Methodology

How every number is calculated, how the model has done, and where it is weak. Reproducible from scripts/build_ratings.py, scripts/backtest_model.py and this site's JSON APIs.

Two questions. How good is each team — MPR, the rating every projection here runs on — and what have they actually done — Resume, the schedule-adjusted record, ranked in its own column. Neither beats the market; see the backtest.
What drives a rating MPR How the MPR is built Resume Margin model Offense/Defense Backtest Against the market Limitations Win probability Postgame replay Data pipeline

#What drives a rating

Every number on this site comes from the same ten things the model measures about a team — none of them the scoreboard, each measured for both teams and both sides of the ball:

Raw stats mostly measure schedule, so every stat is re-rated against the opponents it was earned against (the same least-squares idea as the MPR, one stat at a time, FBS-vs-FBS only). Drives are dropped once the outcome is decided; only the final margin stays unfiltered — it is what the model predicts.

Three inputs are deliberately left out, each of which would move the ratings closer to everyone else's and say less: the market's closing spread (read the line and you reproduce the line), pregame Elo (mostly restates who has been winning) and recruiting rankings (slow-moving, about the roster rather than the games). Roster talent appears only in the MPR's preseason prior (30%). Venue is not one of the ten — the model learns it from the neutral-site flag.

#MPR

The Rankings page's # column is the MPR (McClintock Power Rating): the rating in points versus an average FBS team on a neutral field, used to sort the board all season. Resume keeps its own column and its own order.

MPR = points vs an average FBS team, neutral field Resume = mean(margin − top-10 expectation), ranked in its own column

#How the MPR is built

MPR is predicted points against an average FBS team on a neutral field, solved over the full schedule — games played and still to come. Subtract two MPRs and add 2.5 for the home team for a projected margin.

A[i][i] += 1 // team i played another game A[i][j] -= 1 // ... against team j b[i] += value // capped predicted margin x = argmin ‖A·x − b‖² // np.linalg.lstsq: points better than an average team

Two calibrations keep the early season honest: the live season's adjusted stats are rescaled onto the spread of the seasons the model trained on (a 1-2 game solve comes out 1.4-3.7x too spread out), and the solved rating is shrunk toward a preseason prior — 70% last season's final rating, 30% roster talent — weighted 100/75/50/25/0% across a team's 1st-5th game. Each adjusted stat is anchored toward last season's value (--prior-stat 1), an anchor that never expires. The whole FCS subdivision is folded into one FCS team: every FCS school's plays pool into a single row that the stat solve and the ratings solve both fit, so beating an FCS opponent is priced against a real opponent rating (about 20 points below an average FBS team). Zero is the average FBS team of 2021-2025, one constant across seasons.

The API's rating field is the MPR; the backtest feed labels the site model's rows resume. Past-season ratings are a reconstruction by the current model, not point-in-time predictions.

#Resume (schedule-adjusted)

Resume is how a team has performed relative to what a top-10 team would have done against the same schedule. Every played game counts, FBS and FCS alike (an FCS opponent priced off the folded pool), so it is thinnest in September.

baseline = mean(MPR of the top 10 teams) // --top-n, default 10 expected(team, g) = baseline − MPR(opp) + 2.5 home / −2.5 away / 0 neutral resume_mov(g) = clamp(MOV(g), −30, 30) − expected(team, g) // wins floored at 7

Resume is the per-team mean of those margins, minus the 2021-2025 average: points per game better than the top-10 expectation, capped at ±30 and floored at 7 — a win is worth at least 7 points of MOV, so a one-point escape still counts. A team with no played games falls back to the MPR, not a fake 0.

It is a performance rating, not a betting signal: not fitted to the closing line, built from final scores, so it necessarily tracks win-loss record — by design, not as predictive edge. Opponents are priced off the MPR ladder, so Resume is derived from the MPR rather than independent evidence of it.

#Margin model

margin_home = f(off_stats_home, def_stats_home, off_stats_away, def_stats_away, neutral_site)

The default is diffridge — a sign-symmetric ridge over per-stat offense/defense differences, which is why the MPR decomposes into Offense and Defense. It trains on every completed game through the target season's own played weeks (prior seasons plus this season's earlier weeks), so every prediction is out-of-sample for the game it predicts; spreads are capped at ±20 before the solve. One build rates every season, so the ladder is comparable.

#Offense / Defense

The stats are solved with the two sides separate: unknown Off_i and Def_i are fit together from each game as Off_i − Def_j = stat, KenPom/SP+ style, so a good offense and a good defense stay different things. The linear, sign-symmetric margin model then decomposes exactly: projected margin ≈ (Off + Def) − (opp Off + opp Def).

The Offense and Defense columns are each side's marginal contribution, in points vs an average FBS team. They are two halves of one decomposition — they do not sum to the MPR. The team page's Projected column is a projected final score, not an O/D edge: the pace model sets the game total and the MPR margin (+2.5 at home) splits it. Neither the margin nor the total beats the market (Backtest).

Unit splits (run/pass boxes, QB line) are descriptive rollups through the same solve, never model factors — near-duplicates of the factors they roll up (collinear at 21-54 on the usual variance-inflation scale), so they only say where a team is good.

#Backtest

The MPR is the number that gets graded, and the hard way: a point-in-time replay of what the weekly retrain does. Before week w, re-adjust every stat on games played through week w−1 only, re-train on every completed game before w, predict that week's games, score them, roll on. Week 1 cannot be predicted — the model needs this season's games to rate anyone — so the replay starts at week 2.

Pooled 2021-26 (3,652 games, the window with CFBD advanced box scores behind it): RMSE 16.1, the typical miss on a final margin in points; straight-up 0.713, the winner picked 71.3% of the time; ATS 49.8% — 49.8% pooled (n=2,235), 49.7% on the 1,232 games the model disagreed with the closing spread by ≥3 points, and 49.2% on the 380 it disagreed by ≥7, scored from 2023 on. Beating −110 juice needs about 52.4%. By the model's own yardstick its disagreements with the line are noise, not value.

#Against the market

That ATS number is the honest yardstick, and it is the normal result rather than a failure: no model beats the closing line reliably, and this one is no exception. The independent Prediction Tracker grades the market's own lines alongside dozens of public systems on these same games, every season. Its 2026 board (through 2026-09-27, 215 games): the market's own opening line ranks first straight-up (0.851) with the lowest absolute error (11.3), while the listed systems' spread records run 0.45-0.58 — scattered around a coin flip. This model lands inside that field at 71.3% straight-up and 49.8% ATS pooled over 2021-26.

A betting signal is a different product — not "better than the line on average" (nobody is), but a repeatable, specific disagreement: the few situations where you zig exactly as Vegas zags, often enough to clear −110. That is a hunt for consistent spots, not a higher overall hit rate — the reason the disagreement buckets above are published.

Vegas knows more than anyone. A modern line is a price built to maximise the book's return, not a poll that splits money evenly, and today's computing makes it a forecast in its own right — more capital, data and talent than any public model. Treat it as the ceiling: a model can be right about a game and still be worth nothing as a bet, and any real edge is a small, repeatable niche, never general superiority.

Loading backtest…

#Limitations

#Win probability, expected wins & record odds

Every projection converts a margin into a probability with a normal distribution over margins:

margin = Power_home − Power_away + 2.5 // 0 neutral; SP+ tables use SP+ instead // non-FBS opponents priced at SP+ −35 win% = Φ(margin / 17) // normal CDF, σ = 17 points

Expected wins is the sum of a team's per-game win probabilities; the team page's record-odds grid gives the chance of being exactly w–l after each game, from those same probabilities. σ = 17 is the site-wide scale even though the least-squares σ against outcomes is 13, so model and SP+ probabilities compare.

#Postgame replay (expected score & win expectancy)

The game page's MPR card and the Games page's Expected column ask a different question: not "who wins on Saturday" but "what does this box score say happened". The same ridge is re-run on one game's own ten factors, fit on prior seasons only.

raw = (off + def)_home − (off + def)_away margin = slope · raw + base (+ neutral adj) // slope/base/neutral fit on the 3 prior seasons, never the season shown total = league_pts_per_drive · filtered_drives // the game's own pace score = (total ± margin) / 2 win% = Φ(margin / 17)

Only real possessions count — kickoffs, kneels, overtime rows and sub-minute end-of-half drives that never passed come out. Over 4,155 games (2021-2026, each season calibrated on the three before it) the expected margin correlates 0.77 with the actual. By expected-margin bucket, the predicted home win rate lands within 5 points of the real one in the three buckets inside a touchdown and is 10-14 points off in the five outside it — always compressed toward 50%, because the panel converts a margin that spreads 15.5 points at a fixed 17-point SD. Read it as the competitive-phase expectation, not a final score (Limitations).

#Data & pipeline

Data flows from the CFBD API into a local DuckDB store and a DuckLake catalog on object storage, through dbt, the ratings build and a sync to Postgres that this site's Worker serves; season-scoped stats see only games played up to that week. All model and backtest work runs on the local stores. Every experiment behind these choices — including the rejected ones — is logged in docs/MODEL_ITERATIONS.md.