Methodology
How every number is calculated, how the model has done, and where it is
weak. Reproducible from scripts/build_ratings.py,
scripts/backtest_model.py and this site's JSON APIs.
#What drives a rating
Every number on this site comes from the same ten things the model measures about a team — none of them the scoreboard, each measured for both teams and both sides of the ball:
- Moving the ball: yards per play, early-down yards per play, success rate (share of plays gaining half the yards needed on 1st down, 70% on 2nd, the full distance on 3rd/4th), third-down conversion, and explosive plays (12+ yards on a run, 19+ on a pass).
- Blocking and protection: offensive-line yards per carry, and behind-the-line plays — tackles for loss plus sacks, penetration from the defense's side.
- Finishing drives: points per trip inside the opponent's 40 (a trip counts once the offense has snapped a play inside the 40, or scores; a punt downed there is not). Moving the ball and finishing are separate skills, kept separate.
- Field position: expected points from where a drive starts.
- Takeaways: turnover-equivalents — turnovers, fourth-down stops and 20-yard-plus punt returns, folded into one number.
Raw stats mostly measure schedule, so every stat is re-rated against the opponents it was earned against (the same least-squares idea as the MPR, one stat at a time, FBS-vs-FBS only). Drives are dropped once the outcome is decided; only the final margin stays unfiltered — it is what the model predicts.
Three inputs are deliberately left out, each of which would move the ratings closer to everyone else's and say less: the market's closing spread (read the line and you reproduce the line), pregame Elo (mostly restates who has been winning) and recruiting rankings (slow-moving, about the roster rather than the games). Roster talent appears only in the MPR's preseason prior (30%). Venue is not one of the ten — the model learns it from the neutral-site flag.
#MPR
The Rankings page's # column is the MPR (McClintock Power Rating): the rating in points versus an average FBS team on a neutral field, used to sort the board all season. Resume keeps its own column and its own order.
#How the MPR is built
MPR is predicted points against an average FBS team on a neutral field, solved over the full schedule — games played and still to come. Subtract two MPRs and add 2.5 for the home team for a projected margin.
Two calibrations keep the early season honest: the live season's adjusted stats are
rescaled onto the spread of the seasons the model trained on (a 1-2 game solve comes out
1.4-3.7x too spread out), and the solved rating is shrunk toward a preseason prior — 70%
last season's final rating, 30% roster talent — weighted 100/75/50/25/0% across a team's
1st-5th game. Each adjusted stat is anchored toward last season's value
(--prior-stat 1), an anchor that never expires. The whole FCS subdivision is
folded into one FCS team: every FCS school's plays pool into a single row
that the stat solve and the ratings solve both fit, so beating an FCS opponent is priced
against a real opponent rating (about 20 points below an average FBS team). Zero
is the average FBS team of 2021-2025, one constant across seasons.
The API's rating field is the MPR; the backtest feed labels the
site model's rows resume. Past-season ratings are a
reconstruction by the current model, not point-in-time predictions.
#Resume (schedule-adjusted)
Resume is how a team has performed relative to what a top-10 team would have done against the same schedule. Every played game counts, FBS and FCS alike (an FCS opponent priced off the folded pool), so it is thinnest in September.
Resume is the per-team mean of those margins, minus the 2021-2025 average: points per game better than the top-10 expectation, capped at ±30 and floored at 7 — a win is worth at least 7 points of MOV, so a one-point escape still counts. A team with no played games falls back to the MPR, not a fake 0.
It is a performance rating, not a betting signal: not fitted to the closing line, built from final scores, so it necessarily tracks win-loss record — by design, not as predictive edge. Opponents are priced off the MPR ladder, so Resume is derived from the MPR rather than independent evidence of it.
#Margin model
The default is diffridge — a sign-symmetric ridge over per-stat
offense/defense differences, which is why the MPR decomposes into Offense
and Defense. It trains on every completed game through the target season's own
played weeks (prior seasons plus this season's earlier weeks), so every prediction is
out-of-sample for the game it predicts; spreads are capped at ±20 before the solve. One
build rates every season, so the ladder is comparable.
#Offense / Defense
The stats are solved with the two sides separate: unknown Off_i and
Def_i are fit together from each game as Off_i − Def_j = stat,
KenPom/SP+ style, so a good offense and a good defense stay different things. The linear,
sign-symmetric margin model then decomposes exactly: projected margin ≈ (Off + Def)
− (opp Off + opp Def).
The Offense and Defense columns are each side's marginal contribution, in points vs an average FBS team. They are two halves of one decomposition — they do not sum to the MPR. The team page's Projected column is a projected final score, not an O/D edge: the pace model sets the game total and the MPR margin (+2.5 at home) splits it. Neither the margin nor the total beats the market (Backtest).
Unit splits (run/pass boxes, QB line) are descriptive rollups through the same solve, never model factors — near-duplicates of the factors they roll up (collinear at 21-54 on the usual variance-inflation scale), so they only say where a team is good.
#Backtest
The MPR is the number that gets graded, and the hard way: a point-in-time replay of what the weekly retrain does. Before week w, re-adjust every stat on games played through week w−1 only, re-train on every completed game before w, predict that week's games, score them, roll on. Week 1 cannot be predicted — the model needs this season's games to rate anyone — so the replay starts at week 2.
Pooled 2021-26 (3,652 games, the window with CFBD advanced box scores behind it): RMSE 16.1, the typical miss on a final margin in points; straight-up 0.713, the winner picked 71.3% of the time; ATS 49.8% — 49.8% pooled (n=2,235), 49.7% on the 1,232 games the model disagreed with the closing spread by ≥3 points, and 49.2% on the 380 it disagreed by ≥7, scored from 2023 on. Beating −110 juice needs about 52.4%. By the model's own yardstick its disagreements with the line are noise, not value.
#Against the market
That ATS number is the honest yardstick, and it is the normal result rather than a failure: no model beats the closing line reliably, and this one is no exception. The independent Prediction Tracker grades the market's own lines alongside dozens of public systems on these same games, every season. Its 2026 board (through 2026-09-27, 215 games): the market's own opening line ranks first straight-up (0.851) with the lowest absolute error (11.3), while the listed systems' spread records run 0.45-0.58 — scattered around a coin flip. This model lands inside that field at 71.3% straight-up and 49.8% ATS pooled over 2021-26.
A betting signal is a different product — not "better than the line on average" (nobody is), but a repeatable, specific disagreement: the few situations where you zig exactly as Vegas zags, often enough to clear −110. That is a hunt for consistent spots, not a higher overall hit rate — the reason the disagreement buckets above are published.
Vegas knows more than anyone. A modern line is a price built to maximise the book's return, not a poll that splits money evenly, and today's computing makes it a forecast in its own right — more capital, data and talent than any public model. Treat it as the ceiling: a model can be right about a game and still be worth nothing as a bet, and any real edge is a small, repeatable niche, never general superiority.
#Limitations
- The factor weights are inside noise. Dropping any one of the ten factor groups moves the mean miss on a margin (MAE) by at most 0.6% of its 13.73-point baseline, and that miss and straight-up accuracy disagree in sign for six of the ten. The factor bars on team pages decompose where a rating came from, not a calibrated weight set.
- The turnover term is the one to distrust in a small sample. Sloppy rate — turnovers plus failed fourth downs — is the least stable of the ten year over year (rank persistence 0.58 from 2025 to 2026, against 0.85 for success rate and 0.53 for behind-the-line plays). It carries 10.0% of the rating's variance; success rate carries the most (16.9%), behind-the-line plays the least (1.9%). Games with a ±3 turnover-equivalent gap miss 14.0 points against 9.8 for the rest.
- Two factors are nearly one. Yards per play and early-down YPP correlate 0.976 (2021-25 adjusted), so their contributions overlap.
- MPR alone mis-ranks records: it carries no win/loss term, so on the 2026 week-4 board 4-win teams sat ~3.8 ranks too low and winless teams ~8.8 too high. Resume ranks those closer to where a record-aware reader expects — the reason it keeps its own column.
- The postgame expected score is a competitive-phase number, not a final score. The stat-implied margin runs about a point heavy in the top bucket — 34.8 projected against 33.8 actual across the 164 games projected at 28 or more — because the factors drop garbage time and the scoreboard does not. A monotone calibration was measured and not shipped: RMSE 15.841 → 15.820, ATS 49.84% → 49.44%, inside noise.
#Win probability, expected wins & record odds
Every projection converts a margin into a probability with a normal distribution over margins:
Expected wins is the sum of a team's per-game win probabilities; the team page's record-odds grid gives the chance of being exactly w–l after each game, from those same probabilities. σ = 17 is the site-wide scale even though the least-squares σ against outcomes is 13, so model and SP+ probabilities compare.
#Postgame replay (expected score & win expectancy)
The game page's MPR card and the Games page's Expected column ask a different question: not "who wins on Saturday" but "what does this box score say happened". The same ridge is re-run on one game's own ten factors, fit on prior seasons only.
Only real possessions count — kickoffs, kneels, overtime rows and sub-minute end-of-half drives that never passed come out. Over 4,155 games (2021-2026, each season calibrated on the three before it) the expected margin correlates 0.77 with the actual. By expected-margin bucket, the predicted home win rate lands within 5 points of the real one in the three buckets inside a touchdown and is 10-14 points off in the five outside it — always compressed toward 50%, because the panel converts a margin that spreads 15.5 points at a fixed 17-point SD. Read it as the competitive-phase expectation, not a final score (Limitations).
#Data & pipeline
Data flows from the CFBD API into a local
DuckDB store and a DuckLake catalog on object storage, through dbt, the ratings build and
a sync to Postgres that this site's Worker serves; season-scoped stats see only games
played up to that week. All model and backtest work runs on the local stores. Every
experiment behind these choices — including the rejected ones — is logged in
docs/MODEL_ITERATIONS.md.