FIFA World Cup 2026 · Final Report
PitchProb logged a frozen, timestamped prediction for every match before kickoff — a public audit trail, no hindsight. Below, watch every contender's title odds rise and fall across the whole tournament, then see exactly how those 103 pre-registered forecasts held up.
Each line is one team's live chance of winning the cup. Spain (gold) opened as the model's favourite, slipped to 17% when the group stage tightened, then climbed relentlessly through the knockouts to lift it. Argentina surged to 40% before the final. Every other qualifier is the grey pack behind them.
103 matches, each with a frozen pre-kickoff forecast and a final result. Standard metrics for probabilistic forecasts — accuracy on the win/draw/win result, and Brier score on the full probability vector (lower is better).
Accuracy only means something against what you'd get for free. The model cleared every naive baseline and landed in the range professional, paywalled models occupy.
Calibration is the real test of a probability model — when it says 60%, does that happen 60% of the time? Pooling all 309 forecasts, one pattern was clear: the model's favourites won more often than it predicted. This World Cup had fewer upsets than the odds implied.
Expected Calibration Error came to 0.083 — decent, not immaculate, and exactly the kind of thing a single 104-match tournament is too small to pin down. The honest read: directionally underconfident in strong teams, on a sample too thin for a stronger claim.
The one genuinely novel piece was the context layer — nudging each match for stadium heat, rest days and travel distance. So: did it work?
No. On this tournament, the hand-tuned context layer did not improve the forecasts.
Across the 80 matches where heat, rest or travel moved the numbers, removing the whole context layer changed accuracy by zero and nudged Brier a hair the wrong way. The adjustments were built from research priors, not fit to data — and it shows. Reporting this is the point: the public track record existed precisely so a null like this couldn't be quietly buried.
The straight answer: not a traditional methods paper — but there's a genuine, publishable artifact inside this if it's framed for what it actually is.