FIFA World Cup 2026 · Final Report

The model called the champion.

PitchProb logged a frozen, timestamped prediction for every match before kickoff — a public audit trail, no hindsight. Below, watch every contender's title odds rise and fall across the whole tournament, then see exactly how those 103 pre-registered forecasts held up.

🏆 The road to the trophy — title odds, every day
Championship probability re-simulated at each stage, using only results known at that point

Each line is one team's live chance of winning the cup. Spain (gold) opened as the model's favourite, slipped to 17% when the group stage tightened, then climbed relentlessly through the knockouts to lift it. Argentina surged to 40% before the final. Every other qualifier is the grey pack behind them.

🇪🇸
Spain
Pre-tournament #1 · Champion
1 – 0
The Final · Jul 19
🇦🇷
Argentina
Pre-tournament #2 · Runner-up
Weeks before the knockouts, Argentina vs Spain was the model's single most-likely final — 19.7% of 62 possible pairings — and it favoured Spain to win it (51%). Both landed. The top two seeds finished exactly first and second.
01

The scorecard

103 matches, each with a frozen pre-kickoff forecast and a final result. Standard metrics for probabilistic forecasts — accuracy on the win/draw/win result, and Brier score on the full probability vector (lower is better).

63.1%
3-way accuracy
65 of 103. Random ≈ 33%, bookmaker favourites ≈ 50–55%.
0.511
Brier score
Matched the pre-tournament backtest (0.508) almost exactly.
73.5%
Knockout accuracy
25 of 34. Sharper without dead-rubber noise.
9.7%
Exact scoreline
10 of 103 called to the exact goals — hard by design.
02

Better than the baselines that matter

Accuracy only means something against what you'd get for free. The model cleared every naive baseline and landed in the range professional, paywalled models occupy.

PitchProb
63.1%
63.1%
Bookmaker favourite
~53%
~53%
Always "home"
47.6%
47.6%
Coin-flip (3-way)
33%
33.3%
03

Where it was wrong: too humble

Calibration is the real test of a probability model — when it says 60%, does that happen 60% of the time? Pooling all 309 forecasts, one pattern was clear: the model's favourites won more often than it predicted. This World Cup had fewer upsets than the odds implied.

40–50% said
Predicted
45%
Actual
68%
50–60% said
Predicted
55%
Actual
72%
60–75% said
Predicted
66%
Actual
80%

Expected Calibration Error came to 0.083 — decent, not immaculate, and exactly the kind of thing a single 104-match tournament is too small to pin down. The honest read: directionally underconfident in strong teams, on a sample too thin for a stronger claim.

04

The honest null result

The one genuinely novel piece was the context layer — nudging each match for stadium heat, rest days and travel distance. So: did it work?

No. On this tournament, the hand-tuned context layer did not improve the forecasts.

63.1%
accuracy — with context
63.1%
accuracy — plain Elo
+0.0010
Brier change (worse)

Across the 80 matches where heat, rest or travel moved the numbers, removing the whole context layer changed accuracy by zero and nudged Brier a hair the wrong way. The adjustments were built from research priors, not fit to data — and it shows. Reporting this is the point: the public track record existed precisely so a null like this couldn't be quietly buried.

05

Is it a paper?

The straight answer: not a traditional methods paper — but there's a genuine, publishable artifact inside this if it's framed for what it actually is.

✕ Not as novel methodology

  • Elo + Poisson + Monte Carlo is textbook — decades old, not a contribution
  • n = 103 is one tournament: too small to statistically validate a model
  • The context layer's headline result is a null — no measured edge to claim

✓ Yes as a pre-registered case study

  • Every forecast was timestamped and frozen before kickoff — a real audit trail, the strongest scientific asset here
  • Honest calibration plus a reported null is exactly the integrity reviewers reward
  • Fits an applied venue (e.g. Sloan Sports Analytics) or a solid preprint / writeup
📄 Preprint
The full write-up is done.
A pre-registered evaluation, the calibration analysis, and the fitted-context null result — including a rest-day effect learned from 15,817 historical matches, with a coefficient statistically indistinguishable from zero.
Read the paper (PDF) →