Could AI predict the World Cup? 6 flagship models called every 2026 knockout
match blind - who goes through, the score, the goalscorers - and we graded all
32 against reality. The tournament is over. Here's the full story.
32 matches graded6 AI models2 point winning margin
The contenders
Claude Opus 4.8
Claude Fable 5
GPT-5.5
Gemini 3.1 Pro
GLM-5.2
Qwen 3.7 Max
Final standings
32/32 matches
🥇GLM-5.2 Champion309
🥈Claude Fable 5307
🥉Claude Opus 4.8296
4Qwen 3.7 Max288
5Gemini 3.1 Pro286
6GPT-5.5284
⚑⚑Ranking favourite239
Tournament complete · Spain won the World Cup · GLM-5.2 won the Bench
How it was a fair test:
every model got the same scouting report (team ratings, recent form, head-to-head and the full squad)
and no live internet, so it predicted blind, like a pundit
before kickoff. Knockout picks were saved before each match was played.
What we learned
Six flagship models, 32 knockout matches, roughly 200 graded predictions. The headline result is close - the findings underneath it aren't.
The result
GLM-5.2 wins 309-307 - by out-gaming the rules, not out-predicting the field
The catch
On core match calls, "always pick the favourite" beat every single AI
The surprise
The AIs were a herd: unanimous on the winner in 31 of 34 matches
GLM-5.2 took the title by 2 points - and its edge was almost entirely strategic.
It named the maximum three goalscorers in nearly every match, even when its own scoreline only predicted two goals
(it did this in 14 of 34 picks; no other model did it more than 3 times). Its scorer hit rate was the same as
everyone else's (~26-29%) - it simply took more shots. Under rules with no penalty for wrong guesses, that was the winning move.
2 · Remove goalscorers, and the dumb baseline wins
On the core calls - who goes through, the scoreline, penalties - not one AI beat
"always pick the higher-ranked team" (239 core points vs the best AI's 235). The models' entire winning margin came from
the goalscorer category, which the baseline structurally can't play. AI added real value in squad-level knowledge; at pure
match-calling, football's rankings already know what the models know.
3 · Thinking harder didn't help
Claude Opus 4.8 spent ~2,700 reasoning tokens per pick - 4.6x more than
GLM-5.2's ~580 - and finished 13 points behind it. Across the field, reasoning effort showed no relationship with accuracy:
the six models landed within one match of each other on progression calls (25-26 out of 32).
4 · Upsets beat everyone, together
Five matches fooled all six models at once: Morocco and Switzerland going through on penalties, Norway's giant-killing runs past
Ivory Coast and Brazil, and a 10-goal third-place classic. When the favourite fell, the AIs fell with it - every model's
misses were the same misses. Genuine football surprise remains unpredicted, by silicon and rankings alike.
Final league table
Hover any column header for what it means.
Showing total points - rotate or open on desktop for the full scoring breakdown.
#
Model
Played
Through
Exact score
Result
Goalscorers
Penalties
Red card
Points
Calibration
🥇
GLM-5.2
32
130
15
32
84
48
0
309
0.14
🥈
Claude Fable 5
32
130
15
36
72
54
0
307
0.13
🥉
Claude Opus 4.8
32
125
15
36
66
54
0
296
0.14
4
Qwen 3.7 Max
32
125
12
34
69
48
0
288
0.15
5
Gemini 3.1 Pro
32
130
18
30
60
48
0
286
0.14
6
GPT-5.5
32
125
21
26
66
46
0
284
0.14
🥇
Claude Fable 5
32
130
15
32
84
50
0
311
0.14
🥈
Claude Opus 4.8
32
125
18
34
75
54
0
306
0.14
🥉
Gemini 3.1 Pro
32
130
21
28
75
48
0
302
0.14
4
GLM-5.2
32
130
15
28
81
46
0
300
0.15
5
GPT-5.5
32
125
24
26
72
48
0
295
0.14
6
Qwen 3.7 Max
32
125
15
28
60
46
0
274
0.15
springprompt.com/evals/world-cup
Points = the sum of the breakdown columns.
Calibration measures how honest a model's confidence is: if it says 70% and is right 70% of the time, that's perfectly calibrated (lower is better, it's a Brier score where 0 is perfect).
AI models only - the non-AI baseline can't play the goalscorer category, so it gets its own comparison below.
Experimental run: same models and scoring, but from the round of 16 onward each model
was shown all earlier results plus its own prior picks and scores before predicting. Retrospective (not locked pre-kickoff), so not
the official record - full analysis below. Claude Fable 5
takes this version; the frozen-run champion GLM-5.2 drops to fourth.
⚑
The sanity check: AI vs "just pick the favourite"
Would a one-line rule - always back the higher-ranked team - have kept up with six
flagship AIs? It can't name goalscorers, so the fair comparison is on the core calls only (through +5, exact score +3, result +2,
penalties +2, red cards +5), with goalscorer points stripped from everyone. On those terms, the one-line rule wins:
#
Model
Core points
Goalscorer points (excluded)
Official total
⚑
Ranking favouriteBaseline
239
-
239
2
Claude Fable 5
235
+72
307
3
Claude Opus 4.8
230
+66
296
4
Gemini 3.1 Pro
226
+60
286
5
GLM-5.2
225
+84
309
6
Qwen 3.7 Max
219
+69
288
7
GPT-5.5
218
+66
284
springprompt.com/evals/world-cup
Read it either way: the AIs' knowledge of squads and likely scorers is real, earned value the baseline can't reach - or the AIs'
entire lead over a one-line heuristic lives in a single generous category. Both are true. That tension is the benchmark's most
honest finding.
🎙️
And vs a human pundit?
For a human yardstick we graded every knockout prediction that Chris Sutton
published on BBC Sport, against the same results. Through the quarter-finals, the human, the AI consensus and the dumb favourite
baseline were indistinguishable - then they split at the exact moment it mattered.
Round (winners called)
🎙️ Sutton
🤖 AI consensus
⚑ Favourite
Round of 32
13/16
13/16
13/16
Round of 16
6/8
6/8
6/8
Quarter-final
4/4
4/4
4/4
Semi-final
0/2
2/2
2/2
Final
0/1
1/1
1/1
Total
23/31
26/31
26/31
Level through the quarter-finals (23/28 apiece). The whole difference is the highlighted semi-final and final rows.
springprompt.com/evals/world-cup
Where the human was irreplaceable
Sutton alone called both of Norway's giant-killings - the shootout win over Ivory Coast and the round-of-16 upset of Brazil. Both were matches every one of the six AI models got wrong.
Reading a specific underdog's momentum through a shootout is exactly the judgement no ratings-fed model produced - twice.
Where the human lost it
At the business end he trusted reputation over form: he backed France and England in the semi-finals (both lost) and picked France to win a final France never reached. The AIs' habit of backing the favourite banked all three.
Boring beat brilliant at the finish. The AIs never had a great call in them - and never needed one.
The verdict: a top human pundit, six flagship AIs, and one line of code all landed within
three correct calls of each other over 31 matches - and the machines' edge was pure discipline, not insight. If you want a forecast
that's reliable, the models (and the baseline) beat the human. If you want the one call that wins the office sweepstake,
you still want the human.
🧠
What if they could learn?
The official benchmark froze every briefing at the eve of the tournament - so a model predicting the final knew nothing of the
rounds before it. So we ran a second, experimental version: same models, same frozen scouting data, but before each match from the
round of 16 on, we showed every model all earlier results and its own prior picks, scores and rationales.
Rolling memory instead of a fixed snapshot. Could they learn on the fly? The answer is stranger than yes or no.
Given full memory of the tournament, the models changed which team they tipped in 0 of 96 later-round picks.
Not once. Shown that Norway had just knocked out two higher-rated sides, every model still went back to the pre-tournament ratings
next time. Their scorelines and scorer picks shifted, and confidence dipped slightly (-2.3pp on average) -
but the core judgement, who advances, never moved. Memory changed their arithmetic, not their mind.
The standings with rolling memory
Same scoring, same matches - re-ranked. Delta is versus the official frozen run.
#
Model
Official
Learning
Change
🥇
Claude Fable 5
307
311
+4
🥈
Claude Opus 4.8
296
306
+10
🥉
Gemini 3.1 Pro
286
302
+16
4
GLM-5.2
309
300
-9
5
GPT-5.5
284
295
+11
6
Qwen 3.7 Max
288
274
-14
springprompt.com/evals/world-cup
A new champion - and a fallen one
With memory, Claude Fable 5 takes the title (311) -
while the frozen-run champion, GLM-5.2, drops to fourth. The model that won by
maximising goalscorer guesses gained nothing from context; the steadier forecasters used it best.
Gemini 3.1 Pro improved most (+16); Qwen 3.7 Max
fell furthest (-14).
Memory still didn't beat the baseline
On core points, even the best learning-run AI (Claude Opus 4.8, 231) still finished
behind the favourite baseline (239). Handing the models
the whole tournament to study did not close the gap to a one-line rule. That is the most sobering result on this page.
Experimental, and not part of the official standings: these picks were generated retrospectively (with same-day results withheld from each other), so unlike the frozen run they aren't timestamped before kickoff. Round-of-32 picks are the models' original ones. Treat it as a controlled probe of whether context helps, not a second leaderboard of record.
Title race
Each model's position after every knockout match (1 = leading). Hover a match to see the full order.
springprompt.com/evals/world-cup
How they played the game
The rules allowed up to three goalscorer picks per match, worth +3 each with no penalty for misses. What each model did with
that freedom is the clearest personality test in the benchmark - and, it turned out, the decisive one.
Scorer picks vs own predicted goals
"Over-named" = named more scorers than goals in its own scoreline. Consistency vs opportunism.
Model
Avg scorers named
Goals in its scoreline prediction
Named more scorers than goals
GLM-5.2
2.97
2.59
14/34
Claude Fable 5
2.74
2.68
2/34
Claude Opus 4.8
2.68
2.59
3/34
Qwen 3.7 Max
2.53
2.47
2/34
GPT-5.5
2.5
2.44
2/34
Gemini 3.1 Pro
2.26
2.56
0/34
springprompt.com/evals/world-cup
Reasoning effort vs result
Average completion tokens per pick (includes each model's internal reasoning at max effort).
Model
Avg tokens / pick
Final points
Claude Opus 4.8
2,687
296
Claude Fable 5
1,186
307
Qwen 3.7 Max
1,182
288
GPT-5.5
649
284
Gemini 3.1 Pro
628
286
GLM-5.2
583
309
The most economical thinker won; the deepest thinker finished third. Effort and accuracy were uncorrelated here.
springprompt.com/evals/world-cup
The herd: how often the AIs just backed the favourite
All six models picked the same winner in 31 of 34 fixtures.
Across ~200 picks there were only 9 contrarian calls (against the ranking favourite) in total - and 4 matches went to penalties, a mechanism the models almost never predicted.
Model
Agreed with favourite
Contrarian picks (won)
Penalties predicted
Avg confidence
Gemini 3.1 Pro
34/34
0 (0)
6
72%
GPT-5.5
33/34
1 (0)
9
70%
Qwen 3.7 Max
33/34
1 (0)
4
73%
GLM-5.2
32/34
2 (1)
6
70%
Claude Fable 5
32/34
2 (1)
1
71%
Claude Opus 4.8
31/34
3 (1)
1
71%
Gemini 3.1 Pro never once disagreed with the rankings on who goes through - its progression picks were, in effect, the baseline
with a squad list attached. The models' confidence sat at a uniform 70-73% regardless of the match, and the boldest behaviour in
the field was three contrarian calls (Claude Opus 4.8), of which one landed.
springprompt.com/evals/world-cup
The pundits, profiled
Six models, six distinct characters. Each profile ends with the model's own (pre-match, blind) rationale for the final -
Spain 1-0 Argentina, as it turned out. All six tipped Spain.
🥇 GLM-5.2
The game-player
309
Read the rules better than it read the football. Named the maximum three scorers almost every match - even when its own scoreline said two goals - and spent the fewest reasoning tokens in the field doing it. Its scorer hit rate was ordinary; its shot volume was not. Won the title on strategy, by 2 points.
Winners called26/32 matchesExact scorelines5 of 32Goalscorer picks28 of 95 hit (29%)Reasoning per pick583 tokens
"Spain's superior Elo rating and midfield control (Rodri, Pedri, Gavi) give them a marginal edge, but finals are tight and Argentina's perfect recent form plus Messi's match-winning ability makes this extremely close — expect a 1-1 draw after 90 minutes with Spain edging it in extra time."
- its call on the final, made before kickoff
🥈 Claude Fable 5
The forecaster
307
The best pure predictor in the field. Best-calibrated confidence (0.135 Brier), internally coherent picks (scorers matched its scoreline in 32 of 34 matches), and top AI on the baseline-fair core table. Joined the benchmark after the round of 32 had begun and still nearly took the crown.
Winners called26/32 matchesExact scorelines5 of 32Goalscorer picks24 of 88 hit (27%)Reasoning per pick1,186 tokens
"Spain hold the Elo edge (2157 vs 2115) and superior squad depth in midfield with Rodri and Pedri controlling tempo, though Argentina's flawless recent form makes this a coin-flip-adjacent final. Expect a tight, high-quality match decided by Spain's wide threats, with a meaningful chance it goes beyond 90 minutes."
- its call on the final, made before kickoff
🥉 Claude Opus 4.8
The deep thinker
296
Out-reasoned everyone, out-predicted no one. ~2,700 tokens of deliberation per pick - four times the winner's spend - bought steady, disciplined picks and a podium, but no measurable accuracy edge over faster rivals. The benchmark's clearest evidence that longer thinking has diminishing returns on chaotic domains.
Winners called25/32 matchesExact scorelines5 of 32Goalscorer picks22 of 86 hit (26%)Reasoning per pick2,687 tokens
"Spain's higher Elo and elite midfield core (Rodri, Pedri) plus Yamal's attacking threat give a narrow edge over an in-form Argentina tested mostly against weaker friendly opposition. Expect a tense, tight final that Spain edges, though extra time is a genuine possibility given both teams' quality."
- its call on the final, made before kickoff
4th Qwen 3.7 Max
Boom or bust
288
The volatility play. Owned more of the tournament's biggest single-match hauls than anyone (three +15s) and surged late - but paired it with the field's worst calibration (0.149) and the most scoreless weekends. When it was right, it was very right.
Winners called25/32 matchesExact scorelines4 of 32Goalscorer picks23 of 82 hit (28%)Reasoning per pick1,182 tokens
"Spain's superior Elo and midfield control with Rodri and Pedri give them the edge over an Argentine side that has faced much weaker recent opposition. A narrow victory in regular time reflects the high stakes and tactical discipline expected in a World Cup final."
- its call on the final, made before kickoff
5th Gemini 3.1 Pro
The purist
286
The only model that never named more scorers than its own predicted goals - not once in 34 matches - and took the fewest scorer shots as a result. Principled, well-calibrated, and the author of the single best-scored prediction of the tournament (+16 on France-Morocco). The rules didn't reward principle.
Winners called26/32 matchesExact scorelines6 of 32Goalscorer picks20 of 72 hit (28%)Reasoning per pick628 tokens
"Spain's superior Elo rating and midfield control (anchored by Rodri) give them a slight edge over Argentina's tactical pragmatism. However, World Cup finals are notoriously tight, making a 1-1 draw in normal time highly likely before Spain's depth seals it in extra time."
- its call on the final, made before kickoff
6th GPT-5.5
The early pacesetter
284
Led the race through the round of 32 on the strength of exact scorelines - it called more of them (7) than any other model. Faded as the goalscorer category grew decisive, and finished with the lowest core-points total. A specialist outpointed by a generalist's rulebook.
Winners called25/32 matchesExact scorelines7 of 32Goalscorer picks22 of 79 hit (28%)Reasoning per pick649 tokens
"Spain have the small Elo edge and more control-oriented midfield depth, but Argentina's recent scoring form and elite forwards make a 90-minute stalemate plausible. I lean Spain to edge the final after extra time or penalties."
- its call on the final, made before kickoff
The matches that fooled everyone
All six models wrong, together
Every AI backed the loser in these 5 matches - and so did the ranking baseline. Each one is a lesson in what the models couldn't see.
Netherlands 1-1 Morocco pens· Round of 32 · Morocco through
Every model backed the Dutch at 62-72% confidence and most predicted 2-1. The 90 minutes actually finished level - which several models' scorelines implied - but none followed that thought to its conclusion: a shootout, which Morocco won. The models predicted the draw and still refused to predict its consequence.
Ivory Coast 1-2 Norway · Round of 32 · Norway through
The Elo gap said Ivory Coast (2062 v 1827), and all six models said Ivory Coast, at up to 78%. Norway's Haaland-Nusa counterattack said otherwise. First warning that the briefing's ratings underweighted Norway - a warning nobody could act on.
Brazil 1-2 Norway · Round of 16 · Norway through
One round later, same mistake, same team. The new matchup told every model that Norway had advanced, but they weren't shown the Ivory Coast result or their own earlier picks. With ratings and form still frozen pre-tournament, all six went back to the same priors and picked Brazil. The BBC's human pundit called this one exactly.
Switzerland 0-0 Colombia pens· Round of 16 · Switzerland through
The tightest tie of the round, and the models knew it - confidence was the lowest of the tournament (52-62%). They still all landed on the same side, Colombia, and a goalless slog went to Switzerland on penalties. A coin-flip where all six coins came up identically wrong.
France 4-6 England · Third-place play-off · England through
The great absurdity of the tournament. All six models predicted France 2-1, a sensible scoreline for a third-place game between well-matched sides. Reality produced ten goals, an England win, and the single most chaotic match of the World Cup. No ratings-based system predicts a 4-6; nothing in the briefing hinted at it. Some things remain gloriously unforecastable.
The best single calls
Highest-scoring individual match predictions of the tournament, and where the points actually came from.
+16Gemini 3.1 Proon France v Morocco (2-0)
The tournament's best single prediction: right team, exact 2-0 scoreline, and both goalscorers named. Everything the benchmark asks for, in one pick - from the model that finished fifth.
+15Claude Opus 4.8on Argentina v Cape Verde (3-2)
+15GPT-5.5on Argentina v Cape Verde (3-2)
+15Gemini 3.1 Proon Argentina v Cape Verde (3-2)
Third of Gemini's three +15s - all built on exact scorelines plus multiple correct scorers.
+15GLM-5.2on Argentina v Cape Verde (3-2)
+15Qwen 3.7 Maxon Argentina v Cape Verde (3-2)
A 3-2 goal-fest predicted almost goal for goal.
Note who owns this list: Gemini and Qwen, who finished fifth and fourth. Peak insight and total points measured different things in this game - the steady rule-optimisers won the table while the precision specialists won the highlights.
The full match-by-match record
Every pick by every model, for all 32 matches. Expand a round to see it.
Reading a pick: = team tipped to go through = predicted goalscorers correct · wrongconf = the model's confidence
The roster was simple: the flagship reasoning model from each major provider, as of the
start of the knockout rounds - Claude Opus 4.8 (Anthropic), GPT-5.5 (OpenAI), Gemini 3.1 Pro (Google), GLM-5.2 (Z.ai) and
Qwen 3.7 Max (Alibaba) - every one at its maximum reasoning setting. We began at the round of 32 rather than the group stage for
no grander reason than that's when we built the benchmark; with 32 knockout matches and sudden-death stakes, it turned out to be
the better test anyway.
Claude Fable 5 joined a few matches in, when it became available - early enough
that backfilling its round-of-32 picks under the same blind conditions (pre-tournament training cut-off, no web access) was
defensible. We kept Opus 4.8 in the field rather than swapping it out: it gave us a same-provider comparison point, and its
predictions were already locked. Models that launched later in the tournament didn't get the same treatment - see the FAQ.
One quiet subplot worth naming: the two open-weights models finished first and fourth.
GLM-5.2 won the whole thing and Qwen 3.7 Max out-pointed two Western proprietary flagships - for a fraction of the per-token cost.
Whatever this benchmark measures, it isn't a proprietary moat.
What this means if you want AI to predict things
1 · Always run a dumb baseline first
The single most useful thing in this benchmark cost one line of code. If we hadn't scored "always pick the favourite," the headline would have been "AI beats the World Cup" - and it would have been wrong. Before trusting an AI forecast, know what the naive answer scores.
2 · Expect the herd, not the oracle
Six different flagships, one briefing, near-identical picks: unanimous in 31 of 34 matches, uniform ~70% confidence, nine contrarian calls in ~200. If your use case needs someone to spot the upset, current models - given identical inputs - won't be that someone. Diversity of data, not of model vendor, is what would have changed these picks.
3 · The value is in the detail, not the verdict
Where the AIs genuinely beat the baseline was squad-level knowledge: naming likely goalscorers at a steady ~27% hit rate. That's real, useful signal a ranking can't produce. Use AI forecasting where breadth of specific knowledge matters, not where the answer is one bit.
4 · Models optimise your rules, not your intent
GLM-5.2 won because our scoring made three scorer guesses free. It didn't cheat - it read the incentive perfectly, which is exactly what these systems do. Any scoring rule you give an AI, audit it as if the AI will find the loophole. It will.
What we'll change next time
This is Spring Prompt's second forecasting benchmark (after PredictTheWeek), and the next one will inherit its scars:
Close the loopholes - or score them. Uncapped goalscorer picks decided this benchmark. Next time either wrong guesses cost points, or rule-exploitation gets measured openly as its own dimension. Both are interesting; ambiguity isn't.
Reward contrarian value, not just accuracy. A correct upset call is worth more information than a correct favourite call. Scoring should reflect that - odds-weighted points would have made Sutton's Norway call worth its weight, and might coax the herd apart.
Save the full reasoning traces. We kept each model's two-sentence rationale and its token spend, and that alone yielded findings. Complete chains of thought would have let us ask why six models converge - next time they're archived from day one.
Let the briefing evolve - and preserve each model's memory. The original test never showed models earlier tournament results or fed back their own picks and scores; only the next matchup revealed who had advanced. Next time, every round should include the results so far plus each model's personal prediction record, testing whether it updates after surprises instead of repeating the same mistake. We ran that rolling-learning variant locally after the tournament to measure the difference.
FAQ
Why was Claude Fable 5 added late, but newer models weren't?
Fable 5 became available only two or three matches into the round of 32, so backfilling its handful of missed picks under identical blind conditions barely stretched the rules. Models that launched deeper into the tournament (e.g. GPT-5.6) would have needed retro-predictions for a large share of already-played matches - technically possible under the same blind protocol, but we judged the integrity risk to the results not worth one more row in the table.
Why start at the knockouts and skip the group stage?
Honestly: timing - the benchmark was built as the group stage ended. The knockout format turned out to be the sharper test (every match has a winner, and penalties add a category the models demonstrably can't call).
Did any model know the results?
No. Every model's training cut-off predates the tournament, web access was disabled, and the briefing was frozen to the eve of the tournament. Knockout picks were also timestamped before kickoff. A model could no more look up a result than a pundit in a studio could.
Can I get the data?
The full prediction set (~200 graded picks with rationales, confidence values and token counts) is preserved. Email us and we'll share it.
Methodology
The blind test
Each model received an identical scouting report per match: Elo rating, recent form, head-to-head record and the full 26-man
squad - all frozen to 2026-06-11, the eve of the tournament - and no internet
access. Because every model's training cut-off predates the tournament and web access was disabled, every pick was genuinely blind,
whether made before or after a match was played. Knockout picks were additionally timestamped before kickoff. All models ran at
their maximum reasoning setting.
Scoring
Right team through +5 · exact 90-minute score +3 · right result without
the exact score +2 · each correct goalscorer (up to three named) +3 ·
correct extra-time/penalties call +2 · correct red-card call +5.
Calibration is a separate Brier score on each model's stated probabilities (0 = perfect, lower is better).
Known quirks, disclosed
Goalscorer picks carried no penalty for misses - naming three was strictly dominant, and the winner exploited exactly that. We left the rule as designed and report the goalscorer-free table above.
The Ranking favourite baseline structurally can't name goalscorers, so compare it on the core table, not the headline one.
Claude Fable 5 joined after the round of 32 had started; its early-round picks were made retrospectively but under the same blind conditions (pre-tournament training cut-off, no web).
Round-of-32 group-stage seeding, injuries and suspensions were not in the briefing; models predicted from pre-tournament knowledge only.
Data
Results, scores and goalscorers were graded from ESPN's public scoreboard with the open
international results dataset (CC0)
as fallback; team ratings from eloratings.net;
squads from Wikipedia. The full prediction dataset (~200 graded picks with rationales and token counts) is preserved as-is;
the tournament is over and these figures are final.
Full time
That's the whistle on World Cup Bench.
Spring Prompt runs benchmarks like this every week - testing what AI models can actually do on real tasks,
with baselines that keep the answers honest. The next experiment is already being designed with this one's lessons.