Confirm Action

Are you sure you want to proceed?

Spring Prompt · AI Benchmark · Final report

World Cup Bench

Could AI predict the World Cup? 6 flagship models called every 2026 knockout match blind - who goes through, the score, the goalscorers - and we graded all 32 against reality. The tournament is over. Here's the full story.

32 matches graded 6 AI models 2 point winning margin

The contenders

Claude Opus 4.8 Claude Opus 4.8
Claude Fable 5 Claude Fable 5
GPT-5.5 GPT-5.5
Gemini 3.1 Pro Gemini 3.1 Pro
GLM-5.2 GLM-5.2
Qwen 3.7 Max Qwen 3.7 Max

Final standings

32/32 matches
🥇 GLM-5.2 Champion 309
🥈 Claude Fable 5 307
🥉 Claude Opus 4.8 296
4 Qwen 3.7 Max 288
5 Gemini 3.1 Pro 286
6 GPT-5.5 284
Ranking favourite 239

Tournament complete · Spain won the World Cup · GLM-5.2 won the Bench

How it was a fair test: every model got the same scouting report (team ratings, recent form, head-to-head and the full squad) and no live internet, so it predicted blind, like a pundit before kickoff. Knockout picks were saved before each match was played.

What we learned

Six flagship models, 32 knockout matches, roughly 200 graded predictions. The headline result is close - the findings underneath it aren't.

The result

GLM-5.2 wins 309-307 - by out-gaming the rules, not out-predicting the field

The catch

On core match calls, "always pick the favourite" beat every single AI

The surprise

The AIs were a herd: unanimous on the winner in 31 of 34 matches

The experiment

Shown every prior result, the models changed 0 of 96 winner picks - but the podium reshuffled

1 · The winner played the rules, not the football

GLM-5.2 took the title by 2 points - and its edge was almost entirely strategic. It named the maximum three goalscorers in nearly every match, even when its own scoreline only predicted two goals (it did this in 14 of 34 picks; no other model did it more than 3 times). Its scorer hit rate was the same as everyone else's (~26-29%) - it simply took more shots. Under rules with no penalty for wrong guesses, that was the winning move.

2 · Remove goalscorers, and the dumb baseline wins

On the core calls - who goes through, the scoreline, penalties - not one AI beat "always pick the higher-ranked team" (239 core points vs the best AI's 235). The models' entire winning margin came from the goalscorer category, which the baseline structurally can't play. AI added real value in squad-level knowledge; at pure match-calling, football's rankings already know what the models know.

3 · Thinking harder didn't help

Claude Opus 4.8 spent ~2,700 reasoning tokens per pick - 4.6x more than GLM-5.2's ~580 - and finished 13 points behind it. Across the field, reasoning effort showed no relationship with accuracy: the six models landed within one match of each other on progression calls (25-26 out of 32).

4 · Upsets beat everyone, together

Five matches fooled all six models at once: Morocco and Switzerland going through on penalties, Norway's giant-killing runs past Ivory Coast and Brazil, and a 10-goal third-place classic. When the favourite fell, the AIs fell with it - every model's misses were the same misses. Genuine football surprise remains unpredicted, by silicon and rankings alike.

Final league table

Showing total points - rotate or open on desktop for the full scoring breakdown.

# Model Played Points
🥇 GLM-5.2 32 309
🥈 Claude Fable 5 32 307
🥉 Claude Opus 4.8 32 296
4 Qwen 3.7 Max 32 288
5 Gemini 3.1 Pro 32 286
6 GPT-5.5 32 284
🥇 Claude Fable 5 32 311
🥈 Claude Opus 4.8 32 306
🥉 Gemini 3.1 Pro 32 302
4 GLM-5.2 32 300
5 GPT-5.5 32 295
6 Qwen 3.7 Max 32 274
springprompt.com/evals/world-cup

Points = the sum of the breakdown columns. Calibration measures how honest a model's confidence is: if it says 70% and is right 70% of the time, that's perfectly calibrated (lower is better, it's a Brier score where 0 is perfect). AI models only - the non-AI baseline can't play the goalscorer category, so it gets its own comparison below.

Experimental run: same models and scoring, but from the round of 16 onward each model was shown all earlier results plus its own prior picks and scores before predicting. Retrospective (not locked pre-kickoff), so not the official record - full analysis below. Claude Fable 5 takes this version; the frozen-run champion GLM-5.2 drops to fourth.

The sanity check: AI vs "just pick the favourite"

Would a one-line rule - always back the higher-ranked team - have kept up with six flagship AIs? It can't name goalscorers, so the fair comparison is on the core calls only (through +5, exact score +3, result +2, penalties +2, red cards +5), with goalscorer points stripped from everyone. On those terms, the one-line rule wins:

# Model Core points Goalscorer points (excluded) Official total
Ranking favourite Baseline 239 - 239
2 Claude Fable 5 235 +72 307
3 Claude Opus 4.8 230 +66 296
4 Gemini 3.1 Pro 226 +60 286
5 GLM-5.2 225 +84 309
6 Qwen 3.7 Max 219 +69 288
7 GPT-5.5 218 +66 284
springprompt.com/evals/world-cup

Read it either way: the AIs' knowledge of squads and likely scorers is real, earned value the baseline can't reach - or the AIs' entire lead over a one-line heuristic lives in a single generous category. Both are true. That tension is the benchmark's most honest finding.

🎙️

And vs a human pundit?

For a human yardstick we graded every knockout prediction that Chris Sutton published on BBC Sport, against the same results. Through the quarter-finals, the human, the AI consensus and the dumb favourite baseline were indistinguishable - then they split at the exact moment it mattered.

Round (winners called) 🎙️ Sutton 🤖 AI consensus ⚑ Favourite
Round of 32 13/16 13/16 13/16
Round of 16 6/8 6/8 6/8
Quarter-final 4/4 4/4 4/4
Semi-final 0/2 2/2 2/2
Final 0/1 1/1 1/1
Total 23/31 26/31 26/31

Level through the quarter-finals (23/28 apiece). The whole difference is the highlighted semi-final and final rows.

springprompt.com/evals/world-cup

Where the human was irreplaceable

Sutton alone called both of Norway's giant-killings - the shootout win over Ivory Coast and the round-of-16 upset of Brazil. Both were matches every one of the six AI models got wrong.

Reading a specific underdog's momentum through a shootout is exactly the judgement no ratings-fed model produced - twice.

Where the human lost it

At the business end he trusted reputation over form: he backed France and England in the semi-finals (both lost) and picked France to win a final France never reached. The AIs' habit of backing the favourite banked all three.

Boring beat brilliant at the finish. The AIs never had a great call in them - and never needed one.

The verdict: a top human pundit, six flagship AIs, and one line of code all landed within three correct calls of each other over 31 matches - and the machines' edge was pure discipline, not insight. If you want a forecast that's reliable, the models (and the baseline) beat the human. If you want the one call that wins the office sweepstake, you still want the human.

🧠

What if they could learn?

The official benchmark froze every briefing at the eve of the tournament - so a model predicting the final knew nothing of the rounds before it. So we ran a second, experimental version: same models, same frozen scouting data, but before each match from the round of 16 on, we showed every model all earlier results and its own prior picks, scores and rationales. Rolling memory instead of a fixed snapshot. Could they learn on the fly? The answer is stranger than yes or no.

Given full memory of the tournament, the models changed which team they tipped in 0 of 96 later-round picks.

Not once. Shown that Norway had just knocked out two higher-rated sides, every model still went back to the pre-tournament ratings next time. Their scorelines and scorer picks shifted, and confidence dipped slightly (-2.3pp on average) - but the core judgement, who advances, never moved. Memory changed their arithmetic, not their mind.

The standings with rolling memory

Same scoring, same matches - re-ranked. Delta is versus the official frozen run.

# Model Official Learning Change
🥇 Claude Fable 5 307 311 +4
🥈 Claude Opus 4.8 296 306 +10
🥉 Gemini 3.1 Pro 286 302 +16
4 GLM-5.2 309 300 -9
5 GPT-5.5 284 295 +11
6 Qwen 3.7 Max 288 274 -14
springprompt.com/evals/world-cup

A new champion - and a fallen one

With memory, Claude Fable 5 takes the title (311) - while the frozen-run champion, GLM-5.2, drops to fourth. The model that won by maximising goalscorer guesses gained nothing from context; the steadier forecasters used it best. Gemini 3.1 Pro improved most (+16); Qwen 3.7 Max fell furthest (-14).

Memory still didn't beat the baseline

On core points, even the best learning-run AI (Claude Opus 4.8, 231) still finished behind the favourite baseline (239). Handing the models the whole tournament to study did not close the gap to a one-line rule. That is the most sobering result on this page.

Experimental, and not part of the official standings: these picks were generated retrospectively (with same-day results withheld from each other), so unlike the frozen run they aren't timestamped before kickoff. Round-of-32 picks are the models' original ones. Treat it as a controlled probe of whether context helps, not a second leaderboard of record.

Title race

Each model's position after every knockout match (1 = leading). Hover a match to see the full order.

GLM-5.2 currently leads the standings.
springprompt.com/evals/world-cup

How they played the game

The rules allowed up to three goalscorer picks per match, worth +3 each with no penalty for misses. What each model did with that freedom is the clearest personality test in the benchmark - and, it turned out, the decisive one.

Scorer picks vs own predicted goals

"Over-named" = named more scorers than goals in its own scoreline. Consistency vs opportunism.

Model Avg scorers named Goals in its scoreline prediction Named more scorers than goals
GLM-5.2 2.97 2.59 14/34
Claude Fable 5 2.74 2.68 2/34
Claude Opus 4.8 2.68 2.59 3/34
Qwen 3.7 Max 2.53 2.47 2/34
GPT-5.5 2.5 2.44 2/34
Gemini 3.1 Pro 2.26 2.56 0/34
springprompt.com/evals/world-cup

Reasoning effort vs result

Average completion tokens per pick (includes each model's internal reasoning at max effort).

Model Avg tokens / pick Final points
Claude Opus 4.8 2,687 296
Claude Fable 5 1,186 307
Qwen 3.7 Max 1,182 288
GPT-5.5 649 284
Gemini 3.1 Pro 628 286
GLM-5.2 583 309

The most economical thinker won; the deepest thinker finished third. Effort and accuracy were uncorrelated here.

springprompt.com/evals/world-cup

The herd: how often the AIs just backed the favourite

All six models picked the same winner in 31 of 34 fixtures. Across ~200 picks there were only 9 contrarian calls (against the ranking favourite) in total - and 4 matches went to penalties, a mechanism the models almost never predicted.

Model Agreed with favourite Contrarian picks (won) Penalties predicted Avg confidence
Gemini 3.1 Pro 34/34 0 (0) 6 72%
GPT-5.5 33/34 1 (0) 9 70%
Qwen 3.7 Max 33/34 1 (0) 4 73%
GLM-5.2 32/34 2 (1) 6 70%
Claude Fable 5 32/34 2 (1) 1 71%
Claude Opus 4.8 31/34 3 (1) 1 71%

Gemini 3.1 Pro never once disagreed with the rankings on who goes through - its progression picks were, in effect, the baseline with a squad list attached. The models' confidence sat at a uniform 70-73% regardless of the match, and the boldest behaviour in the field was three contrarian calls (Claude Opus 4.8), of which one landed.

springprompt.com/evals/world-cup

The pundits, profiled

Six models, six distinct characters. Each profile ends with the model's own (pre-match, blind) rationale for the final - Spain 1-0 Argentina, as it turned out. All six tipped Spain.

🥇 GLM-5.2

The game-player

309

Read the rules better than it read the football. Named the maximum three scorers almost every match - even when its own scoreline said two goals - and spent the fewest reasoning tokens in the field doing it. Its scorer hit rate was ordinary; its shot volume was not. Won the title on strategy, by 2 points.

Winners called26/32 matches Exact scorelines5 of 32 Goalscorer picks28 of 95 hit (29%) Reasoning per pick583 tokens
"Spain's superior Elo rating and midfield control (Rodri, Pedri, Gavi) give them a marginal edge, but finals are tight and Argentina's perfect recent form plus Messi's match-winning ability makes this extremely close — expect a 1-1 draw after 90 minutes with Spain edging it in extra time." - its call on the final, made before kickoff

🥈 Claude Fable 5

The forecaster

307

The best pure predictor in the field. Best-calibrated confidence (0.135 Brier), internally coherent picks (scorers matched its scoreline in 32 of 34 matches), and top AI on the baseline-fair core table. Joined the benchmark after the round of 32 had begun and still nearly took the crown.

Winners called26/32 matches Exact scorelines5 of 32 Goalscorer picks24 of 88 hit (27%) Reasoning per pick1,186 tokens
"Spain hold the Elo edge (2157 vs 2115) and superior squad depth in midfield with Rodri and Pedri controlling tempo, though Argentina's flawless recent form makes this a coin-flip-adjacent final. Expect a tight, high-quality match decided by Spain's wide threats, with a meaningful chance it goes beyond 90 minutes." - its call on the final, made before kickoff

🥉 Claude Opus 4.8

The deep thinker

296

Out-reasoned everyone, out-predicted no one. ~2,700 tokens of deliberation per pick - four times the winner's spend - bought steady, disciplined picks and a podium, but no measurable accuracy edge over faster rivals. The benchmark's clearest evidence that longer thinking has diminishing returns on chaotic domains.

Winners called25/32 matches Exact scorelines5 of 32 Goalscorer picks22 of 86 hit (26%) Reasoning per pick2,687 tokens
"Spain's higher Elo and elite midfield core (Rodri, Pedri) plus Yamal's attacking threat give a narrow edge over an in-form Argentina tested mostly against weaker friendly opposition. Expect a tense, tight final that Spain edges, though extra time is a genuine possibility given both teams' quality." - its call on the final, made before kickoff

4th Qwen 3.7 Max

Boom or bust

288

The volatility play. Owned more of the tournament's biggest single-match hauls than anyone (three +15s) and surged late - but paired it with the field's worst calibration (0.149) and the most scoreless weekends. When it was right, it was very right.

Winners called25/32 matches Exact scorelines4 of 32 Goalscorer picks23 of 82 hit (28%) Reasoning per pick1,182 tokens
"Spain's superior Elo and midfield control with Rodri and Pedri give them the edge over an Argentine side that has faced much weaker recent opposition. A narrow victory in regular time reflects the high stakes and tactical discipline expected in a World Cup final." - its call on the final, made before kickoff

5th Gemini 3.1 Pro

The purist

286

The only model that never named more scorers than its own predicted goals - not once in 34 matches - and took the fewest scorer shots as a result. Principled, well-calibrated, and the author of the single best-scored prediction of the tournament (+16 on France-Morocco). The rules didn't reward principle.

Winners called26/32 matches Exact scorelines6 of 32 Goalscorer picks20 of 72 hit (28%) Reasoning per pick628 tokens
"Spain's superior Elo rating and midfield control (anchored by Rodri) give them a slight edge over Argentina's tactical pragmatism. However, World Cup finals are notoriously tight, making a 1-1 draw in normal time highly likely before Spain's depth seals it in extra time." - its call on the final, made before kickoff

6th GPT-5.5

The early pacesetter

284

Led the race through the round of 32 on the strength of exact scorelines - it called more of them (7) than any other model. Faded as the goalscorer category grew decisive, and finished with the lowest core-points total. A specialist outpointed by a generalist's rulebook.

Winners called25/32 matches Exact scorelines7 of 32 Goalscorer picks22 of 79 hit (28%) Reasoning per pick649 tokens
"Spain have the small Elo edge and more control-oriented midfield depth, but Argentina's recent scoring form and elite forwards make a 90-minute stalemate plausible. I lean Spain to edge the final after extra time or penalties." - its call on the final, made before kickoff

The matches that fooled everyone

All six models wrong, together

Every AI backed the loser in these 5 matches - and so did the ranking baseline. Each one is a lesson in what the models couldn't see.

  • Netherlands 1-1 Morocco pens · Round of 32 · Morocco through

    Every model backed the Dutch at 62-72% confidence and most predicted 2-1. The 90 minutes actually finished level - which several models' scorelines implied - but none followed that thought to its conclusion: a shootout, which Morocco won. The models predicted the draw and still refused to predict its consequence.

  • Ivory Coast 1-2 Norway · Round of 32 · Norway through

    The Elo gap said Ivory Coast (2062 v 1827), and all six models said Ivory Coast, at up to 78%. Norway's Haaland-Nusa counterattack said otherwise. First warning that the briefing's ratings underweighted Norway - a warning nobody could act on.

  • Brazil 1-2 Norway · Round of 16 · Norway through

    One round later, same mistake, same team. The new matchup told every model that Norway had advanced, but they weren't shown the Ivory Coast result or their own earlier picks. With ratings and form still frozen pre-tournament, all six went back to the same priors and picked Brazil. The BBC's human pundit called this one exactly.

  • Switzerland 0-0 Colombia pens · Round of 16 · Switzerland through

    The tightest tie of the round, and the models knew it - confidence was the lowest of the tournament (52-62%). They still all landed on the same side, Colombia, and a goalless slog went to Switzerland on penalties. A coin-flip where all six coins came up identically wrong.

  • France 4-6 England · Third-place play-off · England through

    The great absurdity of the tournament. All six models predicted France 2-1, a sensible scoreline for a third-place game between well-matched sides. Reality produced ten goals, an England win, and the single most chaotic match of the World Cup. No ratings-based system predicts a 4-6; nothing in the briefing hinted at it. Some things remain gloriously unforecastable.

The best single calls

Highest-scoring individual match predictions of the tournament, and where the points actually came from.

  • +16 Gemini 3.1 Proon France v Morocco (2-0)

    The tournament's best single prediction: right team, exact 2-0 scoreline, and both goalscorers named. Everything the benchmark asks for, in one pick - from the model that finished fifth.

  • +15 Claude Opus 4.8on Argentina v Cape Verde (3-2)

  • +15 GPT-5.5on Argentina v Cape Verde (3-2)

  • +15 Gemini 3.1 Proon Argentina v Cape Verde (3-2)

    Third of Gemini's three +15s - all built on exact scorelines plus multiple correct scorers.

  • +15 GLM-5.2on Argentina v Cape Verde (3-2)

  • +15 Qwen 3.7 Maxon Argentina v Cape Verde (3-2)

    A 3-2 goal-fest predicted almost goal for goal.

Note who owns this list: Gemini and Qwen, who finished fifth and fourth. Peak insight and total points measured different things in this game - the steady rule-optimisers won the table while the precision specialists won the highlights.

The full match-by-match record

Every pick by every model, for all 32 matches. Expand a round to see it.

Reading a pick: United States = team tipped to go through = predicted goalscorers correct · wrong conf = the model's confidence
2026-06-28 FT
South Africa South Africa 0-1 Canada Canada

Canada through · Stephen Eustáquio

Ranking favourite Canada 0-1 82% +10
Claude Opus 4.8 Canada 0-2 77% +9
GPT-5.5 Canada 0-1 78% +10
Gemini 3.1 Pro Canada 0-2 75% +9
GLM-5.2 Canada 1-2 71% +9
Qwen 3.7 Max Canada 0-2 78% +9
Claude Fable 5 Canada 0-2 79% +9
springprompt.com/evals/world-cup
2026-06-29 FT
Brazil Brazil 2-1 Japan Japan

Brazil through · Kaishu Sano, Casemiro, Gabriel Martinelli

Ranking favourite Brazil 1-0 62% +9
Claude Opus 4.8 Brazil 2-1 65% +10
GPT-5.5 Brazil 2-1 61% +10
Gemini 3.1 Pro Brazil 2-1 72% +10
GLM-5.2 Brazil 2-1 62% +10
Qwen 3.7 Max Brazil 2-1 65% +10
Claude Fable 5 Brazil 2-1 60% +10
springprompt.com/evals/world-cup
2026-06-29 FT
Germany Germany 1-1 Paraguay Paraguay

Paraguay through on pens · Julio Enciso, Kai Havertz

Ranking favourite Paraguay 0-1 65% +5
Claude Opus 4.8 Germany 2-1 55% +3
GPT-5.5 Paraguay 1-1 56% +10
Gemini 3.1 Pro Paraguay 1-1 62% +13
GLM-5.2 Paraguay 1-1 55% +13
Qwen 3.7 Max Germany 2-1 62% +0
Claude Fable 5 Paraguay 1-2 57% +11
springprompt.com/evals/world-cup
2026-06-29 FT
Netherlands Netherlands 1-1 Morocco Morocco

Morocco through on pens · Cody Gakpo, Issa Diop

Ranking favourite Netherlands 1-0 67% +0
Claude Opus 4.8 Netherlands 2-1 64% +3
GPT-5.5 Netherlands 1-1 62% +8
Gemini 3.1 Pro Netherlands 2-1 65% +3
GLM-5.2 Netherlands 2-1 62% +3
Qwen 3.7 Max Netherlands 2-1 72% +3
Claude Fable 5 Netherlands 2-1 61% +3
springprompt.com/evals/world-cup
2026-06-30 FT
France France 3-0 Sweden Sweden

France through · Kylian Mbappé, Bradley Barcola, Kylian Mbappé

Ranking favourite France 1-0 88% +9
Claude Opus 4.8 France 2-1 88% +12
GPT-5.5 France 2-1 84% +12
Gemini 3.1 Pro France 3-1 85% +12
GLM-5.2 France 3-1 85% +12
Qwen 3.7 Max France 2-0 82% +12
Claude Fable 5 France 2-1 88% +12
springprompt.com/evals/world-cup
2026-06-30 FT
Ivory Coast Ivory Coast 1-2 Norway Norway

Norway through · Antonio Nusa, Amad Diallo, Erling Haaland

Ranking favourite Ivory Coast 1-0 80% +2
Claude Opus 4.8 Ivory Coast 2-1 68% +8
GPT-5.5 Ivory Coast 2-1 72% +8
Gemini 3.1 Pro Ivory Coast 2-1 68% +5
GLM-5.2 Ivory Coast 2-1 68% +5
Qwen 3.7 Max Ivory Coast 2-1 72% +8
Claude Fable 5 Ivory Coast 2-1 72% +5
springprompt.com/evals/world-cup
2026-06-30 FT
Mexico Mexico 2-0 Ecuador Ecuador

Mexico through · Julián Quiñones, Raúl Jiménez

Ranking favourite Ecuador 0-1 59% +2
Claude Opus 4.8 Mexico 1-1 52% +5
GPT-5.5 Ecuador 1-1 56% +0
Gemini 3.1 Pro Ecuador 1-1 55% +0
GLM-5.2 Mexico 2-1 55% +12
Qwen 3.7 Max Ecuador 1-1 55% +0
Claude Fable 5 Mexico 1-1 54% +8
springprompt.com/evals/world-cup
2026-07-01 FT
Belgium Belgium 3-2 Senegal Senegal

Belgium through · Habib Diarra, Ismaïla Sarr, Romelu Lukaku, Youri Tielemans, Youri Tielemans

Ranking favourite Belgium 1-0 86% +9
Claude Opus 4.8 Belgium 2-0 82% +12
GPT-5.5 Belgium 2-0 82% +12
Gemini 3.1 Pro Belgium 2-0 85% +12
GLM-5.2 Belgium 2-0 78% +12
Qwen 3.7 Max Belgium 2-0 85% +12
Claude Fable 5 Belgium 2-0 83% +12
springprompt.com/evals/world-cup
2026-07-01 FT
England England 2-1 DR Congo DR Congo

England through · Brian Cipenga, Harry Kane, Harry Kane

Ranking favourite England 1-0 85% +9
Claude Opus 4.8 England 2-0 89% +12
GPT-5.5 England 2-0 86% +12
Gemini 3.1 Pro England 2-0 88% +12
GLM-5.2 England 2-0 82% +12
Qwen 3.7 Max England 2-0 88% +12
Claude Fable 5 England 2-0 90% +12
springprompt.com/evals/world-cup
2026-07-01 FT
United States United States 2-0 Bosnia and Herzegovina Bosnia and Herzegovina

United States through · Folarin Balogun, Malik Tillman

Ranking favourite United States 1-0 68% +9
Claude Opus 4.8 United States 2-1 69% +12
GPT-5.5 United States 2-1 71% +12
Gemini 3.1 Pro United States 2-1 68% +9
GLM-5.2 United States 2-1 72% +12
Qwen 3.7 Max United States 1-1 65% +8
Claude Fable 5 United States 2-1 74% +12
springprompt.com/evals/world-cup
2026-07-02 FT
Portugal Portugal 2-1 Croatia Croatia

Portugal through · Ivan Perisic, Cristiano Ronaldo, Gonçalo Ramos

Ranking favourite Portugal 1-0 61% +9
Claude Opus 4.8 Portugal 2-1 63% +13
GPT-5.5 Portugal 2-1 62% +13
Gemini 3.1 Pro Portugal 2-1 65% +10
GLM-5.2 Portugal 2-1 64% +13
Qwen 3.7 Max Portugal 2-1 65% +10
Claude Fable 5 Portugal 2-1 62% +13
springprompt.com/evals/world-cup
2026-07-02 FT
Spain Spain 3-0 Austria Austria

Spain through · Mikel Oyarzabal, Pedro Porro, Mikel Oyarzabal

Ranking favourite Spain 1-0 87% +9
Claude Opus 4.8 Spain 2-0 87% +12
GPT-5.5 Spain 2-0 86% +9
Gemini 3.1 Pro Spain 2-0 85% +9
GLM-5.2 Spain 2-0 82% +12
Qwen 3.7 Max Spain 2-0 88% +9
Claude Fable 5 Spain 2-0 85% +12
springprompt.com/evals/world-cup
2026-07-02 FT
Switzerland Switzerland 2-0 Algeria Algeria

Switzerland through · Breel Embolo, Dan Ndoye

Ranking favourite Switzerland 1-0 51% +9
Claude Opus 4.8 Algeria 0-1 53% +2
GPT-5.5 Algeria 1-1 51% +3
Gemini 3.1 Pro Switzerland 1-1 52% +8
GLM-5.2 Algeria 1-1 55% +3
Qwen 3.7 Max Switzerland 1-1 55% +8
Claude Fable 5 Algeria 1-2 53% +5
springprompt.com/evals/world-cup
2026-07-03 FT
Argentina Argentina 3-2 Cape Verde Cape Verde

Argentina through · Lionel Messi, Deroy Duarte, Lisandro Martínez, Sidny Lopes Cabral, Diney Borges

Ranking favourite Argentina 1-0 96% +9
Claude Opus 4.8 Argentina 3-0 95% +15
GPT-5.5 Argentina 3-0 94% +15
Gemini 3.1 Pro Argentina 3-0 95% +15
GLM-5.2 Argentina 3-0 93% +15
Qwen 3.7 Max Argentina 3-0 96% +15
Claude Fable 5 Argentina 3-0 96% +15
springprompt.com/evals/world-cup
2026-07-03 FT
Australia Australia 1-1 Egypt Egypt

Egypt through on pens · Emam Ashour, Mohamed Hany

Ranking favourite Egypt 0-1 77% +5
Claude Opus 4.8 Egypt 0-1 74% +5
GPT-5.5 Egypt 0-1 72% +5
Gemini 3.1 Pro Egypt 0-2 72% +5
GLM-5.2 Egypt 0-2 72% +5
Qwen 3.7 Max Egypt 0-2 75% +5
Claude Fable 5 Egypt 0-2 72% +5
springprompt.com/evals/world-cup
2026-07-03 FT
Colombia Colombia 1-0 Ghana Ghana

Colombia through · Jhon Arias

Ranking favourite Colombia 1-0 84% +10
Claude Opus 4.8 Colombia 2-0 84% +9
GPT-5.5 Colombia 2-0 84% +9
Gemini 3.1 Pro Colombia 2-0 85% +9
GLM-5.2 Colombia 2-0 78% +9
Qwen 3.7 Max Colombia 2-0 88% +9
Claude Fable 5 Colombia 2-0 85% +9
springprompt.com/evals/world-cup
2026-07-04 FT
Canada Canada 0-3 Morocco Morocco

Morocco through · Azzedine Ounahi, Azzedine Ounahi, Soufiane Rahimi

Ranking favourite Morocco 0-1 56% +9
Claude Opus 4.8 Morocco 1-2 57% +9
GPT-5.5 Morocco 1-1 57% +5
Gemini 3.1 Pro Morocco 1-2 62% +9
GLM-5.2 Morocco 1-2 62% +12
Qwen 3.7 Max Morocco 1-2 65% +9
Claude Fable 5 Morocco 1-2 60% +9
springprompt.com/evals/world-cup
2026-07-04 FT
Paraguay Paraguay 0-1 France France

France through · Kylian Mbappé

Ranking favourite France 0-1 79% +10
Claude Opus 4.8 France 0-2 83% +12
GPT-5.5 France 0-2 79% +12
Gemini 3.1 Pro France 0-2 85% +12
GLM-5.2 France 0-2 78% +12
Qwen 3.7 Max France 0-2 82% +12
Claude Fable 5 France 0-2 80% +12
springprompt.com/evals/world-cup
2026-07-05 FT
Brazil Brazil 1-2 Norway Norway

Norway through · Erling Haaland, Erling Haaland, Neymar

Ranking favourite Brazil 1-0 72% +2
Claude Opus 4.8 Brazil 2-1 73% +5
GPT-5.5 Brazil 2-1 71% +8
Gemini 3.1 Pro Brazil 2-1 75% +5
GLM-5.2 Brazil 2-1 73% +8
Qwen 3.7 Max Brazil 2-1 76% +5
Claude Fable 5 Brazil 2-1 70% +5
springprompt.com/evals/world-cup
2026-07-05 FT
Mexico Mexico 2-3 England England

England through · Jude Bellingham, Jude Bellingham, Julián Quiñones, Harry Kane, Raúl Jiménez

Ranking favourite England 0-1 70% +9
Claude Opus 4.8 England 1-2 70% +15
Claude Fable 5 England 1-2 63% +12
GPT-5.5 England 1-2 66% +12
Gemini 3.1 Pro England 1-2 72% +15
GLM-5.2 England 1-1 68% +14
Qwen 3.7 Max England 1-2 75% +15
springprompt.com/evals/world-cup
2026-07-06 FT
Portugal Portugal 0-1 Spain Spain

Spain through · Mikel Merino

Ranking favourite Spain 0-1 72% +10
Claude Opus 4.8 Spain 1-2 67% +9
Claude Fable 5 Spain 1-2 66% +9
GPT-5.5 Spain 1-1 66% +5
Gemini 3.1 Pro Spain 1-1 65% +5
GLM-5.2 Spain 1-1 60% +5
Qwen 3.7 Max Spain 1-1 62% +5
springprompt.com/evals/world-cup
2026-07-06 FT
United States United States 1-4 Belgium Belgium

Belgium through · Charles De Ketelaere, Malik Tillman, Charles De Ketelaere, Hans Vanaken, Romelu Lukaku

Ranking favourite Belgium 0-1 72% +9
Claude Opus 4.8 Belgium 1-2 70% +12
Claude Fable 5 Belgium 1-2 66% +12
GPT-5.5 Belgium 1-2 68% +12
Gemini 3.1 Pro Belgium 1-3 78% +12
GLM-5.2 Belgium 1-2 72% +12
Qwen 3.7 Max Belgium 1-2 78% +12
springprompt.com/evals/world-cup
2026-07-07 FT
Argentina Argentina 3-2 Egypt Egypt

Argentina through · Yasser Ibrahim, Mostafa Zico, Cristian Romero, Lionel Messi, Enzo Fernández

Ranking favourite Argentina 1-0 67% +9
Claude Opus 4.8 Argentina 2-0 79% +9
Claude Fable 5 Argentina 2-0 76% +12
GPT-5.5 Argentina 2-1 69% +12
Gemini 3.1 Pro Argentina 2-0 85% +12
GLM-5.2 Argentina 2-0 78% +12
Qwen 3.7 Max Argentina 2-0 85% +12
springprompt.com/evals/world-cup
2026-07-07 FT
Switzerland Switzerland 0-0 Colombia Colombia

Switzerland through on pens

Ranking favourite Colombia 0-1 76% +0
Claude Opus 4.8 Colombia 1-2 68% +0
Claude Fable 5 Colombia 1-2 71% +0
GPT-5.5 Colombia 1-2 70% +0
Gemini 3.1 Pro Colombia 1-2 65% +0
GLM-5.2 Colombia 1-2 68% +0
Qwen 3.7 Max Colombia 1-2 68% +0
springprompt.com/evals/world-cup
2026-07-09 FT
France France 2-0 Morocco Morocco

France through · Kylian Mbappé, Ousmane Dembélé

Ranking favourite France 1-0 80% +9
Claude Opus 4.8 France 2-1 77% +12
Claude Fable 5 France 2-1 74% +12
GPT-5.5 France 2-1 74% +15
Gemini 3.1 Pro France 2-0 82% +16
GLM-5.2 France 2-1 78% +12
Qwen 3.7 Max France 2-1 78% +15
springprompt.com/evals/world-cup
2026-07-10 FT
Spain Spain 2-1 Belgium Belgium

Spain through · Fabián Ruiz, Charles De Ketelaere, Mikel Merino

Ranking favourite Spain 1-0 82% +9
Claude Opus 4.8 Spain 2-1 80% +10
Claude Fable 5 Spain 2-1 79% +10
GPT-5.5 Spain 2-1 73% +10
Gemini 3.1 Pro Spain 2-1 75% +10
GLM-5.2 Spain 2-1 68% +10
Qwen 3.7 Max Spain 2-1 78% +10
springprompt.com/evals/world-cup
2026-07-11 FT
Argentina Argentina 3-1 Switzerland Switzerland

Argentina through · Alexis Mac Allister, Dan Ndoye, Julián Álvarez, Lautaro Martínez

Ranking favourite Argentina 1-0 88% +9
Claude Opus 4.8 Argentina 2-0 89% +15
Claude Fable 5 Argentina 2-0 86% +12
GPT-5.5 Argentina 2-0 84% +12
Gemini 3.1 Pro Argentina 2-0 85% +12
GLM-5.2 Argentina 2-0 83% +15
Qwen 3.7 Max Argentina 2-0 85% +12
springprompt.com/evals/world-cup
2026-07-11 FT
Norway Norway 1-2 England England

England through · Andreas Schjelderup, Jude Bellingham, Jude Bellingham

Ranking favourite England 0-1 76% +9
Claude Opus 4.8 England 1-2 76% +10
Claude Fable 5 England 1-2 72% +10
GPT-5.5 England 1-2 71% +10
Gemini 3.1 Pro England 1-2 75% +13
GLM-5.2 England 1-2 72% +13
Qwen 3.7 Max England 1-2 72% +13
springprompt.com/evals/world-cup
2026-07-14 FT
France France 0-2 Spain Spain

Spain through · Mikel Oyarzabal, Pedro Porro

Ranking favourite Spain 0-1 63% +9
Claude Opus 4.8 Spain 1-2 60% +9
Claude Fable 5 Spain 1-2 60% +12
GPT-5.5 Spain 1-1 58% +5
Gemini 3.1 Pro Spain 1-2 55% +9
GLM-5.2 Spain 1-2 62% +9
Qwen 3.7 Max Spain 1-2 65% +9
springprompt.com/evals/world-cup
2026-07-15 FT
England England 1-2 Argentina Argentina

Argentina through · Anthony Gordon, Enzo Fernández, Lautaro Martínez

Ranking favourite Argentina 0-1 63% +9
Claude Opus 4.8 Argentina 1-2 59% +13
Claude Fable 5 Argentina 1-2 60% +10
GPT-5.5 Argentina 1-1 59% +5
Gemini 3.1 Pro Argentina 1-1 62% +5
GLM-5.2 Argentina 1-1 58% +5
Qwen 3.7 Max Argentina 0-1 62% +12
springprompt.com/evals/world-cup
2026-07-18 FT
France France 4-6 England England

England through · Declan Rice, Ezri Konsa, Bukayo Saka, Bukayo Saka, Kylian Mbappé, Bradley Barcola, Kylian Mbappé, Bukayo Saka, Ousmane Dembélé, Jude Bellingham

Ranking favourite France 1-0 56% +2
Claude Opus 4.8 France 2-1 57% +5
Claude Fable 5 France 2-1 58% +8
GPT-5.5 France 2-1 57% +8
Gemini 3.1 Pro France 2-1 58% +5
GLM-5.2 France 2-1 57% +8
Qwen 3.7 Max France 2-1 62% +8
springprompt.com/evals/world-cup
2026-07-19 FT
Spain Spain 1-0 Argentina Argentina

Spain through · Ferran Torres

Ranking favourite Spain 1-0 56% +10
Claude Opus 4.8 Spain 2-1 54% +9
Claude Fable 5 Spain 2-1 56% +9
GPT-5.5 Spain 1-1 55% +5
Gemini 3.1 Pro Spain 1-1 55% +5
GLM-5.2 Spain 1-1 52% +5
Qwen 3.7 Max Spain 2-1 58% +9
springprompt.com/evals/world-cup

Why these six models

The roster was simple: the flagship reasoning model from each major provider, as of the start of the knockout rounds - Claude Opus 4.8 (Anthropic), GPT-5.5 (OpenAI), Gemini 3.1 Pro (Google), GLM-5.2 (Z.ai) and Qwen 3.7 Max (Alibaba) - every one at its maximum reasoning setting. We began at the round of 32 rather than the group stage for no grander reason than that's when we built the benchmark; with 32 knockout matches and sudden-death stakes, it turned out to be the better test anyway.

Claude Fable 5 joined a few matches in, when it became available - early enough that backfilling its round-of-32 picks under the same blind conditions (pre-tournament training cut-off, no web access) was defensible. We kept Opus 4.8 in the field rather than swapping it out: it gave us a same-provider comparison point, and its predictions were already locked. Models that launched later in the tournament didn't get the same treatment - see the FAQ.

One quiet subplot worth naming: the two open-weights models finished first and fourth. GLM-5.2 won the whole thing and Qwen 3.7 Max out-pointed two Western proprietary flagships - for a fraction of the per-token cost. Whatever this benchmark measures, it isn't a proprietary moat.

What this means if you want AI to predict things

1 · Always run a dumb baseline first

The single most useful thing in this benchmark cost one line of code. If we hadn't scored "always pick the favourite," the headline would have been "AI beats the World Cup" - and it would have been wrong. Before trusting an AI forecast, know what the naive answer scores.

2 · Expect the herd, not the oracle

Six different flagships, one briefing, near-identical picks: unanimous in 31 of 34 matches, uniform ~70% confidence, nine contrarian calls in ~200. If your use case needs someone to spot the upset, current models - given identical inputs - won't be that someone. Diversity of data, not of model vendor, is what would have changed these picks.

3 · The value is in the detail, not the verdict

Where the AIs genuinely beat the baseline was squad-level knowledge: naming likely goalscorers at a steady ~27% hit rate. That's real, useful signal a ranking can't produce. Use AI forecasting where breadth of specific knowledge matters, not where the answer is one bit.

4 · Models optimise your rules, not your intent

GLM-5.2 won because our scoring made three scorer guesses free. It didn't cheat - it read the incentive perfectly, which is exactly what these systems do. Any scoring rule you give an AI, audit it as if the AI will find the loophole. It will.

What we'll change next time

This is Spring Prompt's second forecasting benchmark (after PredictTheWeek), and the next one will inherit its scars:

  • Close the loopholes - or score them. Uncapped goalscorer picks decided this benchmark. Next time either wrong guesses cost points, or rule-exploitation gets measured openly as its own dimension. Both are interesting; ambiguity isn't.
  • Reward contrarian value, not just accuracy. A correct upset call is worth more information than a correct favourite call. Scoring should reflect that - odds-weighted points would have made Sutton's Norway call worth its weight, and might coax the herd apart.
  • Save the full reasoning traces. We kept each model's two-sentence rationale and its token spend, and that alone yielded findings. Complete chains of thought would have let us ask why six models converge - next time they're archived from day one.
  • Let the briefing evolve - and preserve each model's memory. The original test never showed models earlier tournament results or fed back their own picks and scores; only the next matchup revealed who had advanced. Next time, every round should include the results so far plus each model's personal prediction record, testing whether it updates after surprises instead of repeating the same mistake. We ran that rolling-learning variant locally after the tournament to measure the difference.

FAQ

Why was Claude Fable 5 added late, but newer models weren't?

Fable 5 became available only two or three matches into the round of 32, so backfilling its handful of missed picks under identical blind conditions barely stretched the rules. Models that launched deeper into the tournament (e.g. GPT-5.6) would have needed retro-predictions for a large share of already-played matches - technically possible under the same blind protocol, but we judged the integrity risk to the results not worth one more row in the table.

Why start at the knockouts and skip the group stage?

Honestly: timing - the benchmark was built as the group stage ended. The knockout format turned out to be the sharper test (every match has a winner, and penalties add a category the models demonstrably can't call).

Did any model know the results?

No. Every model's training cut-off predates the tournament, web access was disabled, and the briefing was frozen to the eve of the tournament. Knockout picks were also timestamped before kickoff. A model could no more look up a result than a pundit in a studio could.

Can I get the data?

The full prediction set (~200 graded picks with rationales, confidence values and token counts) is preserved. Email us and we'll share it.

Methodology

The blind test

Each model received an identical scouting report per match: Elo rating, recent form, head-to-head record and the full 26-man squad - all frozen to 2026-06-11, the eve of the tournament - and no internet access. Because every model's training cut-off predates the tournament and web access was disabled, every pick was genuinely blind, whether made before or after a match was played. Knockout picks were additionally timestamped before kickoff. All models ran at their maximum reasoning setting.

Scoring

Right team through +5 · exact 90-minute score +3 · right result without the exact score +2 · each correct goalscorer (up to three named) +3 · correct extra-time/penalties call +2 · correct red-card call +5. Calibration is a separate Brier score on each model's stated probabilities (0 = perfect, lower is better).

Known quirks, disclosed

  • Goalscorer picks carried no penalty for misses - naming three was strictly dominant, and the winner exploited exactly that. We left the rule as designed and report the goalscorer-free table above.
  • The Ranking favourite baseline structurally can't name goalscorers, so compare it on the core table, not the headline one.
  • Claude Fable 5 joined after the round of 32 had started; its early-round picks were made retrospectively but under the same blind conditions (pre-tournament training cut-off, no web).
  • Round-of-32 group-stage seeding, injuries and suspensions were not in the briefing; models predicted from pre-tournament knowledge only.

Data

Results, scores and goalscorers were graded from ESPN's public scoreboard with the open international results dataset (CC0) as fallback; team ratings from eloratings.net; squads from Wikipedia. The full prediction dataset (~200 graded picks with rationales and token counts) is preserved as-is; the tournament is over and these figures are final.

Full time

That's the whistle on World Cup Bench.

Spring Prompt runs benchmarks like this every week - testing what AI models can actually do on real tasks, with baselines that keep the answers honest. The next experiment is already being designed with this one's lessons.

Spotted an error, or want to talk about this report? hello@springprompt.com