Back to Blog
Evals World Cup Bench AI Benchmarks

Six Models, One Mind: What the World Cup Revealed About How AIs Actually Think

Ellis Crosby
3 min read
Dark stadium graphic reading: Six models. One mind. 31 of 34 unanimous.
Six models from five labs, one briefing, near-identical picks: the World Cup Bench herd.

Here is a fact that should unsettle anyone building products on the assumption that different AI models offer different opinions.

For World Cup Bench, we had six frontier models predict every knockout match of the 2026 World Cup: two Claude models, GPT-5.5, Gemini 3.1 Pro, GLM-5.2 and Qwen 3.7 Max. Five different labs, three countries, wildly different architectures and training runs, all reasoning at maximum effort.

In 31 of the 34 fixtures they predicted, all six picked the same winner.

📌
Data note: Figures are from the final World Cup Bench dataset: 32 graded knockout matches, roughly 200 predictions, plus a 96-prediction rolling-memory experiment. Full tables and methodology on the report page.

The herd, measured

Every model received an identical scouting report, frozen to the eve of the tournament: Elo ratings, form, head-to-head, squads. Identical inputs are part of the explanation. They are not all of it, because the models also had genuinely different information they could have weighted differently: squad depth, playing styles, tournament history. They did not.

Gemini 3.1 Pro never once disagreed with the ranking favourite on who goes through. In 34 out of 34 fixtures, its pick was the higher-rated team. Its progression picks were, in effect, an Elo table with a squad list attached.

Across all six models and roughly 200 predictions, there were nine contrarian picks in total. Three landed. The models' stated confidence sat in a narrow band, 70 to 73 percent on average, almost regardless of the matchup.

The herd was also, to be fair, competent. The AI consensus called 26 of 31 knockout results correctly, beating the BBC pundit we graded on the same matches, who managed 23. But every single collective failure was identical. When Morocco went through on penalties, all six models had backed the Netherlands. When Norway knocked out Brazil, all six had backed Brazil. There was no wisdom-of-crowds effect, because there was no crowd. There was one opinion, printed six times.

We gave them memory. Nothing happened.

The official benchmark froze every model's knowledge before the tournament, so a model predicting the final knew nothing about the rounds before it. An obvious objection: maybe the models herd because they cannot learn. Show them the tournament as it unfolds and surely they adapt.

So we ran the experiment. From the round of 16 onward, before every match, each model was shown all earlier results plus its own previous picks, its scores, and its own written rationales. Rolling memory instead of a frozen snapshot. Same models, same scoring.

Across 96 later-round predictions, the number of times any model changed which team it tipped was zero.

Not one. Shown that Norway had just eliminated two higher-rated teams in consecutive rounds, every model went straight back to the pre-tournament ratings and picked against Norway again where the ratings said so. Confidence dipped by about two percentage points on average. Scorelines and goalscorer picks shifted at the margins. The core judgement never moved once.

The points did move, interestingly. With memory, Claude Fable 5 overtook GLM-5.2 to top the experimental table, Gemini 3.1 Pro gained sixteen points, and the frozen-run champion dropped to fourth. Memory made the good forecasters modestly better at the details. It changed nobody's mind about anything. You can flip between both leaderboards on the report page.

Why this matters beyond football

Nobody serious is betting a product on AI World Cup predictions. But plenty of teams are doing structurally identical things: asking a model to assess risk, forecast demand, triage candidates, or call which experiment to run, and some of them "get a second opinion" by asking a different vendor's model.

This benchmark says: that second opinion is likely the first opinion with different branding. Given the same inputs, frontier models converge hard, converge confidently, and do not update their core judgements even when handed direct evidence that their reasoning pattern just failed twice in a row.

If you need reliability, that convergence is fine, even useful. The herd matched the rankings and the rankings are good. If you need someone to catch the outlier, the upset, the tail risk, the thing the consensus is wrong about, you will not get it by polling more models. You will get it by changing what the models see: different data, different framings, deliberately adversarial prompts. Diversity of inputs, not of vendors.

The full report has the herd table, the learning experiment, the human pundit comparison and every match-by-match pick. For the raw dataset, email us.

Ellis Crosby

Related Articles

Ready to Optimize Your AI Prompts?

Start testing and improving your prompts with Spring Prompt's professional tools.

Join waitlist