Benchmarks / BulletBench

Measured by Spring Prompt

BulletBench

When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?

Last updated 8 Oct 2026

Results dated
1 Oct 2026 to 7 Oct 2026
Results
23 configurations of 22 models
Unit
ladder Elo
Licence
Spring Prompt original
Sample
12 games per configuration

Ladder Elo: Gemini 3.5 Flash-Lite

Top 15 of 22 results · ladder Elo, higher is better · lines show the 95% range · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Gemini 3.5 Flash-LiteGoogle 892
  2. 2≈ Gemini 3.1 Flash-Lite (minimal reasoning)Google 780
  3. 3≈ Gemini 3.5 Flash (minimal reasoning)Google 727
  4. 4≈ Mercury DecideDecision modelInception 614
  5. 5≈ d1Decision modelLiquid AI 572
  6. 6≈ Seed-2.0-Mini (minimal reasoning)ByteDance 571
  7. 6≈ Mistral Small 4 (no reasoning)Mistral AI 571
  8. 6≈ GPT-5.4 mini (no reasoning)OpenAI 571
  9. 6≈ Inkling Small (no reasoning)Thinking Machines 571
  10. 6≈ Jev 1.13Decision modelTypeSafe 571
  11. 11≈ Clef FlashDecision modelCloudflare 529
  12. 11≈ ClefDecision modelCloudflare 529
  13. 11≈ Mistral Medium 3.5 (no reasoning)Mistral AI 529
  14. 11≈ GPT-6 Luna DecisionsDecision modelOpenAI 529
  15. 15 Decider V1 27BDecision modelPerplexity 420

Lightning: strength against thinking time

Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.

Named: the six bestOther models (hover for names)

-200300800 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1 s increment Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Mercury Decide d1 Seed-2.0-Mini (minimal reasoning)

Full results

BulletBench Lightning 10+1: ladder Elo, ladder Elo, higher is better
#ModelLightning 10+1 · 95% range
ladder Elo, higher is better
Bullet 60s
ladder Elo
Blitz 3+2
ladder Elo
Move time
seconds
Lost on time
% of games
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1≈ Gemini 3.5 Flash-LiteGoogle
892
641–1,081
7807070.9 s0.0%$0.0070 ––
2≈ Gemini 3.1 Flash-Lite (minimal reasoning)Google
780
589–940
912–0.8 s0.0%$0.0058 ––
3≈ Gemini 3.5 Flash (minimal reasoning)Google
727
567–890
810–1.2 s33.3%$0.0213 1.3 s12
4≈ Mercury DecideDecision modelInception
614
470–766
572–0.4 s0.0%$0 0.5 s12
5≈ d1Decision modelLiquid AI
572
428–718
584–0.6 s0.0%$0.0005 0.7 s12
6≈ Seed-2.0-Mini (minimal reasoning)ByteDance
571
420–710
453–0.3 s0.0%$0.0019 0.4 s12
6≈ Mistral Small 4 (no reasoning)Mistral AI
571
420–699
367–0.6 s0.0%$0.0041 1.2 s12
6≈ GPT-5.4 mini (no reasoning)OpenAI
571
412–721
197–0.9 s0.0%$0.0136 ––
6≈ Inkling Small (no reasoning)Thinking Machines
571
428–698
529–0.6 s0.0%$0.0077 1.6 s12
6≈ Jev 1.13Decision modelTypeSafe
571
429–719
572–0.5 s0.0%$0.0014 0.7 s12
11≈ Clef FlashDecision modelCloudflare
529
388–670
529–0.5 s0.0%$0.0024 0.7 s12
11≈ ClefDecision modelCloudflare
529
373–653
529–0.7 s0.0%$0.0072 1.0 s12
11≈ Mistral Medium 3.5 (no reasoning)Mistral AI · best of 2 settings
529
361–669
572–0.5 s0.0%$0.0491 0.8 s12
11≈ GPT-6 Luna DecisionsDecision modelOpenAI
529
407–643
489–0.5 s0.0%$0.0025 0.5 s12
15 Decider V1 27BDecision modelPerplexity
420
367–488
385–0.6 s0.0%$0.0009 0.7 s12
16 Mistral Large 4Mistral AI
249
0–486
–5550.8 s58.3%$0.0117 1.9 s12
17 Gemini 3.8 Flash (low reasoning)Google
197
0–439
1,068–1.8 s83.3%$0.0041 2.1 s12
17 GPT-5.4 nano (no reasoning)OpenAI
197
0–425
33–1.0 s75.0%$0.0025 ––
19 Claude Haiku 4.5Anthropic
33
0–248
4454911.3 s83.3%$0.0095 1.8 s12
19 Claude Sonnet 5.5 (low reasoning)Anthropic
33
0–270
––2.1 s91.7%$0.0080 2.5 s12
19 Gemini 3.6 Flash (minimal reasoning)Google
33
0–248
628–1.5 s91.7%$0.0050 1.9 s12
19 Kev 4BDecision modelJared Palmer
33
0–270
––1.0 s66.7%$0.0006 1.4 s12

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. ≈ marks results whose 95% range overlaps the leader's: they cannot be told apart from it. Each model is shown at its best setting; show every setting. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only. Equal ratings are not a glitch: the ladder's ratings come in steps (see Honest caveats below).

Too slow for the clock

These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Each shows its median time per move and the time that 1 move in 10 took or exceeded: a median under the increment still loses on time when the slow tail is long. Being too slow is the result, not a fault.

Bullet 60s

  • Claude Fable 5.1 median 4.2 s, 1 in 10 moves 6.3 s+
  • Claude Fable 5.1 (low reasoning) median 4.3 s, 1 in 10 moves 9.3 s+
  • Claude Haiku 5.5 median 4.7 s
  • Claude Haiku 5.5 (low reasoning) median 3.6 s
  • Claude Opus 5.5 median 3.1 s, 1 in 10 moves 5.8 s+
  • Claude Opus 5.5 (low reasoning) median 3.1 s, 1 in 10 moves 5.3 s+
  • Claude Sonnet 5.5 median 3.0 s, 1 in 10 moves 9.4 s+
  • Claude Sonnet 5.5 (low reasoning) median 2.6 s, 1 in 10 moves 5.0 s+
  • DeepSeek V4 Pro 0423 median 5.0 s, 1 in 10 moves 27.8 s+
  • DeepSeek V4.1 Flash (low reasoning) median 1.9 s, 1 in 10 moves 10.6 s+
  • Gemini 3.1 Pro Preview median 4.9 s, 1 in 10 moves 8.4 s+
  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+
  • Muse Spark 1.3 median 7.0 s, 1 in 10 moves 28.1 s+
  • Mistral Large 4 median 0.8 s, 1 in 10 moves 2.9 s+
  • Kimi K3 median 3.2 s, 1 in 10 moves 33.8 s+
  • GPT-6 Luna median 4.6 s, 1 in 10 moves 10.8 s+
  • GPT-6 Sol median 4.2 s, 1 in 10 moves 9.4 s+
  • Qwen3.8 Flash median 3.8 s, 1 in 10 moves 16.9 s+
  • Qwen3.8 Max (0902) median 4.9 s, 1 in 10 moves 22.8 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 4.7 s, 1 in 10 moves 10.3 s+
  • Grok 4.7 median 4.7 s, 1 in 10 moves 17.7 s+
  • Grok 4.7 (low reasoning) median 5.4 s, 1 in 10 moves 15.5 s+
  • MiMo-V2.6-Flash median 4.0 s, 1 in 10 moves 11.7 s+
  • GLM 4.7 Flash median 9.7 s, 1 in 10 moves 23.1 s+
  • GLM 5.3 Flash (low reasoning) median 2.2 s, 1 in 10 moves 5.3 s+

Lightning 10+1

  • Claude Fable 5.1 median 4.4 s, 1 in 10 moves 6.7 s+
  • Claude Fable 5.1 (low reasoning) median 4.1 s, 1 in 10 moves 6.7 s+
  • Claude Haiku 5.5 median 1.9 s
  • Claude Haiku 5.5 (low reasoning) median 2.0 s
  • Claude Haiku 5.5 (no reasoning) median 1.8 s
  • Claude Opus 5.5 median 3.0 s, 1 in 10 moves 3.5 s+
  • Claude Opus 5.5 (low reasoning) median 2.8 s, 1 in 10 moves 3.3 s+
  • Claude Sonnet 5.5 median 2.1 s, 1 in 10 moves 2.7 s+
  • DeepSeek V4 Pro 0423 median 4.6 s, 1 in 10 moves 6.6 s+
  • DeepSeek V4.1 Flash (low reasoning) median 0.8 s, 1 in 10 moves 1.7 s+
  • Gemini 3.1 Pro Preview median 2.7 s, 1 in 10 moves 3.4 s+
  • Gemini 3.1 Pro Preview (low reasoning) median 2.5 s, 1 in 10 moves 3.0 s+
  • Gemini 3.8 Flash median 2.3 s, 1 in 10 moves 2.9 s+
  • Muse Spark 1.3 median 4.6 s, 1 in 10 moves 9.5 s+
  • Muse Spark 1.3 (minimal reasoning) median 2.2 s, 1 in 10 moves 4.9 s+
  • Kimi K3 median 1.4 s, 1 in 10 moves 6.5 s+
  • Kimi K3 (low reasoning) median 1.7 s, 1 in 10 moves 4.2 s+
  • GPT-6 Astra median 2.0 s, 1 in 10 moves 3.4 s+
  • GPT-6 Astra (low reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Luna median 2.9 s, 1 in 10 moves 4.2 s+
  • GPT-6 Luna (no reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Sol median 1.6 s, 1 in 10 moves 3.7 s+
  • GPT-6 Sol (no reasoning) median 1.6 s, 1 in 10 moves 2.0 s+
  • Qwen3.8 Flash median 2.5 s, 1 in 10 moves 3.8 s+
  • Qwen3.8 Max (0902) median 3.1 s, 1 in 10 moves 4.6 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 2.9 s, 1 in 10 moves 4.9 s+
  • Grok 4.7 median 2.5 s, 1 in 10 moves 5.5 s+
  • Grok 4.7 (low reasoning) median 2.3 s, 1 in 10 moves 4.3 s+
  • MiMo-V2.6-Flash median 3.1 s, 1 in 10 moves 4.6 s+
  • GLM 4.7 Flash median 4.4 s, 1 in 10 moves 6.0 s+
  • GLM 5.3 median 1.8 s, 1 in 10 moves 6.2 s+
  • GLM 5.3 (low reasoning) median 1.2 s, 1 in 10 moves 4.0 s+
  • GLM 5.3 Flash (low reasoning) median 0.9 s, 1 in 10 moves 2.6 s+

Blitz 3+2

  • DeepSeek V4 Pro 0423 median 6.4 s, 1 in 10 moves 29.7 s+
  • Muse Spark 1.3 median 19.4 s, 1 in 10 moves 60.0 s+
  • Mistral Large 4 (high reasoning) median 26.1 s, 1 in 10 moves 69.5 s+
  • Kimi K3 median 6.8 s, 1 in 10 moves 32.1 s+
  • GPT-6 Sol median 8.2 s, 1 in 10 moves 22.9 s+
  • Qwen3.8 Max (0902) median 7.4 s, 1 in 10 moves 88.9 s+
  • Grok 4.7 median 9.9 s, 1 in 10 moves 32.9 s+

Decision models

Decision models answer a typed question instead of writing text. Each move is one question whose options are the legal moves, so they cannot play an illegal move, and they are not told their clock. They play the same clock, ladder and engine and are in the table above; here they are ranked among themselves.

BulletBench Lightning 10+1: ladder Elo, ladder Elo, higher is better
#ModelLightning 10+1 · 95% range
ladder Elo, higher is better
Bullet 60s
ladder Elo
Move time
seconds
Lost on time
% of games
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1≈ Mercury DecideInception
614
470–766
5720.4 s0.0%$0 0.5 s12
2≈ d1Liquid AI
572
428–718
5840.6 s0.0%$0.0005 0.7 s12
3≈ Jev 1.13TypeSafe
571
429–719
5720.5 s0.0%$0.0014 0.7 s12
4≈ Clef FlashCloudflare
529
388–670
5290.5 s0.0%$0.0024 0.7 s12
4≈ ClefCloudflare
529
373–653
5290.7 s0.0%$0.0072 1.0 s12
4≈ GPT-6 Luna DecisionsOpenAI
529
407–643
4890.5 s0.0%$0.0025 0.5 s12
7≈ Decider V1 27BPerplexity
420
367–488
3850.6 s0.0%$0.0009 0.7 s12
8 Kev 4BJared Palmer
33
0–270
–1.0 s66.7%$0.0006 1.4 s12

Swipe the table sideways for more columns.

Too slow for the clock. These lost their first four games on time and are not rated for that clock.

Bullet 60s

  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+

Real games: the clock at Lightning 10+1

Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.

0 s5 s10 s15 s 0481216202428 Move Clock left Won by checkmate after 28 moves Gemini 3.5 Flash Lite Lost on time after 3 moves Claude Opus 5.5
Gemini 3.5 Flash LiteWon by checkmate after 28 moves
  1. e40.6 s10.4 s left
  2. Nf30.9 s10.4 s left
  3. Bc40.6 s10.9 s left
  4. O-O0.9 s10.9 s left
  5. c30.6 s11.4 s left
  6. d40.9 s11.4 s left
  7. Nxe51.0 s11.4 s left
  8. Nxc60.6 s11.8 s left
  9. Qxg40.9 s11.9 s left
  10. Qxc8+0.9 s12.0 s left
  11. … 18 more moves
Claude Opus 5.5Lost on time after 3 moves
  1. e43.5 s7.5 s left
  2. Nf33.5 s5.0 s left
  3. d43.3 s2.6 s left

More from the results

Bullet: strength against thinking time

The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.

Named: the six bestOther models (hover for names)

-50005001,0001,500 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1.5 s a move Gemini 3.8 Flash (low reasoning) Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Gemini 3.5 Flash-Lite Gemini 3.6 Flash (minimal reasoning) d1

How games are lost

Share of games lost on time, and moves with no legal reply, by clock.

ModelOn time, LightningOn time, BulletInvalid moves, LightningInvalid moves, Bullet
Gemini 3.5 Flash-LiteGoogle0.0%0.0%0.0%0.0%
Gemini 3.1 Flash-Lite (minimal reasoning)Google0.0%8.3%0.2%0.0%
Gemini 3.5 Flash (minimal reasoning)Google33.3%25.0%0.6%0.5%
Mercury DecideInception0.0%0.0%0.0%0.0%
d1Liquid AI0.0%0.0%0.0%0.0%
Seed-2.0-Mini (minimal reasoning)ByteDance0.0%0.0%0.8%0.4%
Mistral Small 4 (no reasoning)Mistral AI0.0%25.0%1.0%0.6%
GPT-5.4 mini (no reasoning)OpenAI0.0%50.0%0.0%0.0%
Inkling Small (no reasoning)Thinking Machines0.0%0.0%5.3%7.2%
Jev 1.13TypeSafe0.0%0.0%0.0%0.0%
Clef FlashCloudflare0.0%0.0%0.0%0.0%
ClefCloudflare0.0%0.0%0.0%0.0%
Mistral Medium 3.5 (no reasoning)Mistral AI0.0%0.0%0.0%0.4%
Mistral Medium 3.5Mistral AI0.0%8.3%0.3%0.8%
GPT-6 Luna DecisionsOpenAI0.0%0.0%0.0%0.0%
Decider V1 27BPerplexity0.0%8.3%0.0%0.0%
Mistral Large 4Mistral AI58.3%–1.8%–
Gemini 3.8 Flash (low reasoning)Google83.3%50.0%0.0%0.0%
GPT-5.4 nano (no reasoning)OpenAI75.0%66.7%4.0%3.1%
Claude Haiku 4.5Anthropic83.3%25.0%0.0%0.3%
Claude Sonnet 5.5 (low reasoning)Anthropic91.7%–0.0%–
Gemini 3.6 Flash (minimal reasoning)Google91.7%41.7%0.0%0.3%
Kev 4BJared Palmer66.7%–0.0%–

Lightning: strength against cost

What one Lightning game cost at the run date's prices, on a log scale. Up and to the left is better; the dashed line joins the strongest model at each cost.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

-200300800 $0.0001$0.001$0.01$0.1 Cost per game, US dollars (log scale) Ladder Elo Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) d1 Seed-2.0-Mini (minimal reasoning) Mistral Small 4 (no reasoning)

Best per budget

  1. d1: 572 at $0.0005
  2. Gemini 3.1 Flash-Lite (minimal reasoning): 780 at $0.0058
  3. Gemini 3.5 Flash-Lite: 892 at $0.0070

Cheapest first: each model here beats every cheaper one on score.

How BulletBench works

  1. 1

    The position

    The model gets the position, the moves so far, its remaining clock and the legal moves.

  2. 2

    The reply

    It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.

  3. 3

    The engine

    Stockfish replies instantly at the model's current ladder level.

  4. 4

    The ladder

    Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.

Why chess on a clock

Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.

The clocks

Lightning 10+1
Ten seconds, plus one second a move. A model can play for ever, but only at about a second an answer.
Bullet 60s
One minute for the whole game, no increment.
Blitz 3+2
Three minutes plus two seconds a move: room for reasoning models, for comparison.

Fair play

Same prompt
Every model gets the same system prompt and position format.
Invalid replies
A reply with no legal move gets two corrective retries, then a random legal move is played and counted.
Settings
Each model runs at its provider's default and, where the provider offers one, its fastest reasoning setting, shown separately.

Honest caveats

A dozen games or fewer per clock (the table gives each configuration's count) give wide intervals, so neighbouring ranks are often within each other's ranges. Ratings also come in steps: the engine plays at fixed levels and the rating is fitted from which levels a model beat, drew or lost to, so models with the same results against the same levels get exactly the same rating. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.

What it measures

  • Decision quality under real time pressure
  • Response latency at the provider's default settings
  • Which reasoning settings are fast enough to use

What it does not measure

  • Any business skill: this is chess against an engine
  • Chess strength without a clock

Method

  • Games against a Stockfish ladder; every second of response time comes off the model's clock
  • Each model at its default setting and, where offered, its fastest reasoning setting
  • Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating