Benchmarks / BulletBench

Measured by Spring Prompt

BulletBench

When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?

Last updated 8 Oct 2026

Results dated
1 Oct 2026 to 8 Oct 2026
Results
31 configurations of 27 models
Unit
% of games
Licence
Spring Prompt original
Sample
8 to 12 games per configuration

Games lost on time: Claude Haiku 5.5

17 results above zero · % of games, lower is better. Choose a model to highlight it.Clear highlight

  1. 11 Gemini 3.1 Flash-Lite (minimal reasoning)Google 8.3%
  2. 11 Decider V1 27BDecision modelPerplexity 8.3%
  3. 13 Claude Haiku 4.5Anthropic 25.0%
  4. 13 Gemini 3.5 Flash (minimal reasoning)Google 25.0%
  5. 13 Mistral Small 4 (no reasoning)Mistral AI 25.0%
  6. 16 Claude Haiku 5.5 (no reasoning)Anthropic 37.5%
  7. 17 Gemini 3.6 Flash (minimal reasoning)Google 41.7%
  8. 18 Gemini 3.8 Flash (low reasoning)Google 50.0%
  9. 18 GPT-5.4 mini (no reasoning)OpenAI 50.0%
  10. 18 GPT-6 Luna (no reasoning)OpenAI 50.0%
  11. 18 GPT-6 Sol (no reasoning)OpenAI 50.0%
  12. 22 Gemini 3.1 Pro Preview (low reasoning)Google 58.3%
  13. 23 GPT-5.4 nano (no reasoning)OpenAI 66.7%
  14. 24 GPT-6 Astra (low reasoning)OpenAI 75.0%
  15. 25 GLM-5.3 (low reasoning)Z.ai 83.3%

10 models had none: Clef, Clef Flash, GPT-6 Luna Decisions, Gemini 3.5 Flash-Lite, Inkling Small (no reasoning), Jev 1.13, Mercury Decide, Mistral Medium 3.5 (no reasoning), Seed-2.0-Mini (minimal reasoning), d1.

Lightning: strength against thinking time

Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.

Named: the six bestOther models (hover for names)

-200300800 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1 s increment Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Mercury Decide d1 Seed-2.0-Mini (minimal reasoning)

Full results

BulletBench Bullet 60s: games lost on time, % of games, lower is better
#ModelGames lost on time, Bullet 60s
% of games, lower is better
Lightning 10+1
ladder Elo
Bullet 60s
ladder Elo
Blitz 3+2
ladder Elo
Move time
seconds
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1≈ Seed-2.0-Mini (minimal reasoning)ByteDance
0.0%
571453–0.3 s$0.0030 0.4 s12
1≈ Clef FlashDecision modelCloudflare
0.0%
529529–0.5 s$0.0025 0.7 s12
1≈ ClefDecision modelCloudflare
0.0%
529529–0.7 s$0.0056 1.0 s12
1≈ Gemini 3.5 Flash-LiteGoogle
0.0%
8927807070.9 s$0.0058 ––
1≈ Mercury DecideDecision modelInception
0.0%
614572–0.4 s$0 0.5 s12
1≈ d1Decision modelLiquid AI
0.0%
572584–0.6 s$0.0007 0.7 s12
1≈ Mistral Medium 3.5 (no reasoning)Mistral AI · best of 2 settings
0.0%
529572–0.5 s$0.0354 0.7 s12
1≈ GPT-6 Luna DecisionsDecision modelOpenAI
0.0%
529489–0.5 s$0.0026 0.7 s12
1≈ Inkling Small (no reasoning)Thinking Machines
0.0%
571529–0.6 s$0.0075 1.5 s12
1≈ Jev 1.13Decision modelTypeSafe
0.0%
571572–0.5 s$0.0013 0.6 s12
11 Gemini 3.1 Flash-Lite (minimal reasoning)Google
8.3%
780912–0.8 s$0.0068 ––
11 Decider V1 27BDecision modelPerplexity
8.3%
420385–0.6 s$0.0010 0.8 s12
13 Claude Haiku 4.5Anthropic
25.0%
334454911.2 s$0.0163 1.7 s12
13 Gemini 3.5 Flash (minimal reasoning)Google
25.0%
727810–1.2 s$0.0289 1.4 s12
13 Mistral Small 4 (no reasoning)Mistral AI
25.0%
571367–0.6 s$0.0039 1.1 s12
16 Claude Haiku 5.5 (no reasoning)Anthropic
37.5%
–310–1.7 s$0.0019 ––
17 Gemini 3.6 Flash (minimal reasoning)Google
41.7%
33628–1.6 s$0.0113 1.9 s12
18 Gemini 3.8 Flash (low reasoning)Google · best of 2 settings
50.0%
1971,068–1.8 s$0.0165 2.2 s12
18 GPT-5.4 mini (no reasoning)OpenAI
50.0%
571197–0.9 s$0.0205 ––
18 GPT-6 Luna (no reasoning)OpenAI
50.0%
–249–1.5 s$0.0016 2.3 s12
18 GPT-6 Sol (no reasoning)OpenAI
50.0%
–445–1.6 s$0.0297 2.1 s12
22 Gemini 3.1 Pro Preview (low reasoning)Google
58.3%
–529–3.1 s$0.0446 4.4 s12
23 GPT-5.4 nano (no reasoning)OpenAI
66.7%
19733–1.1 s$0.0050 ––
24 GPT-6 Astra (low reasoning)OpenAI · best of 2 settings
75.0%
–323–2.3 s$0.10 5.7 s12
25 GLM-5.3 (low reasoning)Z.ai · best of 2 settings
83.3%
–33–1.6 s$0.0220 ––
26 Muse Spark 1.3 (minimal reasoning)Meta
91.7%
–33–4.2 s$0.0142 11.4 s12
26 Kimi K3 (low reasoning)Moonshot AI
91.7%
–33–2.6 s$0.0540 11.3 s12

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only. Equal ratings are not a glitch: the ladder's ratings come in steps (see Honest caveats below).

Too slow for the clock

These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Each shows its median time per move and the time that 1 move in 10 took or exceeded: a median under the increment still loses on time when the slow tail is long. Being too slow is the result, not a fault.

Bullet 60s

  • Claude Fable 5.1 median 4.2 s, 1 in 10 moves 6.3 s+
  • Claude Fable 5.1 (low reasoning) median 4.3 s, 1 in 10 moves 9.3 s+
  • Claude Haiku 5.5 median 4.7 s
  • Claude Haiku 5.5 (low reasoning) median 3.6 s
  • Claude Opus 5.5 median 3.1 s, 1 in 10 moves 5.8 s+
  • Claude Opus 5.5 (low reasoning) median 3.1 s, 1 in 10 moves 5.3 s+
  • Claude Sonnet 5.5 median 3.0 s, 1 in 10 moves 9.4 s+
  • Claude Sonnet 5.5 (low reasoning) median 2.6 s, 1 in 10 moves 5.0 s+
  • DeepSeek V4 Pro 0423 median 5.0 s, 1 in 10 moves 27.8 s+
  • DeepSeek V4.1 Flash (low reasoning) median 1.9 s, 1 in 10 moves 10.6 s+
  • Gemini 3.1 Pro Preview median 4.9 s, 1 in 10 moves 8.4 s+
  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+
  • Muse Spark 1.3 median 7.0 s, 1 in 10 moves 28.1 s+
  • Mistral Large 4 median 0.8 s, 1 in 10 moves 2.9 s+
  • Kimi K3 median 3.2 s, 1 in 10 moves 33.8 s+
  • GPT-6 Luna median 4.6 s, 1 in 10 moves 10.8 s+
  • GPT-6 Sol median 4.2 s, 1 in 10 moves 9.4 s+
  • Qwen3.8 Flash median 3.8 s, 1 in 10 moves 16.9 s+
  • Qwen3.8 Max (0902) median 4.9 s, 1 in 10 moves 22.8 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 4.7 s, 1 in 10 moves 10.3 s+
  • Grok 4.7 median 4.7 s, 1 in 10 moves 17.7 s+
  • Grok 4.7 (low reasoning) median 5.4 s, 1 in 10 moves 15.5 s+
  • MiMo-V2.6-Flash median 4.0 s, 1 in 10 moves 11.7 s+
  • GLM 4.7 Flash median 9.7 s, 1 in 10 moves 23.1 s+
  • GLM 5.3 Flash (low reasoning) median 2.2 s, 1 in 10 moves 5.3 s+

Lightning 10+1

  • Claude Fable 5.1 median 4.4 s, 1 in 10 moves 6.7 s+
  • Claude Fable 5.1 (low reasoning) median 4.1 s, 1 in 10 moves 6.7 s+
  • Claude Haiku 5.5 median 1.9 s
  • Claude Haiku 5.5 (low reasoning) median 2.0 s
  • Claude Haiku 5.5 (no reasoning) median 1.8 s
  • Claude Opus 5.5 median 3.0 s, 1 in 10 moves 3.5 s+
  • Claude Opus 5.5 (low reasoning) median 2.8 s, 1 in 10 moves 3.3 s+
  • Claude Sonnet 5.5 median 2.1 s, 1 in 10 moves 2.7 s+
  • DeepSeek V4 Pro 0423 median 4.6 s, 1 in 10 moves 6.6 s+
  • DeepSeek V4.1 Flash (low reasoning) median 0.8 s, 1 in 10 moves 1.7 s+
  • Gemini 3.1 Pro Preview median 2.7 s, 1 in 10 moves 3.4 s+
  • Gemini 3.1 Pro Preview (low reasoning) median 2.5 s, 1 in 10 moves 3.0 s+
  • Gemini 3.8 Flash median 2.3 s, 1 in 10 moves 2.9 s+
  • Muse Spark 1.3 median 4.6 s, 1 in 10 moves 9.5 s+
  • Muse Spark 1.3 (minimal reasoning) median 2.2 s, 1 in 10 moves 4.9 s+
  • Kimi K3 median 1.4 s, 1 in 10 moves 6.5 s+
  • Kimi K3 (low reasoning) median 1.7 s, 1 in 10 moves 4.2 s+
  • GPT-6 Astra median 2.0 s, 1 in 10 moves 3.4 s+
  • GPT-6 Astra (low reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Luna median 2.9 s, 1 in 10 moves 4.2 s+
  • GPT-6 Luna (no reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Sol median 1.6 s, 1 in 10 moves 3.7 s+
  • GPT-6 Sol (no reasoning) median 1.6 s, 1 in 10 moves 2.0 s+
  • Qwen3.8 Flash median 2.5 s, 1 in 10 moves 3.8 s+
  • Qwen3.8 Max (0902) median 3.1 s, 1 in 10 moves 4.6 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 2.9 s, 1 in 10 moves 4.9 s+
  • Grok 4.7 median 2.5 s, 1 in 10 moves 5.5 s+
  • Grok 4.7 (low reasoning) median 2.3 s, 1 in 10 moves 4.3 s+
  • MiMo-V2.6-Flash median 3.1 s, 1 in 10 moves 4.6 s+
  • GLM 4.7 Flash median 4.4 s, 1 in 10 moves 6.0 s+
  • GLM 5.3 median 1.8 s, 1 in 10 moves 6.2 s+
  • GLM 5.3 (low reasoning) median 1.2 s, 1 in 10 moves 4.0 s+
  • GLM 5.3 Flash (low reasoning) median 0.9 s, 1 in 10 moves 2.6 s+

Blitz 3+2

  • DeepSeek V4 Pro 0423 median 6.4 s, 1 in 10 moves 29.7 s+
  • Muse Spark 1.3 median 19.4 s, 1 in 10 moves 60.0 s+
  • Mistral Large 4 (high reasoning) median 26.1 s, 1 in 10 moves 69.5 s+
  • Kimi K3 median 6.8 s, 1 in 10 moves 32.1 s+
  • GPT-6 Sol median 8.2 s, 1 in 10 moves 22.9 s+
  • Qwen3.8 Max (0902) median 7.4 s, 1 in 10 moves 88.9 s+
  • Grok 4.7 median 9.9 s, 1 in 10 moves 32.9 s+

Decision models

Decision models answer a typed question instead of writing text. Each move is one question whose options are the legal moves, so they cannot play an illegal move, and they are not told their clock. They play the same clock, ladder and engine and are in the table above; here they are ranked among themselves.

BulletBench Bullet 60s: games lost on time, % of games, lower is better
#ModelGames lost on time, Bullet 60s
% of games, lower is better
Lightning 10+1
ladder Elo
Bullet 60s
ladder Elo
Move time
seconds
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1≈ Clef FlashCloudflare
0.0%
5295290.5 s$0.0025 0.7 s12
1≈ ClefCloudflare
0.0%
5295290.7 s$0.0056 1.0 s12
1≈ Mercury DecideInception
0.0%
6145720.4 s$0 0.5 s12
1≈ d1Liquid AI
0.0%
5725840.6 s$0.0007 0.7 s12
1≈ GPT-6 Luna DecisionsOpenAI
0.0%
5294890.5 s$0.0026 0.7 s12
1≈ Jev 1.13TypeSafe
0.0%
5715720.5 s$0.0013 0.6 s12
7 Decider V1 27BPerplexity
8.3%
4203850.6 s$0.0010 0.8 s12

Swipe the table sideways for more columns.

Too slow for the clock. These lost their first four games on time and are not rated for that clock.

Bullet 60s

  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+

Real games: the clock at Lightning 10+1

Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.

0 s5 s10 s15 s 0481216202428 Move Clock left Won by checkmate after 28 moves Gemini 3.5 Flash Lite Lost on time after 3 moves Claude Opus 5.5
Gemini 3.5 Flash LiteWon by checkmate after 28 moves
  1. e40.6 s10.4 s left
  2. Nf30.9 s10.4 s left
  3. Bc40.6 s10.9 s left
  4. O-O0.9 s10.9 s left
  5. c30.6 s11.4 s left
  6. d40.9 s11.4 s left
  7. Nxe51.0 s11.4 s left
  8. Nxc60.6 s11.8 s left
  9. Qxg40.9 s11.9 s left
  10. Qxc8+0.9 s12.0 s left
  11. … 18 more moves
Claude Opus 5.5Lost on time after 3 moves
  1. e43.5 s7.5 s left
  2. Nf33.5 s5.0 s left
  3. d43.3 s2.6 s left

More from the results

Bullet: strength against thinking time

The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.

Named: the six bestOther models (hover for names)

-50005001,0001,500 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1.5 s a move Gemini 3.8 Flash (low reasoning) Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Gemini 3.5 Flash-Lite Gemini 3.6 Flash (minimal reasoning) d1

How games are lost

Share of games lost on time, and moves with no legal reply, by clock.

ModelOn time, LightningOn time, BulletInvalid moves, LightningInvalid moves, Bullet
Gemini 3.5 Flash-LiteGoogle0.0%0.0%0.0%0.0%
Gemini 3.1 Flash-Lite (minimal reasoning)Google0.0%8.3%0.2%0.0%
Gemini 3.5 Flash (minimal reasoning)Google33.3%25.0%0.6%0.5%
Mercury DecideInception0.0%0.0%0.0%0.0%
d1Liquid AI0.0%0.0%0.0%0.0%
Seed-2.0-Mini (minimal reasoning)ByteDance0.0%0.0%0.8%0.4%
Mistral Small 4 (no reasoning)Mistral AI0.0%25.0%1.0%0.6%
GPT-5.4 mini (no reasoning)OpenAI0.0%50.0%0.0%0.0%
Inkling Small (no reasoning)Thinking Machines0.0%0.0%5.3%7.2%
Jev 1.13TypeSafe0.0%0.0%0.0%0.0%
Clef FlashCloudflare0.0%0.0%0.0%0.0%
ClefCloudflare0.0%0.0%0.0%0.0%
Mistral Medium 3.5 (no reasoning)Mistral AI0.0%0.0%0.0%0.4%
Mistral Medium 3.5Mistral AI0.0%8.3%0.3%0.8%
GPT-6 Luna DecisionsOpenAI0.0%0.0%0.0%0.0%
Decider V1 27BPerplexity0.0%8.3%0.0%0.0%
Mistral Large 4Mistral AI58.3%–1.8%–
Gemini 3.8 Flash (low reasoning)Google83.3%50.0%0.0%0.0%
GPT-5.4 nano (no reasoning)OpenAI75.0%66.7%4.0%3.1%
Claude Haiku 4.5Anthropic83.3%25.0%0.0%0.3%
Claude Sonnet 5.5 (low reasoning)Anthropic91.7%–0.0%–
Gemini 3.6 Flash (minimal reasoning)Google91.7%41.7%0.0%0.3%
Kev 4BJared Palmer66.7%–0.0%–

Lightning: strength against cost

What one Lightning game cost at the run date's prices, on a log scale. Up and to the left is better; the dashed line joins the strongest model at each cost.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

-200300800 $0.0001$0.001$0.01$0.1 Cost per game, US dollars (log scale) Ladder Elo Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) d1 Seed-2.0-Mini (minimal reasoning) Mistral Small 4 (no reasoning)

Best per budget

  1. d1: 572 at $0.0005
  2. Gemini 3.1 Flash-Lite (minimal reasoning): 780 at $0.0058
  3. Gemini 3.5 Flash-Lite: 892 at $0.0070

Cheapest first: each model here beats every cheaper one on score.

How BulletBench works

  1. 1

    The position

    The model gets the position, the moves so far, its remaining clock and the legal moves.

  2. 2

    The reply

    It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.

  3. 3

    The engine

    Stockfish replies instantly at the model's current ladder level.

  4. 4

    The ladder

    Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.

Why chess on a clock

Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.

The clocks

Lightning 10+1
Ten seconds, plus one second a move. A model can play for ever, but only at about a second an answer.
Bullet 60s
One minute for the whole game, no increment.
Blitz 3+2
Three minutes plus two seconds a move: room for reasoning models, for comparison.

Fair play

Same prompt
Every model gets the same system prompt and position format.
Invalid replies
A reply with no legal move gets two corrective retries, then a random legal move is played and counted.
Settings
Each model runs at its provider's default and, where the provider offers one, its fastest reasoning setting, shown separately.

Honest caveats

A dozen games or fewer per clock (the table gives each configuration's count) give wide intervals, so neighbouring ranks are often within each other's ranges. Ratings also come in steps: the engine plays at fixed levels and the rating is fitted from which levels a model beat, drew or lost to, so models with the same results against the same levels get exactly the same rating. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.

What it measures

  • Decision quality under real time pressure
  • Response latency at the provider's default settings
  • Which reasoning settings are fast enough to use

What it does not measure

  • Any business skill: this is chess against an engine
  • Chess strength without a clock

Method

  • Games against a Stockfish ladder; every second of response time comes off the model's clock
  • Each model at its default setting and, where offered, its fastest reasoning setting
  • Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating