Benchmarks / BulletBench

Measured by Spring Prompt

BulletBench

When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?

Last updated 8 Oct 2026

Results dated
1 Oct 2026 to 7 Oct 2026
Results
23 configurations of 22 models
Unit
milliseconds
Licence
Spring Prompt original
Sample
12 games per configuration

Median move time: Clef Flash

Top 15 of 22 results · milliseconds, lower is better. Choose a model to highlight it.Clear highlight

  1. 1 Seed-2.0-Mini (minimal reasoning)ByteDance 0.3 s
  2. 2 Mercury DecideDecision modelInception 0.4 s
  3. 3 Mistral Medium 3.5Mistral AI 0.5 s
  4. 3 Jev 1.13Decision modelTypeSafe 0.5 s
  5. 3 GPT-6 Luna DecisionsDecision modelOpenAI 0.5 s
  6. 3 Clef FlashDecision modelCloudflare 0.5 s
  7. 7 Mistral Small 4 (no reasoning)Mistral AI 0.6 s
  8. 7 Inkling Small (no reasoning)Thinking Machines 0.6 s
  9. 7 Decider V1 27BDecision modelPerplexity 0.6 s
  10. 7 d1Decision modelLiquid AI 0.6 s
  11. 11 ClefDecision modelCloudflare 0.7 s
  12. 12 Gemini 3.1 Flash-Lite (minimal reasoning)Google 0.8 s
  13. 12 Mistral Large 4Mistral AI 0.8 s
  14. 14 Gemini 3.5 Flash-LiteGoogle 0.9 s
  15. 14 GPT-5.4 mini (no reasoning)OpenAI 0.9 s

Lightning: strength against thinking time

Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.

Named: the six bestOther models (hover for names)

-200300800 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1 s increment Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Mercury Decide d1 Seed-2.0-Mini (minimal reasoning)

Full results

BulletBench Lightning 10+1: median move time, milliseconds, lower is better
#ModelMedian move time, Lightning 10+1
milliseconds, lower is better
Lightning 10+1
ladder Elo
Bullet 60s
ladder Elo
Blitz 3+2
ladder Elo
Lost on time
% of games
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1 Seed-2.0-Mini (minimal reasoning)ByteDance
0.3 s
571453–0.0%$0.0019 0.4 s12
2 Mercury DecideDecision modelInception
0.4 s
614572–0.0%$0 0.5 s12
3 Mistral Medium 3.5Mistral AI · best of 2 settings
0.5 s
5293855550.0%$0.0297 0.8 s12
3 Jev 1.13Decision modelTypeSafe
0.5 s
571572–0.0%$0.0014 0.7 s12
3 GPT-6 Luna DecisionsDecision modelOpenAI
0.5 s
529489–0.0%$0.0025 0.5 s12
3 Clef FlashDecision modelCloudflare
0.5 s
529529–0.0%$0.0024 0.7 s12
7 Mistral Small 4 (no reasoning)Mistral AI
0.6 s
571367–0.0%$0.0041 1.2 s12
7 Inkling Small (no reasoning)Thinking Machines
0.6 s
571529–0.0%$0.0077 1.6 s12
7 Decider V1 27BDecision modelPerplexity
0.6 s
420385–0.0%$0.0009 0.7 s12
7 d1Decision modelLiquid AI
0.6 s
572584–0.0%$0.0005 0.7 s12
11 ClefDecision modelCloudflare
0.7 s
529529–0.0%$0.0072 1.0 s12
12 Gemini 3.1 Flash-Lite (minimal reasoning)Google
0.8 s
780912–0.0%$0.0058 ––
12 Mistral Large 4Mistral AI
0.8 s
249–55558.3%$0.0117 1.9 s12
14 Gemini 3.5 Flash-LiteGoogle
0.9 s
8927807070.0%$0.0070 ––
14 GPT-5.4 mini (no reasoning)OpenAI
0.9 s
571197–0.0%$0.0136 ––
16 Kev 4BDecision modelJared Palmer
1.0 s
33––66.7%$0.0006 1.4 s12
16 GPT-5.4 nano (no reasoning)OpenAI
1.0 s
19733–75.0%$0.0025 ––
18 Gemini 3.5 Flash (minimal reasoning)Google
1.2 s
727810–33.3%$0.0213 1.3 s12
19 Claude Haiku 4.5Anthropic
1.3 s
3344549183.3%$0.0095 1.8 s12
20 Gemini 3.6 Flash (minimal reasoning)Google
1.5 s
33628–91.7%$0.0050 1.9 s12
21 Gemini 3.8 Flash (low reasoning)Google
1.8 s
1971,068–83.3%$0.0041 2.1 s12
22 Claude Sonnet 5.5 (low reasoning)Anthropic
2.1 s
33––91.7%$0.0080 2.5 s12

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only. Equal ratings are not a glitch: the ladder's ratings come in steps (see Honest caveats below).

Too slow for the clock

These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Each shows its median time per move and the time that 1 move in 10 took or exceeded: a median under the increment still loses on time when the slow tail is long. Being too slow is the result, not a fault.

Bullet 60s

  • Claude Fable 5.1 median 4.2 s, 1 in 10 moves 6.3 s+
  • Claude Fable 5.1 (low reasoning) median 4.3 s, 1 in 10 moves 9.3 s+
  • Claude Haiku 5.5 median 4.7 s
  • Claude Haiku 5.5 (low reasoning) median 3.6 s
  • Claude Opus 5.5 median 3.1 s, 1 in 10 moves 5.8 s+
  • Claude Opus 5.5 (low reasoning) median 3.1 s, 1 in 10 moves 5.3 s+
  • Claude Sonnet 5.5 median 3.0 s, 1 in 10 moves 9.4 s+
  • Claude Sonnet 5.5 (low reasoning) median 2.6 s, 1 in 10 moves 5.0 s+
  • DeepSeek V4 Pro 0423 median 5.0 s, 1 in 10 moves 27.8 s+
  • DeepSeek V4.1 Flash (low reasoning) median 1.9 s, 1 in 10 moves 10.6 s+
  • Gemini 3.1 Pro Preview median 4.9 s, 1 in 10 moves 8.4 s+
  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+
  • Muse Spark 1.3 median 7.0 s, 1 in 10 moves 28.1 s+
  • Mistral Large 4 median 0.8 s, 1 in 10 moves 2.9 s+
  • Kimi K3 median 3.2 s, 1 in 10 moves 33.8 s+
  • GPT-6 Luna median 4.6 s, 1 in 10 moves 10.8 s+
  • GPT-6 Sol median 4.2 s, 1 in 10 moves 9.4 s+
  • Qwen3.8 Flash median 3.8 s, 1 in 10 moves 16.9 s+
  • Qwen3.8 Max (0902) median 4.9 s, 1 in 10 moves 22.8 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 4.7 s, 1 in 10 moves 10.3 s+
  • Grok 4.7 median 4.7 s, 1 in 10 moves 17.7 s+
  • Grok 4.7 (low reasoning) median 5.4 s, 1 in 10 moves 15.5 s+
  • MiMo-V2.6-Flash median 4.0 s, 1 in 10 moves 11.7 s+
  • GLM 4.7 Flash median 9.7 s, 1 in 10 moves 23.1 s+
  • GLM 5.3 Flash (low reasoning) median 2.2 s, 1 in 10 moves 5.3 s+

Lightning 10+1

  • Claude Fable 5.1 median 4.4 s, 1 in 10 moves 6.7 s+
  • Claude Fable 5.1 (low reasoning) median 4.1 s, 1 in 10 moves 6.7 s+
  • Claude Haiku 5.5 median 1.9 s
  • Claude Haiku 5.5 (low reasoning) median 2.0 s
  • Claude Haiku 5.5 (no reasoning) median 1.8 s
  • Claude Opus 5.5 median 3.0 s, 1 in 10 moves 3.5 s+
  • Claude Opus 5.5 (low reasoning) median 2.8 s, 1 in 10 moves 3.3 s+
  • Claude Sonnet 5.5 median 2.1 s, 1 in 10 moves 2.7 s+
  • DeepSeek V4 Pro 0423 median 4.6 s, 1 in 10 moves 6.6 s+
  • DeepSeek V4.1 Flash (low reasoning) median 0.8 s, 1 in 10 moves 1.7 s+
  • Gemini 3.1 Pro Preview median 2.7 s, 1 in 10 moves 3.4 s+
  • Gemini 3.1 Pro Preview (low reasoning) median 2.5 s, 1 in 10 moves 3.0 s+
  • Gemini 3.8 Flash median 2.3 s, 1 in 10 moves 2.9 s+
  • Muse Spark 1.3 median 4.6 s, 1 in 10 moves 9.5 s+
  • Muse Spark 1.3 (minimal reasoning) median 2.2 s, 1 in 10 moves 4.9 s+
  • Kimi K3 median 1.4 s, 1 in 10 moves 6.5 s+
  • Kimi K3 (low reasoning) median 1.7 s, 1 in 10 moves 4.2 s+
  • GPT-6 Astra median 2.0 s, 1 in 10 moves 3.4 s+
  • GPT-6 Astra (low reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Luna median 2.9 s, 1 in 10 moves 4.2 s+
  • GPT-6 Luna (no reasoning) median 2.1 s, 1 in 10 moves 3.1 s+
  • GPT-6 Sol median 1.6 s, 1 in 10 moves 3.7 s+
  • GPT-6 Sol (no reasoning) median 1.6 s, 1 in 10 moves 2.0 s+
  • Qwen3.8 Flash median 2.5 s, 1 in 10 moves 3.8 s+
  • Qwen3.8 Max (0902) median 3.1 s, 1 in 10 moves 4.6 s+
  • Qwen3.8 Max (0902) (minimal reasoning) median 2.9 s, 1 in 10 moves 4.9 s+
  • Grok 4.7 median 2.5 s, 1 in 10 moves 5.5 s+
  • Grok 4.7 (low reasoning) median 2.3 s, 1 in 10 moves 4.3 s+
  • MiMo-V2.6-Flash median 3.1 s, 1 in 10 moves 4.6 s+
  • GLM 4.7 Flash median 4.4 s, 1 in 10 moves 6.0 s+
  • GLM 5.3 median 1.8 s, 1 in 10 moves 6.2 s+
  • GLM 5.3 (low reasoning) median 1.2 s, 1 in 10 moves 4.0 s+
  • GLM 5.3 Flash (low reasoning) median 0.9 s, 1 in 10 moves 2.6 s+

Blitz 3+2

  • DeepSeek V4 Pro 0423 median 6.4 s, 1 in 10 moves 29.7 s+
  • Muse Spark 1.3 median 19.4 s, 1 in 10 moves 60.0 s+
  • Mistral Large 4 (high reasoning) median 26.1 s, 1 in 10 moves 69.5 s+
  • Kimi K3 median 6.8 s, 1 in 10 moves 32.1 s+
  • GPT-6 Sol median 8.2 s, 1 in 10 moves 22.9 s+
  • Qwen3.8 Max (0902) median 7.4 s, 1 in 10 moves 88.9 s+
  • Grok 4.7 median 9.9 s, 1 in 10 moves 32.9 s+

Decision models

Decision models answer a typed question instead of writing text. Each move is one question whose options are the legal moves, so they cannot play an illegal move, and they are not told their clock. They play the same clock, ladder and engine and are in the table above; here they are ranked among themselves.

BulletBench Lightning 10+1: median move time, milliseconds, lower is better
#ModelMedian move time, Lightning 10+1
milliseconds, lower is better
Lightning 10+1
ladder Elo
Bullet 60s
ladder Elo
Lost on time
% of games
Cost per game
US dollars
Slowest 10% of moves
seconds, 90th percentile
Games
at this clock
1 Mercury DecideInception
0.4 s
6145720.0%$0 0.5 s12
2 Jev 1.13TypeSafe
0.5 s
5715720.0%$0.0014 0.7 s12
2 GPT-6 Luna DecisionsOpenAI
0.5 s
5294890.0%$0.0025 0.5 s12
2 Clef FlashCloudflare
0.5 s
5295290.0%$0.0024 0.7 s12
5 Decider V1 27BPerplexity
0.6 s
4203850.0%$0.0009 0.7 s12
5 d1Liquid AI
0.6 s
5725840.0%$0.0005 0.7 s12
7 ClefCloudflare
0.7 s
5295290.0%$0.0072 1.0 s12
8 Kev 4BJared Palmer
1.0 s
33–66.7%$0.0006 1.4 s12

Swipe the table sideways for more columns.

Too slow for the clock. These lost their first four games on time and are not rated for that clock.

Bullet 60s

  • Kev 4B median 1.9 s, 1 in 10 moves 15.4 s+

Real games: the clock at Lightning 10+1

Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.

0 s5 s10 s15 s 0481216202428 Move Clock left Won by checkmate after 28 moves Gemini 3.5 Flash Lite Lost on time after 3 moves Claude Opus 5.5
Gemini 3.5 Flash LiteWon by checkmate after 28 moves
  1. e40.6 s10.4 s left
  2. Nf30.9 s10.4 s left
  3. Bc40.6 s10.9 s left
  4. O-O0.9 s10.9 s left
  5. c30.6 s11.4 s left
  6. d40.9 s11.4 s left
  7. Nxe51.0 s11.4 s left
  8. Nxc60.6 s11.8 s left
  9. Qxg40.9 s11.9 s left
  10. Qxc8+0.9 s12.0 s left
  11. … 18 more moves
Claude Opus 5.5Lost on time after 3 moves
  1. e43.5 s7.5 s left
  2. Nf33.5 s5.0 s left
  3. d43.3 s2.6 s left

More from the results

Bullet: strength against thinking time

The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.

Named: the six bestOther models (hover for names)

-50005001,0001,500 0.1 s1 s10 s Median time per move (log scale) Ladder Elo 1.5 s a move Gemini 3.8 Flash (low reasoning) Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) Gemini 3.5 Flash-Lite Gemini 3.6 Flash (minimal reasoning) d1

How games are lost

Share of games lost on time, and moves with no legal reply, by clock.

ModelOn time, LightningOn time, BulletInvalid moves, LightningInvalid moves, Bullet
Gemini 3.5 Flash-LiteGoogle0.0%0.0%0.0%0.0%
Gemini 3.1 Flash-Lite (minimal reasoning)Google0.0%8.3%0.2%0.0%
Gemini 3.5 Flash (minimal reasoning)Google33.3%25.0%0.6%0.5%
Mercury DecideInception0.0%0.0%0.0%0.0%
d1Liquid AI0.0%0.0%0.0%0.0%
Seed-2.0-Mini (minimal reasoning)ByteDance0.0%0.0%0.8%0.4%
Mistral Small 4 (no reasoning)Mistral AI0.0%25.0%1.0%0.6%
GPT-5.4 mini (no reasoning)OpenAI0.0%50.0%0.0%0.0%
Inkling Small (no reasoning)Thinking Machines0.0%0.0%5.3%7.2%
Jev 1.13TypeSafe0.0%0.0%0.0%0.0%
Clef FlashCloudflare0.0%0.0%0.0%0.0%
ClefCloudflare0.0%0.0%0.0%0.0%
Mistral Medium 3.5 (no reasoning)Mistral AI0.0%0.0%0.0%0.4%
Mistral Medium 3.5Mistral AI0.0%8.3%0.3%0.8%
GPT-6 Luna DecisionsOpenAI0.0%0.0%0.0%0.0%
Decider V1 27BPerplexity0.0%8.3%0.0%0.0%
Mistral Large 4Mistral AI58.3%–1.8%–
Gemini 3.8 Flash (low reasoning)Google83.3%50.0%0.0%0.0%
GPT-5.4 nano (no reasoning)OpenAI75.0%66.7%4.0%3.1%
Claude Haiku 4.5Anthropic83.3%25.0%0.0%0.3%
Claude Sonnet 5.5 (low reasoning)Anthropic91.7%–0.0%–
Gemini 3.6 Flash (minimal reasoning)Google91.7%41.7%0.0%0.3%
Kev 4BJared Palmer66.7%–0.0%–

Lightning: strength against cost

What one Lightning game cost at the run date's prices, on a log scale. Up and to the left is better; the dashed line joins the strongest model at each cost.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

-200300800 $0.0001$0.001$0.01$0.1 Cost per game, US dollars (log scale) Ladder Elo Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite (minimal reasoning) Gemini 3.5 Flash (minimal reasoning) d1 Seed-2.0-Mini (minimal reasoning) Mistral Small 4 (no reasoning)

Best per budget

  1. d1: 572 at $0.0005
  2. Gemini 3.1 Flash-Lite (minimal reasoning): 780 at $0.0058
  3. Gemini 3.5 Flash-Lite: 892 at $0.0070

Cheapest first: each model here beats every cheaper one on score.

How BulletBench works

  1. 1

    The position

    The model gets the position, the moves so far, its remaining clock and the legal moves.

  2. 2

    The reply

    It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.

  3. 3

    The engine

    Stockfish replies instantly at the model's current ladder level.

  4. 4

    The ladder

    Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.

Why chess on a clock

Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.

The clocks

Lightning 10+1
Ten seconds, plus one second a move. A model can play for ever, but only at about a second an answer.
Bullet 60s
One minute for the whole game, no increment.
Blitz 3+2
Three minutes plus two seconds a move: room for reasoning models, for comparison.

Fair play

Same prompt
Every model gets the same system prompt and position format.
Invalid replies
A reply with no legal move gets two corrective retries, then a random legal move is played and counted.
Settings
Each model runs at its provider's default and, where the provider offers one, its fastest reasoning setting, shown separately.

Honest caveats

A dozen games or fewer per clock (the table gives each configuration's count) give wide intervals, so neighbouring ranks are often within each other's ranges. Ratings also come in steps: the engine plays at fixed levels and the rating is fitted from which levels a model beat, drew or lost to, so models with the same results against the same levels get exactly the same rating. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.

What it measures

  • Decision quality under real time pressure
  • Response latency at the provider's default settings
  • Which reasoning settings are fast enough to use

What it does not measure

  • Any business skill: this is chess against an engine
  • Chess strength without a clock

Method

  • Games against a Stockfish ladder; every second of response time comes off the model's clock
  • Each model at its default setting and, where offered, its fastest reasoning setting
  • Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating