4 models had none: Claude Haiku 4.5, Gemini 3.5 Flash-Lite, Mistral Large 4, Mistral Medium 3.5.
0%20%40%60%80%
Lightning: strength against thinking time
Each configuration's median time per move at 10 seconds plus 1 a move. Right of the line, a model thinks longer than the increment it earns and the clock drains.
Named: the six bestOther models (hover for names)
Full results
BulletBench Blitz 3+2: games lost on time, % of games, lower is better
#
Model
Games lost on time, Blitz 3+2 % of games, lower is better
Ranks follow the score as shown, so equal numbers share a rank. Each model runs at its provider's default reasoning setting and, where the provider offers a faster one, again at its fastest setting, listed separately. Blitz 3+2 was run at default settings only. Equal ratings are not a glitch: the ladder's ratings come in steps (see Honest caveats below).
Too slow for the clock
These configurations lost their first four games on time, so they were stopped there and are not rated for that clock. Each shows its median time per move and the time that 1 move in 10 took or exceeded: a median under the increment still loses on time when the slow tail is long. Being too slow is the result, not a fault.
Bullet 60s
Claude Fable 5.1median 4.2 s, 1 in 10 moves 6.3 s+
Claude Fable 5.1 (low reasoning)median 4.3 s, 1 in 10 moves 9.3 s+
Claude Haiku 5.5median 4.7 s
Claude Haiku 5.5 (low reasoning)median 3.6 s
Claude Opus 5.5median 3.1 s, 1 in 10 moves 5.8 s+
Claude Opus 5.5 (low reasoning)median 3.1 s, 1 in 10 moves 5.3 s+
Claude Sonnet 5.5median 3.0 s, 1 in 10 moves 9.4 s+
Claude Sonnet 5.5 (low reasoning)median 2.6 s, 1 in 10 moves 5.0 s+
DeepSeek V4 Pro 0423median 5.0 s, 1 in 10 moves 27.8 s+
Mistral Large 4 (high reasoning)median 26.1 s, 1 in 10 moves 69.5 s+
Kimi K3median 6.8 s, 1 in 10 moves 32.1 s+
GPT-6 Solmedian 8.2 s, 1 in 10 moves 22.9 s+
Qwen3.8 Max (0902)median 7.4 s, 1 in 10 moves 88.9 s+
Grok 4.7median 9.9 s, 1 in 10 moves 32.9 s+
Decision models
Decision models answer a typed question instead of writing text. Each move is one question whose options are the legal moves, so they cannot play an illegal move, and they are not told their clock. They play the same clock, ladder and engine and are in the table above; here they are ranked among themselves.
No decision model was rated on this measure.
Too slow for the clock. These lost their first four games on time and are not rated for that clock.
Bullet 60s
Kev 4Bmedian 1.9 s, 1 in 10 moves 15.4 s+
Real games: the clock at Lightning 10+1
Each model's clock after every move of one game. A fast model earns back more than it spends; a slow one runs out within a few moves, whatever the position.
Gemini 3.5 Flash LiteWon by checkmate after 28 moves
e40.6 s10.4 s left
Nf30.9 s10.4 s left
Bc40.6 s10.9 s left
O-O0.9 s10.9 s left
c30.6 s11.4 s left
d40.9 s11.4 s left
Nxe51.0 s11.4 s left
Nxc60.6 s11.8 s left
Qxg40.9 s11.9 s left
Qxc8+0.9 s12.0 s left
… 18 more moves
Claude Opus 5.5Lost on time after 3 moves
e43.5 s7.5 s left
Nf33.5 s5.0 s left
d43.3 s2.6 s left
More from the results
Bullet: strength against thinking time
The same at 60 seconds for the whole game, no increment. A 40-move game leaves 1.5 seconds a move.
Named: the six bestOther models (hover for names)
How games are lost
Share of games lost on time, and moves with no legal reply, by clock.
Model
On time, Lightning
On time, Bullet
Invalid moves, Lightning
Invalid moves, Bullet
Gemini 3.5 Flash-LiteGoogle
0.0%
0.0%
0.0%
0.0%
Gemini 3.1 Flash-Lite (minimal reasoning)Google
0.0%
8.3%
0.2%
0.0%
Gemini 3.5 Flash (minimal reasoning)Google
33.3%
25.0%
0.6%
0.5%
Mercury DecideInception
0.0%
0.0%
0.0%
0.0%
d1Liquid AI
0.0%
0.0%
0.0%
0.0%
Seed-2.0-Mini (minimal reasoning)ByteDance
0.0%
0.0%
0.8%
0.4%
Mistral Small 4 (no reasoning)Mistral AI
0.0%
25.0%
1.0%
0.6%
GPT-5.4 mini (no reasoning)OpenAI
0.0%
50.0%
0.0%
0.0%
Inkling Small (no reasoning)Thinking Machines
0.0%
0.0%
5.3%
7.2%
Jev 1.13TypeSafe
0.0%
0.0%
0.0%
0.0%
Clef FlashCloudflare
0.0%
0.0%
0.0%
0.0%
ClefCloudflare
0.0%
0.0%
0.0%
0.0%
Mistral Medium 3.5 (no reasoning)Mistral AI
0.0%
0.0%
0.0%
0.4%
Mistral Medium 3.5Mistral AI
0.0%
8.3%
0.3%
0.8%
GPT-6 Luna DecisionsOpenAI
0.0%
0.0%
0.0%
0.0%
Decider V1 27BPerplexity
0.0%
8.3%
0.0%
0.0%
Mistral Large 4Mistral AI
58.3%
–
1.8%
–
Gemini 3.8 Flash (low reasoning)Google
83.3%
50.0%
0.0%
0.0%
GPT-5.4 nano (no reasoning)OpenAI
75.0%
66.7%
4.0%
3.1%
Claude Haiku 4.5Anthropic
83.3%
25.0%
0.0%
0.3%
Claude Sonnet 5.5 (low reasoning)Anthropic
91.7%
–
0.0%
–
Gemini 3.6 Flash (minimal reasoning)Google
91.7%
41.7%
0.0%
0.3%
Kev 4BJared Palmer
66.7%
–
0.0%
–
Lightning: strength against cost
What one Lightning game cost at the run date's prices, on a log scale. Up and to the left is better; the dashed line joins the strongest model at each cost.
Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost
Best per budget
d1: 572 at $0.0005
Gemini 3.1 Flash-Lite (minimal reasoning): 780 at $0.0058
Gemini 3.5 Flash-Lite: 892 at $0.0070
Cheapest first: each model here beats every cheaper one on score.
How BulletBench works
1
The position
The model gets the position, the moves so far, its remaining clock and the legal moves.
2
The reply
It answers with one move. The whole API round trip, including any hidden reasoning, comes off its clock.
3
The engine
Stockfish replies instantly at the model's current ladder level.
4
The ladder
Win and the next game is a level up; lose and it is a level down; draw and it stays. The results give a rating with a 95% interval.
Why chess on a clock
Many products need an answer in about a second: routing, triage, autocomplete, live agents. BulletBench asks which models can still think usefully at that speed. Chess gives a hard, objective score, and the clock makes slow thinking a real cost instead of a free extra.
The clocks
Lightning 10+1
Ten seconds, plus one second a move. A model can play for ever, but only at about a second an answer.
Bullet 60s
One minute for the whole game, no increment.
Blitz 3+2
Three minutes plus two seconds a move: room for reasoning models, for comparison.
Fair play
Same prompt
Every model gets the same system prompt and position format.
Invalid replies
A reply with no legal move gets two corrective retries, then a random legal move is played and counted.
Settings
Each model runs at its provider's default and, where the provider offers one, its fastest reasoning setting, shown separately.
Honest caveats
A dozen games or fewer per clock (the table gives each configuration's count) give wide intervals, so neighbouring ranks are often within each other's ranges. Ratings also come in steps: the engine plays at fixed levels and the rating is fitted from which levels a model beat, drew or lost to, so models with the same results against the same levels get exactly the same rating. Latency depends on the provider's servers on the day. The ladder's Elo anchors come from v1's Stockfish; this edition runs Stockfish 19, so ratings are not comparable with v1's.
What it measures
Decision quality under real time pressure
Response latency at the provider's default settings
Which reasoning settings are fast enough to use
What it does not measure
Any business skill: this is chess against an engine
Chess strength without a clock
Method
Games against a Stockfish ladder; every second of response time comes off the model's clock
Each model at its default setting and, where offered, its fastest reasoning setting
Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating