BulletBench
How smart is an AI per second? We make AI models play speed chess against a chess computer, on a real clock: every second a model spends producing its move drains its time, and when the clock hits zero it loses - even from a winning position. Fast and good wins. Slow genius flags.
- Games played
- 12,096
- Configurations
- 66
- Model makers
- 11
- Fastest reply
- 0.5s
Bullet (60s) standings
12,096 gamesChess rating at the 60-second bullet format. Full board below.
Spring Prompt
Which LLM is best for quick decisions?
Chess rating with 60 seconds of thinking per game (Bullet) · whiskers show the 95% confidence interval
BulletBench
Takeaway
Gemini 3.5 Flash (minimal) is the only class of model that can sustain ~1-second chess: ladder Elo 781 at the 10+1 lightning control, 1.1s per move.
Takeaway
All 23 frontier heavyweight configurations lost essentially every lightning game on time - including positions they were winning on the board.
The fastest LLMs, ranked: who can actually think fast?
This is the leaderboard for low-latency AI: the two fastest formats mirror jobs where a model must be smart right now (routing requests, classifying, and making quick decisions in real-time agents). Higher ratings mean stronger play on our calibrated Stockfish ladder; 400 is the random-move anchor, and these are not FIDE ratings. Click any column header to sort. Priority/Fast rows are ranked in the same table but remain provisional at 8 games per control.
Bullet - 60 seconds of thinking for an entire game.
Lightning - 10 seconds plus 1 per move: it can play forever, but only at about a second an answer.
| Model | Bullet rating 60s whole game | Lightning rating 10s + 1s per move | Response time median secs per move | Thinking per move tokens generated | Time losses % games lost on time | Cost avg $ per game |
|---|---|---|---|---|---|---|
|
Qwen3.7 Flash (off) Alibaba via OpenRouter · off reasoning |
466
403–513
|
438
363–504
|
0.5s | 3 | 7% 13/192 | <$0.01 |
|
Qwen 3.5 Flash (off) Alibaba via OpenRouter · off reasoning |
459
396–516
|
459
405–516
|
0.5s | 3 | 3% 6/192 | <$0.01 |
|
Ministral 3 3B (off) Mistral via OpenRouter · off reasoning |
58
0–159
|
161
57–254
|
0.5s | 4 | 12% 23/192 | <$0.01 |
|
Mercury 2 (Inception) (off) Inception via OpenRouter · off reasoning |
147
0–250
|
313
229–379
|
0.5s | 67 | 21% 41/192 | <$0.01 |
|
Gemma 4 31B · Cerebras (off) Google · off reasoning |
570
511–629
|
625
564–676
|
0.5s | 4 | 0% 0/192 | $0.02 |
|
Inkling Small (off) Thinking Machines via OpenRouter · off reasoning |
445
394–513
|
480
421–527
|
0.6s | 6 | 2% 4/192 | <$0.01 |
|
Ministral 14B (off) Mistral via OpenRouter · off reasoning |
322
237–396
|
409
325–478
|
0.6s | 4 | 12% 24/192 | <$0.01 |
|
Mistral Small 4 (high) Mistral via OpenRouter · high reasoning heavy |
347
244–425
|
500
451–561
|
0.6s | 4 | 7% 13/192 | <$0.01 |
|
Nemotron 3 Nano 30B-A3B (off) NVIDIA via OpenRouter · off reasoning |
35
0–150
|
246
143–321
|
0.6s | 6 | 32% 61/192 | <$0.01 |
|
Gemma 4 26B (A4B) (off) Google via OpenRouter · off reasoning |
493
428–551
|
506
436–575
|
0.7s | 4 | 11% 22/192 | <$0.01 |
|
Mistral Small 4 (off) Mistral via OpenRouter · off reasoning |
200
58–304
|
459
387–519
|
0.7s | 4 | 12% 23/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (minimal) Google · minimal reasoning |
722
664–792
|
744
663–826
|
0.7s | 3 | 1% 1/192 | <$0.01 |
|
Nova Micro (off) Amazon via OpenRouter · off reasoning |
97
0–229
|
295
178–391
|
0.8s | 4 | 14% 27/192 | $0.00 |
|
Qwen 3.5 9B (off) Alibaba via OpenRouter · off reasoning |
115
0–238
|
≤0
|
1.0s | 4 | 78% 149/192 | <$0.01 |
|
Gemini 3.5 Flash Lite (minimal) New Google · minimal reasoning |
672
611–745
|
776
718–835
|
1.0s | 3 | 5% 10/192 | <$0.01 |
|
Gemini 3.5 Flash Lite (low) New Google · low reasoning |
716
653–783
|
749
690–805
|
1.0s | 29 | 10% 20/192 | <$0.01 |
|
Claude Haiku 4.5 (off) Anthropic · off reasoning |
330
226–411
|
200
46–300
|
1.0s | 13 | 38% 72/192 | $0.01 |
|
Nova 2.0 Lite (off) Amazon via OpenRouter · off reasoning |
304
204–382
|
212
58–305
|
1.0s | 5 | 30% 58/192 | $0.00 |
|
Qwen 3.7 Max (off) Alibaba via OpenRouter · off reasoning heavy |
539
475–590
|
539
462–605
|
1.1s | 3 | 15% 28/192 | $0.03 |
|
Gemini 3.5 Flash (minimal) Google · minimal reasoning |
808
727–890
|
781
718–856
|
1.1s | 3 | 28% 54/192 | $0.02 |
|
Gemini 3.6 Flash (minimal) New Google · minimal reasoning |
808
746–878
|
694
623–779
|
1.2s | 3 | 24% 47/192 | $0.03 |
|
Gemini 3 Flash (minimal) Google · minimal reasoning |
792
699–882
|
619
530–716
|
1.3s | 3 | 39% 74/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (low) Google · low reasoning |
613
534–702
|
402
297–485
|
1.3s | 159 | 49% 94/192 | <$0.01 |
|
DeepSeek V4 Flash 0731 (off) DeepSeek via OpenRouter · off reasoning |
161
0–291
|
0
0–0
|
1.4s | 3 | 77% 148/192 | <$0.01 |
|
Gemini 3 Flash (low) Google · low reasoning |
649
567–745
|
35
0–175
|
1.4s | 85 | 70% 135/192 | <$0.01 |
|
Claude Opus 4.8 (low) Anthropic · low reasoning heavy |
304
168–416
|
0
0–84
|
1.6s | 6 | 69% 132/192 | $0.06 |
|
Gemini 3.5 Flash (low) Google · low reasoning |
744
678–821
|
175
23–291
|
1.7s | 131 | 71% 136/192 | $0.01 |
|
Gemini 3.6 Flash (low) New Google · low reasoning |
694
620–773
|
224
56–345
|
1.8s | 120 | 71% 137/192 | $0.01 |
|
Gemini 3.5 Flash Lite (medium) New Google · medium reasoning heavy |
506
424–583
|
0
0–0
|
1.9s | 325 | 80% 154/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (medium) Google · medium reasoning heavy |
589
505–672
|
0
0–85
|
1.9s | 459 | 77% 147/192 | <$0.01 |
|
GPT-5.4 mini (low) OpenAI · low reasoning |
147
2–267
|
0
0–100
|
2.1s | 524 | 80% 154/192 | $0.02 |
|
Gemini 3.5 Flash (medium) Google · medium reasoning heavy |
601
507–669
|
0
0–0
|
2.3s | 320 | 81% 156/192 | <$0.01 |
|
GPT-5.6 Terra (low) OpenAI · low reasoning |
79
0–214
|
0
0–14
|
2.4s | 348 | 95% 183/192 | $0.04 |
|
GPT-5.6 Terra (medium) OpenAI · medium reasoning heavy |
200
64–310
|
≤0
|
2.4s | 355 | 95% 182/192 | $0.04 |
|
Gemini 3.5 Flash (high) Google · high reasoning heavy |
416
333–491
|
≤0
|
2.5s | 378 | 89% 171/192 | <$0.01 |
|
Gemini 3.6 Flash (high) New Google · high reasoning heavy |
513
430–589
|
0
0–0
|
2.5s | 284 | 85% 163/192 | <$0.01 |
|
Gemini 3.6 Flash (medium) New Google · medium reasoning heavy |
513
442–583
|
0
0–0
|
2.6s | 297 | 84% 162/192 | <$0.01 |
|
Gemini 3.5 Flash Lite (high) New Google · high reasoning heavy |
339
240–452
|
≤0
|
2.6s | 559 | 92% 176/192 | <$0.01 |
|
GPT-5.6 Sol (low) OpenAI · low reasoning |
35
0–171
|
0
0–14
|
2.7s | 207 | 97% 186/192 | $0.05 |
|
Gemini 3.1 Flash Lite (high) Google · high reasoning heavy |
79
0–209
|
≤0
|
2.8s | 983 | 97% 187/192 | <$0.01 |
|
GPT-5.6 Sol (medium) OpenAI · medium reasoning heavy |
0
0–126
|
0
0–14
|
2.8s | 228 | 97% 187/192 | $0.05 |
|
GPT-5.6 Luna (low) OpenAI · low reasoning |
147
0–266
|
≤0
|
2.8s | 534 | 95% 183/192 | $0.02 |
|
GPT-5.6 Terra (low) Fast · pilot OpenAI · low · Priority · 8/control |
191
0–495
Provisional 8/96
|
≤0
Provisional 8/96
|
2.8s | 357 | 94% 15/16 | $0.06 |
|
GPT-5.6 Sol (low) Fast · pilot OpenAI · low · Priority · 8/control |
≤0
Provisional 8/96
|
≤0
Provisional 8/96
|
2.9s | 212 | 100% 16/16 | $0.10 |
|
GPT-5.6 Luna (medium) OpenAI · medium reasoning heavy |
0
0–83
|
≤0
|
2.9s | 670 | 99% 190/192 | $0.02 |
|
GPT-5.6 Luna (low) Fast · pilot OpenAI · low · Priority · 8/control |
191
0–495
Provisional 8/96
|
≤0
Provisional 8/96
|
2.9s | 443 | 94% 15/16 | <$0.01 |
|
Mercury 2 (Inception) (high) Inception via OpenRouter · high reasoning heavy |
0
0–80
|
≤0
|
3.7s | 3,496 | 96% 184/192 | $0.01 |
|
Mercury 2 (Inception) (medium) Inception via OpenRouter · medium reasoning heavy |
0
0–9
|
≤0
|
3.8s | 3,459 | 96% 185/192 | $0.01 |
|
Gemini 3 Flash (high) Google · high reasoning heavy |
79
0–209
|
≤0
|
3.9s | 764 | 97% 187/192 | <$0.01 |
|
Mercury 2 (Inception) (low) Inception via OpenRouter · low reasoning |
≤0
|
≤0
|
3.9s | 3,487 | 99% 190/192 | $0.01 |
|
Gemini 3 Flash (medium) Google · medium reasoning heavy |
35
0–198
|
≤0
|
3.9s | 699 | 98% 188/192 | <$0.01 |
|
Grok 4.5 (high) xAI via OpenRouter · high reasoning heavy |
0
0–84
|
≤0
|
4.2s | 598 | 98% 189/192 | $0.02 |
|
Claude Fable 5 (low) Anthropic · low reasoning heavy |
79
0–196
|
≤0
|
4.3s | 70 | 97% 187/192 | $0.08 |
|
Grok 4.5 (low) xAI via OpenRouter · low reasoning heavy |
0
0–118
|
≤0
|
4.3s | 587 | 98% 188/192 | $0.02 |
|
GPT-5.4 nano (low) OpenAI · low reasoning |
0
0–97
|
≤0
|
4.6s | 1,196 | 96% 184/192 | <$0.01 |
|
Grok 4.5 (medium) xAI via OpenRouter · medium reasoning heavy |
0
0–62
|
≤0
|
4.6s | 591 | 98% 189/192 | $0.02 |
|
Gemini 3.1 Pro (low) Google · low reasoning heavy |
35
0–172
|
≤0
|
4.7s | 461 | 98% 188/192 | $0.03 |
|
Gemini 3.1 Pro (medium) Google · medium reasoning heavy |
35
0–162
|
≤0
|
4.9s | 466 | 98% 188/192 | $0.03 |
|
GPT-5.5 (medium) OpenAI · medium reasoning heavy |
≤0
|
≤0
|
4.9s | 479 | 100% 192/192 | $0.06 |
|
Gemini 3.1 Pro (high) Google · high reasoning heavy |
35
0–172
|
≤0
|
5.1s | 470 | 98% 188/192 | $0.03 |
|
Qwen 3.7 Max (default/on) Alibaba via OpenRouter · default/on reasoning heavy |
≤0
|
≤0
|
9.2s | 953 | 100% 192/192 | $0.01 |
|
Qwen 3.5 9B (on) Alibaba via OpenRouter · on reasoning |
≤0
|
≤0
|
12.3s | 2,300 | 100% 192/192 | <$0.01 |
|
Gemma 4 26B (A4B) (on) Google via OpenRouter · on reasoning |
≤0
|
≤0
|
13.5s | 1,340 | 100% 192/192 | <$0.01 |
|
Nemotron 3 Nano 30B-A3B (on) NVIDIA via OpenRouter · on reasoning |
≤0
|
≤0
|
13.8s | 3,590 | 100% 192/192 | <$0.01 |
|
Qwen 3.5 Flash (on) Alibaba via OpenRouter · on reasoning |
≤0
|
≤0
|
17.6s | 6,179 | 100% 192/192 | <$0.01 |
|
Nova 2.0 Lite (on) Amazon via OpenRouter · on reasoning |
≤0
|
≤0
|
17.6s | 2,506 | 100% 192/192 | $0.00 |
Rating cells: greener = stronger play, redder = clock death; the small figures are the 95% confidence interval. A rating at or near zero means the model lost essentially every game on the clock - not that it plays worse than random. Ranking ties are broken by score, then fewer timeout losses, then faster median move time.
Beat the Bench
Think you can outplay the leader?
Play a 60-second bullet game against the model at the top of this board - currently Gemini 3.6 Flash (minimal). Same clock rules as the benchmark: every second the AI spends on the API comes off its own time. You play White.
New · paired Fast/Priority study
Does GPT-5.6 Fast mode matter under a real clock?
Yes, directionally. Across all three clocks, Priority processing cut pooled median move latency by 15–27%. The clearest gains appeared in the 180-second games, where every model lost fewer games on time.
144 games
72 Standard + 72 verified Priority
priority service tier, with zero downgrades.
GPT-5.6 Luna
27.4%
lower pooled latency · 1.38× faster
Time losses: 22/24 → 19/24
GPT-5.6 Terra
22.5%
lower pooled latency · 1.29× faster
Time losses: 23/24 → 18/24
GPT-5.6 Sol
14.8%
lower pooled latency · 1.17× faster
Time losses: 22/24 → 19/24
The strongest signal: 180-second games
At the longest tested clock, Priority was faster and reduced time losses for Luna, Terra, and Sol.
| Model | Median move | Time losses | Score |
|---|---|---|---|
| GPT-5.6 Luna | 7.374s → 4.747s | 6/8 → 4/8 | 25.0% → 37.5% |
| GPT-5.6 Terra | 5.961s → 4.354s | 7/8 → 3/8 | 0.0% → 31.2% |
| GPT-5.6 Sol | 5.474s → 4.256s | 7/8 → 3/8 | 12.5% → 50.0% |
Low reasoning · identical openings, colors, engine levels, and benchmark seeds in each pair · run completed 2026-07-31
Download the aggregate JSON →Watch the games
Full games with per-move API latency. Step through and watch the clock do its work.
Gemini 3.5 Flash (low) as white
Lightning (10+1)
· vs Stockfish L2 (~1000)
19 plies · 18.1s thinking · $0.01
Move
-
Gemini 3.6 Flash (low) as black
Lightning (10+1)
· vs Stockfish L2 (~1000)
24 plies · 20.3s thinking · $0.01
Move
-
Gemini 3.6 Flash (low) as white
Bullet (60)
· vs Stockfish L2 (~1000)
33 plies · 36.1s thinking · $0.01
Move
-
Ministral 3 3B (off) as black
Bullet (60)
· vs Stockfish L0 (~400)
199 plies · 60.3s thinking · $0.01
Move
-
Choosing a fast LLM: quick answers
Which LLM is best for quick decisions?
Right now: Gemini 3.6 Flash (minimal) makes the best decisions under a 60-second clock, and Gemini 3.5 Flash (minimal) leads when every answer must come back in about a second. The big reasoning models lose on time long before their intelligence becomes usable - see the live fast board for the full ranking, updated with every run.
What is the fastest LLM right now?
Qwen3.7 Flash (off) is the fastest model we measure, answering in about 0.5 seconds per decision. But raw speed isn't the whole story - several sub-second models play barely better than random. BulletBench exists to measure whether fast answers are also good answers.
Are reasoning models suitable for low-latency use cases?
It depends on the budget. In these fast controls, higher reasoning settings frequently increase response time and timeout losses enough to erase the quality gain. Compare each model's explicit reasoning rows in the fast board. Minimal or low reasoning is usually the safer starting point for hard real-time limits.
How should I pick a model for routing or classification?
Weigh three numbers together: response time, quality under time pressure (the ratings above), and cost per decision. A model that's 0.5s slower but markedly smarter often wins; a model that's cheap and fast but near-random loses you more than it saves. Cross-reference with our overall model leaderboard and best-models-by-task rankings to check a candidate's general ability.
How it works
The clock is real
This published edition uses two clocks: Lightning (10s plus 1s per move) and Bullet (60s for the whole game). The wall-clock latency of every API response - including hidden reasoning - is deducted. Hit zero and it's a loss on time. The remaining clock is stated in every prompt, so models that pace themselves are rewarded.
The opponent ladder
A 9-level Stockfish ladder from random mover (anchor 400) to full strength (anchor 2800). Levels step adaptively - win and face a stronger engine, lose and drop down - and a maximum-likelihood performance rating is fitted from all games. "Ladder Elo" is internally consistent, not FIDE-calibrated.
Fair-play rules
Legal moves are listed in the prompt (we measure decision quality, not notation trivia). Illegal replies get two corrective retries - on the clock - then a random legal move is played and counted. Provider transport errors pause the clock: they measure infrastructure flakiness, not model speed.
Honest caveats
Latency includes provider serving infrastructure (that's the point for routing decisions, but infra changes can move results). Chess knowledge is part of what's measured - this is fast applied intelligence, one domain among several we test. Preview-endpoint models may be slower than their GA versions.
Routing latency-sensitive AI workloads?
BulletBench measures decision quality under a real clock. For task selection, browse the predicted-fit pages we publish only when external evidence clears our review gates.