BulletBench
How smart is an AI per second? We make AI models play speed chess against a chess computer, on a real clock: every second a model spends producing its move drains its time, and when the clock hits zero it loses - even from a winning position. Fast and good wins. Slow genius flags.
- Games played
- 11,136
- Configurations
- 58
- Model makers
- 10
- Fastest reply
- 0.1s
Bullet (60s) standings
11,136 gamesChess rating at the 60-second bullet format. Full board below.
Spring Prompt
Which LLM is best for quick decisions?
Chess rating with 60 seconds of thinking per game (Bullet) · whiskers show the 95% confidence interval
BulletBench
Takeaway
Gemini 3.5 Flash (minimal) is the only class of model that can sustain ~1-second chess: ladder Elo 834 at the 10+1 lightning control, 0.6s per move.
Takeaway
All 14 frontier heavyweight configurations lost essentially every lightning game on time - including positions they were winning on the board.
The fastest LLMs, ranked: who can actually think fast?
This is the leaderboard for low-latency AI: the two fastest formats mirror jobs where a model must be smart right now (routing requests, classifying, and making quick decisions in real-time agents). Higher ratings mean stronger play on our calibrated Stockfish ladder; 400 is the random-move anchor, and these are not FIDE ratings. Click any column header to sort.
Bullet - 60 seconds of thinking for an entire game.
Lightning - 10 seconds plus 1 per move: it can play forever, but only at about a second an answer.
| Model | Bullet rating 60s whole game | Lightning rating 10s + 1s per move | Response time median secs per move | Thinking per move tokens generated | Time losses % games lost on time | Cost avg $ per game |
|---|---|---|---|---|---|---|
|
Celeris-1 (off) Celeris · off reasoning |
519
468–568
|
500
438–561
|
0.1s | 8 | 0% 0/192 | <$0.01 |
|
Mercury 2 (Inception) (off) Inception via OpenRouter · off reasoning |
266
185–343
|
387
314–440
|
0.3s | 93 | 8% 16/192 | $0.01 |
|
Inkling Small (off) Thinking Machines via OpenRouter · off reasoning |
438
371–496
|
486
437–537
|
0.3s | 7 | 2% 3/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (minimal) Google · minimal reasoning |
666
604–729
|
689
623–755
|
0.3s | 3 | 0% 0/192 | <$0.01 |
|
Gemini 3.5 Flash Lite (minimal) New Google · minimal reasoning |
677
618–745
|
631
576–696
|
0.4s | 2 | 0% 0/192 | <$0.01 |
|
Gemini 3.5 Flash Lite (low) New Google · low reasoning |
677
614–734
|
649
591–707
|
0.4s | 30 | 0% 0/192 | <$0.01 |
|
Ministral 14B (off) Mistral via OpenRouter · off reasoning |
486
428–535
|
480
418–542
|
0.4s | 4 | 0% 0/192 | <$0.01 |
|
Ministral 3 3B (off) Mistral via OpenRouter · off reasoning |
79
0–173
|
161
54–255
|
0.4s | 4 | 3% 6/192 | <$0.01 |
|
Nova Micro (off) Amazon via OpenRouter · off reasoning |
379
286–447
|
355
250–435
|
0.4s | 4 | 0% 0/192 | <$0.01 |
|
Mistral Small 4 (high) Mistral via OpenRouter · high reasoning heavy |
339
212–415
|
466
409–522
|
0.5s | 4 | 8% 15/192 | <$0.01 |
|
Mistral Small 4 (off) Mistral via OpenRouter · off reasoning |
371
282–440
|
486
428–535
|
0.5s | 4 | 6% 12/192 | <$0.01 |
|
Gemini 3.6 Flash (minimal) New Google · minimal reasoning |
771
705–841
|
818
739–900
|
0.6s | 3 | 1% 1/192 | $0.03 |
|
Gemma 4 26B (A4B) (off) Google via OpenRouter · off reasoning |
506
440–567
|
416
347–486
|
0.6s | 4 | 20% 39/192 | <$0.01 |
|
Nova 2.0 Lite (off) Amazon via OpenRouter · off reasoning |
276
185–355
|
431
386–485
|
0.6s | 5 | 7% 14/192 | <$0.01 |
|
Gemini 3.5 Flash (minimal) Google · minimal reasoning |
860
787–952
|
834
766–910
|
0.6s | 3 | 1% 1/192 | $0.03 |
|
Gemini 3 Flash (minimal) Google · minimal reasoning |
755
690–822
|
786
730–854
|
0.6s | 3 | 2% 3/192 | $0.01 |
|
Gemini 3 Flash (low) Google · low reasoning |
792
717–865
|
727
654–798
|
0.6s | 37 | 4% 7/192 | $0.01 |
|
Qwen 3.5 Flash (off) Alibaba via OpenRouter · off reasoning |
409
332–468
|
486
429–535
|
0.7s | 3 | 4% 7/192 | <$0.01 |
|
Qwen3.7 Flash (off) Alibaba via OpenRouter · off reasoning |
322
249–405
|
493
443–543
|
0.7s | 3 | 7% 14/192 | <$0.01 |
|
Qwen 3.5 9B (off) Alibaba via OpenRouter · off reasoning |
355
259–448
|
276
170–354
|
0.7s | 4 | 23% 45/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (low) Google · low reasoning |
689
627–743
|
677
625–740
|
0.9s | 159 | 6% 11/192 | $0.01 |
|
Gemini 3.7 Flash (low) Google · low reasoning |
834
760–922
|
765
670–862
|
0.9s | 64 | 32% 62/192 | $0.02 |
|
DeepSeek V4 Flash 0731 (off) DeepSeek via OpenRouter · off reasoning |
200
83–294
|
0
0–89
|
1.0s | 12 | 70% 134/192 | <$0.01 |
|
Qwen 3.7 Max (off) Alibaba via OpenRouter · off reasoning heavy |
438
342–514
|
452
374–520
|
1.1s | 3 | 28% 54/192 | $0.02 |
|
Gemini 3.7 Flash (medium) Google · medium reasoning heavy |
683
603–764
|
35
0–171
|
1.1s | 137 | 76% 146/192 | $0.01 |
|
Gemini 3.6 Flash (low) New Google · low reasoning |
755
689–817
|
583
512–656
|
1.2s | 123 | 44% 84/192 | $0.05 |
|
Gemini 3.5 Flash (low) Google · low reasoning |
865
797–946
|
722
638–796
|
1.2s | 138 | 36% 69/192 | $0.06 |
|
Gemini 3.5 Flash Lite (medium) New Google · medium reasoning heavy |
564
495–625
|
295
179–380
|
1.4s | 333 | 58% 112/192 | $0.02 |
|
DeepSeek V4 Flash (off) DeepSeek via OpenRouter · off reasoning |
161
41–273
|
0
0–83
|
1.6s | 3 | 64% 123/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (medium) Google · medium reasoning heavy |
379
275–462
|
115
0–233
|
1.7s | 458 | 80% 153/192 | $0.02 |
|
Gemini 3.5 Flash (medium) Google · medium reasoning heavy |
631
560–708
|
115
0–254
|
1.8s | 338 | 76% 146/192 | $0.07 |
|
Gemini 3.6 Flash (high) New Google · high reasoning heavy |
519
444–604
|
0
0–0
|
2.0s | 312 | 84% 162/192 | $0.05 |
|
Gemini 3.6 Flash (medium) New Google · medium reasoning heavy |
552
476–623
|
0
0–89
|
2.0s | 305 | 83% 159/192 | $0.05 |
|
GPT-5.4 mini (low) OpenAI · low reasoning |
147
0–250
|
0
0–0
|
2.1s | 555 | 75% 144/192 | $0.02 |
|
Gemini 3.5 Flash (high) Google · high reasoning heavy |
558
473–627
|
0
0–0
|
2.1s | 396 | 82% 157/192 | $0.07 |
|
Gemini 3.5 Flash Lite (high) New Google · high reasoning heavy |
387
309–460
|
0
0–80
|
2.1s | 580 | 85% 164/192 | $0.03 |
|
GPT-5.6 Terra (low) OpenAI · low reasoning |
79
0–214
|
0
0–14
|
2.4s | 348 | 95% 183/192 | $0.04 |
|
GPT-5.6 Sol (low) OpenAI · low reasoning |
35
0–171
|
0
0–14
|
2.7s | 207 | 97% 186/192 | $0.05 |
|
GPT-5.6 Luna (low) OpenAI · low reasoning |
0
0–140
|
≤0
|
2.8s | 589 | 97% 187/192 | <$0.01 |
|
Gemini 3.1 Flash Lite (high) Google · high reasoning heavy |
0
0–0
|
≤0
|
2.8s | 986 | 99% 190/192 | $0.02 |
|
GPT-5.6 Luna (medium) OpenAI · medium reasoning heavy |
0
0–146
|
0
0–14
|
2.9s | 712 | 98% 188/192 | <$0.01 |
|
Gemini 3 Flash (medium) Google · medium reasoning heavy |
212
67–320
|
≤0
|
3.2s | 746 | 94% 181/192 | $0.02 |
|
Gemini 3.1 Pro (low) Google · low reasoning heavy |
355
252–438
|
≤0
|
3.2s | 185 | 91% 175/192 | $0.03 |
|
Gemini 3.1 Pro (medium) Google · medium reasoning heavy |
79
0–210
|
≤0
|
3.6s | 251 | 97% 187/192 | $0.04 |
|
Mercury 2 (Inception) (high) Inception via OpenRouter · high reasoning heavy |
0
0–0
|
≤0
|
3.9s | 3,467 | 96% 184/192 | $0.01 |
|
Mercury 2 (Inception) (medium) Inception via OpenRouter · medium reasoning heavy |
0
0–129
|
≤0
|
4.0s | 3,468 | 95% 182/192 | $0.01 |
|
Mercury 2 (Inception) (low) Inception via OpenRouter · low reasoning |
35
0–186
|
≤0
|
4.0s | 3,457 | 96% 184/192 | $0.01 |
|
GPT-5.4 nano (low) OpenAI · low reasoning |
0
0–158
|
≤0
|
4.1s | 1,220 | 93% 179/192 | <$0.01 |
|
Gemini 3.1 Pro (high) Google · high reasoning heavy |
0
0–64
|
≤0
|
4.2s | 346 | 99% 190/192 | $0.04 |
|
Grok 4.5 (high) xAI via OpenRouter · high reasoning heavy |
0
0–18
|
≤0
|
5.3s | 536 | 99% 191/192 | $0.02 |
|
Grok 4.5 (low) xAI via OpenRouter · low reasoning heavy |
≤0
|
≤0
|
5.4s | 547 | 100% 192/192 | $0.02 |
|
Grok 4.5 (medium) xAI via OpenRouter · medium reasoning heavy |
0
0–18
|
≤0
|
5.4s | 552 | 99% 191/192 | $0.02 |
|
Inkling Small (low) Thinking Machines via OpenRouter · low reasoning |
≤0
|
≤0
|
7.3s | 2,885 | 100% 192/192 | $0.01 |
|
Qwen3.7 Flash (on) Alibaba via OpenRouter · on reasoning |
≤0
|
≤0
|
9.1s | 2,426 | 100% 192/192 | <$0.01 |
|
Qwen 3.7 Max (default/on) Alibaba via OpenRouter · default/on reasoning heavy |
≤0
|
≤0
|
9.4s | 946 | 100% 192/192 | $0.01 |
|
DeepSeek V4 Flash 0731 (low) DeepSeek via OpenRouter · low reasoning |
≤0
|
≤0
|
10.5s | 1,722 | 100% 192/192 | <$0.01 |
|
Nova 2.0 Lite (on) Amazon via OpenRouter · on reasoning |
≤0
|
≤0
|
15.9s | 2,663 | 100% 192/192 | $0.02 |
|
Qwen 3.5 Flash (on) Alibaba via OpenRouter · on reasoning |
≤0
|
≤0
|
17.9s | 2,893 | 100% 192/192 | <$0.01 |
Rating cells: greener = stronger play, redder = clock death; the small figures are the 95% confidence interval. A rating at or near zero means the model lost essentially every game on the clock - not that it plays worse than random. Ranking ties are broken by score, then fewer timeout losses, then faster median move time.
Beat the Bench
Think you can outplay the leader?
Play a 60-second bullet game against the model at the top of this board - currently Gemini 3.5 Flash (low). Same clock rules as the benchmark: every second the AI spends on the API comes off its own time. You play White.
New · paired Fast/Priority study
Does GPT-5.6 Fast mode matter under a real clock?
Yes, directionally. Across all three clocks, Priority processing cut pooled median move latency by 15–27%. The clearest gains appeared in the 180-second games, where every model lost fewer games on time.
144 games
72 Standard + 72 verified Priority
priority service tier, with zero downgrades.
GPT-5.6 Luna
27.4%
lower pooled latency · 1.38× faster
Time losses: 22/24 → 19/24
GPT-5.6 Terra
22.5%
lower pooled latency · 1.29× faster
Time losses: 23/24 → 18/24
GPT-5.6 Sol
14.8%
lower pooled latency · 1.17× faster
Time losses: 22/24 → 19/24
The strongest signal: 180-second games
At the longest tested clock, Priority was faster and reduced time losses for Luna, Terra, and Sol.
| Model | Median move | Time losses | Score |
|---|---|---|---|
| GPT-5.6 Luna | 7.374s → 4.747s | 6/8 → 4/8 | 25.0% → 37.5% |
| GPT-5.6 Terra | 5.961s → 4.354s | 7/8 → 3/8 | 0.0% → 31.2% |
| GPT-5.6 Sol | 5.474s → 4.256s | 7/8 → 3/8 | 12.5% → 50.0% |
Low reasoning · identical openings, colors, engine levels, and benchmark seeds in each pair · run completed 2026-07-31
Provisional supporting study
Watch the games
Full games with per-move API latency. Step through and watch the clock do its work.
Gemini 3.6 Flash (minimal) as white
Lightning (10+1)
· vs Stockfish L2 (~1000)
13 plies · 5.4s thinking · $0.01
Move
-
Gemini 3.5 Flash (low) as white
Bullet (60)
· vs Stockfish L2 (~1000)
21 plies · 15.7s thinking · $0.02
Move
-
Gemini 3.5 Flash (medium) as white
Bullet (60)
· vs Stockfish L2 (~1000)
29 plies · 29.9s thinking · $0.06
Move
-
Ministral 3 3B (off) as black
Bullet (60)
· vs Stockfish L0 (~400)
199 plies · 60.0s thinking · $0.01
Move
-
Choosing a fast LLM: quick answers
Which LLM is best for quick decisions?
Right now: Gemini 3.5 Flash (low) makes the best decisions under a 60-second clock, and Gemini 3.5 Flash (minimal) leads when every answer must come back in about a second. The big reasoning models lose on time long before their intelligence becomes usable - see the live fast board for the full ranking, updated with every run.
What is the fastest LLM right now?
Celeris-1 (off) is the fastest model we measure, answering in about 0.1 seconds per decision. But raw speed isn't the whole story - several sub-second models play barely better than random. BulletBench exists to measure whether fast answers are also good answers.
Are reasoning models suitable for low-latency use cases?
It depends on the budget. In these fast controls, higher reasoning settings frequently increase response time and timeout losses enough to erase the quality gain. Compare each model's explicit reasoning rows in the fast board. Minimal or low reasoning is usually the safer starting point for hard real-time limits.
How should I pick a model for routing or classification?
Weigh three numbers together: response time, quality under time pressure (the ratings above), and cost per decision. A model that's 0.5s slower but markedly smarter often wins; a model that's cheap and fast but near-random loses you more than it saves. Cross-reference with our overall model leaderboard and best-models-by-task rankings to check a candidate's general ability.
How it works
The clock is real
This published edition uses two clocks: Lightning (10s plus 1s per move) and Bullet (60s for the whole game). The wall-clock latency of every API response - including hidden reasoning - is deducted. Hit zero and it's a loss on time. The remaining clock is stated in every prompt, so models that pace themselves are rewarded.
The opponent ladder
A 9-level Stockfish ladder from random mover (anchor 400) to full strength (anchor 2800). Levels step adaptively - win and face a stronger engine, lose and drop down - and a maximum-likelihood performance rating is fitted from all games. "Ladder Elo" is internally consistent, not FIDE-calibrated.
Fair-play rules
Legal moves are listed in the prompt (we measure decision quality, not notation trivia). Illegal replies get two corrective retries - on the clock - then a random legal move is played and counted. Provider transport errors pause the clock: they measure infrastructure flakiness, not model speed.
Honest caveats
Latency includes provider serving infrastructure (that's the point for routing decisions, but infra changes can move results). Chess knowledge is part of what's measured - this is fast applied intelligence, one domain among several we test. Preview-endpoint models may be slower than their GA versions.
Routing latency-sensitive AI workloads?
BulletBench measures decision quality under a real clock. For task selection, browse the predicted-fit pages we publish only when external evidence clears our review gates.