Confirm Action

Are you sure you want to proceed?

Spring Prompt · Live AI Benchmark

BulletBench

How smart is an AI per second? We make AI models play speed chess against a chess computer, on a real clock: every second a model spends producing its move drains its time, and when the clock hits zero it loses - even from a winning position. Fast and good wins. Slow genius flags.

Games played
12,096
Configurations
66
Model makers
11
Fastest reply
0.5s

Bullet (60s) standings

12,096 games
🥇 Gemini 3.6 Flash (minimal) 1.2s/mv 808
🥈 Gemini 3.5 Flash (minimal) 1.2s/mv 808
🥉 Gemini 3 Flash (minimal) 1.2s/mv 792
4 Gemini 3.5 Flash (low) 1.8s/mv 744
5 Gemini 3.1 Flash Lite (minimal) 0.7s/mv 722
6 Gemini 3.5 Flash Lite (low) 1.0s/mv 716

Chess rating at the 60-second bullet format. Full board below.

Core leaderboard complete: every Standard Lightning and Bullet cell has 96 games. Priority/Fast pilots: 8 games per control · clearly marked provisional

Spring Prompt

Which LLM is best for quick decisions?

Chess rating with 60 seconds of thinking per game (Bullet) · whiskers show the 95% confidence interval

BulletBench

Gemini 3.6 Flash (minimal)
808
Gemini 3.5 Flash (minimal)
808
Gemini 3 Flash (minimal)
792
Gemini 3.5 Flash (low)
744
Gemini 3.1 Flash Lite (minimal)
722
Gemini 3.5 Flash Lite (low)
716
Gemini 3.6 Flash (low)
694
Gemini 3.5 Flash Lite (minimal)
672
Gemini 3 Flash (low)
649
Gemini 3.1 Flash Lite (low)
613
Gemini 3.5 Flash (medium) · heavy
601
Gemini 3.1 Flash Lite (medium) · heavy
589
Gemma 4 31B · Cerebras (off)
570
Qwen 3.7 Max (off) · heavy
539
Gemini 3.6 Flash (medium) · heavy
513
springprompt.com/evals/bullet-chess

Takeaway

Gemini 3.5 Flash (minimal) is the only class of model that can sustain ~1-second chess: ladder Elo 781 at the 10+1 lightning control, 1.1s per move.

Takeaway

All 23 frontier heavyweight configurations lost essentially every lightning game on time - including positions they were winning on the board.

The fastest LLMs, ranked: who can actually think fast?

This is the leaderboard for low-latency AI: the two fastest formats mirror jobs where a model must be smart right now (routing requests, classifying, and making quick decisions in real-time agents). Higher ratings mean stronger play on our calibrated Stockfish ladder; 400 is the random-move anchor, and these are not FIDE ratings. Click any column header to sort. Priority/Fast rows are ranked in the same table but remain provisional at 8 games per control.

Bullet - 60 seconds of thinking for an entire game.

Lightning - 10 seconds plus 1 per move: it can play forever, but only at about a second an answer.

Model Bullet rating 60s whole game Lightning rating 10s + 1s per move Response time median secs per move Thinking per move tokens generated Time losses % games lost on time Cost avg $ per game

Qwen3.7 Flash (off)

Alibaba via OpenRouter · off reasoning

466 403–513
438 363–504
0.5s 3 7% 13/192 <$0.01

Qwen 3.5 Flash (off)

Alibaba via OpenRouter · off reasoning

459 396–516
459 405–516
0.5s 3 3% 6/192 <$0.01

Ministral 3 3B (off)

Mistral via OpenRouter · off reasoning

58 0–159
161 57–254
0.5s 4 12% 23/192 <$0.01

Mercury 2 (Inception) (off)

Inception via OpenRouter · off reasoning

147 0–250
313 229–379
0.5s 67 21% 41/192 <$0.01

Gemma 4 31B · Cerebras (off)

Google · off reasoning

570 511–629
625 564–676
0.5s 4 0% 0/192 $0.02

Inkling Small (off)

Thinking Machines via OpenRouter · off reasoning

445 394–513
480 421–527
0.6s 6 2% 4/192 <$0.01

Ministral 14B (off)

Mistral via OpenRouter · off reasoning

322 237–396
409 325–478
0.6s 4 12% 24/192 <$0.01

Mistral Small 4 (high)

Mistral via OpenRouter · high reasoning heavy

347 244–425
500 451–561
0.6s 4 7% 13/192 <$0.01

Nemotron 3 Nano 30B-A3B (off)

NVIDIA via OpenRouter · off reasoning

35 0–150
246 143–321
0.6s 6 32% 61/192 <$0.01

Gemma 4 26B (A4B) (off)

Google via OpenRouter · off reasoning

493 428–551
506 436–575
0.7s 4 11% 22/192 <$0.01

Mistral Small 4 (off)

Mistral via OpenRouter · off reasoning

200 58–304
459 387–519
0.7s 4 12% 23/192 <$0.01

Gemini 3.1 Flash Lite (minimal)

Google · minimal reasoning

722 664–792
744 663–826
0.7s 3 1% 1/192 <$0.01

Nova Micro (off)

Amazon via OpenRouter · off reasoning

97 0–229
295 178–391
0.8s 4 14% 27/192 $0.00

Qwen 3.5 9B (off)

Alibaba via OpenRouter · off reasoning

115 0–238
≤0
1.0s 4 78% 149/192 <$0.01

Gemini 3.5 Flash Lite (minimal) New

Google · minimal reasoning

672 611–745
776 718–835
1.0s 3 5% 10/192 <$0.01

Gemini 3.5 Flash Lite (low) New

Google · low reasoning

716 653–783
749 690–805
1.0s 29 10% 20/192 <$0.01

Claude Haiku 4.5 (off)

Anthropic · off reasoning

330 226–411
200 46–300
1.0s 13 38% 72/192 $0.01

Nova 2.0 Lite (off)

Amazon via OpenRouter · off reasoning

304 204–382
212 58–305
1.0s 5 30% 58/192 $0.00

Qwen 3.7 Max (off)

Alibaba via OpenRouter · off reasoning heavy

539 475–590
539 462–605
1.1s 3 15% 28/192 $0.03

Gemini 3.5 Flash (minimal)

Google · minimal reasoning

808 727–890
781 718–856
1.1s 3 28% 54/192 $0.02

Gemini 3.6 Flash (minimal) New

Google · minimal reasoning

808 746–878
694 623–779
1.2s 3 24% 47/192 $0.03

Gemini 3 Flash (minimal)

Google · minimal reasoning

792 699–882
619 530–716
1.3s 3 39% 74/192 <$0.01

Gemini 3.1 Flash Lite (low)

Google · low reasoning

613 534–702
402 297–485
1.3s 159 49% 94/192 <$0.01

DeepSeek V4 Flash 0731 (off)

DeepSeek via OpenRouter · off reasoning

161 0–291
0 0–0
1.4s 3 77% 148/192 <$0.01

Gemini 3 Flash (low)

Google · low reasoning

649 567–745
35 0–175
1.4s 85 70% 135/192 <$0.01

Claude Opus 4.8 (low)

Anthropic · low reasoning heavy

304 168–416
0 0–84
1.6s 6 69% 132/192 $0.06

Gemini 3.5 Flash (low)

Google · low reasoning

744 678–821
175 23–291
1.7s 131 71% 136/192 $0.01

Gemini 3.6 Flash (low) New

Google · low reasoning

694 620–773
224 56–345
1.8s 120 71% 137/192 $0.01

Gemini 3.5 Flash Lite (medium) New

Google · medium reasoning heavy

506 424–583
0 0–0
1.9s 325 80% 154/192 <$0.01

Gemini 3.1 Flash Lite (medium)

Google · medium reasoning heavy

589 505–672
0 0–85
1.9s 459 77% 147/192 <$0.01

GPT-5.4 mini (low)

OpenAI · low reasoning

147 2–267
0 0–100
2.1s 524 80% 154/192 $0.02

Gemini 3.5 Flash (medium)

Google · medium reasoning heavy

601 507–669
0 0–0
2.3s 320 81% 156/192 <$0.01

GPT-5.6 Terra (low)

OpenAI · low reasoning

79 0–214
0 0–14
2.4s 348 95% 183/192 $0.04

GPT-5.6 Terra (medium)

OpenAI · medium reasoning heavy

200 64–310
≤0
2.4s 355 95% 182/192 $0.04

Gemini 3.5 Flash (high)

Google · high reasoning heavy

416 333–491
≤0
2.5s 378 89% 171/192 <$0.01

Gemini 3.6 Flash (high) New

Google · high reasoning heavy

513 430–589
0 0–0
2.5s 284 85% 163/192 <$0.01

Gemini 3.6 Flash (medium) New

Google · medium reasoning heavy

513 442–583
0 0–0
2.6s 297 84% 162/192 <$0.01

Gemini 3.5 Flash Lite (high) New

Google · high reasoning heavy

339 240–452
≤0
2.6s 559 92% 176/192 <$0.01

GPT-5.6 Sol (low)

OpenAI · low reasoning

35 0–171
0 0–14
2.7s 207 97% 186/192 $0.05

Gemini 3.1 Flash Lite (high)

Google · high reasoning heavy

79 0–209
≤0
2.8s 983 97% 187/192 <$0.01

GPT-5.6 Sol (medium)

OpenAI · medium reasoning heavy

0 0–126
0 0–14
2.8s 228 97% 187/192 $0.05

GPT-5.6 Luna (low)

OpenAI · low reasoning

147 0–266
≤0
2.8s 534 95% 183/192 $0.02

GPT-5.6 Terra (low)

Fast · pilot OpenAI · low · Priority · 8/control

191 0–495 Provisional 8/96
≤0 Provisional 8/96
2.8s 357 94% 15/16 $0.06

GPT-5.6 Sol (low)

Fast · pilot OpenAI · low · Priority · 8/control

≤0 Provisional 8/96
≤0 Provisional 8/96
2.9s 212 100% 16/16 $0.10

GPT-5.6 Luna (medium)

OpenAI · medium reasoning heavy

0 0–83
≤0
2.9s 670 99% 190/192 $0.02

GPT-5.6 Luna (low)

Fast · pilot OpenAI · low · Priority · 8/control

191 0–495 Provisional 8/96
≤0 Provisional 8/96
2.9s 443 94% 15/16 <$0.01

Mercury 2 (Inception) (high)

Inception via OpenRouter · high reasoning heavy

0 0–80
≤0
3.7s 3,496 96% 184/192 $0.01

Mercury 2 (Inception) (medium)

Inception via OpenRouter · medium reasoning heavy

0 0–9
≤0
3.8s 3,459 96% 185/192 $0.01

Gemini 3 Flash (high)

Google · high reasoning heavy

79 0–209
≤0
3.9s 764 97% 187/192 <$0.01

Mercury 2 (Inception) (low)

Inception via OpenRouter · low reasoning

≤0
≤0
3.9s 3,487 99% 190/192 $0.01

Gemini 3 Flash (medium)

Google · medium reasoning heavy

35 0–198
≤0
3.9s 699 98% 188/192 <$0.01

Grok 4.5 (high)

xAI via OpenRouter · high reasoning heavy

0 0–84
≤0
4.2s 598 98% 189/192 $0.02

Claude Fable 5 (low)

Anthropic · low reasoning heavy

79 0–196
≤0
4.3s 70 97% 187/192 $0.08

Grok 4.5 (low)

xAI via OpenRouter · low reasoning heavy

0 0–118
≤0
4.3s 587 98% 188/192 $0.02

GPT-5.4 nano (low)

OpenAI · low reasoning

0 0–97
≤0
4.6s 1,196 96% 184/192 <$0.01

Grok 4.5 (medium)

xAI via OpenRouter · medium reasoning heavy

0 0–62
≤0
4.6s 591 98% 189/192 $0.02

Gemini 3.1 Pro (low)

Google · low reasoning heavy

35 0–172
≤0
4.7s 461 98% 188/192 $0.03

Gemini 3.1 Pro (medium)

Google · medium reasoning heavy

35 0–162
≤0
4.9s 466 98% 188/192 $0.03

GPT-5.5 (medium)

OpenAI · medium reasoning heavy

≤0
≤0
4.9s 479 100% 192/192 $0.06

Gemini 3.1 Pro (high)

Google · high reasoning heavy

35 0–172
≤0
5.1s 470 98% 188/192 $0.03

Qwen 3.7 Max (default/on)

Alibaba via OpenRouter · default/on reasoning heavy

≤0
≤0
9.2s 953 100% 192/192 $0.01

Qwen 3.5 9B (on)

Alibaba via OpenRouter · on reasoning

≤0
≤0
12.3s 2,300 100% 192/192 <$0.01

Gemma 4 26B (A4B) (on)

Google via OpenRouter · on reasoning

≤0
≤0
13.5s 1,340 100% 192/192 <$0.01

Nemotron 3 Nano 30B-A3B (on)

NVIDIA via OpenRouter · on reasoning

≤0
≤0
13.8s 3,590 100% 192/192 <$0.01

Qwen 3.5 Flash (on)

Alibaba via OpenRouter · on reasoning

≤0
≤0
17.6s 6,179 100% 192/192 <$0.01

Nova 2.0 Lite (on)

Amazon via OpenRouter · on reasoning

≤0
≤0
17.6s 2,506 100% 192/192 $0.00
springprompt.com/evals/bullet-chess

Rating cells: greener = stronger play, redder = clock death; the small figures are the 95% confidence interval. A rating at or near zero means the model lost essentially every game on the clock - not that it plays worse than random. Ranking ties are broken by score, then fewer timeout losses, then faster median move time.

Beat the Bench

Think you can outplay the leader?

Play a 60-second bullet game against the model at the top of this board - currently Gemini 3.6 Flash (minimal). Same clock rules as the benchmark: every second the AI spends on the API comes off its own time. You play White.

♟ Start a challenge →

New · paired Fast/Priority study

Does GPT-5.6 Fast mode matter under a real clock?

Yes, directionally. Across all three clocks, Priority processing cut pooled median move latency by 15–27%. The clearest gains appeared in the 180-second games, where every model lost fewer games on time.

144 games

72 Standard + 72 verified Priority

Directional study, not the main leaderboard: 8 matched games per model and clock, compared with 96 games per cell in the published rankings. Priority responses reported the priority service tier, with zero downgrades.

GPT-5.6 Luna

27.4%

lower pooled latency · 1.38× faster

Standard 5.863s Priority 4.259s

Time losses: 22/24 → 19/24

GPT-5.6 Terra

22.5%

lower pooled latency · 1.29× faster

Standard 4.969s Priority 3.850s

Time losses: 23/24 → 18/24

GPT-5.6 Sol

14.8%

lower pooled latency · 1.17× faster

Standard 4.095s Priority 3.487s

Time losses: 22/24 → 19/24

The strongest signal: 180-second games

At the longest tested clock, Priority was faster and reduced time losses for Luna, Terra, and Sol.

Model Median move Time losses Score
GPT-5.6 Luna 7.374s 4.747s 6/8 4/8 25.0% 37.5%
GPT-5.6 Terra 5.961s 4.354s 7/8 3/8 0.0% 31.2%
GPT-5.6 Sol 5.474s 4.256s 7/8 3/8 12.5% 50.0%
What did not work: all Standard and Priority runs timed out at 10+1. Fast processing helped some individual latency cells, but it did not make these models viable at roughly one second per move.
Price and claim check: we observed 1.17–1.38× pooled speedups, below the advertised “up to 2.5×.” Leaderboard costs use OpenAI’s July 30 rates: Luna 80% lower, Terra 20% lower, and Sol unchanged. OpenAI pricing announcement.

Low reasoning · identical openings, colors, engine levels, and benchmark seeds in each pair · run completed 2026-07-31

Download the aggregate JSON →

Watch the games

Full games with per-move API latency. Step through and watch the clock do its work.

Gemini 3.5 Flash (low) as white

Lightning (10+1) · vs Stockfish L2 (~1000)
19 plies · 18.1s thinking · $0.01

win · checkmate

Move

-

Gemini 3.6 Flash (low) as black

Lightning (10+1) · vs Stockfish L2 (~1000)
24 plies · 20.3s thinking · $0.01

win · checkmate

Move

-

Gemini 3.6 Flash (low) as white

Bullet (60) · vs Stockfish L2 (~1000)
33 plies · 36.1s thinking · $0.01

win · checkmate

Move

-

Ministral 3 3B (off) as black

Bullet (60) · vs Stockfish L0 (~400)
199 plies · 60.3s thinking · $0.01

loss · time

Move

-

Choosing a fast LLM: quick answers

Which LLM is best for quick decisions?

Right now: Gemini 3.6 Flash (minimal) makes the best decisions under a 60-second clock, and Gemini 3.5 Flash (minimal) leads when every answer must come back in about a second. The big reasoning models lose on time long before their intelligence becomes usable - see the live fast board for the full ranking, updated with every run.

What is the fastest LLM right now?

Qwen3.7 Flash (off) is the fastest model we measure, answering in about 0.5 seconds per decision. But raw speed isn't the whole story - several sub-second models play barely better than random. BulletBench exists to measure whether fast answers are also good answers.

Are reasoning models suitable for low-latency use cases?

It depends on the budget. In these fast controls, higher reasoning settings frequently increase response time and timeout losses enough to erase the quality gain. Compare each model's explicit reasoning rows in the fast board. Minimal or low reasoning is usually the safer starting point for hard real-time limits.

How should I pick a model for routing or classification?

Weigh three numbers together: response time, quality under time pressure (the ratings above), and cost per decision. A model that's 0.5s slower but markedly smarter often wins; a model that's cheap and fast but near-random loses you more than it saves. Cross-reference with our overall model leaderboard and best-models-by-task rankings to check a candidate's general ability.

How it works

The clock is real

This published edition uses two clocks: Lightning (10s plus 1s per move) and Bullet (60s for the whole game). The wall-clock latency of every API response - including hidden reasoning - is deducted. Hit zero and it's a loss on time. The remaining clock is stated in every prompt, so models that pace themselves are rewarded.

The opponent ladder

A 9-level Stockfish ladder from random mover (anchor 400) to full strength (anchor 2800). Levels step adaptively - win and face a stronger engine, lose and drop down - and a maximum-likelihood performance rating is fitted from all games. "Ladder Elo" is internally consistent, not FIDE-calibrated.

Fair-play rules

Legal moves are listed in the prompt (we measure decision quality, not notation trivia). Illegal replies get two corrective retries - on the clock - then a random legal move is played and counted. Provider transport errors pause the clock: they measure infrastructure flakiness, not model speed.

Honest caveats

Latency includes provider serving infrastructure (that's the point for routing decisions, but infra changes can move results). Chess knowledge is part of what's measured - this is fast applied intelligence, one domain among several we test. Preview-endpoint models may be slower than their GA versions.

Routing latency-sensitive AI workloads?

BulletBench measures decision quality under a real clock. For task selection, browse the predicted-fit pages we publish only when external evidence clears our review gates.