Confirm Action

Are you sure you want to proceed?

Spring Prompt · Live AI Benchmark

BulletBench

How smart is an AI per second? We make AI models play speed chess against a chess computer, on a real clock: every second a model spends producing its move drains its time, and when the clock hits zero it loses - even from a winning position. Fast and good wins. Slow genius flags.

Games played
11,136
Configurations
58
Model makers
10
Fastest reply
0.1s

Bullet (60s) standings

11,136 games
🥇 Gemini 3.5 Flash (low) 1.2s/mv 865
🥈 Gemini 3.5 Flash (minimal) 0.6s/mv 860
🥉 Gemini 3.7 Flash (low) 1.0s/mv 834
4 Gemini 3 Flash (low) 0.7s/mv 792
5 Gemini 3.6 Flash (minimal) 0.6s/mv 771
6 Gemini 3 Flash (minimal) 0.6s/mv 755

Chess rating at the 60-second bullet format. Full board below.

Published fast edition: every Lightning and Bullet cell is complete. 96 games per model per control · updated 2026-08-14

Spring Prompt

Which LLM is best for quick decisions?

Chess rating with 60 seconds of thinking per game (Bullet) · whiskers show the 95% confidence interval

BulletBench

Gemini 3.5 Flash (low)
865
Gemini 3.5 Flash (minimal)
860
Gemini 3.7 Flash (low)
834
Gemini 3 Flash (low)
792
Gemini 3.6 Flash (minimal)
771
Gemini 3 Flash (minimal)
755
Gemini 3.6 Flash (low)
755
Gemini 3.1 Flash Lite (low)
689
Gemini 3.7 Flash (medium) · heavy
683
Gemini 3.5 Flash Lite (minimal)
677
Gemini 3.5 Flash Lite (low)
677
Gemini 3.1 Flash Lite (minimal)
666
Gemini 3.5 Flash (medium) · heavy
631
Gemini 3.5 Flash Lite (medium) · heavy
564
Gemini 3.5 Flash (high) · heavy
558
springprompt.com/evals/bullet-chess

Takeaway

Gemini 3.5 Flash (minimal) is the only class of model that can sustain ~1-second chess: ladder Elo 834 at the 10+1 lightning control, 0.6s per move.

Takeaway

All 14 frontier heavyweight configurations lost essentially every lightning game on time - including positions they were winning on the board.

The fastest LLMs, ranked: who can actually think fast?

This is the leaderboard for low-latency AI: the two fastest formats mirror jobs where a model must be smart right now (routing requests, classifying, and making quick decisions in real-time agents). Higher ratings mean stronger play on our calibrated Stockfish ladder; 400 is the random-move anchor, and these are not FIDE ratings. Click any column header to sort.

Bullet - 60 seconds of thinking for an entire game.

Lightning - 10 seconds plus 1 per move: it can play forever, but only at about a second an answer.

Model Bullet rating 60s whole game Lightning rating 10s + 1s per move Response time median secs per move Thinking per move tokens generated Time losses % games lost on time Cost avg $ per game

Celeris-1 (off)

Celeris · off reasoning

519 468–568
500 438–561
0.1s 8 0% 0/192 <$0.01

Mercury 2 (Inception) (off)

Inception via OpenRouter · off reasoning

266 185–343
387 314–440
0.3s 93 8% 16/192 $0.01

Inkling Small (off)

Thinking Machines via OpenRouter · off reasoning

438 371–496
486 437–537
0.3s 7 2% 3/192 <$0.01

Gemini 3.1 Flash Lite (minimal)

Google · minimal reasoning

666 604–729
689 623–755
0.3s 3 0% 0/192 <$0.01

Gemini 3.5 Flash Lite (minimal) New

Google · minimal reasoning

677 618–745
631 576–696
0.4s 2 0% 0/192 <$0.01

Gemini 3.5 Flash Lite (low) New

Google · low reasoning

677 614–734
649 591–707
0.4s 30 0% 0/192 <$0.01

Ministral 14B (off)

Mistral via OpenRouter · off reasoning

486 428–535
480 418–542
0.4s 4 0% 0/192 <$0.01

Ministral 3 3B (off)

Mistral via OpenRouter · off reasoning

79 0–173
161 54–255
0.4s 4 3% 6/192 <$0.01

Nova Micro (off)

Amazon via OpenRouter · off reasoning

379 286–447
355 250–435
0.4s 4 0% 0/192 <$0.01

Mistral Small 4 (high)

Mistral via OpenRouter · high reasoning heavy

339 212–415
466 409–522
0.5s 4 8% 15/192 <$0.01

Mistral Small 4 (off)

Mistral via OpenRouter · off reasoning

371 282–440
486 428–535
0.5s 4 6% 12/192 <$0.01

Gemini 3.6 Flash (minimal) New

Google · minimal reasoning

771 705–841
818 739–900
0.6s 3 1% 1/192 $0.03

Gemma 4 26B (A4B) (off)

Google via OpenRouter · off reasoning

506 440–567
416 347–486
0.6s 4 20% 39/192 <$0.01

Nova 2.0 Lite (off)

Amazon via OpenRouter · off reasoning

276 185–355
431 386–485
0.6s 5 7% 14/192 <$0.01

Gemini 3.5 Flash (minimal)

Google · minimal reasoning

860 787–952
834 766–910
0.6s 3 1% 1/192 $0.03

Gemini 3 Flash (minimal)

Google · minimal reasoning

755 690–822
786 730–854
0.6s 3 2% 3/192 $0.01

Gemini 3 Flash (low)

Google · low reasoning

792 717–865
727 654–798
0.6s 37 4% 7/192 $0.01

Qwen 3.5 Flash (off)

Alibaba via OpenRouter · off reasoning

409 332–468
486 429–535
0.7s 3 4% 7/192 <$0.01

Qwen3.7 Flash (off)

Alibaba via OpenRouter · off reasoning

322 249–405
493 443–543
0.7s 3 7% 14/192 <$0.01

Qwen 3.5 9B (off)

Alibaba via OpenRouter · off reasoning

355 259–448
276 170–354
0.7s 4 23% 45/192 <$0.01

Gemini 3.1 Flash Lite (low)

Google · low reasoning

689 627–743
677 625–740
0.9s 159 6% 11/192 $0.01

Gemini 3.7 Flash (low)

Google · low reasoning

834 760–922
765 670–862
0.9s 64 32% 62/192 $0.02

DeepSeek V4 Flash 0731 (off)

DeepSeek via OpenRouter · off reasoning

200 83–294
0 0–89
1.0s 12 70% 134/192 <$0.01

Qwen 3.7 Max (off)

Alibaba via OpenRouter · off reasoning heavy

438 342–514
452 374–520
1.1s 3 28% 54/192 $0.02

Gemini 3.7 Flash (medium)

Google · medium reasoning heavy

683 603–764
35 0–171
1.1s 137 76% 146/192 $0.01

Gemini 3.6 Flash (low) New

Google · low reasoning

755 689–817
583 512–656
1.2s 123 44% 84/192 $0.05

Gemini 3.5 Flash (low)

Google · low reasoning

865 797–946
722 638–796
1.2s 138 36% 69/192 $0.06

Gemini 3.5 Flash Lite (medium) New

Google · medium reasoning heavy

564 495–625
295 179–380
1.4s 333 58% 112/192 $0.02

DeepSeek V4 Flash (off)

DeepSeek via OpenRouter · off reasoning

161 41–273
0 0–83
1.6s 3 64% 123/192 <$0.01

Gemini 3.1 Flash Lite (medium)

Google · medium reasoning heavy

379 275–462
115 0–233
1.7s 458 80% 153/192 $0.02

Gemini 3.5 Flash (medium)

Google · medium reasoning heavy

631 560–708
115 0–254
1.8s 338 76% 146/192 $0.07

Gemini 3.6 Flash (high) New

Google · high reasoning heavy

519 444–604
0 0–0
2.0s 312 84% 162/192 $0.05

Gemini 3.6 Flash (medium) New

Google · medium reasoning heavy

552 476–623
0 0–89
2.0s 305 83% 159/192 $0.05

GPT-5.4 mini (low)

OpenAI · low reasoning

147 0–250
0 0–0
2.1s 555 75% 144/192 $0.02

Gemini 3.5 Flash (high)

Google · high reasoning heavy

558 473–627
0 0–0
2.1s 396 82% 157/192 $0.07

Gemini 3.5 Flash Lite (high) New

Google · high reasoning heavy

387 309–460
0 0–80
2.1s 580 85% 164/192 $0.03

GPT-5.6 Terra (low)

OpenAI · low reasoning

79 0–214
0 0–14
2.4s 348 95% 183/192 $0.04

GPT-5.6 Sol (low)

OpenAI · low reasoning

35 0–171
0 0–14
2.7s 207 97% 186/192 $0.05

GPT-5.6 Luna (low)

OpenAI · low reasoning

0 0–140
≤0
2.8s 589 97% 187/192 <$0.01

Gemini 3.1 Flash Lite (high)

Google · high reasoning heavy

0 0–0
≤0
2.8s 986 99% 190/192 $0.02

GPT-5.6 Luna (medium)

OpenAI · medium reasoning heavy

0 0–146
0 0–14
2.9s 712 98% 188/192 <$0.01

Gemini 3 Flash (medium)

Google · medium reasoning heavy

212 67–320
≤0
3.2s 746 94% 181/192 $0.02

Gemini 3.1 Pro (low)

Google · low reasoning heavy

355 252–438
≤0
3.2s 185 91% 175/192 $0.03

Gemini 3.1 Pro (medium)

Google · medium reasoning heavy

79 0–210
≤0
3.6s 251 97% 187/192 $0.04

Mercury 2 (Inception) (high)

Inception via OpenRouter · high reasoning heavy

0 0–0
≤0
3.9s 3,467 96% 184/192 $0.01

Mercury 2 (Inception) (medium)

Inception via OpenRouter · medium reasoning heavy

0 0–129
≤0
4.0s 3,468 95% 182/192 $0.01

Mercury 2 (Inception) (low)

Inception via OpenRouter · low reasoning

35 0–186
≤0
4.0s 3,457 96% 184/192 $0.01

GPT-5.4 nano (low)

OpenAI · low reasoning

0 0–158
≤0
4.1s 1,220 93% 179/192 <$0.01

Gemini 3.1 Pro (high)

Google · high reasoning heavy

0 0–64
≤0
4.2s 346 99% 190/192 $0.04

Grok 4.5 (high)

xAI via OpenRouter · high reasoning heavy

0 0–18
≤0
5.3s 536 99% 191/192 $0.02

Grok 4.5 (low)

xAI via OpenRouter · low reasoning heavy

≤0
≤0
5.4s 547 100% 192/192 $0.02

Grok 4.5 (medium)

xAI via OpenRouter · medium reasoning heavy

0 0–18
≤0
5.4s 552 99% 191/192 $0.02

Inkling Small (low)

Thinking Machines via OpenRouter · low reasoning

≤0
≤0
7.3s 2,885 100% 192/192 $0.01

Qwen3.7 Flash (on)

Alibaba via OpenRouter · on reasoning

≤0
≤0
9.1s 2,426 100% 192/192 <$0.01

Qwen 3.7 Max (default/on)

Alibaba via OpenRouter · default/on reasoning heavy

≤0
≤0
9.4s 946 100% 192/192 $0.01

DeepSeek V4 Flash 0731 (low)

DeepSeek via OpenRouter · low reasoning

≤0
≤0
10.5s 1,722 100% 192/192 <$0.01

Nova 2.0 Lite (on)

Amazon via OpenRouter · on reasoning

≤0
≤0
15.9s 2,663 100% 192/192 $0.02

Qwen 3.5 Flash (on)

Alibaba via OpenRouter · on reasoning

≤0
≤0
17.9s 2,893 100% 192/192 <$0.01
springprompt.com/evals/bullet-chess

Rating cells: greener = stronger play, redder = clock death; the small figures are the 95% confidence interval. A rating at or near zero means the model lost essentially every game on the clock - not that it plays worse than random. Ranking ties are broken by score, then fewer timeout losses, then faster median move time.

Beat the Bench

Think you can outplay the leader?

Play a 60-second bullet game against the model at the top of this board - currently Gemini 3.5 Flash (low). Same clock rules as the benchmark: every second the AI spends on the API comes off its own time. You play White.

♟ Start a challenge →

New · paired Fast/Priority study

Does GPT-5.6 Fast mode matter under a real clock?

Yes, directionally. Across all three clocks, Priority processing cut pooled median move latency by 15–27%. The clearest gains appeared in the 180-second games, where every model lost fewer games on time.

144 games

72 Standard + 72 verified Priority

Directional study, not the main leaderboard: 8 matched games per model and clock, compared with 96 games per cell in the published rankings. Priority responses reported the priority service tier, with zero downgrades.

GPT-5.6 Luna

27.4%

lower pooled latency · 1.38× faster

Standard 5.863s Priority 4.259s

Time losses: 22/24 → 19/24

GPT-5.6 Terra

22.5%

lower pooled latency · 1.29× faster

Standard 4.969s Priority 3.850s

Time losses: 23/24 → 18/24

GPT-5.6 Sol

14.8%

lower pooled latency · 1.17× faster

Standard 4.095s Priority 3.487s

Time losses: 22/24 → 19/24

The strongest signal: 180-second games

At the longest tested clock, Priority was faster and reduced time losses for Luna, Terra, and Sol.

Model Median move Time losses Score
GPT-5.6 Luna 7.374s 4.747s 6/8 4/8 25.0% 37.5%
GPT-5.6 Terra 5.961s 4.354s 7/8 3/8 0.0% 31.2%
GPT-5.6 Sol 5.474s 4.256s 7/8 3/8 12.5% 50.0%
What did not work: all Standard and Priority runs timed out at 10+1. Fast processing helped some individual latency cells, but it did not make these models viable at roughly one second per move.
Price and claim check: we observed 1.17–1.38× pooled speedups, below the advertised “up to 2.5×.” Leaderboard costs use OpenAI’s July 30 rates: Luna 80% lower, Terra 20% lower, and Sol unchanged. OpenAI pricing announcement.

Low reasoning · identical openings, colors, engine levels, and benchmark seeds in each pair · run completed 2026-07-31

Provisional supporting study

Watch the games

Full games with per-move API latency. Step through and watch the clock do its work.

Gemini 3.6 Flash (minimal) as white

Lightning (10+1) · vs Stockfish L2 (~1000)
13 plies · 5.4s thinking · $0.01

win · checkmate

Move

-

Gemini 3.5 Flash (low) as white

Bullet (60) · vs Stockfish L2 (~1000)
21 plies · 15.7s thinking · $0.02

win · checkmate

Move

-

Gemini 3.5 Flash (medium) as white

Bullet (60) · vs Stockfish L2 (~1000)
29 plies · 29.9s thinking · $0.06

win · checkmate

Move

-

Ministral 3 3B (off) as black

Bullet (60) · vs Stockfish L0 (~400)
199 plies · 60.0s thinking · $0.01

loss · time

Move

-

Choosing a fast LLM: quick answers

Which LLM is best for quick decisions?

Right now: Gemini 3.5 Flash (low) makes the best decisions under a 60-second clock, and Gemini 3.5 Flash (minimal) leads when every answer must come back in about a second. The big reasoning models lose on time long before their intelligence becomes usable - see the live fast board for the full ranking, updated with every run.

What is the fastest LLM right now?

Celeris-1 (off) is the fastest model we measure, answering in about 0.1 seconds per decision. But raw speed isn't the whole story - several sub-second models play barely better than random. BulletBench exists to measure whether fast answers are also good answers.

Are reasoning models suitable for low-latency use cases?

It depends on the budget. In these fast controls, higher reasoning settings frequently increase response time and timeout losses enough to erase the quality gain. Compare each model's explicit reasoning rows in the fast board. Minimal or low reasoning is usually the safer starting point for hard real-time limits.

How should I pick a model for routing or classification?

Weigh three numbers together: response time, quality under time pressure (the ratings above), and cost per decision. A model that's 0.5s slower but markedly smarter often wins; a model that's cheap and fast but near-random loses you more than it saves. Cross-reference with our overall model leaderboard and best-models-by-task rankings to check a candidate's general ability.

How it works

The clock is real

This published edition uses two clocks: Lightning (10s plus 1s per move) and Bullet (60s for the whole game). The wall-clock latency of every API response - including hidden reasoning - is deducted. Hit zero and it's a loss on time. The remaining clock is stated in every prompt, so models that pace themselves are rewarded.

The opponent ladder

A 9-level Stockfish ladder from random mover (anchor 400) to full strength (anchor 2800). Levels step adaptively - win and face a stronger engine, lose and drop down - and a maximum-likelihood performance rating is fitted from all games. "Ladder Elo" is internally consistent, not FIDE-calibrated.

Fair-play rules

Legal moves are listed in the prompt (we measure decision quality, not notation trivia). Illegal replies get two corrective retries - on the clock - then a random legal move is played and counted. Provider transport errors pause the clock: they measure infrastructure flakiness, not model speed.

Honest caveats

Latency includes provider serving infrastructure (that's the point for routing decisions, but infra changes can move results). Chess knowledge is part of what's measured - this is fast applied intelligence, one domain among several we test. Preview-endpoint models may be slower than their GA versions.

Routing latency-sensitive AI workloads?

BulletBench measures decision quality under a real clock. For task selection, browse the predicted-fit pages we publish only when external evidence clears our review gates.