Benchmarks / BulletBench

Measured by Spring Prompt

BulletBench

When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?

Results dated
29 Sep 2026 to 30 Sep 2026
Models
16
Unit
milliseconds
Licence
Spring Prompt original
BulletBench 180+2: median move time, milliseconds, lower is better
#ModelMedian move time
milliseconds, lower is better
Ladder Elo
ladder Elo
1 Mistral Medium 3.5Mistral
0.5 s
555
2 Gemini 3.5 Flash LiteGoogle
0.8 s
707
3 Claude Haiku 4.5Anthropic
1.3 s
491
4 Gemini 3.8 FlashGoogle
4.0 s
1,126
5 Claude Fable 5.1Anthropic
5.4 s
358
6 GPT-6 AstraOpenAI
5.6 s
839
7 Gemini 3.1 Pro PreviewGoogle
6.2 s
683
8 DeepSeek V4 Pro 0423DeepSeek
6.4 s
0
9 Kimi K3Moonshot AI
6.8 s
0
10 GPT-6 LunaOpenAI
7.0 s
557
11 Qwen3.8 Max (0902)Alibaba
7.4 s
0
12 Claude Opus 5.5Anthropic
7.5 s
557
13 Claude Sonnet 5.5Anthropic
8.1 s
557
14 GPT-6 SolOpenAI
8.2 s
0
15 Grok 4.7xAI
10.1 s
147
16 Muse Spark 1.3Meta
20.7 s
0

Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

What it measures

  • Decision quality under real time pressure
  • Response latency at the provider's default settings

What it does not measure

  • Any business skill: this is chess against an engine
  • Chess strength without a clock

Method

  • Games against a Stockfish ladder at 3 minutes plus 2 seconds a move
  • Every second of response time comes off the model's clock
  • Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating