Benchmarks / BulletBench

Measured by Spring Prompt

BulletBench

When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?

Results dated
29 Sep 2026 to 30 Sep 2026
Models
16
Unit
ladder Elo
Licence
Spring Prompt original
BulletBench 180+2: ladder Elo, ladder Elo, higher is better
#ModelLadder Elo · 95% range
ladder Elo, higher is better
Median move time
milliseconds
1 Gemini 3.8 FlashGoogle
1,126
949–1,347
4.0 s
1 GPT-6 AstraOpenAI
839
595–1,072
5.6 s
1 Gemini 3.1 Pro PreviewGoogle
683
487–1,130
6.2 s
2 Gemini 3.5 Flash LiteGoogle
707
458–933
0.8 s
2 Claude Opus 5.5Anthropic
557
370–727
7.5 s
2 Claude Sonnet 5.5Anthropic
557
370–727
8.1 s
2 GPT-6 LunaOpenAI
557
370–727
7.0 s
2 Mistral Medium 3.5Mistral
555
396–765
0.5 s
2 Claude Haiku 4.5Anthropic
491
299–659
1.3 s
2 Claude Fable 5.1Anthropic
358
0–619
5.4 s
6 Grok 4.7xAI
147
0–387
10.1 s
12 Qwen3.8 Max (0902)Alibaba
0
7.4 s
12 DeepSeek V4 Pro 0423DeepSeek
0
6.4 s
12 Muse Spark 1.3Meta
0
20.7 s
12 Kimi K3Moonshot AI
0
6.8 s
12 GPT-6 SolOpenAI
0
8.2 s

Models share a rank when their ranges overlap. Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

What it measures

  • Decision quality under real time pressure
  • Response latency at the provider's default settings

What it does not measure

  • Any business skill: this is chess against an engine
  • Chess strength without a clock

Method

  • Games against a Stockfish ladder at 3 minutes plus 2 seconds a move
  • Every second of response time comes off the model's clock
  • Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating