1,126
949–1,347
Benchmarks / BulletBench
Measured by Spring Prompt
BulletBench
When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?
- Results dated
- 29 Sep 2026 to 30 Sep 2026
- Models
- 16
- Unit
- ladder Elo
- Licence
- Spring Prompt original
| # | Model | Ladder Elo · 95% range ladder Elo, higher is better | Median move time milliseconds |
|---|---|---|---|
| 1 | Gemini 3.8 FlashGoogle | 4.0 s | |
| 1 | GPT-6 AstraOpenAI |
839 595–1,072 |
5.6 s |
| 1 | Gemini 3.1 Pro PreviewGoogle |
683 487–1,130 |
6.2 s |
| 2 | Gemini 3.5 Flash LiteGoogle |
707 458–933 |
0.8 s |
| 2 | Claude Opus 5.5Anthropic |
557 370–727 |
7.5 s |
| 2 | Claude Sonnet 5.5Anthropic |
557 370–727 |
8.1 s |
| 2 | GPT-6 LunaOpenAI |
557 370–727 |
7.0 s |
| 2 | Mistral Medium 3.5Mistral |
555 396–765 |
0.5 s |
| 2 | Claude Haiku 4.5Anthropic |
491 299–659 |
1.3 s |
| 2 | Claude Fable 5.1Anthropic |
358 0–619 |
5.4 s |
| 6 | Grok 4.7xAI |
147 0–387 |
10.1 s |
| 12 | Qwen3.8 Max (0902)Alibaba |
0
|
7.4 s |
| 12 | DeepSeek V4 Pro 0423DeepSeek |
0
|
6.4 s |
| 12 | Muse Spark 1.3Meta |
0
|
20.7 s |
| 12 | Kimi K3Moonshot AI |
0
|
6.8 s |
| 12 | GPT-6 SolOpenAI |
0
|
8.2 s |
Models share a rank when their ranges overlap. Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.
What it measures
- Decision quality under real time pressure
- Response latency at the provider's default settings
What it does not measure
- Any business skill: this is chess against an engine
- Chess strength without a clock
Method
- Games against a Stockfish ladder at 3 minutes plus 2 seconds a move
- Every second of response time comes off the model's clock
- Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating