Benchmarks / BulletBench
Measured by Spring Prompt
BulletBench
When every second of thinking time comes off the clock, which models are fast enough to still make good decisions?
- Results dated
- 29 Sep 2026 to 30 Sep 2026
- Models
- 16
- Unit
- milliseconds
- Licence
- Spring Prompt original
| # | Model | Median move time milliseconds, lower is better | Ladder Elo ladder Elo |
|---|---|---|---|
| 1 | Mistral Medium 3.5Mistral |
0.5 s
|
555 |
| 2 | Gemini 3.5 Flash LiteGoogle |
0.8 s
|
707 |
| 3 | Claude Haiku 4.5Anthropic |
1.3 s
|
491 |
| 4 | Gemini 3.8 FlashGoogle |
4.0 s
|
1,126 |
| 5 | Claude Fable 5.1Anthropic |
5.4 s
|
358 |
| 6 | GPT-6 AstraOpenAI |
5.6 s
|
839 |
| 7 | Gemini 3.1 Pro PreviewGoogle |
6.2 s
|
683 |
| 8 | DeepSeek V4 Pro 0423DeepSeek |
6.4 s
|
0 |
| 9 | Kimi K3Moonshot AI |
6.8 s
|
0 |
| 10 | GPT-6 LunaOpenAI |
7.0 s
|
557 |
| 11 | Qwen3.8 Max (0902)Alibaba |
7.4 s
|
0 |
| 12 | Claude Opus 5.5Anthropic |
7.5 s
|
557 |
| 13 | Claude Sonnet 5.5Anthropic |
8.1 s
|
557 |
| 14 | GPT-6 SolOpenAI |
8.2 s
|
0 |
| 15 | Grok 4.7xAI |
10.1 s
|
147 |
| 16 | Muse Spark 1.3Meta |
20.7 s
|
0 |
Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.
What it measures
- Decision quality under real time pressure
- Response latency at the provider's default settings
What it does not measure
- Any business skill: this is chess against an engine
- Chess strength without a clock
Method
- Games against a Stockfish ladder at 3 minutes plus 2 seconds a move
- Every second of response time comes off the model's clock
- Ladder Elo with a 95% interval; it orders models, it is not a FIDE rating