Models / Gemini 3.6 Flash

Google

Gemini 3.6 Flash

31 published results from 3 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Google
Sources
3
Our benchmarks
1
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
$0.0113
Cost per game · rank 8 of 23
Unit
US dollars, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
41.7%
Games lost on time · rank 10 of 23
Unit
% of games, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.3%
Invalid moves · rank 15 of 23
Unit
% of moves, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
628
Ladder Elo · rank 2 of 23
Unit
ladder Elo, higher is better
Range
429 to 849
Sample
12 games
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
1.6 s
Median move time · rank 14 of 23
Unit
milliseconds, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · BulletBench
$0.0050
Cost per game · rank 5 of 14
Unit
US dollars, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not your cost: prices are those charged on the run date.
Measured by Spring Prompt · BulletBench
91.7%
Games lost on time · rank 13 of 14
Unit
% of games, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
0.0%
Invalid moves · rank 1 of 14
Unit
% of moves, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
33
Ladder Elo · rank 10 of 14
Unit
ladder Elo, higher is better
Range
0 to 248
Sample
12 games
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
1.5 s
Median move time · rank 12 of 14
Unit
milliseconds, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning), minimal reasoning
Measured
1 Oct 2026
Not shown
Not general response speed: replies are one short move.

Reported by others

-0.08
Confirmed task success · rank 28 of 46
Unit
IPS effect estimate, higher is better
Range
-0.10 to -0.05
Sample
17636 observations
Configuration
Gemini 3.6 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.09
Praise over complaint · rank 32 of 46
Unit
IPS effect estimate, higher is better
Range
-0.11 to -0.06
Sample
7408 observations
Configuration
Gemini 3.6 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.04
Steerability · rank 24 of 46
Unit
IPS effect estimate, higher is better
Range
-0.06 to -0.02
Sample
24888 observations
Configuration
Gemini 3.6 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
0.00 to 0.00
Sample
1249005 observations
Configuration
Gemini 3.6 Flash
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,453
Overall · rank 17 of 44
Unit
Arena rating, higher is better
Range
1,443 to 1,464
Sample
2975 votes
Configuration
Gemini 3.6 Flash (high)
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,470
Overall · rank 14 of 177
Unit
Arena rating, higher is better
Range
1,466 to 1,474
Sample
33361 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,472
Business, management and finance · rank 13 of 402
Unit
Arena rating, higher is better
Range
1,464 to 1,480
Sample
6456 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,469
Creative writing · rank 6 of 407
Unit
Arena rating, higher is better
Range
1,461 to 1,477
Sample
6873 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,500
Expert prompts · rank 14 of 359
Unit
Arena rating, higher is better
Range
1,490 to 1,510
Sample
3963 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,473
Instruction following · rank 13 of 409
Unit
Arena rating, higher is better
Range
1,466 to 1,479
Sample
12339 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,482
Overall · rank 11 of 409
Unit
Arena rating, higher is better
Range
1,477 to 1,486
Sample
33739 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Writing, literature and language · rank 8 of 408
Unit
Arena rating, higher is better
Range
1,465 to 1,480
Sample
9055 votes
Configuration
Gemini 3.6 Flash (high)
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by UGI Leaderboard
11.0%
Requested-length error · rank 91 of 370
Unit
% off the requested word count, lower is better
Configuration
Gemini 3.6 Flash (high)
Measured
3 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
12.0%
Requested-length error · rank 98 of 370
Unit
% off the requested word count, lower is better
Configuration
Gemini 3.6 Flash (medium reasoning)
Measured
3 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
9.0%
Requested-length error · rank 67 of 370
Unit
% off the requested word count, lower is better
Configuration
Gemini 3.6 Flash (minimal reasoning)
Measured
3 Sep 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.32
Style adherence · rank 285 of 370
Unit
score from 0 to 1, higher is better
Configuration
Gemini 3.6 Flash (high)
Measured
3 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.33
Style adherence · rank 266 of 370
Unit
score from 0 to 1, higher is better
Configuration
Gemini 3.6 Flash (medium reasoning)
Measured
3 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.30
Style adherence · rank 326 of 370
Unit
score from 0 to 1, higher is better
Configuration
Gemini 3.6 Flash (minimal reasoning)
Measured
3 Sep 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
69.4
Writing score · rank 29 of 370
Unit
score out of 100, higher is better
Configuration
Gemini 3.6 Flash (high)
Measured
3 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
69.8
Writing score · rank 24 of 370
Unit
score out of 100, higher is better
Configuration
Gemini 3.6 Flash (medium reasoning)
Measured
3 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
65.7
Writing score · rank 58 of 370
Unit
score out of 100, higher is better
Configuration
Gemini 3.6 Flash (minimal reasoning)
Measured
3 Sep 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.

Compare Gemini 3.6 Flash with