Models / Gemini 3.1 Pro Preview

Google

Gemini 3.1 Pro Preview

54 published results from 8 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Google
Sources
8
Our benchmarks
3
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
683
Ladder Elo · rank 1 of 16
Unit
ladder Elo, higher is better
Range
487 to 1,130
Sample
8 games
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
6.2 s
Median move time · rank 7 of 16
Unit
milliseconds, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · CatalogBench
96.6%
Content quality · rank 8 of 18
Unit
% of checks, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
3.6%
Publish-ready listings · rank 9 of 18
Unit
% of products, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Reliably publish-ready · rank 10 of 18
Unit
% of products, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
7.21
Unsupported claims · rank 13 of 18
Unit
claims per product, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
97.0%
Channel compliance · rank 15 of 18
Unit
% of products, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
100.0%
Conflicts caught · rank 1 of 18
Unit
% of conflicts, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
96.5%
Content quality · rank 9 of 18
Unit
% of checks, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
$0.0355
Cost per product · rank 14 of 18
Unit
US dollars, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · CatalogBench
94.8%
Decision accuracy · rank 10 of 18
Unit
% of decisions, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
0.0%
Failed outputs · rank 1 of 18
Unit
% of products, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a provider outage: provider errors are retried before a product counts as failed.
Measured by Spring Prompt · CatalogBench
95.4%
Field accuracy · rank 1 of 18
Unit
% of missing fields, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.5%
Invented values · rank 10 of 18
Unit
% of filled values, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
42.3%
Publish-ready listings · rank 10 of 18
Unit
% of products, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
25.0%
Reliably publish-ready · rank 9 of 18
Unit
% of products, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · CatalogBench
1.64
Unsupported claims · rank 12 of 18
Unit
claims per product, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
30 Sep 2026
Not shown
Not your catalogue: invented products with generated images, so a model's result on your own feed can differ.
Measured by Spring Prompt · ROASBench
$0.52
Cost of a run · rank 12 of 17
Unit
US dollars, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
43.5
Overall score · rank 9 of 17
Unit
score out of 100, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.42
Return on ad spend · rank 9 of 17
Unit
profit per £1 spent, higher is better
Configuration
Gemini 3.1 Pro Preview, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

-0.08
Confirmed task success · rank 32 of 46
Unit
IPS effect estimate, higher is better
Range
-0.10 to -0.06
Sample
69883 observations
Configuration
Gemini 3.1 Pro Preview
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.01
Praise over complaint · rank 17 of 46
Unit
IPS effect estimate, higher is better
Range
-0.03 to 0.02
Sample
26441 observations
Configuration
Gemini 3.1 Pro Preview
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.04
Steerability · rank 29 of 46
Unit
IPS effect estimate, higher is better
Range
-0.05 to -0.03
Sample
97403 observations
Configuration
Gemini 3.1 Pro Preview
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 1 of 46
Unit
IPS effect estimate, higher is better
Range
-0.01 to 0.00
Sample
2578396 observations
Configuration
Gemini 3.1 Pro Preview
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,459
Overall · rank 17 of 44
Unit
Arena rating, higher is better
Range
1,454 to 1,464
Sample
49457 votes
Configuration
Gemini 3.1 Pro Preview
Measured
13 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,208
Overall · rank 5 of 34
Unit
Arena rating, higher is better
Range
1,202 to 1,213
Sample
113282 votes
Configuration
Gemini 3.1 Pro Preview
Measured
24 Aug 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,472
Overall · rank 13 of 177
Unit
Arena rating, higher is better
Range
1,469 to 1,474
Sample
118511 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,477
Business, management and finance · rank 11 of 402
Unit
Arena rating, higher is better
Range
1,472 to 1,482
Sample
23398 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Creative writing · rank 4 of 407
Unit
Arena rating, higher is better
Range
1,475 to 1,486
Sample
21305 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,509
Expert prompts · rank 12 of 359
Unit
Arena rating, higher is better
Range
1,503 to 1,516
Sample
12336 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,479
Instruction following · rank 11 of 409
Unit
Arena rating, higher is better
Range
1,475 to 1,484
Sample
40860 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,487
Overall · rank 8 of 409
Unit
Arena rating, higher is better
Range
1,484 to 1,490
Sample
119196 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Writing, literature and language · rank 6 of 408
Unit
Arena rating, higher is better
Range
1,475 to 1,485
Sample
29992 votes
Configuration
Gemini 3.1 Pro Preview
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
60.6%
Average · rank 9 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
44.1%
Front end (SWE-Bench Multimodal) · rank 4 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
25.0%
Greenfield (Commit0) · rank 14 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
81.8%
Information gathering (GAIA) · rank 4 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
76.8%
Issue resolution (SWE-Bench) · rank 6 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
75.1%
Testing (SWT-Bench) · rank 9 of 34
Unit
% resolved, higher is better
Configuration
gemini-3.1-pro-preview
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by tau2-bench
9.3%
Consistency (pass^4) · rank 17 of 21
Unit
% of tasks, higher is better
Configuration
gemini-3.1-pro-preview
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by tau2-bench
26.0%
Task success (pass^1) · rank 15 of 21
Unit
% of tasks, higher is better
Configuration
gemini-3.1-pro-preview
Measured
5 May 2026
Not shown
Not your policies or customers: the tasks and policies are tau2's own, and the customer is simulated by gpt-5.2.
Reported by UGI Leaderboard
20.0%
Requested-length error · rank 179 of 370
Unit
% off the requested word count, lower is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
19.0%
Requested-length error · rank 163 of 370
Unit
% off the requested word count, lower is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
19.0%
Requested-length error · rank 163 of 370
Unit
% off the requested word count, lower is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.33
Style adherence · rank 249 of 370
Unit
score from 0 to 1, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.34
Style adherence · rank 209 of 370
Unit
score from 0 to 1, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.33
Style adherence · rank 259 of 370
Unit
score from 0 to 1, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
68.2
Writing score · rank 38 of 370
Unit
score out of 100, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
72.2
Writing score · rank 12 of 370
Unit
score out of 100, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
70.2
Writing score · rank 21 of 370
Unit
score out of 100, higher is better
Configuration
gemini-3.1-pro-preview
Measured
19 Feb 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
99.4%
Answer rate · rank 58 of 108
Unit
% of documents, higher is better
Configuration
Gemini 3.1 Pro Preview
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
10.4%
Hallucination rate · rank 61 of 108
Unit
% of summaries, lower is better
Configuration
Gemini 3.1 Pro Preview
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare Gemini 3.1 Pro Preview with