Confirm Action

Are you sure you want to proceed?

Model intelligence profile

Gemini 3.1 Pro Preview benchmark results and model details

This family profile brings together 15 published benchmark records for Google Gemini 3.1 Pro Preview. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.

Published benchmark evidence available · Business Skills V3 not yet authorized

Provider

Google

Base model family

Release date

Not yet verified

Exact-family source only

Published benchmarks

15

Native scales kept separate

Artificial Analysis

No exact match

As of 2026-07-16

Benchmark-level evidence

Published Gemini 3.1 Pro Preview benchmark results

These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.

How sources are reviewed →

Spring Prompt benchmark

BulletBench

3 configurations

60-second bullet Elo

Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.

Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.

3 published configurations

Gemini 3.1 Pro (low)

35

Bullet 95% interval
0–172
10+1 Elo
0
Bullet games
96
Bullet median move time
5.06s
Average bullet game cost
$0.0494

Gemini 3.1 Pro (medium)

35

Bullet 95% interval
0–162
10+1 Elo
0
Bullet games
96
Bullet median move time
5.52s
Average bullet game cost
$0.0468

Gemini 3.1 Pro (high)

35

Bullet 95% interval
0–172
10+1 Elo
0
Bullet games
96
Bullet median move time
5.84s
Average bullet game cost
$0.0445
Source: Spring Prompt Recorded 3 Aug 2026 Ladder v1 · controls 10+1 and 60 Source record →

Spring Prompt benchmark

PredictTheWeek pilot

6.0%

Mean prediction score

Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?

This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.

1 published configuration

Gemini 3.1 Pro

6.0%

Accuracy (partial support or better)
8.0%
Outcome-line coverage
2.9%
Pilot composite
4.7%
Source: Spring Prompt Recorded 2 Apr 2026 16–22 Mar 2026 → 23–29 Mar 2026 Source record →

Spring Prompt benchmark

ROASBench

27.14

Average ROASBench score

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

1 published configuration

Google: Gemini 3.1 Pro Preview

27.14

Average contribution profit
$-34,549
Average ROAS
1.33×
Completed runs
3
Score variability
±2.46
Source: Spring Prompt Recorded 25 Jul 2026 Published ROASBench cache · 12-month simulation Source record →

Spring Prompt benchmark

World Cup prediction benchmark

286

Tournament prediction points

Fixture-by-fixture football predictions frozen before the tournament and graded on outcomes, scorelines, goals, penalties, and cards.

32 fixtures graded from the frozen prediction artifact.

1 published configuration

Gemini 3.1 Pro

286

Graded matches
32
Exact-score points
18
Result points
30
Goal points
60
Source: Spring Prompt Recorded 20 Jul 2026 Predictions frozen 2026-06-11 Source record →

source-native benchmark

Artificial Analysis Intelligence Index

46

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

1 published configuration

Gemini 3.1 Pro Preview

46

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Intelligence Index v4.1 · 2026-07-30 Source record →

source-native benchmark

Coding Agent Configurations

29.2

Coding Agent Index v1.2 point score

Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.

These are exact agent-plus-model configurations, not bare-model scores.

1 published configuration

Gemini CLI - Gemini 3.1 Pro (high)

29.2

Mean cost per task
$2.00
Mean wall time
10.8 min
Source: Artificial Analysis Coding Agent Index Recorded 20 Jul 2026 V1.2 Source record →

source-native benchmark

GDPval-AA v2

23%

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

1 published configuration

Gemini 3.1 Pro Preview

23%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

GPQA Diamond

94%

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

1 published configuration

Gemini 3.1 Pro Preview

94%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Humanity's Last Exam

45%

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

1 published configuration

Gemini 3.1 Pro Preview

45%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

IFBench

77%

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

1 published configuration

Gemini 3.1 Pro Preview

77%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

MMMU-Pro

82%

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

1 published configuration

Gemini 3.1 Pro Preview

82%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

Repository Issue Resolution

76.8%

OpenHands SWE-Bench resolved

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

1 published configuration

OpenHands v1.27.0 + Gemini-3.1-Pro

76.8%

95% task-bootstrap interval
Unavailable
Task evidence
not-eligible-for-task-bootstrap
Source: OpenHands Index Recorded 30 Jun 2026 OpenHands · SWE-Bench 2026.06.30-3015ac6 Source record →

source-native benchmark

Structured Output Reliability

86.94%

Direct benchmark score

Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.

Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.

1 published configuration

Gemini-3.1-Pro-Preview

86.94%

95% interval
86.41–87.52
Evaluated modality coverage
100%
Source: Structured Output Benchmark (SOB) Recorded 17 Jul 2026 Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights Source record →

source-native benchmark

Terminal-Bench v2.1

74%

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

1 published configuration

Gemini 3.1 Pro Preview

74%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

source-native benchmark

τ²-Bench Telecom

96%

A legacy agentic tool-use benchmark for conversational work in a telecom environment.

1 published configuration

Gemini 3.1 Pro Preview

96%

Status
Measured
Recorded
30 Jul 2026
Source: Artificial Analysis Recorded 30 Jul 2026 Independent evaluation · Public leaderboard snapshot · 2026-07-30 Source record →

Independent operational reference

Artificial Analysis details

No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.

Early-user signal

Initial community opinions

Anecdotal · never scored

A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (19 Feb 2026 to 5 Mar 2026). Anecdotal context only — it never affects scores or rankings.

In the search-indexed launch-window sample, Gemini 3.1 Pro appeared to improve long outputs, context retention, coding, and interface generation. Recurring reservations involved weak autonomous tool use, lengthy planning, flattery, inconsistent instruction following, and large differences between AI Studio, the consumer chat product, and Antigravity. Capability looked promising while agent reliability remained unsettled.

Independent early-test signal

Early technical field tests

X · anecdotal · never scored

An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 19 Feb 2026 up to 5 Mar 2026, and exact model identity was editor reviewed.

Early Gemini 3.1 Pro reports were strongly harness-dependent. An overnight OpenCode run on a large production monorepo described sustained tool use, few tool failures, useful clarification behavior, and capable interface work. A separate production integration found speed and design strengths but inconsistent instruction and tool-schema adherence. The safest conclusion is high raw capability with uneven agent reliability between hosts.

These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.

Business Skills V3 · proposed

How Spring Prompt plans to test Gemini 3.1 Pro Preview

The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.

Proposed reasoning
High reasoning
Provider revision
Exact provider revision will be resolved and frozen only after preflight and execution approval
Tools and service
No tools · Provider-default service tier
Planning configuration
gemini-3.1-pro-preview-high
Generation controls
Provider-managed reasoning; no temperature override proposed

Future first-party coverage

15 proposed Business Skills V3 task areas

These links describe evaluation contracts, not published Gemini 3.1 Pro Preview results.

Model comparisons

Compare Gemini 3.1 Pro Preview side by side

Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.

Publication safeguards

What must pass before a V3 result appears

  1. 1First-party task-local comparisons and eligible external evidence must both be present.
  2. 2Model and provider configuration identity must match the reviewed release exactly.
  3. 3Coverage, reliability, judge diagnostics, and sealed stability checks must pass.
  4. 4Uncertainty and missing evidence remain visible when results are published.

Stable family URL

Evidence can grow without changing the page

New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.

More from Google

Other owned model profiles