Provider
Base model family
Model intelligence profile
This family profile brings together 15 published benchmark records for Google Gemini 3.1 Pro Preview. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
Base model family
Release date
Not yet verified
Exact-family source only
Published benchmarks
15
Native scales kept separate
Artificial Analysis
No exact match
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
3 configurations
60-second bullet Elo
Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.
Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.
Gemini 3.1 Pro (low)
35
Gemini 3.1 Pro (medium)
35
Gemini 3.1 Pro (high)
35
Spring Prompt benchmark
6.0%
Mean prediction score
Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?
This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.
Gemini 3.1 Pro
6.0%
Spring Prompt benchmark
27.14
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Google: Gemini 3.1 Pro Preview
27.14
Spring Prompt benchmark
286
Tournament prediction points
Fixture-by-fixture football predictions frozen before the tournament and graded on outcomes, scorelines, goals, penalties, and cards.
32 fixtures graded from the frozen prediction artifact.
Gemini 3.1 Pro
286
source-native benchmark
46
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Gemini 3.1 Pro Preview
46
source-native benchmark
29.2
Coding Agent Index v1.2 point score
Artificial Analysis Coding Agent Index v1.2 point scores and task-specific operational measurements for exact agent and model configurations.
These are exact agent-plus-model configurations, not bare-model scores.
Gemini CLI - Gemini 3.1 Pro (high)
29.2
source-native benchmark
23%
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Gemini 3.1 Pro Preview
23%
source-native benchmark
94%
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Gemini 3.1 Pro Preview
94%
source-native benchmark
45%
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Gemini 3.1 Pro Preview
45%
source-native benchmark
77%
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Gemini 3.1 Pro Preview
77%
source-native benchmark
82%
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Gemini 3.1 Pro Preview
82%
source-native benchmark
76.8%
OpenHands SWE-Bench resolved
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
OpenHands v1.27.0 + Gemini-3.1-Pro
76.8%
source-native benchmark
86.94%
Direct benchmark score
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Gemini-3.1-Pro-Preview
86.94%
source-native benchmark
74%
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Gemini 3.1 Pro Preview
74%
source-native benchmark
96%
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Gemini 3.1 Pro Preview
96%
Independent operational reference
No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.
Early-user signal
A paraphrased editorial synthesis of 4 Reddit discussions created in the 14 days after launch (19 Feb 2026 to 5 Mar 2026). Anecdotal context only — it never affects scores or rankings.
In the search-indexed launch-window sample, Gemini 3.1 Pro appeared to improve long outputs, context retention, coding, and interface generation. Recurring reservations involved weak autonomous tool use, lengthy planning, flattery, inconsistent instruction following, and large differences between AI Studio, the consumer chat product, and Antigravity. Capability looked promising while agent reliability remained unsettled.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 19 Feb 2026 up to 5 Mar 2026, and exact model identity was editor reviewed.
Early Gemini 3.1 Pro reports were strongly harness-dependent. An overnight OpenCode run on a large production monorepo described sustained tool use, few tool failures, useful clarification behavior, and capable interface work. A separate production integration found speed and design strengths but inconsistent instruction and tool-schema adherence. The safest conclusion is high raw capability with uneven agent reliability between hosts.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published Gemini 3.1 Pro Preview results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from Google