Confirm Action

Are you sure you want to proceed?

Independent model comparison

Gemini 3 Flash Preview vs Gemini 3.1 Pro Preview: evidence side by side

Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.

Exact comparison coverage

4 benchmark records share an exact reviewed comparison key.

Evidence strength

Limited exact coverage

The page is useful for inspection but does not yet contain enough pair-native evidence for Search publication.

Models at a glance

Key differences

Detail Gemini 3 Flash Preview Gemini 3.1 Pro Preview
Provider

Google

Google

Release date Shown only where an exact reviewed release date is available.

Not verified

Not verified

Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted.

4 shared records

4 shared records

Price and speed appear only when an exact reviewed operational reference is attached to that model identity.

Primary comparison evidence

Exact matched benchmark evidence

4 exact matches

A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.

PredictTheWeek pilot

Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?

Gemini 3 Flash Preview

12.0%

Mean prediction score

This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.

Configuration: Gemini 3 Flash

Spring Prompt · as of 2 Apr 2026 · 16–22 Mar 2026 → 23–29 Mar 2026

Gemini 3.1 Pro Preview

6.0%

Mean prediction score

This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.

Configuration: Gemini 3.1 Pro

Spring Prompt · as of 2 Apr 2026 · 16–22 Mar 2026 → 23–29 Mar 2026

Repository Issue Resolution

Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.

Gemini 3 Flash Preview

74.6%

OpenHands SWE-Bench resolved

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

Configuration: OpenHands v1.8.3 + Gemini-3-Flash

OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6

Gemini 3.1 Pro Preview

76.8%

OpenHands SWE-Bench resolved

Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.

Configuration: OpenHands v1.27.0 + Gemini-3.1-Pro

OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6

ROASBench

A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.

Gemini 3 Flash Preview

16.29

Average ROASBench score

Configuration: Google: Gemini 3 Flash Preview

Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation

Gemini 3.1 Pro Preview

27.14

Average ROASBench score

Configuration: Google: Gemini 3.1 Pro Preview

Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation

Structured Output Reliability

Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.

Gemini 3 Flash Preview

83.27%

Direct benchmark score

Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.

Configuration: Gemini-3-Flash-Preview

Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights

Gemini 3.1 Pro Preview

86.94%

Direct benchmark score

Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.

Configuration: Gemini-3.1-Pro-Preview

Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights

Secondary context · not head-to-head

Independent signals for each model

The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.

Anecdotal model-level reports

Initial community opinions from Reddit

Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.

Gemini 3 Flash Preview

No editor-reviewed Reddit opinions paragraph is active for this model snapshot.

Gemini 3.1 Pro Preview

Initial community opinions

In the search-indexed launch-window sample, Gemini 3.1 Pro appeared to improve long outputs, context retention, coding, and interface generation. Recurring reservations involved weak autonomous tool use, lengthy planning, flattery, inconsistent instruction following, and large differences between AI Studio, the consumer chat product, and Antigravity. Capability looked promising while agent reliability remained unsettled.

Source venue: Reddit · window: 2026-02-19 to 2026-03-05 · observed items: 4

Independent early-test reports

Early technical field tests from X

Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.

Gemini 3 Flash Preview

No editor-reviewed X field-test paragraph is active for this model snapshot.

Gemini 3.1 Pro Preview

Early technical field tests

Early Gemini 3.1 Pro reports were strongly harness-dependent. An overnight OpenCode run on a large production monorepo described sustained tool use, few tool failures, useful clarification behavior, and capable interface work. A separate production integration found speed and design strengths but inconsistent instruction and tool-schema adherence. The safest conclusion is high raw capability with uneven agent reliability between hosts.

Source venue: X · window: 2026-02-19 to 2026-03-05 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed

X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.

How to read this comparison

Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.

Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.