Gemini 3 Flash Preview
- Benchmark records
- 4
- Snapshot
- 2026-07-17
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
3 benchmark records share an exact reviewed comparison key.
Evidence strength
Limited exact coverage
The page is useful for inspection but does not yet contain enough pair-native evidence for Search publication.
Models at a glance
| Detail | Gemini 3 Flash Preview | Gemini 3.5 Flash |
|---|---|---|
| Provider |
|
|
| Release date Shown only where an exact reviewed release date is available. |
Not verified |
Not verified |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
3 shared records |
3 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Gemini 3 Flash Preview
74.6%
OpenHands SWE-Bench resolved
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
Configuration: OpenHands v1.8.3 + Gemini-3-Flash
OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6
Gemini 3.5 Flash
78.6%
OpenHands SWE-Bench resolved
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
Configuration: OpenHands v1.28.0 + Gemini-3.5-Flash
OpenHands Index · as of 30 Jun 2026 · OpenHands · SWE-Bench 2026.06.30-3015ac6
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Gemini 3 Flash Preview
16.29
Average ROASBench score
Configuration: Google: Gemini 3 Flash Preview
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
Gemini 3.5 Flash
20.74
Average ROASBench score
Configuration: Google: Gemini 3.5 Flash · High
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Gemini 3 Flash Preview
83.27%
Direct benchmark score
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Configuration: Gemini-3-Flash-Preview
Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights
Gemini 3.5 Flash
85.64%
Direct benchmark score
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Configuration: Gemini-3.5-Flash
Structured Output Benchmark (SOB) · as of 17 Jul 2026 · Sob-v1@da785a8521c8954283b2989d01e54d80c4e023c6:upstream-provider-configs:temperature-0-where-supported:max-output-2048:reasoning-disabled-or-minimum-where-required:official-modality-weights
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
No editor-reviewed model-level signals are attached to this comparison snapshot.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.