Provider
Base model family
Model intelligence profile
This family profile brings together 12 published benchmark records for Google Gemini 3.5 Flash. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
Base model family
Release date
Not yet verified
Exact-family source only
Published benchmarks
12
Native scales kept separate
Artificial Analysis
No exact match
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
4 configurations
60-second bullet Elo
Chess decision quality under a shared 60-second game clock and a 10+1 per-move clock.
Ladder Elo: anchored to a Stockfish skill ladder (random mover = 400). Internally consistent ordering, not FIDE-calibrated.
Gemini 3.5 Flash (minimal)
808
Gemini 3.5 Flash (low)
744
Gemini 3.5 Flash (medium)
601
Gemini 3.5 Flash (high)
416
Spring Prompt benchmark
20.74
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Google: Gemini 3.5 Flash · High
20.74
source-native benchmark
50
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Gemini 3.5 Flash (high)
50
source-native benchmark
42%
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Gemini 3.5 Flash (high)
42%
source-native benchmark
92%
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Gemini 3.5 Flash (high)
92%
source-native benchmark
41%
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Gemini 3.5 Flash (high)
41%
source-native benchmark
76%
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Gemini 3.5 Flash (high)
76%
source-native benchmark
84%
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Gemini 3.5 Flash (high)
84%
source-native benchmark
78.6%
OpenHands SWE-Bench resolved
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
OpenHands v1.28.0 + Gemini-3.5-Flash
78.6%
source-native benchmark
85.64%
Direct benchmark score
Directly measured ability to return accurate values in the requested structured schema across the benchmark's evaluated text, image and audio modalities.
Point order reproduces the source's direct Overall score. Rank ranges come from overlap of marginal record-cluster bootstrap intervals for Overall; they are not simultaneous confidence intervals for rank.
Gemini-3.5-Flash
85.64%
source-native benchmark
79%
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Gemini 3.5 Flash (high)
79%
source-native benchmark
95%
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Gemini 3.5 Flash (high)
95%
Independent operational reference
No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.
Early-user signal
No editor-reviewed community-opinions paragraph is active for this model. An active paragraph requires at least two retained Reddit discussions from the minimum 14-day launch window—implemented exactly as release date through day 14, end-exclusive—plus paraphrase-only review and a sealed publication artifact. Community reports stay separate from every benchmark score and model comparison.
Independent early-test signal
No independent field-test write-up has cleared review for this model yet — insufficient independent tests.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published Gemini 3.5 Flash results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from Google