Gemini 3.5 Flash
- Benchmark records
- 11
- Snapshot
- 2026-07-30
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
7 benchmark records share an exact reviewed comparison key.
Evidence strength
Editorial review pending
Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.
Models at a glance
OpenAI
| Detail | Gemini 3.5 Flash | GPT-5.6 Luna |
|---|---|---|
| Provider |
|
OpenAI |
| Release date Shown only where an exact reviewed release date is available. |
Not verified |
Not verified |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
7 shared records |
7 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Gemini 3.5 Flash
50
Artificial Analysis Intelligence Index
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
GPT-5.6 Luna
46
Artificial Analysis Intelligence Index
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Gemini 3.5 Flash
42%
GDPval-AA v2
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
GPT-5.6 Luna
48%
GDPval-AA v2
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Gemini 3.5 Flash
92%
GPQA Diamond
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
GPT-5.6 Luna
89%
GPQA Diamond
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Gemini 3.5 Flash
41%
Humanity's Last Exam
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
GPT-5.6 Luna
32%
Humanity's Last Exam
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Gemini 3.5 Flash
84%
MMMU-Pro
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
GPT-5.6 Luna
78%
MMMU-Pro
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
Gemini 3.5 Flash
20.74
Average ROASBench score
Configuration: Google: Gemini 3.5 Flash · High
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
GPT-5.6 Luna
4 configurations
Average ROASBench score
Configuration: 4 configurations compared separately
Spring Prompt · Reviewed public-catalogue SQLite projection · 12-month simulation
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Gemini 3.5 Flash
79%
Terminal-Bench v2.1
Configuration: Gemini 3.5 Flash (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
GPT-5.6 Luna
70%
Terminal-Bench v2.1
Configuration: GPT-5.6 Luna (high)
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
No editor-reviewed model-level signals are attached to this comparison snapshot.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.