Gemini 3.1 Flash Lite
- Benchmark records
- 8
- Snapshot
- 2026-07-30
Independent model comparison
Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.
Exact comparison coverage
8 benchmark records share an exact reviewed comparison key.
Evidence strength
Editorial review pending
Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.
Models at a glance
| Detail | Gemini 3.1 Flash Lite | Gemini 3.1 Pro Preview |
|---|---|---|
| Provider |
|
|
| Release date Shown only where an exact reviewed release date is available. |
Not verified |
Not verified |
| Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted. |
8 shared records |
8 shared records |
Price and speed appear only when an exact reviewed operational reference is attached to that model identity.
Primary comparison evidence
A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.
A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.
Gemini 3.1 Flash Lite
25
Artificial Analysis Intelligence Index
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Gemini 3.1 Pro Preview
46
Artificial Analysis Intelligence Index
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30
Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.
Gemini 3.1 Flash Lite
7%
GDPval-AA v2
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
23%
GDPval-AA v2
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.
Gemini 3.1 Flash Lite
82%
GPQA Diamond
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
94%
GPQA Diamond
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A broad expert-level benchmark of difficult academic reasoning and knowledge questions.
Gemini 3.1 Flash Lite
16%
Humanity's Last Exam
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
45%
Humanity's Last Exam
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.
Gemini 3.1 Flash Lite
77%
IFBench
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
77%
IFBench
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.
Gemini 3.1 Flash Lite
76%
MMMU-Pro
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
82%
MMMU-Pro
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.
Gemini 3.1 Flash Lite
31%
Terminal-Bench v2.1
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
74%
Terminal-Bench v2.1
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
A legacy agentic tool-use benchmark for conversational work in a telecom environment.
Gemini 3.1 Flash Lite
31%
τ²-Bench Telecom
Configuration: Gemini 3.1 Flash-Lite
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Gemini 3.1 Pro Preview
96%
τ²-Bench Telecom
Configuration: Gemini 3.1 Pro Preview
Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30
Secondary context · not head-to-head
The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.
Anecdotal model-level reports
Recurring themes from each model's declared observation window. They are not a representative survey and are not direct comparisons between these models.
No editor-reviewed Reddit opinions paragraph is active for this model snapshot.
Initial community opinions
In the search-indexed launch-window sample, Gemini 3.1 Pro appeared to improve long outputs, context retention, coding, and interface generation. Recurring reservations involved weak autonomous tool use, lengthy planning, flattery, inconsistent instruction following, and large differences between AI Studio, the consumer chat product, and Antigravity. Capability looked promising while agent reliability remained unsettled.
Source venue: Reddit · window: 2026-02-19 to 2026-03-05 · observed items: 4
Independent early-test reports
Editor-paraphrased reports from independent authors in each model's fixed 14-day launch window. These selectively surfaced anecdotes are not direct head-to-head tests unless they also appear in the dedicated direct-report section above.
No editor-reviewed X field-test paragraph is active for this model snapshot.
Early technical field tests
Early Gemini 3.1 Pro reports were strongly harness-dependent. An overnight OpenCode run on a large production monorepo described sustained tool use, few tool failures, useful clarification behavior, and capable interface work. A separate production integration found speed and design strengths but inconsistent instruction and tool-schema adherence. The safest conclusion is high raw capability with uneven agent reliability between hosts.
Source venue: X · window: 2026-02-19 to 2026-03-05 · 2 tests · 2 independent authors · 2 with a method or artifact · identity reviewed
X field tests remain separate from Reddit community opinions, Artificial Analysis facts, and every scored benchmark.
Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.
Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.