Confirm Action

Are you sure you want to proceed?

Independent model comparison

Gemini 3.1 Flash Lite vs Gemini 3.5 Flash: evidence side by side

Compare the models using only evidence that can be identified and attributed. Exact benchmark matches appear first; independent facts and user reports remain separate and do not create an overall ranking.

Exact comparison coverage

8 benchmark records share an exact reviewed comparison key.

Evidence strength

Editorial review pending

Substantial matched evidence exists, but this pair has not passed the pair-native publication gate.

Models at a glance

Key differences

Detail Gemini 3.1 Flash Lite Gemini 3.5 Flash
Provider

Google

Google

Release date Shown only where an exact reviewed release date is available.

Not verified

Not verified

Exact benchmark coverage Only records carrying the same reviewed comparison identity are counted.

8 shared records

8 shared records

Price and speed appear only when an exact reviewed operational reference is attached to that model identity.

Primary comparison evidence

Exact matched benchmark evidence

8 exact matches

A shared benchmark name is not enough. A result appears side by side only when both model records carry the same reviewed comparison key. Duplicate or unversioned records are omitted rather than guessed into alignment.

Artificial Analysis Intelligence Index

A composite index of language-model performance across agentic work, coding, scientific reasoning, knowledge, and long-context reasoning.

Gemini 3.1 Flash Lite

25

Artificial Analysis Intelligence Index

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30

Gemini 3.5 Flash

50

Artificial Analysis Intelligence Index

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Intelligence Index v4.1 · 2026-07-30

GDPval-AA v2

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

Gemini 3.1 Flash Lite

7%

GDPval-AA v2

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

42%

GDPval-AA v2

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

GPQA Diamond

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

Gemini 3.1 Flash Lite

82%

GPQA Diamond

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

92%

GPQA Diamond

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Humanity's Last Exam

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

Gemini 3.1 Flash Lite

16%

Humanity's Last Exam

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

41%

Humanity's Last Exam

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

IFBench

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

Gemini 3.1 Flash Lite

77%

IFBench

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

76%

IFBench

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

MMMU-Pro

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

Gemini 3.1 Flash Lite

76%

MMMU-Pro

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

84%

MMMU-Pro

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Terminal-Bench v2.1

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

Gemini 3.1 Flash Lite

31%

Terminal-Bench v2.1

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

79%

Terminal-Bench v2.1

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

τ²-Bench Telecom

A legacy agentic tool-use benchmark for conversational work in a telecom environment.

Gemini 3.1 Flash Lite

31%

τ²-Bench Telecom

Configuration: Gemini 3.1 Flash-Lite

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Gemini 3.5 Flash

95%

τ²-Bench Telecom

Configuration: Gemini 3.5 Flash (high)

Artificial Analysis · as of 30 Jul 2026 · Independent evaluation · Public leaderboard snapshot · 2026-07-30

Secondary context · not head-to-head

Independent signals for each model

The sources below evaluated or discussed each model independently. Putting them in adjacent columns makes them easier to inspect, but it does not make their metrics, samples, or observations directly comparable.

No editor-reviewed model-level signals are attached to this comparison snapshot.

How to read this comparison

Spring Prompt does not name an overall winner from unrelated benchmark scores, third-party metrics, Reddit opinions, or X field-test reports.

Different benchmarks measure different constructs and may use different populations, prompts, harnesses, model revisions, tools, and score scales. This page does not average those values or infer a winner from source coverage. Test both models on your own production work before making a consequential choice.