Compare / Gemini 3.5 Flash Lite vs GPT-6.1 Sol

Gemini 3.5 Flash Lite vs GPT-6.1 Sol

51 results from 3 sources that measured both models. Each row is in the source's own units; there is no overall winner.

Gemini 3.5 Flash Lite
Google · profile
GPT-6.1 Sol
OpenAI · profile
Results compared
51
Shared sources
3
BenchmarkMetricGemini 3.5 Flash LiteGPT-6.1 Sol
CatalogBenchmeasured by usChannel rules broken% of products 1.2%0.0%
CatalogBenchmeasured by usClaims to check% of products 7.7%0.0%
CatalogBenchmeasured by usContent quality% of checks 94.8%98.8%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 21.4%5.4%
CatalogBenchmeasured by usPublish-ready listings% of products 10.7%72.0%
CatalogBenchmeasured by usReliably publish-ready% of products 3.6%67.9%
CatalogBenchmeasured by usUnsupported claims% of products 78.6%3.6%
CatalogBenchmeasured by usUnsupported claimsclaims per product 3.440.07
CatalogBenchmeasured by usWrong attributes% of products 35.1%19.1%
CatalogBenchmeasured by usWrong category or variant% of products 10.7%1.8%
CatalogBenchmeasured by usChannel compliance% of products 98.2%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 1.8%0.0%
CatalogBenchmeasured by usClaims to check% of products 1.8%0.0%
CatalogBenchmeasured by usConflicts caught% of conflicts 84.8%100.0%
CatalogBenchmeasured by usContent quality% of checks 95.5%98.6%
CatalogBenchmeasured by usCost per productUS dollars $0.0020$0.0084
CatalogBenchmeasured by usDecision accuracy% of decisions 93.4%98.7%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 92.5%94.4%
CatalogBenchmeasured by usInvented values% of filled values 0.8%2.2%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 20.2%4.8%
CatalogBenchmeasured by usPublish-ready listings% of products 35.1%78.0%
CatalogBenchmeasured by usReliably publish-ready% of products 17.9%75.0%
CatalogBenchmeasured by usUnsupported claimsclaims per product 1.180.00
CatalogBenchmeasured by usUnsupported claims% of products 33.3%0.0%
CatalogBenchmeasured by usWrong attributes% of products 34.5%16.7%
CatalogBenchmeasured by usWrong category or variant% of products 9.5%1.8%
DeckBenchmeasured by usAccurate decks% of tasks 0.0%100.0%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%0.0%
DeckBenchmeasured by usClean layout% of tasks 33.3%83.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0126$0.15
DeckBenchmeasured by usDeck ratingrating 6441,412
DeckBenchmeasured by usDesign quality% of the maximum 60.2%71.4%
DeckBenchmeasured by usDraft figure quoted% of tasks 50.0%0.0%
DeckBenchmeasured by usFindings missing% of tasks 33.3%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 26.9%74.6%
DeckBenchmeasured by usLayout defects% of tasks 66.7%16.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%50.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 83.3%0.0%
DeckBenchmeasured by usSlides needing work% of tasks 16.7%50.0%
DeckBenchmeasured by usUnsupported claims% of tasks 16.7%0.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 0.0%0.0%
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 22.251.8
Artificial AnalysisreportedHumanity's Last Exam% of questions 18.8%52.9%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 76.0%84.0%
Artificial AnalysisreportedSciCode% of problems 41.3%55.8%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 1.0%56.1%

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.