Compare / Mistral Large 4 vs GPT-6.1 Sol

Mistral Large 4 vs GPT-6.1 Sol

53 results from 3 sources that measured both models. Each row is in the source's own units; there is no overall winner.

Mistral Large 4
Mistral · profile
GPT-6.1 Sol
OpenAI · profile
Results compared
53
Shared sources
3
BenchmarkMetricMistral Large 4GPT-6.1 Sol
CatalogBenchmeasured by usChannel rules broken% of products 1.2%0.0%
CatalogBenchmeasured by usClaims to check% of products 3.3%0.0%
CatalogBenchmeasured by usContent quality% of checks 95.9%98.8%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 3.0%5.4%
CatalogBenchmeasured by usPublish-ready listings% of products 10.7%72.0%
CatalogBenchmeasured by usReliably publish-ready% of products 1.8%67.9%
CatalogBenchmeasured by usUnsupported claims% of products 80.4%3.6%
CatalogBenchmeasured by usUnsupported claimsclaims per product 3.690.07
CatalogBenchmeasured by usWrong attributes% of products 17.3%19.1%
CatalogBenchmeasured by usWrong category or variant% of products 1.8%1.8%
CatalogBenchmeasured by usChannel compliance% of products 100.0%100.0%
CatalogBenchmeasured by usChannel rules broken% of products 3.0%0.0%
CatalogBenchmeasured by usClaims to check% of products 0.7%0.0%
CatalogBenchmeasured by usConflicts caught% of conflicts 100.0%100.0%
CatalogBenchmeasured by usContent quality% of checks 96.5%98.6%
CatalogBenchmeasured by usCost per productUS dollars $0.0033$0.0084
CatalogBenchmeasured by usDecision accuracy% of decisions 93.1%98.7%
CatalogBenchmeasured by usFailed outputs% of products 0.0%0.0%
CatalogBenchmeasured by usField accuracy% of missing fields 94.4%94.4%
CatalogBenchmeasured by usInvented values% of filled values 3.5%2.2%
CatalogBenchmeasured by usMissing UK information% of products 0.0%0.0%
CatalogBenchmeasured by usNot findable% of products 4.2%4.8%
CatalogBenchmeasured by usPublish-ready listings% of products 39.9%78.0%
CatalogBenchmeasured by usReliably publish-ready% of products 10.7%75.0%
CatalogBenchmeasured by usUnsupported claimsclaims per product 0.910.00
CatalogBenchmeasured by usUnsupported claims% of products 33.3%0.0%
CatalogBenchmeasured by usWrong attributes% of products 19.6%16.7%
CatalogBenchmeasured by usWrong category or variant% of products 1.2%1.8%
DeckBenchmeasured by usAccurate decks% of tasks 0.0%100.0%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%0.0%
DeckBenchmeasured by usClean layout% of tasks 33.3%83.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0148$0.15
DeckBenchmeasured by usDeck ratingrating 5151,412
DeckBenchmeasured by usDesign quality% of the maximum 30.5%71.4%
DeckBenchmeasured by usDraft figure quoted% of tasks 16.7%0.0%
DeckBenchmeasured by usFindings missing% of tasks 0.0%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 10.7%74.6%
DeckBenchmeasured by usLayout defects% of tasks 66.7%16.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%50.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 83.3%0.0%
DeckBenchmeasured by usSlides needing work% of tasks 100.0%50.0%
DeckBenchmeasured by usUnsupported claims% of tasks 50.0%0.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 33.3%0.0%
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 38.451.8
Artificial AnalysisreportedHumanity's Last Exam% of questions 35.0%52.9%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 81.3%84.0%
Artificial AnalysisreportedOutput speedtokens per second 10657.2
Artificial AnalysisreportedSciCode% of problems 54.2%55.8%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 26.8%56.1%
Artificial AnalysisreportedTime to first answer tokenseconds 19.81.81

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.