Compare / DeepSeek V4 Pro 0423 vs GPT-6.1 Sol

DeepSeek V4 Pro 0423 vs GPT-6.1 Sol

22 results from 3 sources that measured both models. Each row is in the source's own units; there is no overall winner.

DeepSeek V4 Pro 0423
DeepSeek · profile
GPT-6.1 Sol
OpenAI · profile
Results compared
22
Shared sources
3
BenchmarkMetricDeepSeek V4 Pro 0423GPT-6.1 Sol
DeckBenchmeasured by usAccurate decks% of tasks 0.0%100.0%
DeckBenchmeasured by usCaveat dropped% of tasks 16.7%0.0%
DeckBenchmeasured by usClean layout% of tasks 33.3%83.3%
DeckBenchmeasured by usCost per deckUS dollars $0.0465$0.15
DeckBenchmeasured by usDeck ratingrating 6151,412
DeckBenchmeasured by usDesign quality% of the maximum 40.9%71.4%
DeckBenchmeasured by usDraft figure quoted% of tasks 33.3%0.0%
DeckBenchmeasured by usFindings missing% of tasks 16.7%0.0%
DeckBenchmeasured by usHead-to-head win rate% of comparisons 18.2%74.6%
DeckBenchmeasured by usLayout defects% of tasks 66.7%16.7%
DeckBenchmeasured by usMisleading metric used% of tasks 0.0%0.0%
DeckBenchmeasured by usPresentable decks% of tasks 0.0%50.0%
DeckBenchmeasured by usRecommendation late or wrong% of tasks 66.7%0.0%
DeckBenchmeasured by usSlides needing work% of tasks 83.3%50.0%
DeckBenchmeasured by usUnsupported claims% of tasks 66.7%0.0%
DeckBenchmeasured by usUnsupported numbers% of tasks 0.0%0.0%
Artificial AnalysisreportedArtificial Analysis Intelligence Indexindex score 36.051.8
Artificial AnalysisreportedHumanity's Last Exam% of questions 41.0%52.9%
Artificial AnalysisreportedLong-context reasoning (AA-LCR)% of questions 80.3%84.0%
Artificial AnalysisreportedSciCode% of problems 51.0%55.8%
Artificial AnalysisreportedTerminal-Bench 4.0% of tasks 14.1%56.1%
SimpleQA Verified (Epoch AI)reportedCorrect answers% of questions 47.0%73.9%

Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.