Provider
DeepSeek
Base model family
Model intelligence profile
This family profile brings together 3 published benchmark records for DeepSeek V3.2. Every result keeps its original benchmark, configuration, scale, and source; unrelated scores are never averaged.
Provider
DeepSeek
Base model family
Release date
Not yet verified
Exact-family source only
Published benchmarks
3
Native scales kept separate
Artificial Analysis
No exact match
As of 2026-07-16
Benchmark-level evidence
These are individual benchmarks, not collection rollups. Scores remain on their original scales, and agent or harness results stay labelled as system configurations.
Spring Prompt benchmark
10.0%
Mean prediction score
Can LLMs anticipate next week’s Guardian agenda from last week’s coverage?
This benchmark is not on a weekly live cadence yet. The table and charts reflect a single multi-model comparison on one forecast window—a pilot snapshot, not an updating leaderboard. The evaluation setup (clustering, prompts, automated judge, and scoring checks) is still being refined; reported scores and details may change as we improve the pipeline.
DeepSeek V3.2
10.0%
Spring Prompt benchmark
19.77
Average ROASBench score
A 12-month performance-marketing simulation scored on business outcomes, planning, behavior, and persona fit.
DeepSeek: DeepSeek V3.2
19.77
source-native benchmark
71.6%
OpenHands SWE-Bench resolved
Direct OpenHands SWE-Bench outcomes for resolving real repository issues with a pinned OpenHands agent and language-model configuration.
Scores are official OpenHands SWE-Bench resolved percentages for exact OpenHands and model configurations, not bare-model scores or SpringPrompt predictions. Twenty-nine exact task sidecars support marginal 95% bootstrap score intervals; five rows explicitly have no interval. No full-cohort rank confidence is claimed.
OpenHands v1.8.3 + DeepSeek-V3.2-Reasoner
71.6%
Independent operational reference
No exact, reviewed Artificial Analysis operational record is attached to this family. We do not inherit price, release, or throughput values from a nearby model name.
Early-user signal
No editor-reviewed community-opinions paragraph is active for this model. An active paragraph requires at least two retained Reddit discussions from the minimum 14-day launch window—implemented exactly as release date through day 14, end-exclusive—plus paraphrase-only review and a sealed publication artifact. Community reports stay separate from every benchmark score and model comparison.
Independent early-test signal
An editorial paraphrase of 2 launch-window field tests from 2 independent authors on X, including 2 reports with a described method or inspectable artifact. The window runs from 1 Dec 2025 up to 15 Dec 2025, and exact model identity was editor reviewed.
Early DeepSeek V3.2 evidence covered two narrow reasoning checks. A procedurally generated, out-of-distribution logic benchmark produced a favorable result within its tested open-model set, while a separate user shared the output from one International Mathematical Olympiad challenge. Both are useful snapshots, but neither demonstrates tool-use reliability or consistent production behavior.
These reports are selectively surfaced and are not a representative sample. They never affect benchmark scores, rankings, winners, or comparison outcomes.
Business Skills V3 · proposed
The setup below is proposed and may change until preflight and execution approval are complete. Existing benchmark evidence above does not authorize or stand in for a Business Skills V3 result.
Future first-party coverage
These links describe evaluation contracts, not published DeepSeek V3.2 results.
Model comparisons
Editorially reviewed comparisons appear first. Every page matches only benchmark records with the same reviewed protocol key and does not manufacture an overall winner.
Publication safeguards
Stable family URL
New reviewed benchmark snapshots, operational facts, community themes, and early field-test syntheses can be added here while the canonical model-family identity remains fixed.
More from DeepSeek