Models / DeepSeek V4 Pro 0423

DeepSeek

DeepSeek V4 Pro 0423

37 published results from 7 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
DeepSeek
Sources
7
Our benchmarks
2
Price
Not yet published

Measured by Spring Prompt

Measured by Spring Prompt · BulletBench
0
Ladder Elo · rank 12 of 16
Unit
ladder Elo, higher is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
30 Sep 2026
Not shown
Not a business skill: chess against an engine on a real clock, used here as a speed-under-pressure signal.
Measured by Spring Prompt · BulletBench
6.4 s
Median move time · rank 8 of 16
Unit
milliseconds, lower is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
30 Sep 2026
Not shown
Not general response speed: replies are one short move.
Measured by Spring Prompt · ROASBench
$0.20
Cost of a run · rank 8 of 17
Unit
US dollars, lower is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
29 Sep 2026
Not shown
Not your cost: prices are those charged through OpenRouter on the run date.
Measured by Spring Prompt · ROASBench
0
Months over budget · rank 1 of 17
Unit
months of 12, lower is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
28.2
Overall score · rank 12 of 17
Unit
score out of 100, higher is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.
Measured by Spring Prompt · ROASBench
£0.26
Return on ad spend · rank 11 of 17
Unit
profit per £1 spent, higher is better
Configuration
DeepSeek V4 Pro 0423, provider default reasoning
Measured
29 Sep 2026
Not shown
Not real ad performance: the market is a deterministic simulation of one invented skincare brand.

Reported by others

-0.04
Confirmed task success · rank 24 of 46
Unit
IPS effect estimate, higher is better
Range
-0.06 to -0.01
Sample
25002 observations
Configuration
DeepSeek V4 Pro 0423
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.02
Praise over complaint · rank 17 of 46
Unit
IPS effect estimate, higher is better
Range
-0.05 to 0.02
Sample
9770 observations
Configuration
DeepSeek V4 Pro 0423
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.03
Steerability · rank 22 of 46
Unit
IPS effect estimate, higher is better
Range
-0.05 to -0.01
Sample
36510 observations
Configuration
DeepSeek V4 Pro 0423
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
-0.00
Tool grounding · rank 30 of 46
Unit
IPS effect estimate, higher is better
Range
-0.00 to -0.00
Sample
1800929 observations
Configuration
DeepSeek V4 Pro 0423
Measured
28 Sep 2026
Not shown
Not a success rate: the estimated effect of choosing this model as the agent's orchestrator, against an even mix of the models tested, inside Arena's own harness.
1,455
Overall · rank 43 of 177
Unit
Arena rating, higher is better
Range
1,452 to 1,458
Sample
57523 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,451
Overall · rank 28 of 408
Unit
Arena rating, higher is better
Range
1,445 to 1,457
Sample
14117 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,460
Business, management and finance · rank 31 of 402
Unit
Arena rating, higher is better
Range
1,453 to 1,466
Sample
11411 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,446
Creative writing · rank 24 of 407
Unit
Arena rating, higher is better
Range
1,439 to 1,453
Sample
9697 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,480
Expert prompts · rank 37 of 359
Unit
Arena rating, higher is better
Range
1,471 to 1,488
Sample
5925 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,452
Instruction following · rank 40 of 409
Unit
, lower is better
Range
1,447 to 1,458
Sample
20066 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
None
1,458
Overall · rank 48 of 409
Unit
Arena rating, higher is better
Range
1,454 to 1,461
Sample
57564 votes
Configuration
DeepSeek V4 Pro 0423
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
25.3%
Consistency (pass^5) · rank 5 of 5
Unit
% of tasks, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
45.6%
Customer support task success (pass@1) · rank 4 of 5
Unit
% of tasks, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
48.4%
Shopping assistant task success (pass@1) · rank 4 of 5
Unit
% of tasks, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
47.2%
Task success (pass@1) · rank 4 of 5
Unit
% of tasks, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
47.6%
Travel task success (pass@1) · rank 4 of 5
Unit
% of tasks, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not the model alone: it runs inside STATE-Bench's agent loop against simulated users, and versions of the benchmark are not comparable.
3.36
User experience · rank 5 of 5
Unit
score from 1 to 5, higher is better
Configuration
deepseek-v4-pro
Measured
25 May 2026
Not shown
Not real customer satisfaction: the ratings come from a model-based judge.
Reported by OpenHands Index
40.7%
Average · rank 29 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
36.8%
Front end (SWE-Bench Multimodal) · rank 11 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
12.5%
Greenfield (Commit0) · rank 25 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
12.7%
Information gathering (GAIA) · rank 33 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
73.2%
Issue resolution (SWE-Bench) · rank 22 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
68.1%
Testing (SWT-Bench) · rank 18 of 34
Unit
% resolved, higher is better
Configuration
deepseek-v4-pro
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by UGI Leaderboard
42.0%
Requested-length error · rank 295 of 370
Unit
% off the requested word count, lower is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
34.0%
Requested-length error · rank 269 of 370
Unit
% off the requested word count, lower is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.35
Style adherence · rank 173 of 370
Unit
score from 0 to 1, higher is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.36
Style adherence · rank 140 of 370
Unit
score from 0 to 1, higher is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
60.2
Writing score · rank 98 of 370
Unit
score out of 100, higher is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
68.4
Writing score · rank 36 of 370
Unit
score out of 100, higher is better
Configuration
deepseek-v4-pro
Measured
26 Apr 2026
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
97.2%
Answer rate · rank 87 of 108
Unit
% of documents, higher is better
Configuration
DeepSeek V4 Pro 0423
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
8.6%
Hallucination rate · rank 44 of 108
Unit
% of summaries, lower is better
Configuration
DeepSeek V4 Pro 0423
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare DeepSeek V4 Pro 0423 with