Models / DeepSeek V3.2 Exp

DeepSeek

DeepSeek V3.2 Exp

38 published results from 4 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
DeepSeek
Sources
4
Our benchmarks
0
Price
Not yet published

Reported by others

1,436
Overall · rank 74 of 177
Unit
Arena rating, higher is better
Range
1,430 to 1,442
Sample
11853 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,446
Overall · rank 44 of 177
Unit
Arena rating, higher is better
Range
1,435 to 1,457
Sample
2514 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,422
Business, management and finance · rank 86 of 402
Unit
Arena rating, higher is better
Range
1,409 to 1,434
Sample
2246 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,406
Business, management and finance · rank 107 of 402
Unit
Arena rating, higher is better
Range
1,392 to 1,420
Sample
1777 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,409
Creative writing · rank 66 of 407
Unit
Arena rating, higher is better
Range
1,394 to 1,423
Sample
1648 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,391
Creative writing · rank 79 of 407
Unit
Arena rating, higher is better
Range
1,374 to 1,408
Sample
1186 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,425
Expert prompts · rank 98 of 359
Unit
Arena rating, higher is better
Range
1,402 to 1,447
Sample
661 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,449
Expert prompts · rank 53 of 359
Unit
Arena rating, higher is better
Range
1,420 to 1,478
Sample
417 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,414
Instruction following · rank 87 of 409
Unit
Arena rating, higher is better
Range
1,404 to 1,424
Sample
3366 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,416
Instruction following · rank 86 of 409
Unit
Arena rating, higher is better
Range
1,404 to 1,427
Sample
2547 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,422
Overall · rank 105 of 409
Unit
Arena rating, higher is better
Range
1,416 to 1,428
Sample
11947 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,425
Overall · rank 99 of 409
Unit
Arena rating, higher is better
Range
1,418 to 1,432
Sample
8942 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,410
Writing, literature and language · rank 77 of 408
Unit
Arena rating, higher is better
Range
1,399 to 1,422
Sample
2646 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,398
Writing, literature and language · rank 91 of 408
Unit
Arena rating, higher is better
Range
1,385 to 1,411
Sample
2087 votes
Configuration
DeepSeek V3.2 Exp
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
93.2%
Irrelevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
67.0%
Irrelevance detection · rank 86 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not general refusal or safety behaviour.
54.2%
Memory · rank 8 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
44.1%
Memory · rank 16 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not long-term personal memory in a product: sessions are BFCL's scripted ones.
37.4%
Multi-turn tasks · rank 31 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
44.9%
Multi-turn tasks · rank 22 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not open-ended agent work: the tools and tasks are BFCL's simulated APIs.
54.1%
Overall accuracy · rank 19 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
56.7%
Overall accuracy · rank 14 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not a neutral average: the weighting is BFCL's. Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
37.5%
Relevance detection · rank 103 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
93.8%
Relevance detection · rank 8 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not whether the call itself was right; only that one was attempted.
34.9%
Single-turn calls (curated) · rank 105 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
85.5%
Single-turn calls (curated) · rank 39 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
53.7%
Single-turn calls (user-contributed) · rank 97 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
76.0%
Single-turn calls (user-contributed) · rank 35 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not reliability on your own tools: BFCL's functions and queries are a fixed test set, and a correct call is judged by its form, not by what it achieved.
69.5%
Web search · rank 16 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
58.0%
Web search · rank 22 of 109
Unit
% correct, higher is better
Configuration
deepseek-v3.2-exp
Measured
16 Dec 2025
Not shown
Not general research quality: questions have short, checkable answers.
Reported by UGI Leaderboard
27.0%
Requested-length error · rank 239 of 370
Unit
% off the requested word count, lower is better
Configuration
deepseek-v3.2-exp
Measured
2 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
25.0%
Requested-length error · rank 228 of 370
Unit
% off the requested word count, lower is better
Configuration
DeepSeek V3.2 Exp
Measured
2 Oct 2025
Not shown
Not other format limits such as character counts or bullet counts.
Reported by UGI Leaderboard
0.35
Style adherence · rank 196 of 370
Unit
score from 0 to 1, higher is better
Configuration
deepseek-v3.2-exp
Measured
2 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
0.35
Style adherence · rank 196 of 370
Unit
score from 0 to 1, higher is better
Configuration
DeepSeek V3.2 Exp
Measured
2 Oct 2025
Not shown
Not brand voice on your own examples: UGI's prompts are private and lean towards creative writing.
Reported by UGI Leaderboard
51.0
Writing score · rank 154 of 370
Unit
score out of 100, higher is better
Configuration
deepseek-v3.2-exp
Measured
2 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
Reported by UGI Leaderboard
54.0
Writing score · rank 133 of 370
Unit
score out of 100, higher is better
Configuration
DeepSeek V3.2 Exp
Measured
2 Oct 2025
Not shown
Not business copy quality: UGI's prompts are private and lean towards creative writing, and models that often refuse get no score.
96.6%
Answer rate · rank 89 of 108
Unit
% of documents, higher is better
Configuration
DeepSeek V3.2 Exp
Measured
22 Sep 2026
Not shown
Not a quality score: a low rate usually means content filters were triggered, and hallucination rates are measured on answered documents only.
5.3%
Hallucination rate · rank 15 of 108
Unit
% of summaries, lower is better
Configuration
DeepSeek V3.2 Exp
Measured
22 Sep 2026
Not shown
Not errors in open questions or other tasks: only summarisation, judged by Vectara's own model (HHEM-2.3), not by people, on news-style documents rather than your data.

Compare DeepSeek V3.2 Exp with