Models / DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp
Fastest of the 22 models from developers based in China, and 2nd most intelligent of the 10 models writing 200 or more tokens a second. 4th most intelligent of the 68 models priced $0.25–1 per million tokens. Among models available today, in the top quarter for reasoning and knowledge, and agents and tool use.
- Price per million tokens
- $0.44 in · $1.32 out
Typical host on OpenRouter, 7 Oct 2026 · compare prices - Speed
- 217 tokens a second
Artificial Analysis, on its usual API - Intelligence Index
- 34.8
Artificial Analysis - Spring Prompt overall
- Not ranked yet
needs our benchmarks and two groups of results - Developer
- DeepSeek, based in China
Weights published - Released
- Not recorded
1,048,576 tokens of context
Where it stands among the models you could choose
Ranked only among models available today (those you can call through OpenRouter), in the groups people choose between: the same price band, similar intelligence, the same speed, developers based in the same place, open weights. Retired models are left out.
- Fastest of the 22 models from developers based in China
- 2nd most intelligent of the 10 models writing 200 or more tokens a second
- Fastest of the 7 models with similar intelligence
- 3rd fastest of the 21 models priced $0.25–1 per million tokens
| Among | Intelligence Index | Spring Prompt overall | Price | Speed |
|---|---|---|---|---|
| Models available today | 35th of 195Claude Opus 5.5 | – | 100th of 221Mistral Nemo | 7th of 63Trinity Large Thinking |
| Models priced $0.25–1 per million tokens | 4th of 68MiMo-V2.6-Pro | – | the group | 3rd of 21Trinity Large Thinking |
| Models with similar intelligence | the group | – | 5th of 17DeepSeek V4 Flash 0423 | 1st of 7leads |
| Models writing 200 or more tokens a second | 2nd of 10DeepSeek V4.1 Flash | – | 8th of 10gpt-oss-20b | the group |
| Models from developers based in China | 10th of 84MiMo-V2.6-Pro | – | 53rd of 96Qwen3.5-9B | 1st of 22leads |
| Open-weights modelsincluding 1 with weights announced but not yet published | 10th of 108MiMo-V2.6-Pro | – | 78th of 117Mistral Nemo | 5th of 41Trinity Large Thinking |
The name under each rank is the leader of that group. Price is the developer's list price (or the typical OpenRouter host where we haven't read one) for three input tokens to one output; intelligence and speed are from Artificial Analysis (data sourced from Artificial Analysis); "similar intelligence" means within 4 points on its Intelligence Index. Developer locations are where each company is based, not where a model is served. Groups under 3 models are not ranked.
How good is it, and for what?
Each use case ranks the available models its benchmarks measured, then those priced $0.25–1 per million tokens. The bar shows where it falls in that field, best to the right. Open a row for the results behind it.
Product listingsTurning a sparse product feed and photos into listings that can go live
- CatalogBench: has not measured this model.
Decks from an analysisTurning a finished analysis into a deck you could present as it is
- DeckBench: has not measured this model.
Marketing planningPlanning a year of ad spend without overspending
- ROASBench: has not measured this model.
Agents and tool useMulti-step tasks with tools: support desks, coding agents, function calls
| Benchmark | DeepSeek V4 Flash Vision Exp | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Terminal-Bench 4.0 · reported | 12.1% | 29th of 102 | Claude Sonnet 5.5 63.6% |
| Artificial AnalysisArtificial Analysis: Terminal-Bench 2.1 · reported | 74.2% | 32nd of 113 | Claude Fable 5.1 91.4% |
| Artificial AnalysisArtificial Analysis: τ-bench banking · reported | 41.0% | 15th of 110 | Grok 4.6 50.7% |
- tau2-bench: has not measured this model.
- Berkeley Function Calling Leaderboard (BFCL) V4: has not measured this model.
- OpenHands Index: has not measured this model.
- Microsoft STATE-Bench: has not measured this model.
- Vending-Bench 2: has not measured this model.
Professional workReal tasks from banking, consulting and law, business documents and freelance projects
- APEX-Agents: has not measured this model.
- GDP.pdf: has not measured this model.
- Remote Labor Index: has not measured this model.
Reasoning and knowledgeHard questions across science, maths and general knowledge
| Benchmark | DeepSeek V4 Flash Vision Exp | Rank among available | Best available |
|---|---|---|---|
| Artificial AnalysisArtificial Analysis: Artificial Analysis Intelligence Index · reported | 34.8 | 35th of 195 | Claude Opus 5.5 57.6 |
| Artificial AnalysisArtificial Analysis: Humanity's Last Exam · reported | 34.5% | 53rd of 194 | Claude Opus 5.5 61.4% |
| Artificial AnalysisArtificial Analysis: GPQA Diamond · reported | 91.3% | 27th of 185 | GPT-6 Astra 96.3% |
WritingWhat people prefer in blind comparisons, and judged writing quality
- Arena (formerly LMArena): has not measured this model.
- UGI Leaderboard: has not measured this model.
Sticking to the factsSummarising without inventing things, and factual answers
- Vectara Hallucination Leaderboard: has not measured this model.
- Arena (formerly LMArena): has not measured this model.
- SimpleQA Verified (Epoch AI): has not measured this model.
Speed under pressureGood decisions against a real clock (fast chess)
- BulletBench: has not measured this model.
Against the alternatives
The models you would most likely weigh it against: the leaders of the groups above.
| DeepSeek V4 Flash Vision Exp | MiMo-V2.6-Pro most intelligent at $0.25–1 per million tokens | DeepSeek V4 Flash 0423 cheapest with similar intelligence | Claude Opus 5.5 most intelligent available today | |
|---|---|---|---|---|
| Price per million tokens | $0.66 | $0.54 | $0.18 | $8.00 |
| Tokens a second | 217 | 40 | – | 97 |
| Intelligence Index | 34.8 | 46.3 | 34.3 | 57.6 |
| Spring Prompt overall | – | – | – | 82 |
| Rank among available models, by use case | ||||
| Product listings | – | – | – | 6th |
| Decks from an analysis | – | – | – | 5th |
| Marketing planning | – | – | – | 5th |
| Agents and tool use | 35th of 161 | 15th | 34th | 2nd |
| Professional work | – | – | – | 3rd |
| Reasoning and knowledge | 41st of 195 | 10th | 40th | 1st |
| Writing | – | 18th | 62nd | 4th |
| Sticking to the facts | – | 9th | 81st | 6th |
| Speed under pressure | – | – | – | – |
Shaded figures are better than DeepSeek V4 Flash Vision Exp's. A dash means no result.
Where it has been measured
- Artificial Analysis11 results
- BulletBenchNot measured
- CatalogBenchNot measured
- DeckBenchNot measured
- ROASBenchNot measured
- APEX-AgentsNot measured
- Arena (formerly LMArena)Not measured
- Berkeley Function Calling Leaderboard (BFCL) V4Not measured
- GDP.pdfNot measured
- Microsoft STATE-BenchNot measured
- OpenHands IndexNot measured
- Remote Labor IndexNot measured
- SimpleQA Verified (Epoch AI)Not measured
- tau2-benchNot measured
- UGI LeaderboardNot measured
- Vectara Hallucination LeaderboardNot measured
- Vending-Bench 2Not measured
Every result
11 published results from 1 source, each in the source's own units, with the configuration that produced it.
Artificial Analysis · reported by Artificial Analysis · 11 results
| Measure | Value | Rank | Configuration | Dated |
|---|---|---|---|---|
| Artificial Analysis: Artificial Analysis Coding Indexindex score, higher is better | 65.0 | 67 of 188 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Artificial Analysis Intelligence Indexindex score, higher is better | 34.8 | 79 of 402 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: GPQA Diamond% of questions, higher is better | 91.3% | 49 of 359 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Humanity's Last Exam% of questions, higher is better | 34.5% | 112 of 400 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Long-context reasoning (AA-LCR)% of questions, higher is better | 81.3% | 47 of 388 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Output speedtokens per second, higher is better | 217 | 8 of 128 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: SciCode% of problems, higher is better | 49.7% | 104 of 186 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Terminal-Bench 2.1% of tasks, higher is better | 74.2% | 65 of 187 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Terminal-Bench 4.0% of tasks, higher is better | 12.1% | 73 of 182 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: Time to first answer tokenseconds, lower is better | 10.0 | 59 of 128 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
| Artificial Analysis: τ-bench banking% of tasks, higher is better | 41.0% | 28 of 175 | DeepSeek V4 Flash Vision (experimental, max) | 7 Oct 2026 |
Not shown: Not your codebase or tools.