Mistral Large 4 Is Fast, Cheap and Europe's Best Model. On Business Work, It Still Makes Things Up.

Mistral released Mistral Large 4 on 6 October. Mistral calls it "le Chonk": a mixture-of-experts model with about a trillion parameters, trained in its own European data centres, with open weights promised by the end of October.
It is a big step up for Mistral. Artificial Analysis now ranks it the most intelligent model from a developer outside the US and China. It is also quick and, during the launch discount, inexpensive.
We ran it through our four benchmarks the same day. Our verdict: it is a good choice for EU-hosted, cost-sensitive work on structured data and tools, but not for anything a customer will read unless the facts are checked first.
Mistral Large 4 at a glance
- Use it where an EU developer and EU hosting matter, and for extraction, tool calls and long documents, where it is accurate, fast and cheap.
- Don't use it unchecked for product copy, sales briefs or decks. At its default setting, 64% of its product listings and 97% of its sales briefs contained claims the source data did not support.
- Against models of similar price, intelligence and speed, it underperforms. Gemini 3.8 Flash and GPT-6 Luna beat it on listings, decks and planning.
- Launch-day reaction was mostly warm, with a loud critical minority: 54% of the posts and comments that took a side were positive and 33% negative. The praise was for its speed and for Europe having a contender; the criticism was that it trails Chinese open models on price and performance.
- Test both reasoning settings. High reasoning halved the invented claims in listings, but the same setting cut its ad-budget planning score from 37.0 to 27.4.
- Price: $0.68 in and $2.09 out per million tokens during a launch discount that Artificial Analysis says lasts two weeks. The list price is $1.36 and $4.18, which is still cheaper than Mistral Medium 3.5.
Specs, price and what Mistral claims
According to Mistral's announcement and documentation:
- Architecture: about 1 trillion total parameters, with 49 billion active per token (52 billion counting embeddings and output layers). It has a 1.6-billion-parameter vision encoder.
- Inputs: text and images in, text out. Mistral's documentation gives a 1-million-token context window. OpenRouter and Artificial Analysis list 524,288 tokens.
- Reasoning: two settings on Mistral's API, none and high.
- Weights: promised by the end of October. The licence has not been published yet.
- A preview: Guillaume Lample says the reinforcement-learning run behind it is "still in flight", and a final version will come with the weights.
- Claimed strengths: coding, cyber-security, finance and vision. Mistral's headline numbers include 61.7% on DeepSWE v1.1, 59.9% on AutomationBench and 93% on Cybench.
VentureBeat points out that the competitive picture is "more complicated than Mistral's charts suggest". On the live DeepSWE leaderboard, other models score higher than Mistral's chart shows, under different configurations.
Mistral Large 4 benchmarks, measured by Spring Prompt
A rank among every model only tells you so much. What matters is how it does against the models you would actually weigh it against. So, as on our model pages, we rank it only among models available today (those you can call through OpenRouter), and within the three groups it belongs to:
- Same price: $1–3 per million tokens, blending three input tokens to one output. Mistral Large 4 costs $1.03 at the launch discount and $2.07 at list price, so it stays in this band.
- Similar intelligence: within 4 points of its 38.4 on the Artificial Analysis Intelligence Index.
- Same speed: 100–200 tokens a second. It writes 106.
How to read this table: each row is one of our benchmarks. Reading across, it shows where Mistral Large 4 ranks among every model available today, then among only the models at its price, its intelligence and its speed. Under each rank is the best model in that group and its score, so you can see how far behind it is.
| Benchmarkwhat it measures | Mistral Large 4its score | Rank among all modelsavailable today | Rank at its price$1–3 per million tokens | Rank at its intelligencewithin 4 points on the AA index | Rank at its speed100–200 tokens a second |
|---|---|---|---|---|---|
| CatalogBenchshare of product listings reliably ready to publish | 10.7%at high reasoning; 7.1% at default | 17th of 19Best: GPT-6 Astra and GPT-6.1 Sol, 75.0% | 3rd of 5Best: Muse Spark 1.3, 48.2% | 4th of 4 (last)Best: GPT-6 Luna, 55.4% | 7th of 7 (last)Best: GPT-6 Luna, 55.4% |
| CatalogBench sales briefshare of sales briefs reliably ready to publish | 1.8%at high reasoning; 0.0% at default | 15th of 19Best: GPT-6 Astra, 69.6% | 3rd of 5Best: Muse Spark 1.3, 10.7% | 4th of 4 (last)Best: GPT-6 Luna, 51.8% | 7th of 7 (last)Best: GPT-6 Luna, 51.8% |
| DeckBenchdeck rating from head-to-head comparisons | 515default setting (the only one run) | 18th of 18 (last)Best: GPT-6 Astra, 1,489 | 4th of 4 (last)Best: Gemini 3.8 Flash, 904 | 4th of 4 (last)Best: GPT-6 Luna, 1,273 | 8th of 8 (last)Best: Claude Sonnet 5.5, 1,356 |
| ROASBenchscore out of 100 for a year of ad-spend planning | 37.0at default; 27.4 at high reasoning | 10th of 18Best: GPT-6 Astra, 56.6 | 2nd of 5Best: Gemini 3.8 Flash, 51.4 | 3rd of 4Best: GPT-6 Luna, 53.4 | 5th of 8Best: GPT-6 Luna, 53.4 |
| BulletBench Lightningchess rating with 10 seconds plus 1 a move | 249 Elodefault setting (the only one run) | 9th of 14Best: Gemini 3.5 Flash-Lite, 892 | 2nd of 5Best: GPT-5.4 Mini, 571 | Not rankedonly 2 of these models measured | 3rd of 5Best: Mistral Small 4, 571 |
Each model is ranked at its better reasoning setting. A group only counts models that benchmark has measured, and a group of fewer than three is not ranked. "Best" is the top model in that group, not necessarily one you would choose instead: see the alternatives.
At its price, Mistral Large 4 is mid-table: second of five at ad planning and at the 10-second chess clock, third of five on listings, and last of four on decks. Against models of similar intelligence it does worse, finishing last on listings and decks. On business work, it does less well than its Intelligence Index score suggests.
Two models keep coming out ahead of it. Gemini 3.8 Flash falls in all three of its groups ($1.50 per million tokens, 40.9 on the Intelligence Index, 192 tokens a second) and beats it on everything except the 10-second chess clock. GPT-6 Luna has similar intelligence and speed at $0.20, and beats it on all four benchmarks both have run.
Product listings: accurate on fields, loose with claims
CatalogBench gives each model a sparse product feed and product photos, and asks for listings that could go live. Mistral Large 4 does the structured part well. At its default setting it filled 94.4% of missing fields correctly (3rd of 20), caught every conflict between feed and image, never failed to return an output, and cost $0.0033 per product (3rd cheapest).
The trouble is what it adds. 64.3% of its listings contained at least one claim the source data did not support. Only 7.1% were reliably ready to publish, against 75.0% for the best models. On the harder sales-brief variant, 97.0% of listings had unsupported claims, an average of 6.7 per product. That is the most of any model we have run.
High reasoning helps here. Unsupported claims fall to 33.3% of listings and 0.91 per product. Reliably publish-ready listings rise to 10.7%. The cost is more failures: 7.1% of products got no usable output, against none at the default setting. Cost per product rises to $0.0125.
Decks: last place
On DeckBench, where models turn a finished analysis into a presentable deck, Mistral Large 4 finished 18th of 18. It won 10.7% of its head-to-head comparisons. Its content wasn't the main problem: it never left out a finding and never used a misleading metric. But every deck needed slide work, two-thirds had layout defects, and a third quoted numbers the analysis didn't contain.
Ad-spend planning: mid-table, and worse with more thinking
On ROASBench, a 12-month simulated ad-budget game, Mistral Large 4 scored 37.0 at its default setting. That puts it 10th of 19 and well ahead of Mistral Medium 3.5 (15.2). It made a simulated $533,401 in contribution profit, but went over budget in three months.
At high reasoning it never went over budget, but its overall score fell to 27.4 and profit fell to $156,831. It played too safe. Simon Willison found the opposite on his SVG test, where high reasoning gave a better drawing with fewer tokens. The reasoning setting matters, and which one is better depends on the task.
Speed: good with seconds to spare, poor under one
Mistral Large 4 makes a typical move in 0.8 seconds. With three minutes per game (Blitz 3+2) it lost no games on time and reached 555 ladder Elo, level with Mistral Medium 3.5. At the 10-second Lightning clock, 58.3% of its games ended on time and its rating fell to 249. If you need answers in well under a second every time, look at Gemini 3.5 Flash-Lite.
Third-party benchmarks
- Artificial Analysis (reported 7 October): Intelligence Index 38.4, against 14 for Medium 3.5 and 9 for Large 3. It reaches 81.3% on long-context reasoning (AA-LCR), 26.8% on Terminal-Bench 4.0 and 35.0% on Humanity's Last Exam. Output speed is 106 tokens a second. It is verbose: Artificial Analysis says it used 200 million tokens to run the index, against a median of 81 million.
- Pending: Arena, tau2-bench, BFCL, OpenHands, Vending-Bench 2, the Vectara Hallucination Leaderboard, SimpleQA Verified and the professional-work sources (APEX-Agents, GDP.pdf, Remote Labor Index).
What people are saying about Mistral Large 4
Reactions from launch day, on X, blogs and Hacker News. Quotes are as posted; each card links to the original.
How launch day read: 54% positive, 12% mixed, 33% negative of the 114 posts and comments that took a side
NegativeMixedPositive
On 7 October we read the top posts on X for "Mistral Large 4" and "le Chonk", and every top-level comment on the Hacker News launch thread. We labelled each one positive, mixed, negative or neutral. News, jokes, questions, and posts from Mistral and its partners count as neutral and are left out of the bars. Labels are by Claude Opus 5.5 against a written rubric, and we checked a sample. X ranks top posts by engagement, so its sample leans to the loudest voices. Every labelled post.
A selection, with how we read each one:
Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China
It is at the frontier of open models, and by far the strongest open-weight model from the US or Europe. The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.
It's certainly not a Fable-class model, but it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier.
Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them. We'll see what the kitten can pull off, but the disparity is pretty stark
Mistral Large 4 beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security tasks. Meanwhile Opus and Astra had ~40% of tasks blocked by their own safety filters.
Mistral Large 4 just lost to DeepSeek V4.1 Flash. Same sunset ocean prompt. Mistral rendered a blank screen. DeepSeek nailed it for 3 cents. […] Europe's flagship model can't beat a Chinese flash model.
@MistralAI Large 4, Le Chonk, looks like a great model for data analytics. We just ran it through our benchmark. It's 10x cheaper ($20 to $2) and improves its grade by two letters. At this rate, it'll ace the exam next year.
Worth a benchmark run but not yet worth a sovereignty promise.
Fans praised its speed, the name and having a strong European open-weights model at all. Critics said it still trails the Chinese open models, questioned the comparisons Mistral chose, and some reported weak first tests.
Mistral Large 4 vs Gemini 3.8 Flash, GPT-6 Luna and DeepSeek V4.1 Flash
| Model | Why compare | Price* | Tokens a second | AA Intelligence | Spring Prompt overall |
|---|---|---|---|---|---|
| Mistral Large 4 | — | $1.03 | 106 | 38.4 | 31 |
| Gemini 3.8 Flash compare | In all three of its groups | $1.50 | 192 | 40.9 | 61 |
| GPT-6 Luna compare | Similar intelligence and speed | $0.20 | 134 | 38.1 | 41 |
| Muse Spark 1.3 compare | Most intelligent at $1–3 | $2.00 | 137 | 48.1 | 66 |
| DeepSeek V4.1 Flash compare | Cheapest with similar intelligence | $0.11 | 213 | 39.5 | 58 |
| Mistral Medium 3.5 compare | The other Mistral you might use | $3.00 | 163 | 14.2 | 21 |
*OpenRouter list price per million tokens, blended at three input tokens to one output, on 7 October. Mistral Large 4's price includes the launch discount. Speed and intelligence are from Artificial Analysis.
If EU hosting isn't a requirement, there is a better choice at every price point. Gemini 3.8 Flash costs about half as much again, is faster and beats it on listings, decks and planning. Once the launch discount ends, Gemini 3.8 Flash will also be the cheaper of the two. GPT-6 Luna matches its intelligence for a fifth of the price, and DeepSeek V4.1 Flash for a tenth. Both score higher on our benchmarks. Muse Spark 1.3 is markedly stronger for about twice the price.
Against Mistral Medium 3.5, Large 4 is cheaper and much better at planning. Medium 3.5 still writes better decks (649 against 515) and more publish-ready listings (17.9% against 10.7%), so don't assume Large 4 replaces it everywhere.
Coming soon. Going by each line's average gap between launches on our trends page, new GPT, GLM, DeepSeek and Muse Spark models are due around October, and Kimi and Qwen Plus are overdue. That is arithmetic, not inside knowledge. Separately, Google has said the full Gemini 4 release will come well before the end of the year. In this price band, the picture may change within weeks.
Is Mistral Large 4 good? Who should use it, and for what
- Teams that need an EU developer and EU hosting. It is now clearly the strongest option from an EU developer. Mistral says the model is served end to end in its EU region. Once the weights are out, you can also host it yourself.
- Extraction, classification and tool calls. It is accurate on structured fields, never failed to return output at its default setting, and ranks in the top quarter for agents and tool use among models available today.
- Long documents. 81.3% on AA-LCR, with a context window of at least 512k tokens.
- Not for customer-facing copy without a fact check. If it writes product listings, sales briefs or decks, check every claim against the source, or use a model that invents less.
- Not where every answer must land in about a second.
As always, these benchmarks aren't your data, so test it on your own prompts before switching. If it would write your catalogue, our catalogue feed diagnostic measures how often it invents attributes on your real feed.
What we'll update
We'll publish a week-one verdict once Arena, tau2-bench and the hallucination leaderboards have results. When the final version arrives with the open weights and licence, we'll run it through our benchmarks again. Until then, the Mistral Large 4 model page has every result as it lands.
Quick answers
How much does Mistral Large 4 cost?
Mistral's list price is $1.36 per million input tokens and $4.18 per million output tokens. During the launch discount, which Artificial Analysis says lasts two weeks from 6 October 2026, it is $0.68 and $2.09.
What is Mistral Large 4's context window?
Mistral's documentation gives 1 million tokens. OpenRouter and Artificial Analysis list 524,288 tokens.
Is Mistral Large 4 open source?
Not yet. It launched on 6 October 2026 as a preview on Mistral's API. Mistral has promised the weights, with a final version of the model, by the end of October 2026, and has not yet published the licence.
How good is Mistral Large 4?
Artificial Analysis gives it 38.4 on its Intelligence Index, the highest of any model from a developer outside the US and China. On Spring Prompt's business benchmarks it is mid-table among models priced $1–3 per million tokens, and behind most models of similar intelligence.
Does Mistral Large 4 make things up?
On CatalogBench at its default setting, 64% of its product listings contained at least one claim the product data did not support, and 97% of its sales briefs did. High reasoning cut the listings figure to 33%.
Mistral Large 4 or Gemini 3.8 Flash?
Gemini 3.8 Flash costs about the same, scores slightly higher on the Intelligence Index (40.9 against 38.4), writes faster (192 against 106 tokens a second) and beat Mistral Large 4 on our listings, decks and ad-planning benchmarks. Choose Mistral Large 4 if you need an EU developer and EU hosting.
Model pages: Mistral Large 4 · Gemini 3.8 Flash · GPT-6 Luna · DeepSeek V4.1 Flash · Muse Spark 1.3 · Mistral Medium 3.5