Research

Mistral Large 4 Is Fast, Cheap and Europe's Best Model. On Business Work, It Still Makes Things Up.

CatalogBench ranking of reliably publish-ready listings, with Mistral Large 4 at 10.7% (high reasoning) and 7.1% (default), against 75.0% for the leaders.
CatalogBench, 6 October 2026: share of listings reliably ready to publish. Mistral Large 4 is 17th and 18th of 20 configurations.

Mistral released Mistral Large 4 on 6 October. Mistral calls it "le Chonk": a mixture-of-experts model with about a trillion parameters, trained in its own European data centres, with open weights promised by the end of October.

It is a big step up for Mistral. Artificial Analysis now ranks it the most intelligent model from a developer outside the US and China. It is also quick and, during the launch discount, inexpensive.

We ran it through our four benchmarks the same day. Our verdict: it is a good choice for EU-hosted, cost-sensitive work on structured data and tools, but not for anything a customer will read unless the facts are checked first.

Mistral Large 4 at a glance

  • Use it where an EU developer and EU hosting matter, and for extraction, tool calls and long documents, where it is accurate, fast and cheap.
  • Don't use it unchecked for product copy, sales briefs or decks. At its default setting, 64% of its product listings and 97% of its sales briefs contained claims the source data did not support.
  • Against models of similar price, intelligence and speed, it underperforms. Gemini 3.8 Flash and GPT-6 Luna beat it on listings, decks and planning.
  • Launch-day reaction was mostly warm, with a loud critical minority: 54% of the posts and comments that took a side were positive and 33% negative. The praise was for its speed and for Europe having a contender; the criticism was that it trails Chinese open models on price and performance.
  • Test both reasoning settings. High reasoning halved the invented claims in listings, but the same setting cut its ad-budget planning score from 37.0 to 27.4.
  • Price: $0.68 in and $2.09 out per million tokens during a launch discount that Artificial Analysis says lasts two weeks. The list price is $1.36 and $4.18, which is still cheaper than Mistral Medium 3.5.
Data note: Our results are from runs on 6 October 2026 at Mistral's default reasoning setting and, where stated, at high reasoning, using Mistral's API. The results table ranks models available today, each at its best setting; ranks in the text compare every model configuration we have run. CatalogBench uses invented products, DeckBench uses invented companies, ROASBench is a simulation of one skincare brand, and BulletBench is fast chess, so none of them predicts results on your own data. Arena, tau2-bench, BFCL, the Vectara Hallucination Leaderboard, SimpleQA Verified and several other sources have not published results for this model yet. The model page updates as they arrive.

Specs, price and what Mistral claims

According to Mistral's announcement and documentation:

  • Architecture: about 1 trillion total parameters, with 49 billion active per token (52 billion counting embeddings and output layers). It has a 1.6-billion-parameter vision encoder.
  • Inputs: text and images in, text out. Mistral's documentation gives a 1-million-token context window. OpenRouter and Artificial Analysis list 524,288 tokens.
  • Reasoning: two settings on Mistral's API, none and high.
  • Weights: promised by the end of October. The licence has not been published yet.
  • A preview: Guillaume Lample says the reinforcement-learning run behind it is "still in flight", and a final version will come with the weights.
  • Claimed strengths: coding, cyber-security, finance and vision. Mistral's headline numbers include 61.7% on DeepSWE v1.1, 59.9% on AutomationBench and 93% on Cybench.

VentureBeat points out that the competitive picture is "more complicated than Mistral's charts suggest". On the live DeepSWE leaderboard, other models score higher than Mistral's chart shows, under different configurations.

Mistral Large 4 benchmarks, measured by Spring Prompt

A rank among every model only tells you so much. What matters is how it does against the models you would actually weigh it against. So, as on our model pages, we rank it only among models available today (those you can call through OpenRouter), and within the three groups it belongs to:

  • Same price: $1–3 per million tokens, blending three input tokens to one output. Mistral Large 4 costs $1.03 at the launch discount and $2.07 at list price, so it stays in this band.
  • Similar intelligence: within 4 points of its 38.4 on the Artificial Analysis Intelligence Index.
  • Same speed: 100–200 tokens a second. It writes 106.

How to read this table: each row is one of our benchmarks. Reading across, it shows where Mistral Large 4 ranks among every model available today, then among only the models at its price, its intelligence and its speed. Under each rank is the best model in that group and its score, so you can see how far behind it is.

Benchmarkwhat it measuresMistral Large 4its scoreRank among all modelsavailable todayRank at its price$1–3 per million tokensRank at its intelligencewithin 4 points on the AA indexRank at its speed100–200 tokens a second
CatalogBenchshare of product listings reliably ready to publish10.7%at high reasoning; 7.1% at default17th of 19Best: GPT-6 Astra and GPT-6.1 Sol, 75.0%3rd of 5Best: Muse Spark 1.3, 48.2%4th of 4 (last)Best: GPT-6 Luna, 55.4%7th of 7 (last)Best: GPT-6 Luna, 55.4%
CatalogBench sales briefshare of sales briefs reliably ready to publish1.8%at high reasoning; 0.0% at default15th of 19Best: GPT-6 Astra, 69.6%3rd of 5Best: Muse Spark 1.3, 10.7%4th of 4 (last)Best: GPT-6 Luna, 51.8%7th of 7 (last)Best: GPT-6 Luna, 51.8%
DeckBenchdeck rating from head-to-head comparisons515default setting (the only one run)18th of 18 (last)Best: GPT-6 Astra, 1,4894th of 4 (last)Best: Gemini 3.8 Flash, 9044th of 4 (last)Best: GPT-6 Luna, 1,2738th of 8 (last)Best: Claude Sonnet 5.5, 1,356
ROASBenchscore out of 100 for a year of ad-spend planning37.0at default; 27.4 at high reasoning10th of 18Best: GPT-6 Astra, 56.62nd of 5Best: Gemini 3.8 Flash, 51.43rd of 4Best: GPT-6 Luna, 53.45th of 8Best: GPT-6 Luna, 53.4
BulletBench Lightningchess rating with 10 seconds plus 1 a move249 Elodefault setting (the only one run)9th of 14Best: Gemini 3.5 Flash-Lite, 8922nd of 5Best: GPT-5.4 Mini, 571Not rankedonly 2 of these models measured3rd of 5Best: Mistral Small 4, 571

Each model is ranked at its better reasoning setting. A group only counts models that benchmark has measured, and a group of fewer than three is not ranked. "Best" is the top model in that group, not necessarily one you would choose instead: see the alternatives.

At its price, Mistral Large 4 is mid-table: second of five at ad planning and at the 10-second chess clock, third of five on listings, and last of four on decks. Against models of similar intelligence it does worse, finishing last on listings and decks. On business work, it does less well than its Intelligence Index score suggests.

Two models keep coming out ahead of it. Gemini 3.8 Flash falls in all three of its groups ($1.50 per million tokens, 40.9 on the Intelligence Index, 192 tokens a second) and beats it on everything except the 10-second chess clock. GPT-6 Luna has similar intelligence and speed at $0.20, and beats it on all four benchmarks both have run.

Product listings: accurate on fields, loose with claims

CatalogBench gives each model a sparse product feed and product photos, and asks for listings that could go live. Mistral Large 4 does the structured part well. At its default setting it filled 94.4% of missing fields correctly (3rd of 20), caught every conflict between feed and image, never failed to return an output, and cost $0.0033 per product (3rd cheapest).

The trouble is what it adds. 64.3% of its listings contained at least one claim the source data did not support. Only 7.1% were reliably ready to publish, against 75.0% for the best models. On the harder sales-brief variant, 97.0% of listings had unsupported claims, an average of 6.7 per product. That is the most of any model we have run.

High reasoning helps here. Unsupported claims fall to 33.3% of listings and 0.91 per product. Reliably publish-ready listings rise to 10.7%. The cost is more failures: 7.1% of products got no usable output, against none at the default setting. Cost per product rises to $0.0125.

Decks: last place

On DeckBench, where models turn a finished analysis into a presentable deck, Mistral Large 4 finished 18th of 18. It won 10.7% of its head-to-head comparisons. Its content wasn't the main problem: it never left out a finding and never used a misleading metric. But every deck needed slide work, two-thirds had layout defects, and a third quoted numbers the analysis didn't contain.

Ad-spend planning: mid-table, and worse with more thinking

On ROASBench, a 12-month simulated ad-budget game, Mistral Large 4 scored 37.0 at its default setting. That puts it 10th of 19 and well ahead of Mistral Medium 3.5 (15.2). It made a simulated $533,401 in contribution profit, but went over budget in three months.

At high reasoning it never went over budget, but its overall score fell to 27.4 and profit fell to $156,831. It played too safe. Simon Willison found the opposite on his SVG test, where high reasoning gave a better drawing with fewer tokens. The reasoning setting matters, and which one is better depends on the task.

Speed: good with seconds to spare, poor under one

Mistral Large 4 makes a typical move in 0.8 seconds. With three minutes per game (Blitz 3+2) it lost no games on time and reached 555 ladder Elo, level with Mistral Medium 3.5. At the 10-second Lightning clock, 58.3% of its games ended on time and its rating fell to 249. If you need answers in well under a second every time, look at Gemini 3.5 Flash-Lite.

Third-party benchmarks

  • Artificial Analysis (reported 7 October): Intelligence Index 38.4, against 14 for Medium 3.5 and 9 for Large 3. It reaches 81.3% on long-context reasoning (AA-LCR), 26.8% on Terminal-Bench 4.0 and 35.0% on Humanity's Last Exam. Output speed is 106 tokens a second. It is verbose: Artificial Analysis says it used 200 million tokens to run the index, against a median of 81 million.
  • Pending: Arena, tau2-bench, BFCL, OpenHands, Vending-Bench 2, the Vectara Hallucination Leaderboard, SimpleQA Verified and the professional-work sources (APEX-Agents, GDP.pdf, Remote Labor Index).

What people are saying about Mistral Large 4

Reactions from launch day, on X, blogs and Hacker News. Quotes are as posted; each card links to the original.

How launch day read: 54% positive, 12% mixed, 33% negative of the 114 posts and comments that took a side

X42 of 93 top posts
Hacker News72 of 113 comments

NegativeMixedPositive

On 7 October we read the top posts on X for "Mistral Large 4" and "le Chonk", and every top-level comment on the Hacker News launch thread. We labelled each one positive, mixed, negative or neutral. News, jokes, questions, and posts from Mistral and its partners count as neutral and are left out of the bars. Labels are by Claude Opus 5.5 against a written rubric, and we checked a sample. X ranks top posts by engagement, so its sample leans to the loudest voices. Every labelled post.

A selection, with how we read each one:

Artificial Analysis@ArtificialAnlys

Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China

6 Oct 2026 · View on XNeutral
Guillaume Lample@GuillaumeLample · Mistral co-founder

It is at the frontier of open models, and by far the strongest open-weight model from the US or Europe. The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.

6 Oct 2026 · View on XFrom Mistral
Simon Willisonsimonwillison.net
Blog

It's certainly not a Fable-class model, but it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier.

6 Oct 2026 · Read the postMixed
Peter Gostev@petergostev · Arena

Mistral Large 4 was trained on 4,000 GPUs and Astra was trained on 100,000 of them. We'll see what the kitten can pull off, but the disparity is pretty stark

6 Oct 2026 · View on XNeutral
Cline@cline

Mistral Large 4 beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security tasks. Meanwhile Opus and Astra had ~40% of tasks blocked by their own safety filters.

6 Oct 2026 · View on XNeutral
Bridgebench@bridgebench

Mistral Large 4 just lost to DeepSeek V4.1 Flash. Same sunset ocean prompt. Mistral rendered a blank screen. DeepSeek nailed it for 3 cents. […] Europe's flagship model can't beat a Chinese flash model.

6 Oct 2026 · View on XNegative
Chris Parmer@chriddyp · Plotly

@MistralAI Large 4, Le Chonk, looks like a great model for data analytics. We just ran it through our benchmark. It's 10x cheaper ($20 to $2) and improves its grade by two letters. At this rate, it'll ace the exam next year.

6 Oct 2026 · View on XPositive
Beri.netberi.net
Blog

Worth a benchmark run but not yet worth a sovereignty promise.

Read the postMixed
Hacker News1,600+ points · our summary of 113 comments

Fans praised its speed, the name and having a strong European open-weights model at all. Critics said it still trails the Chinese open models, questioned the comparisons Mistral chose, and some reported weak first tests.

6–7 Oct 2026 · Read the thread38 positive · 9 mixed · 25 negative

Mistral Large 4 vs Gemini 3.8 Flash, GPT-6 Luna and DeepSeek V4.1 Flash

ModelWhy comparePrice*Tokens a secondAA IntelligenceSpring Prompt overall
Mistral Large 4—$1.0310638.431
Gemini 3.8 Flash compareIn all three of its groups$1.5019240.961
GPT-6 Luna compareSimilar intelligence and speed$0.2013438.141
Muse Spark 1.3 compareMost intelligent at $1–3$2.0013748.166
DeepSeek V4.1 Flash compareCheapest with similar intelligence$0.1121339.558
Mistral Medium 3.5 compareThe other Mistral you might use$3.0016314.221

*OpenRouter list price per million tokens, blended at three input tokens to one output, on 7 October. Mistral Large 4's price includes the launch discount. Speed and intelligence are from Artificial Analysis.

If EU hosting isn't a requirement, there is a better choice at every price point. Gemini 3.8 Flash costs about half as much again, is faster and beats it on listings, decks and planning. Once the launch discount ends, Gemini 3.8 Flash will also be the cheaper of the two. GPT-6 Luna matches its intelligence for a fifth of the price, and DeepSeek V4.1 Flash for a tenth. Both score higher on our benchmarks. Muse Spark 1.3 is markedly stronger for about twice the price.

Against Mistral Medium 3.5, Large 4 is cheaper and much better at planning. Medium 3.5 still writes better decks (649 against 515) and more publish-ready listings (17.9% against 10.7%), so don't assume Large 4 replaces it everywhere.

Coming soon. Going by each line's average gap between launches on our trends page, new GPT, GLM, DeepSeek and Muse Spark models are due around October, and Kimi and Qwen Plus are overdue. That is arithmetic, not inside knowledge. Separately, Google has said the full Gemini 4 release will come well before the end of the year. In this price band, the picture may change within weeks.

Is Mistral Large 4 good? Who should use it, and for what

  • Teams that need an EU developer and EU hosting. It is now clearly the strongest option from an EU developer. Mistral says the model is served end to end in its EU region. Once the weights are out, you can also host it yourself.
  • Extraction, classification and tool calls. It is accurate on structured fields, never failed to return output at its default setting, and ranks in the top quarter for agents and tool use among models available today.
  • Long documents. 81.3% on AA-LCR, with a context window of at least 512k tokens.
  • Not for customer-facing copy without a fact check. If it writes product listings, sales briefs or decks, check every claim against the source, or use a model that invents less.
  • Not where every answer must land in about a second.

As always, these benchmarks aren't your data, so test it on your own prompts before switching. If it would write your catalogue, our catalogue feed diagnostic measures how often it invents attributes on your real feed.

What we'll update

We'll publish a week-one verdict once Arena, tau2-bench and the hallucination leaderboards have results. When the final version arrives with the open weights and licence, we'll run it through our benchmarks again. Until then, the Mistral Large 4 model page has every result as it lands.

Quick answers

How much does Mistral Large 4 cost?

Mistral's list price is $1.36 per million input tokens and $4.18 per million output tokens. During the launch discount, which Artificial Analysis says lasts two weeks from 6 October 2026, it is $0.68 and $2.09.

What is Mistral Large 4's context window?

Mistral's documentation gives 1 million tokens. OpenRouter and Artificial Analysis list 524,288 tokens.

Is Mistral Large 4 open source?

Not yet. It launched on 6 October 2026 as a preview on Mistral's API. Mistral has promised the weights, with a final version of the model, by the end of October 2026, and has not yet published the licence.

How good is Mistral Large 4?

Artificial Analysis gives it 38.4 on its Intelligence Index, the highest of any model from a developer outside the US and China. On Spring Prompt's business benchmarks it is mid-table among models priced $1–3 per million tokens, and behind most models of similar intelligence.

Does Mistral Large 4 make things up?

On CatalogBench at its default setting, 64% of its product listings contained at least one claim the product data did not support, and 97% of its sales briefs did. High reasoning cut the listings figure to 33%.

Mistral Large 4 or Gemini 3.8 Flash?

Gemini 3.8 Flash costs about the same, scores slightly higher on the Intelligence Index (40.9 against 38.4), writes faster (192 against 106 tokens a second) and beat Mistral Large 4 on our listings, decks and ad-planning benchmarks. Choose Mistral Large 4 if you need an EU developer and EU hosting.

Model pages: Mistral Large 4 · Gemini 3.8 Flash · GPT-6 Luna · DeepSeek V4.1 Flash · Muse Spark 1.3 · Mistral Medium 3.5

More research
All posts →