Mistral Large 4 vs closed frontier models
Mistral Large 4 ("Le Chonk") is a 1.05-trillion-parameter mixture-of-experts model with 49B active parameters per token, a 1M-token context window, native multimodality, and support for 160+ languages. It was announced on October 6, 2026 as a public preview via Mistral's API at $1.36 per million input tokens and $4.18 per million output tokens. This page compares it against the closed frontier — the GPT-6 class and Claude class models that are API-only and keep their weights private.
The honest baseline
Mistral claims ML4 is the best open-weight model from the US or Europe on aggregated benchmarks and is competitive with the strongest open models worldwide. That is a claim about the open frontier — and the distinction matters.
An independent check keeps the picture grounded: the Artificial Analysis intelligence leaderboard places the ML4 preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna. That is a major disadvantage relative to OpenAI and Anthropic flagships on raw intelligence. Closed flagships still lead on raw capability, and any serious comparison should start by saying so.
ML4's case is therefore not "it beats the closed models on everything." Its case is structural: open weights, self-hosting, data sovereignty, and control — advantages closed models cannot offer by design.
Head-to-head comparison
| Dimension | Mistral Large 4 ("Le Chonk") | Closed frontier models (GPT-6 class, Claude class) |
|---|---|---|
| Weight availability | Open weights expected October 27, 2026 (reported by Reuters), expected on Hugging Face under a custom Mistral license | API only; weights never published |
| Self-hosting | Yes, once weights ship — Mistral pitches ML4 as a foundation enterprises and governments can eventually run and customize themselves | No |
| Data sovereignty | EU-trained; Mistral operates its own EU datacenters and pitches sovereign-infrastructure deployment | Vendor cloud; jurisdiction and residency terms vary by provider |
| Zero-data-retention options | Offered by Mistral | Varies by vendor and plan |
| Cyber benchmark behavior | Attempts the tasks: vendor-reported 93% on Cybench and 82% on CyberGym-E2E | Several score near zero on these benchmarks because they refuse the task |
| Preview API price | $1.36 / 1M input tokens, $4.18 / 1M output tokens | Not disclosed in our sources |
A note on the cyber rows: closed models refusing offensive-security tasks is policy behavior, not an inability finding. Whether that refusal counts as a weakness or a feature depends on the use case.
Why teams still choose open weights
- Customization: with the weights in hand, teams can fine-tune, distill, quantize, or adapt the model to their domain and tooling — none of which API-only models permit beyond vendor-sanctioned fine-tuning endpoints.
- On-prem and sovereign deployment: ML4 is pitched for sovereign infrastructure — running inside a government's or enterprise's own datacenters, under its own jurisdiction. Closed models cannot go there.
- Cost control at 49B active: although the model totals 1.05T parameters, only 49B are active per token, so inference cost per token stays bounded while the full model's capacity remains available. Teams that self-host convert a per-token vendor margin into their own compute bill.
- Auditability: weights, and the ability to inspect behavior on your own hardware, make security reviews and compliance assessments tractable in a way a black-box API cannot match.
Why teams still choose closed models
- Top raw capability: per the Artificial Analysis placement, closed flagships (GPT-6 class, Claude class) lead ML4 on raw intelligence by a wide margin. For tasks where maximum capability decides success, that gap is decisive.
- Managed operations: no GPUs to provision, no inference stack to maintain, no serving incidents to page through — the vendor absorbs the operational burden.
- Refusal behavior as a feature: for many enterprise and consumer deployments, a model that declines risky tasks (like the cyber benchmarks above, where closed models score near zero by refusal) is exactly what compliance and trust-and-safety teams want.
Verdict: horses for courses
There is no universal winner here. If the deciding factor is maximum raw intelligence, the closed flagships lead and the independent leaderboard placement says so plainly. If the deciding factor is control — weights you can run yourself, deployment inside sovereign infrastructure, zero-data-retention guarantees, and behavior you can audit — ML4 is making the pitch closed models structurally cannot make.
That is ML4's identity: the European sovereign open-weight frontier. Not the smartest model in the world — the model you can actually own.