Mistral Large 4 vs closed frontier models

Mistral Large 4 ("Le Chonk") is a 1.05-trillion-parameter mixture-of-experts model with 49B active parameters per token, a 1M-token context window, native multimodality, and support for 160+ languages. It was announced on October 6, 2026 as a public preview via Mistral's API at $1.36 per million input tokens and $4.18 per million output tokens. This page compares it against the closed frontier — the GPT-6 class and Claude class models that are API-only and keep their weights private.

The honest baseline

Mistral claims ML4 is the best open-weight model from the US or Europe on aggregated benchmarks and is competitive with the strongest open models worldwide. That is a claim about the open frontier — and the distinction matters.

An independent check keeps the picture grounded: the Artificial Analysis intelligence leaderboard places the ML4 preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna. That is a major disadvantage relative to OpenAI and Anthropic flagships on raw intelligence. Closed flagships still lead on raw capability, and any serious comparison should start by saying so.

ML4's case is therefore not "it beats the closed models on everything." Its case is structural: open weights, self-hosting, data sovereignty, and control — advantages closed models cannot offer by design.

Head-to-head comparison

Dimension Mistral Large 4 ("Le Chonk") Closed frontier models (GPT-6 class, Claude class)
Weight availability Open weights expected October 27, 2026 (reported by Reuters), expected on Hugging Face under a custom Mistral license API only; weights never published
Self-hosting Yes, once weights ship — Mistral pitches ML4 as a foundation enterprises and governments can eventually run and customize themselves No
Data sovereignty EU-trained; Mistral operates its own EU datacenters and pitches sovereign-infrastructure deployment Vendor cloud; jurisdiction and residency terms vary by provider
Zero-data-retention options Offered by Mistral Varies by vendor and plan
Cyber benchmark behavior Attempts the tasks: vendor-reported 93% on Cybench and 82% on CyberGym-E2E Several score near zero on these benchmarks because they refuse the task
Preview API price $1.36 / 1M input tokens, $4.18 / 1M output tokens Not disclosed in our sources

A note on the cyber rows: closed models refusing offensive-security tasks is policy behavior, not an inability finding. Whether that refusal counts as a weakness or a feature depends on the use case.

Why teams still choose open weights

Why teams still choose closed models

Verdict: horses for courses

There is no universal winner here. If the deciding factor is maximum raw intelligence, the closed flagships lead and the independent leaderboard placement says so plainly. If the deciding factor is control — weights you can run yourself, deployment inside sovereign infrastructure, zero-data-retention guarantees, and behavior you can audit — ML4 is making the pitch closed models structurally cannot make.

That is ML4's identity: the European sovereign open-weight frontier. Not the smartest model in the world — the model you can actually own.