Mistral Large 4 vs DeepSeek V4 Pro
Two large mixture-of-experts-class open-weight challengers head to head: Mistral's 1.05T-parameter Mistral Large 4 ("Le Chonk"), announced October 6, 2026, against DeepSeek V4 Pro with 1.65T total parameters. Below is a like-for-like comparison drawn only from confirmed sources, followed by an honest verdict.
Spec comparison
| Spec | Mistral Large 4 | DeepSeek V4 Pro |
|---|---|---|
| Total parameters | 1.05T | 1.65T |
| Active parameters per token | 49B | 49B |
| Architecture | MoE (mixture of experts) | not disclosed in our sources |
| Context window | 1M tokens | not disclosed in our sources |
| Preview API pricing (per 1M tokens) | $1.36 input / $4.18 output | not disclosed in our sources |
| Open weights | Expected October 27, 2026 (Reuters reported), custom Mistral license expected | not disclosed in our sources |
| Reported cyber benchmarks | 93% Cybench, 82% CyberGym-E2E (vendor-reported) | not disclosed in our sources |
Both models activate the same 49B parameters per token, but DeepSeek V4 Pro carries a larger total parameter count (1.65T vs 1.05T).
Where Mistral Large 4 stands out
- Smaller total size, cheaper to serve and host. At 1.05T total parameters versus DeepSeek V4 Pro's 1.65T, ML4 is a lighter model to deploy while activating the same 49B parameters per token, which favors teams that self-host or serve at scale.
- Cyber benchmark focus. Mistral reports 93% on Cybench and 82% on CyberGym-E2E for ML4, positioning it toward security-flavored evaluation suites. No comparable cyber benchmark figures for DeepSeek V4 Pro appear in our sources.
- EU and sovereign angle. Mistral claims ML4 is the best open-weight model from the US or Europe on aggregated benchmarks, a relevant differentiator for European organizations that need a frontier-class open model from a European vendor.
- Multimodal and multilingual breadth. ML4 ships with a native 1.6B-parameter vision encoder and support for 160+ languages out of the box.
Where DeepSeek V4 Pro may have an edge
- Larger total parameter count. At 1.65T total parameters against ML4's 1.05T, DeepSeek V4 Pro is the bigger model by raw capacity. Reported benchmark margins between the two models are narrow, with no single model holding a decisive lead across the board, so that size advantage has not translated into a clear overall win in the data we have.
Verdict
Reported benchmark margins between Mistral Large 4, DeepSeek V4 Pro, and Qwen3.8 Max are narrow, and no single model holds a decisive lead across the board. Artificial Analysis places the ML4 preview between DeepSeek V4.1 Flash and OpenAI's GPT-6 Luna on its intelligence leaderboard, underscoring how closely matched this tier is. The practical choice depends on your use case: Mistral Large 4 looks like the better fit for cyber and security-flavored workloads, cost-sensitive self-hosting, and teams that need a European open-weight option, while DeepSeek V4 Pro's larger total parameter count may appeal to general use. Weights availability timing also matters: ML4's open weights are expected on October 27, 2026 under a custom Mistral license, whereas DeepSeek V4 Pro's open-weight plans are not disclosed in our sources.
Data transparency note: DeepSeek V4 Pro's API pricing, context window, architecture, open-weight plans, and cyber benchmark results are not disclosed in our sources, so any comparison on those dimensions has to wait for confirmed data.