Mistral Large 4 vs Qwen3.8 Max

Mistral Large 4 (ML4, "Le Chonk"), announced October 6, 2026, is Mistral AI's open-weight flagship at 1.05T total parameters. Qwen3.8 Max is the 2.4T-total-parameter model from the Chinese side of the field. The two are the heavyweights of the open race, but reported benchmark margins between ML4, DeepSeek V4 Pro, and Qwen3.8 Max are narrow: no single model holds a decisive lead across the board.

Head-to-head specs

SpecMistral Large 4Qwen3.8 Max
Total parameters1.05T2.4T
Active parameters per token49B (MoE)Not disclosed in our sources
ArchitectureMoENot disclosed in our sources
Context window1M tokensNot disclosed in our sources
Preview API pricing$1.36 / 1M input tokens, $4.18 / 1M output tokensNot disclosed in our sources
Open weightsExpected October 27, 2026 (Reuters reported); custom Mistral license expectedNot disclosed in our sources
Training infrastructureTrained from scratch on ~3,800 Nvidia Grace Blackwell GPUs in Mistral's own European datacentersNot disclosed in our sources
Reported cyber benchmarks93% Cybench, 82% CyberGym-E2E (vendor-reported)Not disclosed in our sources

Where Mistral Large 4 stands out

Where Qwen3.8 Max may have an edge

Qwen3.8 Max's 2.4T total parameter count is the largest of the three models in this comparison set (ML4 at 1.05T, DeepSeek V4 Pro at 1.65T). Scale alone does not decide capability — the reported margins across the three are narrow — but it is the biggest open-class model by total parameter count. Beyond that, our sources do not disclose its active parameter count, pricing, or benchmark figures, so any deeper claim would be speculation.

Verdict

The margins are narrow and no single model holds a decisive lead across the board, so this is not a one-sided race. ML4 offers the smaller, cheaper-to-serve footprint (less than half Qwen3.8 Max's total parameters), a strong cyber benchmark story backed by Mensch's claim of leading Chinese models on cyber, EU-trained provenance for sovereignty-sensitive deployments, and broad multilingual coverage. Qwen3.8 Max counters with the largest total parameter count of the three.

The decision depends on your workload and deployment constraints: cost-efficient hosting and EU data residency point toward Mistral Large 4; raw scale points toward Qwen3.8 Max. If pricing and benchmarks for Qwen3.8 Max are published, revisit this comparison — the current gap in disclosed figures is the biggest unknown.