Mistral Large 4 Benchmarks

Mistral's reported headline figures for Mistral Large 4 ("Le Chonk") are 93% on Cybench and 82% on CyberGym-E2E, reflecting a deliberate cyber-defense focus. Below: the headline results, a chart of the two reported scores, what the numbers mean, honest caveats, and one independent cross-check.

Headline results

BenchmarkScoreWhat it measures
Cybench 93% Offensive-security-style benchmark: the model attempts realistic cybersecurity tasks that Mistral reports, rather than academic question sets.
CyberGym-E2E 82% End-to-end cyber operations benchmark: the model carries out multi-step security workflows rather than single-shot answers.

Scores reported by Mistral in the October 6, 2026 announcement.

Reported scores at a glance

Vendor-reported scores from Mistral's October 6, 2026 announcement. No competitor figures plotted — margins between ML4, DeepSeek V4 Pro and Qwen3.8 Max are reported as narrow with no decisive leader.

What the numbers mean

Both headline benchmarks sit in the cybersecurity domain, which Mistral names as one of ML4's focus areas alongside coding, finance, manufacturing, visual grounding, agentic workflows, and multimodal understanding. The scores describe how often the model successfully completes realistic offensive-security-style tasks, not conventional academic accuracy.

One nuance matters for interpreting comparisons: Mistral says several closed frontier models score near zero on these benchmarks because they refuse the task rather than attempting it. That means a gap between ML4 and a closed model on Cybench or CyberGym-E2E may reflect differing safety refusal policies, not a difference in underlying capability.

Mistral claims ML4 is the best open-weight model from the US or Europe on aggregated benchmarks, and competitive with the strongest open models worldwide. CEO Arthur Mensch said it is above Chinese models on certain aspects, including cyber.

Caveats

  • All figures are vendor-reported, from Mistral's October 6, 2026 announcement — not independent testing.
  • The tested artifact is a preview checkpoint, not the final release.
  • Mistral says reinforcement learning is still ongoing, and capabilities may improve before the final weights release.
  • Scrutiny of the cybersecurity framing may intensify once the weights are public.

Independent testing

One independent data point is available so far: Artificial Analysis's intelligence leaderboard places the ML4 preview between DeepSeek V4.1 Flash and OpenAI's entry-level GPT-6 Luna — a major disadvantage versus OpenAI/Anthropic flagships, per that firm's testing.

On the parameter side, ML4 is a 1.05T total / 49B active mixture-of-experts model. For context, DeepSeek V4 Pro is 1.65T total with 49B active, and Qwen3.8 Max is 2.4T total; reported benchmark margins between the three are narrow, with no single model leading decisively across the board.

Model-by-model comparisons