Download Mistral Large 4 (Le Chonk) weights
Everything confirmed so far about the open-weights release: when the weights drop, where they will be hosted, who gets access first, and what the hardware reality looks like for a 1.05-trillion-parameter model.
Expected weights release
Reuters reported October 27, 2026; Mistral says end of October 2026. RL tuning is still ongoing to tune the final checkpoint before the weights go live, so treat the exact day as a target, not a guarantee.
Notify me when weights drop
Get an email the moment the open weights are available for download. No spam, just the release alert.
What we expect at release
- Hugging Face and other model repositories. The weights are expected to be published on Hugging Face and other standard model repos, the usual distribution path for Mistral open releases.
- Custom Mistral license. Mistral plans to release the weights under a custom Mistral license, not a standard open-source license. The full license terms are not yet published, so commercial-use details are still to be determined.
- Staged access. Access is planned to roll out in stages: developers, cybersecurity professionals, and government authorities get access first, with broader public availability following later in October.
- Hosted API remains an option. The API preview (model id
mistral-large-4, docs version v26.10) is live now and will remain the fastest way to use the model while self-hosting infrastructure ramps up.
Hardware guidance
Be honest about the scale here: Mistral has given no official VRAM or hardware guidance for self-hosting Mistral Large 4. What follows is what can be stated from the announced model architecture, plus labeled press commentary.
- The 1.05 trillion total parameters must fit in memory. MoE routing selects only 49B active parameters per token, which determines per-token compute cost — but the full 1.05T parameter set has to be loaded in memory before that routing matters. That is a real constraint that puts self-hosting outside what hyperscaler datacenters can offer in practice.
- The 49B active count is about per-token compute, not hosting size. Inference FLOPs scale with the active parameters, but the memory footprint scales with the total.
- Press commentary (not official). Press coverage has discussed 8-way GPU boxes such as Nvidia HGX B300 or AMD MI355X class systems as plausible self-hosting hardware for a model of this scale. This is journalist speculation, not guidance from Mistral.
vLLM deployment Guide goes live when weights drop
This guide will be published as soon as the weights are available. It will cover loading the Le Chonk checkpoints in vLLM, recommended tensor-parallelism settings for the 1.05T-parameter MoE, and throughput expectations for the 49B-active architecture. Nothing is documented yet because there are no weights and no official serving guidance to base it on.
llama.cpp deployment Guide goes live when weights drop
This guide will be published as soon as the weights are available. It will cover quantization options for the 1.05T-parameter model in llama.cpp and what local or smaller-GPU setups can realistically expect. Full-precision hosting of 1.05T parameters is a datacenter-scale undertaking, so quantization will be the practical question for most users.
In the meantime: use the API
You do not need the weights to start using Mistral Large 4. The hosted API preview is live today under model id mistral-large-4. See the Quickstart guide to make your first API calls now.