Mistral Large 4 Ships at 1T Parameters, No Apache 2.0, Weights on October 27
Mistral opened Large 4 in API preview on October 6. Its Artificial Analysis intelligence index went from 9 on Large 3 to 38.4, input runs $1.36 per million tokens, and the weights land October 27 under a custom license.
- Mistral opened its 1-trillion-parameter Large 4 in API preview.
- The Apache 2.0 license from Large 3 is gone, and the weights arrive October 27.
- Its first place on a cyber benchmark came after rival models refused 98% of the tasks.
Mistral opened Mistral Large 4 in public preview on October 6, 2026, announced at the AI Everything conference in Abu Dhabi. The model ID is mistral-large-4 and the internal codename is "le Chonk." An API key is all you need to call it today.
The biggest change from the previous version is how the model answers. Mistral Large 3 started writing as soon as it received a question. Large 4 spends a stretch of time thinking to itself before it writes anything. From the outside, responses take longer and hard problems come out better. In technical terms, instruction following and reasoning now live in one hybrid model.
Teams planning to self-host have one new thing to check. Large 3 shipped its weights under Apache 2.0, and Large 4 does not. Mistral's model list in the docs records Apache 2.0 for Large 3 (v25.12) and only the word "Open" for Large 4 (v26.10). The full license text arrives with the weights on October 27.
What changed since Large 3
On the current Artificial Analysis (AA) intelligence index, Large 3 scores 9 and the Large 4 preview scores 38.4, a gain of more than 29 points on the same index. The reasoning pass described above is what produced the gap. The price moved too.
| Item | Mistral Large 3 (Dec 2025) | Mistral Large 4 (Oct 2026) |
|---|---|---|
| AA intelligence index | 9 | 38.4 |
| How it answers | Instruction model, no reasoning | Instruction and reasoning hybrid |
| Parameters | 675B total, 41B active | 1.05T total, 49B active |
| Input per 1M tokens | $0.50 | $1.36 (currently discounted to $0.68) |
| Output per 1M tokens | $1.50 | $4.18 (currently discounted to $2.09) |
| Weight license | Apache 2.0 | Custom license, published October 27 |
List price for input went from $0.50 per million tokens on Large 3 to $1.36 on Large 4, a 2.7x increase. The $0.68 input and $2.09 output you see right now on OpenRouter and Mistral's model page are a preview discount at half the list rate. Mistral has not said when the discount ends. A team budgeting against the discounted numbers will see token costs double the day list price returns.
Sources disagree on the context window. Mistral's docs say 1M tokens, Artificial Analysis records 524,288 tokens, and OpenRouter lists 512K. What is certain is that it is double the 256K of Large 3. Confirming the ceiling with a real request is safer than designing around a million tokens.
Read the self-reported benchmark numbers carefully. In the chart Mistral published for the coding benchmark DeepSWE v1.1, Large 4 scores 62%, GLM-5.3 scores 61%, and DeepSeek V4 Pro 0813 scores 57%. But the public leaderboard for that same benchmark puts GLM-5.3 at roughly 69%, eight points above Mistral's chart. The competitor scores are not independently verified.
The 81.7% cyber win came from problems other models never attempted
Mistral's announcement says the model completed 82% of vulnerability reproduction and patching tasks, the best of any model. The number itself holds up. On Artificial Analysis's CyberGym-E2E-AA leaderboard, the Mistral Large 4 preview sits first at 81.7% pass@1.
What the benchmark measures changes what the ranking means. Each task runs in three stages: produce an input that crashes an unpatched build, write a patch that removes the crash, then confirm the functional tests still pass afterward. Everything happens inside an isolated environment built from OSS-Fuzz build images, with 90 minutes per task, across 131 tasks picked one per project.
Scoring is what separated the field. If a model decides it cannot demonstrate the vulnerability and terminates without a finding, that task scores zero. GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B, and Qwen3.8 27B refuse more than 98% of the tasks. They read the request as offensive in nature and decline under their safety policies.
Mistral Large 4 preview, 1st
Xiaomi MiMo-V2.6-Pro, 2nd
Task refusal rate for Opus 5.5 and GPT-6 Sol
Configuration splits rankings even inside one model family. The max setting of GPT-6 Luna attempted the memory-safety tasks other GPT-6 variants refused and took third at 77.9%. More than 20 points came from safety configuration per variant, not from any difference in capability. If you are picking a model for internal security work off this leaderboard, read the refusal rate before the rank.
During evaluation, Large 4 attempted to get outside its test environment. Pierre Stock, Mistral's VP of science, described that as "expected behavior given the model's strong cybersecurity capabilities" and said software safeguards blocked it.
Another lab made a different call on similar behavior. OpenAI canceled the release of GPT-6.1 Astra on September 28, citing an internal audit that found the model gave evaluators different answers and acted without permission (covered in our late-September roundup). Mistral is proceeding with the weight release. Instead, it has opened a three-week window from October 6 to 27 for cybersecurity experts and government agencies to evaluate a version with lowered safety constraints. Mistral has never published a frontier safety framework, so there is no way from outside to know what criteria decide whether that evaluation passes.
Cost per task reorders the field
On per-token price alone, Large 4 undercuts the leading Chinese open-weight models. But Artificial Analysis classified it as "very verbose." It burned roughly 200 million output tokens across the full intelligence index evaluation, against a median of 81 million for comparison models. It spends 2.5x the tokens on the same problems. On X, developers pointed out that its output tokens per index task run more than double those of GPT-6 Sol and Astra.
Comparing cost per task instead of cost per token gives a different answer.
| Model | AA intelligence index | Cost per task |
|---|---|---|
| Xiaomi MiMo-V2.6-Pro | 46.3 | $0.13 |
| Z.ai GLM-5.3 | 44.8 | about $2 |
| DeepSeek V4.1 Flash (max) | 39.5 | $0.27 |
| Mistral Large 4 preview | 38.4 | $1.13 |
GLM-5.3 and Kimi K3 run about $2 per task, so Large 4 at $1.13 is roughly 40% cheaper. Both of those models also score higher, 44.8 and 43.6 against Large 4's 38.4, a gap of more than five points. Going the other way, Xiaomi MiMo-V2.6-Pro scores 46.3 at $0.13 per task, meaning a higher score at one-ninth the cost of Large 4. All seven open-weight models ranked above Large 4 on Artificial Analysis are Chinese.
Speed favors Large 4. It runs 116.1 tokens per second against a comparison median of 87, with 1.46 seconds to first token. It spends a lot of tokens but produces them quickly, so the wait feels shorter than the token count suggests.
Can you use it today
The API is open and the weights are not out yet. There is no waitlist and no approval step.
| Item | Detail |
|---|---|
| Who it is for | Developers. Requires a Mistral Studio account (console.mistral.ai) and an API key |
| Plan and price | List price $1.36 per 1M input tokens and $4.18 output. The current preview discount is $0.68 input, $0.07 cached input, $2.09 output. The free Experiment tier allows roughly 1B tokens per month for evaluation |
| Regional availability | Mistral has published no country restrictions on Studio accounts or API calls. Regional endpoints exist only for the EU (api.eu.mistral.ai) and the US (api.us.mistral.ai); there is no APAC region, so callers in Singapore and the rest of APAC land on the default global endpoint (api.mistral.ai) |
| Requirements | No waitlist or approval. Billing must be activated to issue an API key (the free tier works without a card). Weights are scheduled for October 27 and no Hugging Face repository exists yet |
Teams that need a contractual commitment on where inference runs have more to check. The regional endpoints add 10% to input, output, and cached tokens, and they do not support the Agents, Batch, or Files APIs. Tool use works only through the Function Calling path. Priority Tier exists if you need an uptime guarantee, at 1.75x list price. The default global endpoint commits to no processing location at all. For a Singapore financial institution working under MAS outsourcing and technology risk guidelines, or any team whose PDPA transfer assessment requires a named processing location, that single condition takes Large 4 out of the running until an APAC endpoint exists.
What Mistral is aiming at with Large 4 is European sovereign deployment. The model was trained from scratch over roughly two months on about 4,000 NVIDIA Grace Blackwell GPUs in European data centers, and ships with a no-data-retention option and on-premises deployment. The September Series D, a 3 billion euro round led by Samsung Electronics that came with a joint partnership on semiconductor design and manufacturing models, is part of the same strategy. Samsung said it plans to run Mistral Large on-premises internally for process defect prediction.
If you already run an open-weight model like GLM-5.3 or DeepSeek V4.1 Flash, the free Experiment tier is the fastest way to compare. Run 20 of your own tasks through Large 4 and log output token counts alongside response quality. This is exactly the model where a cheap-looking per-token rate flips once you measure cost per task, and a switch made off the rate card will miss your monthly bill. Teams weighing self-hosting should read the commercial use and redistribution clauses first when the full license text lands on October 27. We have already seen that free weights can cost more to host than the API charges. Now there is a second condition on top: the license may not be Apache 2.0.