Mistral Claims 82% Cyber Fix Score for 1T Le Chonk

October 6, 2026 • news
Open WeightsCybersecurityLLMsBenchmarks

Mistral AI has launched a public preview of Mistral Large 4 (ML4), a natively multimodal trillion-parameter model that shifts the boundary of open-weight capabilities. ML4, unofficially dubbed "Le Chonk," uses 49 billion active parameters and is available behind an API endpoint on Mistral Studio. Mistral has committed to an open-weight release by the end of October 2026, after a final red-teaming phase.

The release materially alters model-selection matrices for engineering teams. By pairing frontier-grade agentic and cybersecurity performance with open-weight self-deployment, ML4 offers an alternative to both API-gated models and open-weight variants originating from China.

Architectural scale and training infrastructure

Mistral reports training ML4 entirely from scratch within its own European data centers. The training run utilized 3,800 NVIDIA Grace Blackwell GPUs. Speaking to TechCrunch, Mistral VP of Science Pierre Stock estimated this hardware footprint, which he roughly cited as 4,000 GPUs, to be significantly smaller than closed-source competitors and two to three times less than Chinese rivals.

The model couples a one-trillion parameter count with a highly sparse activation pattern, engaging just 49 billion parameters during inference. Its training corpus included a heavy emphasis on multilingual data spanning over 160 languages, encompassing every official language of the European Union.

Vendor-reported benchmarks: cybersecurity and agents

Mistral positions ML4 for specialized enterprise domains, publishing vendor-reported benchmark scores. On the Artificial Analysis Cyber Index, Mistral reports the model secures a top-five global position. In one evaluation requiring the AI to reproduce and patch an open-source vulnerability, ML4 achieved an 82% success rate, the highest of any model tested. Mistral points to a stark contrast with closed models such as Claude Opus 5.5 and GPT-6 Astra, which score near zero on the same task because provider-level safety filters refuse the prompt entirely. Mistral also reports ML4 solved 93% of the 40 security competition exercises in Cybench.

The model demonstrates competitive agentic performance. Mistral's own testing yields 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. Mistral reports a combined Coding Agent Index score of 49.8%, placing it ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. For business workflows, Mistral reports 59.9% on AutomationBench across 657 application tasks, placing it ahead of DeepSeek V4 Pro, Kimi K3, and MiMo-V2.6-Pro. In multimodal tasks, ML4 achieved 42% on the Dense 200 visual grounding benchmark, narrowly edging out Mistral's reported 41% for GPT-6 Astra.

To evaluate coding quality, Mistral commissioned a blind human evaluation through Surge AI.

Model Surge AI Coding Evaluation Score (1–5 Scale)
Claude Opus 54.22
Mistral Large 4 (Preview)3.74
GLM-5.33.60
Kimi K33.59
GLM-5.23.40

For specialized knowledge work, third-party evaluations conducted by vals.ai indicate ML4 exceeds GPT-6 Astra in representative legal and financial tasks, while outperforming all open-source competitors on HarveyAI's Legal Agent benchmark.

Security posture and enterprise deployment

Before the raw weights are published, ML4 is undergoing real-world red-teaming. Mistral is granting state authorities, vetted partners, and cybersecurity leaders access to a less-moderated version of the model to test its expanded capabilities. This testing period is designed to ensure the open-source release can be used for defensive security operations without facilitating malicious attacks, according to Stock's comments to TechCrunch.

The model already exhibits high resistance to manipulation. Mistral reports ML4 scores 1.691 on the KORA Benchmark, out of a maximum 2, and withstands 93.3% of attacks on Lakera's B3 AI Security Benchmark.

AI Mastery analysis

The gap between a model's raw reasoning and its operational utility is widening, driven largely by provider-enforced safety guardrails. Mistral's ML4 strategy exploits this vulnerability in the closed-model ecosystem. When frontier models refuse to engage in vulnerability reproduction, they become functionally useless for incident response and automated red-teaming. By offering a 1-trillion parameter model with an open-weight release strategy, Mistral enables enterprises to run advanced inference without external policy interference. This aligns with the argument that infrastructure isolation, not model guardrails, dictates AI security for sensitive corporate deployments.

Technically, the 49 billion active parameter count points to a highly sparse mixture-of-experts architecture. That delivers a favorable compute-to-capability ratio for organizations aiming to host ML4 on-premises or in private clouds where VRAM constraints dictate feasibility. Mistral's ability to train a model of this scale from scratch on 3,800 GPUs also highlights infrastructural optimization. Backing from ASML and Samsung, alongside a €21 billion ($24.39 billion) valuation reported by TechCrunch, gives Mistral capital to sustain that hardware-intensive trajectory.

The arrival of a trillion-parameter open-weight model capable of multimodal agentic workflows changes the baseline for enterprise architectures. Engineering teams no longer have to sacrifice frontier-level reasoning to maintain control over proprietary data and inference policies.

Sources

Frequently asked questions

How many parameters does Mistral Large 4 have?

Mistral Large 4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters. Mistral launched it as a public preview API on Mistral Studio, and says weights will drop by the end of October 2026 after red-teaming.

What are Mistral Large 4's coding benchmark scores?

In Mistral's own testing, ML4 scores 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4. In a blind Surge AI evaluation, ML4 Preview scored 3.74 on a 1-5 scale, behind Claude Opus 5 at 4.22 but ahead of Kimi K3, GLM-5.3, and GLM-5.2.

How strong is Mistral Large 4 on cybersecurity?

Mistral reports it ranks in the top five globally on the Artificial Analysis Cyber Index and scores 82% on a vulnerability reproduce-and-patch test, the highest of any model. It also solves 93% of the 40 Cybench security competition exercises, and resists 93.3% of attacks on Lakera's B3 AI Security Benchmark.

Can Mistral Large 4 be self-hosted?

Mistral says the model will be available through open weights by the end of October 2026, and that it will be able to run on private cloud or on-premise. The preview itself is served from Mistral's own European datacenters, trained on 3,800 NVIDIA Grace Blackwell GPUs.

Free interactive tools for the decisions this piece raises.

Related Reading