JetBrains Mellum2.1: 2.5B Active Params Reach 47 SWE-bench Verified
In this article
JetBrains has released Mellum2.1, a 12-billion-parameter mixture-of-experts (MoE) thinking model optimized for autonomous coding agents and local execution. It ships under the Apache 2.0 license on Hugging Face and activates 2.5 billion parameters per token, keeping compute bounded while sharply improving agentic coding over Mellum2.
The upgrade is unusual because the base architecture did not change. JetBrains concentrated nearly all new work in post-training, primarily reinforcement learning (RL), and left the underlying model untouched. The result is a self-hostable reasoning agent that can explore a repository, edit files, and check its own changes without a proprietary cloud API.
Architecture and post-training environment
Mellum2.1 keeps Mellum2's architecture: 28 layers, 64 experts, and a router that activates eight experts per token. That holds the active footprint at 2.5 billion parameters out of 12 billion total. Attention uses grouped-query attention with 32 query heads and 4 KV heads; three of every four layers use a 1,024-token sliding window. Context length is 131,072 tokens, vocabulary size is 98,304 tokens, and weights ship in bfloat16.
The main change is the RL pipeline. JetBrains says RL moved from a short final stage to the dominant part of training. The company ran millions of sandboxed runs across thousands of environments, giving the model a shell and file-editing tools inside real repositories. It was rewarded when the repository's native test suite passed. Open RL datasets for math, competitive programming, science, tool use, and software engineering were filtered for broken tests and unverifiable answers before training.
Benchmark performance and evaluation pipelines
In JetBrains' self-reported evaluations, run across one shared pipeline in thinking mode, the repository-driven RL phase produced large agentic gains. SWE-bench Verified rose from 2.0 to 47.0, SWE-bench Pro from 0.0 to 28.0, and Terminal-Bench 2.1 from 0.6 to 17.4. Agentic runs used the open-source Pi v0.73.1 harness with a 114,000-token context.
JetBrains reports Mellum2.1 scored 82.0 on LiveCodeBench v6, ahead of Qwen3.5-9B at 75.4 and Gemma 4 E4B at 69.4. It also reached 62.3 on BFCL v4 and 91.5 on HumanEval+. Qwen3.5-9B still leads on some broad measures: JetBrains measured 77.8 on GPQA Diamond and 86.7 on AIME 25/26, versus Mellum2.1's 64.6 and 83.3. Because evaluating models across varying harness configurations changes results, JetBrains notes that Qwen's own card lists 65.6 on LiveCodeBench v6 and 81.7 on GPQA Diamond, while JetBrains measured 75.4 and 77.8 in its shared pipeline. JetBrains also reports safety improved: HarmBench fell from 21.5 to 8.5, where lower is better.
| Model | Active Params | Context Limit | SWE-bench Verified | GPQA Diamond |
|---|---|---|---|---|
| Mellum2.1 | 2.5B | 131,072 | 47.0 | 64.6 |
| Mellum2 Thinking | 2.5B | 131,072 | 2.0 | 51.0 |
| Qwen3.5-9B | 9B | 262,144 native | 50.0 | 77.8 |
| Gemma 4 E4B | 4.5B Effective | 128,000 | 23.0 | 53.1 |
Throughput and local deployment
Because post-training did not touch the architecture, Mellum2.1 retains Mellum2's speed. JetBrains says that under heavy load on a single NVIDIA H200 GPU, the model serves almost twice as many tokens as Qwen3.5-9B, and that multi-token prediction (MTP) makes a single request about 1.6 times faster. An MTP head for speculative decoding in vLLM is listed as coming soon.
Operators running the full model in vLLM must parse thinking outputs with --reasoning-parser qwen3; tool calling adds --enable-auto-tool-choice and --tool-call-parser hermes. JetBrains recommends temperature 0.6, top_p 0.95, and top_k 20.
MarkTechPost reports a GGUF repository already lists five builds. The recommended Q4_K_M build is 8.1 GB with an 88.0% top-token match against the 24.3 GB bfloat16 reference. The smallest build, MXFP4_MOE, is 7.0 GB with an 85.6% match.
AI Mastery analysis
Mellum2.1 shows parameter count is no longer the binding constraint for autonomous software engineering, provided post-training mirrors the deployment environment. By running millions of iterations inside real repositories with shells and test suites, JetBrains taught an efficient 2.5B-active-parameter model the recursive loop of exploring, editing, and validating code. The result is SWE-bench Verified performance that approaches the best open model in its comparison set without dense scale.
That efficiency matters for agentic loops. Autonomous coding is iterative and latency-sensitive; a 2.5B active footprint keeps time-to-first-token and generation cost low enough for tight local feedback loops.
The tradeoff is breadth. Mellum2.1 trails Qwen3.5-9B on GPQA Diamond and AIME, making it a specialized operational tool rather than a general oracle. Open coding models are moving toward repository-aware agents, and JetBrains has set a credible local, privacy-first baseline.
Sources
Frequently asked questions
How many active parameters does Mellum2.1 use per token?
Mellum2.1 is a 12B mixture-of-experts model with 2.5B active parameters per token. Its router selects 8 of 64 experts, keeping compute bounded while the full model retains a 131,072-token context.
What changed between Mellum2 and Mellum2.1?
The base architecture remained unchanged. JetBrains moved reinforcement learning from a short final stage to the main part of post-training, running millions of sandboxed runs in real repositories and rewarding the model when tests passed. SWE-bench Verified rose from 2.0 to 47.0.
How does Mellum2.1 compare with Qwen3.5-9B?
In JetBrains' shared pipeline, Mellum2.1 scores 82.0 on LiveCodeBench v6 versus Qwen3.5-9B's 75.4 and wins BFCL v4 with 62.3. Qwen3.5-9B still leads on SWE-bench Verified at 50.0 and GPQA Diamond at 77.8.
Can Mellum2.1 run locally on small quantized builds?
Yes. MarkTechPost reports GGUF builds starting at 7.0 GB for MXFP4_MOE and 8.1 GB for the recommended Q4_K_M, which retains an 88.0% top-token match against the 24.3 GB bfloat16 reference.
How do you run Mellum2.1 in vLLM?
Use `--reasoning-parser qwen3` to parse thinking outputs. Tool calling adds `--enable-auto-tool-choice` and `--tool-call-parser hermes`; JetBrains recommends temperature 0.6, top_p 0.95, and top_k 20.
Related Reading
Cohere's 218B MoE Scores 83.6 on WMT26, Beats DeepL and Google Translate
North Small Translate activates 25B of 218B parameters per token, scores 83.6 on WMT26 across 50 languages, and runs on a single B200 in 4-bit mode.
IFM K2 Horizon: Six Apache 2.0 Models, 0.9B to 375B, With Self-Audit
IFM releases six open-weight models from 0.9B to 375B, plus training corpus, code, and a self-published reward-hacking audit that corrects 70.2% to 66.9%.

Qwen 3.8 27B Is Strong but Overthinks by Default
Alibaba's 17 GB Qwen 3.8 27B excels at vision, tool use, and coding agents — but its xhigh reasoning default burns tokens on trivial prompts.