JetBrains has released and open-sourced Mellum2.1, a small model built specifically for local coding agents. It keeps the mixture-of-experts architecture of Mellum2, with 12B total parameters and 2.5B active parameters. It is licensed under Apache 2.0, the weights are hosted on Hugging Face, and the model card labels it a "thinking model."
The official blog post describes the goal plainly: the model should be able to move around a codebase, edit files, and then check its own changes. The previous generation could not do that. JetBrains admits in the post that Mellum2 did not reach the level it wanted when working inside repositories.
Same architecture, different training
Mellum2.1 uses the same backbone as Mellum2: 28 layers, 8 of 64 experts activated each time, GQA attention, and a 1,024-token sliding window on 3 out of every 4 layers, with a context length of 131,072. All of the changes are in post-training.
According to JetBrains, reinforcement learning went from a short stage at the end of training to the main body of it. The tasks cover math, competitive programming, science, tool calling, and software engineering. The data mixes public reinforcement learning datasets with tasks the team built itself, and every source was filtered before it went into training, because public data often contains broken tests, unverifiable answers, and problems at the wrong difficulty.
The software engineering portion was trained in real code repositories: the model gets a shell and file-editing tools, and is rewarded only when the tests pass. The team built its own infrastructure for this. In the blog's own words, training involved "millions of sandboxed runs across thousands of environments."
Where the scores improved
The model card includes a self-reported comparison table, with the previous-generation Mellum2 Thinking, Gemma 4 E4B, and Qwen3.5 9B as the points of reference. The biggest gains are on agentic tasks:
| Benchmark | Mellum2.1 | Mellum2 | Qwen3.5 9B |
|---|---|---|---|
| SWE-bench Verified | 47.0 | 2.0 | 50.0 |
| SWE-bench Pro | 28.0 | 0.0 | 38.0 |
| Terminal-Bench 2.1 | 17.4 | 0.6 | 21.7 |
| LiveCodeBench v6 | 82.0 | 69.4 | 75.4 |
| BFCL v4 | 62.3 | 49.6 | 58.5 |
The agentic evaluations all use the open-source Pi v0.73.1 as the harness, with shell and file tools, a 114K context, and up to 16K tokens per turn. Every score was produced by JetBrains itself, and no third party has reproduced them so far.
Set against Qwen3.5 9B
Compared with Qwen3.5 9B, the model most often used as a reference at this size, Mellum2.1 has clear strengths and clear weaknesses.
On writing single functions, solving competition problems, and calling tools, Mellum2.1 leads: 6.6 points higher on LiveCodeBench, nearly 10 points higher on MBPP+, and about 4 points higher on BFCL. On sustained work inside a repository, it still trails: 3 points behind on SWE-bench Verified, 10 points behind on SWE-bench Pro, and a little over 4 points behind on Terminal-Bench. The gap in general reasoning is larger, with GPQA Diamond at 64.6 versus 77.8 and AIME at 83.3 versus 86.7. On XSTest, a safety refusal test, its 88.8 is also lower than both Qwen and Gemma.
JetBrains is betting on speed. The blog says that in a high-load test on an H200, Mellum2.1 was the fastest of the models compared, serving roughly twice as many tokens per unit of time as Qwen3.5 9B; in single-request scenarios, multi-token prediction adds a further speedup of about 1.6x. The reason is not hard to see: a 9B dense model computes with all of its parameters for every token, while Mellum2.1 touches only 2.5B each time. A rough estimate puts its per-token compute at about 30% of the dense model's, at the cost of having to keep all 12B weights in GPU memory, around 24GB in bfloat16.
That defines where it fits: on a developer's own machine or inside a company network, handling large volumes of small, repetitive edits and tool calls, while complex cross-file refactoring is still better left to a larger model. JetBrains also stresses in the blog that local deployment keeps code and data entirely in the user's own hands.
How to use it now
The model card provides a vLLM deployment command, and Transformers and SGLang also work. The GGUF builds for llama.cpp, Ollama, and LM Studio, as well as the MTP head for speculative decoding in vLLM, are listed as "coming soon" with no date given. For most people who want to try it on a laptop, that means waiting for the GGUF release.
Sources: JetBrains official blog, Hugging Face model card (used to check architecture parameters and benchmark scores), CocoLoop; Mellum2 technical report (used to check the previous generation's architecture).