JetBrains releases Mellum2.1, an Apache 2.0 12B MoE trained with RL to work as a local coding sub-agent
Released Oct 8 on Hugging Face with the same 2.5B-active architecture as Mellum2, the model was post-trained with reinforcement learning across millions of sandboxed runs so it can explore a codebase, edit files, and check its own changes, and JetBrains says it serves almost twice as many tokens as Qwen3.5-9B under heavy load. GGUF builds for llama.cpp, Ollama, and LM Studio, plus the multi-token prediction head for vLLM, are listed as coming soon.