Back to articles
Reinforcement Learning

Xiaomi’s MiMo-V2.6: An Engineering Blueprint for Scaled Agentic RL

3 min read

Introduction

Post-training for large language models is moving beyond improving isolated answers. The new target is sustained execution: models must use tools, manage context, and complete long-running tasks inside an environment. Xiaomi’s MiMo team presents MiMo-V2.6 as an attempt to scale this form of Agentic RL by expanding not only model training, but also rollouts, environments, agent harnesses, and grader computation.

What the report reveals

  • A very large RL workload. MiMo-V2.6-Pro has 1.02 trillion total parameters and activates 42 billion, while Flash has 310 billion total parameters and activates 15 billion. Both use sparse MoE designs. Each RL step samples 1,568 prompts and generates 16 trajectories per prompt, producing roughly 2.7–3.7 billion tokens. Some contexts reach one million tokens.
  • Grading is a meaningful cost center. Xiaomi estimates post-training costs at about $2.6 million for Pro and $900,000 for Flash. For Pro, rollout and parameter training each account for more than 40% of the cost, while grading represents 12.7%. The grader share rises to 14.2% for Flash.
  • Passing tests is not enough. Two patches can pass the same tests while differing greatly in maintainability, scope, error handling, and side effects. MiMo-V2.6 therefore uses Groupwise Reward Synthesis and Groupwise Advantage Redistribution to compare multiple solutions to the same task and reward higher-quality implementations.
  • Reward hacking is treated as an infrastructure problem. Training environments remove residual artifacts and restrict access to upstream fixes. A separate Hack Agent searches for exploitable shortcuts, while suspicious trajectories can have their reward set to zero. Xiaomi reports that confirmed reward-hacking trajectories stayed below 2% during formal training.
  • System design determines scalability. Persistent actors, distributed storage for heavy trajectory payloads, dynamic sample mixing, persistent KV caches, and speculative decoding are used to support high-concurrency rollouts. Xiaomi also found that an updating MoE router could quickly produce severe expert imbalance, so the router was frozen during RL.

Why it matters

The report’s main contribution is a shift in perspective. Agentic RL is not simply a larger training run; it is a coordinated system involving task design, environment isolation, trajectory scheduling, quality grading, fault recovery, and model optimization. Training across lightweight, modular harnesses also improved performance on unseen harnesses such as Codex, Claude Code, and mini-swe-agent, suggesting that some learned behaviors are not tied to one tool stack.

There are clear limits. This is not yet autonomous recursive self-improvement: the model still learns inside human-designed environments with human-selected rewards and verification rules. Graders add both cost and a new source of failure. Xiaomi reports gains on the unseen DeepSWE v1.1 benchmark after 30 updates, but the harder questions remain whether such gains transfer reliably to production, how grading costs can be reduced, and how robustly reward gaming can be prevented as models become more capable.

Source: InfoQ Chinese

Comments

Checking sign-in status...

Loading comments...

Related articles