Xiaomi fixes MiMo's repeated tool calls for $90,000

Xiaomi's MiMo team published a technical blog post on September 27 walking back through a complaint developers kept raising after MiMo-V2.6 shipped: inside coding agents, the model would fire off the same or near-identical tool calls over and over, burning through context and compute while the task went nowhere. The fixed weights are already open-sourced under a filename carrying the MOPD suffix. The API platform switched to the new version starting 6 a.m. Beijing time on September 25, and MiMo Desktop users had their remaining quota for the period reset as compensation.

Separating "too many calls" from "repeated calls"

The blog doesn't treat every case of excessive calling as a defect. Xiaomi splits it into three buckets: parallel calls, a normal optimization where several calls each fetch different information; call sprawl, where the number of calls simply exceeds what the task needs; and call repetition, where neither the environment state nor the model's known information has changed, yet it keeps firing the same action. It's this third category the team set out to fix.

By Xiaomi's internal evaluation, the reply-level repetition rate blew past its 0.05% tolerance line, and the gap between environments was wide:

EnvironmentFlashPro
OpenCode1.02%0.54%
Claude Code0.27%0.10%
MiMo Desktop0.19%0.19%
MiMo Code0.11%0.07%

The Flash model in OpenCode posted the worst numbers by far — roughly one in every 100 replies spinning in place.

The root cause traces back to reinforcement learning

The team replayed checkpoints across the reinforcement-learning run and found that the share of single-turn samples firing more than ten calls rose steadily from 11.1% at step 0 to 24.6% at step 20; in the MiMo Code environment it was worse still, climbing from 30.6% to 41.7%. Training had originally built in a penalty, but only for turns exceeding 32 calls. From step 15 onward, samples triggering that penalty grew noticeably more common — a sign that 32 was far too loose a threshold to stop the model from settling into a habit of "a few extra calls never hurts."

The blog offers one striking figure: before the fix, at the point of a model's 12th tool call, the probability it would choose to keep calling was 94.56%, against just 5.43% for stopping. In rough terms, it almost never stopped on its own.

Two fixes, an order of magnitude apart

The most direct path was to lower the penalty threshold from 32 calls down to 8 and rerun the existing MixRL pipeline. In internal testing, the repetition rate fell from 13.45% to 3.83% — but it meant rolling back 20 training steps and retraining from there, which Xiaomi estimated at roughly $2.31 million, with no guarantee the results would hold on a different dataset.

The approach Xiaomi ultimately shipped is called MOPD, multi-teacher online distillation. It works by taking the repetition samples collected internally and training a separate, dedicated reinforcement-learning teacher whose only job is learning when to stop — just 12 training steps on about 7,000 examples. The main model is then rolled back 5 steps and continues training under that teacher's guidance. All in, the bill came to about $90,000 — roughly 4% of the first option's cost.

The payoff: the stop probability at the 12th call rose from 5.43% to 92.17%. Per the blog's cumulative stop-probability curve, before the fix a model needed to reach its 59th call before the odds of stopping crossed 50%; after the fix, it already sits at 99.87% by the 8th call. Xiaomi says repetition rates fell sharply across every platform, the improvement held consistently across different context lengths, and overall benchmark scores didn't slip.

Placed back on the V2.6 timeline

MiMo-V2.6 launched and was open-sourced on September 22, and at the time Xiaomi's headline message was about post-training scale: according to figures reported in the media, Pro and Flash each ran 30 reinforcement-learning steps, for a combined cost of roughly $3.5 million. This post-mortem, arriving five days later, effectively lays out a side effect of that very training run in public — down to disclosing the $2.31 million road not taken.

Public post-launch post-mortems from domestic model teams are uncommon, and it's rarer still to see the cost of a rejected fix listed alongside the one that shipped. For developers plugging Chinese models into OpenCode or Claude Code, the more practical takeaway is this: if you ran into MiMo repeatedly re-reading the same file or rerunning the same command in the past few days, the API is already serving the fixed version; anyone self-hosting needs to switch to the weights carrying the MOPD suffix. All of the repetition figures cited here come from Xiaomi's own internal evaluation — independent third-party retesting hasn't been published yet.

Sources: Xiaomi MiMo team technical blog, CocoLoop; per-environment repetition and sprawl rates, along with both fix-cost estimates, reflect Xiaomi's internal evaluation.