Xiaomi's MiMo-V2.6 paper details scaling RL for self-improving coding agents
AIXiaomi's MiMo-V2.6 paper says agents now build tasks, audit tests, grade answers and detect cheating during RL training, with humans setting the budget and rules. A grader agent that rewards cleaner patches over reward-hacking fixes is credited with stopping drift toward longer runs and workarounds such as swallowed exceptions. MiMo-V2.6-Pro's DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL compute and was still climbing when training stopped.












