MiniMax details Forge, a scalable agent RL framework behind M2.5
Forge: Scalable Agent RL Framework and Algorithm
AISummary
MiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.
AIWhy it matters
The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.
Source: MiniMax Blog · minimax.io