We scaled Beam with the largest publicly documented RL run we’re aware of, which ran on 10.5k GB300s for 4 weeks.
Algorithmic advances coupled with distributed infra enabled an RL system that scales.
Across our eval suite, capabilities continued to improve as we increased RL - with no signs of plateau.
