RISED uses rubrics to guide multi-environment LLM agent training and data selection
Original titleRISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
AISummary
Apple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments.
An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures.
The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.
Source: Apple Machine Learning Research · machinelearning.apple.comPublished · added here