RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback
Original titleRLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback
AISummary
Apple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights.
On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation.
A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.
Source: Apple Machine Learning Research · machinelearning.apple.comPublished · added here