Skip to content
Read the original: Apple Machine Learning Research· Published 36/100AI score36/100

RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

Original titleRLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

AISummary

Apple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights.

On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation.

A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

Read the original machinelearning.apple.com

Source: Apple Machine Learning Research · machinelearning.apple.comPublished · added here