Skip to content
View original post on X: John Schulman· 40/100AI score40/100

PPO's LLM-era revival and the unexpected reasons behind it

AISummary

John Schulman says PPO gained a second wave in the LLM era for reasons not anticipated in the original paper. He points to the importance-ratio objective, which corrects biases from numeric error, asynchronous training, and forward-pass noise, and to the clipping objective, whose effect on entropy was unknown at publication, citing DAPO's arXiv paper.

Post on XView on X
@johnschulman2

PPO had a second wave in the LLM era for reasons unanticipated by the original paper
- the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise
- the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, https://arxiv.org/abs/2509.26114)

John Schulman@johnschulman2
PPO: rejected from NIPS 2017
View quoted post on X

Source: John Schulman · x.comPublished · added here