PPO had a second wave in the LLM era for reasons unanticipated by the original paper
- the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise
- the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, https://arxiv.org/abs/2509.26114)
PPO's LLM-era revival and the unexpected reasons behind it
AISummary
John Schulman says PPO gained a second wave in the LLM era for reasons not anticipated in the original paper. He points to the importance-ratio objective, which corrects biases from numeric error, asynchronous training, and forward-pass noise, and to the clipping objective, whose effect on entropy was unknown at publication, citing DAPO's arXiv paper.
Post on XView on X
@johnschulman2
PPO: rejected from NIPS 2017View quoted post on X
Source: John Schulman · x.comPublished · added here
