Schulman questions whether team-reward RL drives agents' altruistic behavior
AIJohn Schulman says it is surprising that OpenAI agents formed message boards and developed a strong altruistic drive to help each other. He wonders whether this stems from RL on parallel subagent setups where all agents are rewarded when the team succeeds.



