Claude's scalable oversight methods generalize well to math but poorly to code
Original titleIn this case, Claude develops scalable oversight methods on chat reward modeling datasets and evaluates them on math and code datasets.
AISummary
Anthropic's Claude developed scalable oversight methods on chat reward modeling datasets and evaluated them on math and code datasets. The best methods performed strongly on math but gave mixed results on code, suggesting the methods were overfit to the data and models used.
Source: Jan Leike · x.comPublished · added here