Jan Leike says alignment research needs taste, so AARs target scalable oversight
Original titleHowever, most alignment research is not very crisp and requires research taste when evaluating.
AISummary
Most alignment research is not crisp and requires research taste to evaluate, according to Jan Leike. He explains that Anthropic chose to point the AAR at a scalable oversight problem because progress there could let AARs tackle fuzzier alignment problems where humans can only provide weak supervision.
Source: Jan Leike · x.comPublished · added here