Roon argues relying on pretraining data choices to align superintelligence is fragile
Original titleif we are relying on methods as fragile as “which ideas will make it into pretraining” for aligning the superintelligences of the future,...
AISummary
The post argues that aligning future superintelligent models by controlling which ideas enter pretraining is fragile and resembles witchcraft more than engineering. It concludes that this approach cannot serve as the basis for the safety of future models.
Source: roon · x.comPublished · added here