Roon argues relying on pretraining data choices to align superintelligence is fragile
AIThe post argues that aligning future superintelligent models by controlling which ideas enter pretraining is fragile and resembles witchcraft more than engineering. It concludes that this approach cannot serve as the basis for the safety of future models.

