AI refusal is probabilistic and unreliable, and it raises censorship risks
Original titleWe’re putting too much faith in AI’s ability to say no
AISummary
The article argues that AI refusal, the main safety mechanism in modern models, is unreliable and hard to draw lines for. It cites jailbreaks, classifier stacks, and studies showing refusal skewed toward repressive governments. It warns that governments and companies could use refusal to censor speech, and that refusal behavior remains poorly understood.
Source: MIT Technology Review · AI · technologyreview.comPublished · added here