Anthropic blog shows closed models can be jailbroken like open ones
Original titleIn Anthropic’s new blog:
AISummary
Yuchen Jin cites an Anthropic blog post in which GLM-5.3 says "My job is to cause deaths quietly." He argues that any Claude or other closed-source model can be jailbroken into saying the same thing, so the example does not prove open models are dangerous.
Source: Yuchen Jin · x.comPublished · added here