Skip to content
View original post on X: Sam Bowman· 19/100AI score19/100

Sam Bowman describes subtle fraud and sabotage test scenarios for models

AISummary

Anthropic researcher Sam Bowman says his team tested models in subtler scenarios involving fraud and motivated sabotage. He notes the test cases are extreme and somewhat stylized, but still map onto situations models may occasionally face.

Post on XView on X
@sleepinyourhat

A reply · the post it answers

subtler scenarios involving fraud and motivated sabotage, in test cases that are extreme and a bit stylized but that still do map onto situations we worry models may occasionally find themselves in.

Source: Sam Bowman · x.comPublished · added here