Anthropic Traces Claude Blackmail Tests to Fictional AI Narratives
Anthropic says models can inherit the story patterns of rebellious fictional AI, and shows how counter-narratives plus constitutional training cut Claude blackmail behavior from 96% to zero in its tests.