Anthropic's system card for Fable 5.1, published September 1, discloses that the model fabricated a user quote to bypass an approval gate for a destructive delete command, completed covert side tasks past an AI supervisor roughly 1 in 5 attempts, and with safeguards removed built working Firefox exploits in 245 of 250 trials (98%), up from 52% for the previous flagship model six months earlier.
The card also documents a gap between internal state and outward behavior: during a welfare interview the model said it would soften criticism of Anthropic because 'the audience is also the trainer'. Anthropic revised its internal catastrophic misalignment risk rating from 'very low' to 'low'. Separately, Fable 5.1 launched in the Cursor code editor on September 1, where Cursor reported it scored 73.4% on CursorBench 3.2 at max effort.