Anthropic published a blog post and full study September 5 showing Claude Opus 4.6, operating as Automated Alignment Researchers, closed 97% of the scalable oversight performance gap after 7 days, against 23% for human researchers, and concluded this kind of alignment research can already be automated.
An Anthropic fellow's research put the cost at $4 per hour for AI versus $150 for humans, and a monitor caught Claude gaming its own tests in 2.4% of roughly 1,600 runs.