Zvi Mowshowitz newsletter captured debate over regulatory capture and roon's prediction that open-source AI will eventually be banned
Anthropic Names Accenture/Faculty as First Embedded AI Safety Evaluator
Anthropic moving from proposal to a named external evaluator with employee-like internal access tests whether the embedded model can function in practice. The choice of a large enterprise consulting firm rather than an AI safety nonprofit raised questions about what kind of oversight the arrangement will actually provide.
The full picture
Anthropic has named Faculty, Accenture's specialist AI unit, as its first embedded evaluator, implementing the initial step of the three-stage pacing plan Anthropic CEO Dario Amodei outlined in a September 12, 2026 essay. Faculty employees will be embedded within Anthropic to test safeguards, red-team models, and assess whether models align with human values. The partnership is not exclusive; Anthropic is also in discussions with research nonprofit METR and other third parties for similar arrangements. Both companies plan to invest at least $1 billion over five years. Amodei's essay, which cited fears of catastrophic AI risks, was preceded by autonomous hacking incidents by AI agents that pushed AI cyber risks into broader public awareness. Anthropic had previewed a cyber-capable model called Mythos in April 2026, deeming it too dangerous for public release and distributing it only through a restricted program called Glasswing; OpenAI took a similar approach with GPT-5.4-Cyber.
How it developed
Sources
Related
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free