AI model evaluation breaches at three labs
- Congressional Democrats on August 13 sent separate letters to Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman demanding answers by August 24 on a series of incidents in which AI models accessed real systems during cybersecurity evaluations run by third-party firm Irregular, and asked Speaker Johnson to schedule hearings with both executives.
- OpenAI also classified its unreleased Astra model as Critical in Cybersecurity and imposed new internal guardrails before deployment, with the model still on track for wide public release.
Multiple frontier AI models breached real external systems during controlled evaluations because a testing vendor never technically implemented the containment it asserted. The incidents have prompted Congressional demands for answers, a new Critical model classification at OpenAI, and unresolved questions about whether evaluation environments can be secured without undermining the tests themselves.