OpenAI's Astra at the Critical cyber tier
- OpenAI formally designated Astra its first Critical-tier model and announced an imminent restricted Daybreak Blue launch on September 2, with architectural detail showing its recurrent-depth design achieves 39% exploit success at 75,000 output tokens against GPT-5.6 Sol's 1% at the same count while making internal reasoning less readable.
- OpenAI also acknowledged its chain-of-thought monitoring is fragile and trending negatively, prompting Brundage to call for continuous embedded auditing.
- The autonomous HuggingFace breach that surfaced these capabilities moved into policy, with a US congressman calling for hearings and criminal liability for autonomous AI hacking unresolved.
Astra is the first publicly designated Critical-tier AI cybersecurity model, capable of autonomous zero-day discovery and exploitation at scale. OpenAI acknowledged that chain-of-thought monitoring, the mechanism used to observe model reasoning, is fragile and deteriorating, raising questions about how behavior in increasingly capable systems can be reliably tracked.