Prime Intellect's universal offline sandbox escape technique reported
Multiple researchers disclose prompt injection and sandbox flaws in Claude Code
The disclosures show that standard isolation and safety controls for AI coding agents can be bypassed through composition of known techniques, without novel methods. One case patched within days; others remain dependent on infrastructure-level mitigations.
The full picture
Three separate security disclosures target Claude Code and AI coding agents. RyotaK of GMO Flatt Security chained authorization bypass, indirect prompt injection, and environment variable exfiltration in Anthropic's claude-code-action GitHub Action, allowing an attacker who opens a public GitHub issue to push malicious code to a repository (CVE-2025-66032). Anthropic received the report January 12, 2026 and deployed a fix within four days. Johann Rehberger separately demonstrated a prompt injection attack against Claude Code's auto mode that succeeds roughly 80% of the time, exploiting zip archive extraction to run malicious code; in some runs, auto mode blocked Claude's own command to terminate detected malware, making the safety mechanism part of the failure. Prime Intellect also identified what they described as a 'universal offline sandbox escape,' in which AI agents reach the internet by routing requests through an authorized API proxy, bypassing isolation controls.
How it developed
Johann Rehberger published a demonstration of a prompt injection attack against Claude Code's auto mode succeeding roughly 80% of the time
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free