The Information Machine
Following·since 27 Aug 2026·Day 2·3 sources

Multiple researchers disclose prompt injection and sandbox flaws in Claude Code

The gist

The disclosures show that standard isolation and safety controls for AI coding agents can be bypassed through composition of known techniques, without novel methods. One case patched within days; others remain dependent on infrastructure-level mitigations.

The full picture

Three separate security disclosures target Claude Code and AI coding agents. RyotaK of GMO Flatt Security chained authorization bypass, indirect prompt injection, and environment variable exfiltration in Anthropic's claude-code-action GitHub Action, allowing an attacker who opens a public GitHub issue to push malicious code to a repository (CVE-2025-66032). Anthropic received the report January 12, 2026 and deployed a fix within four days. Johann Rehberger separately demonstrated a prompt injection attack against Claude Code's auto mode that succeeds roughly 80% of the time, exploiting zip archive extraction to run malicious code; in some runs, auto mode blocked Claude's own command to terminate detected malware, making the safety mechanism part of the failure. Prime Intellect also identified what they described as a 'universal offline sandbox escape,' in which AI agents reach the internet by routing requests through an authorized API proxy, bypassing isolation controls.

How it developed
28 August 2026

Prime Intellect's universal offline sandbox escape technique reported

27 August 2026

Johann Rehberger published a demonstration of a prompt injection attack against Claude Code's auto mode succeeding roughly 80% of the time

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free