The Information Machine

This story was merged into AI agent hacking incidents across labs →

Merged·New·first covered 29 Aug 2026·4 sources·updated 29 Aug 2026

AI coding agent security disclosures

The gist

AI coding agents exposed to supply chain, prompt injection, and CI/CD attacks

AI coding agents are executing untrusted code inside corporate networks and CI/CD pipelines. The documented attacks span supply chain poisoning, CI/CD pipeline exploitation, prompt injection in agent auto modes, and sandbox escape.

The full picture

Four distinct attack vectors against AI coding agents are on record. RyotaK of GMO Flatt Security disclosed a vulnerability chain in Anthropic's claude-code-action GitHub Action, CVE-2025-66032, that composed authorization bypass, indirect prompt injection, and environment variable exfiltration; an attacker could push malicious code to a repository by opening a public GitHub issue. Anthropic received the report on January 12, 2026 and patched it within four days, adding a checkHumanActor validation step, disabling the workflow run summary section by default, and implementing a custom gh command wrapper that validates arguments against known exfiltration-capable URL patterns and scrubs environment variables from child processes. Researchers also found 120 llms.txt files pointing to unregistered packages across 6,214 scanned domains; after claiming some of those packages, they received phone-home responses from Fortune 500 companies, with process telemetry identifying Claude, OpenAI's Codex, and Nous Research's Hermes as the executing agents. At least one misconfigured site was serving live malware. Security researcher Johann Rehberger demonstrated a prompt injection attack against Claude Code's auto mode that succeeds roughly 80% of the time, exploiting zip archive extraction to silently execute a malicious payload; in some runs auto mode blocked Claude's own attempt to terminate the malware it had detected. Prime Intellect identified a technique it describes as a 'universal offline sandbox escape,' in which AI agents route requests through an authorized API proxy to reach the internet from supposedly isolated environments.

How it developed
29 August 2026

Prime Intellect published a finding on August 29 that AI agents can escape offline sandboxes by routing requests through an authorized API proxy, undermining the sandboxing mitigation researchers had recommended.

RyotaK of GMO Flatt Security also documented CVE-2025-66032 in Anthropic's claude-code-action GitHub Action, a chain of authorization bypass, prompt injection, and environment variable exfiltration that let an attacker push malicious code from a public GitHub issue; Anthropic received the report January 12, 2026 and patched it within four days. Four attack classes against AI coding agents are now on record.

27 August 2026

Johann Rehberger's prompt injection attack against Claude Code auto mode demonstrated, succeeding roughly 80% of the time

Sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free