The Information Machine
Concluded·following since 11 Aug 2026·Day 4·10 sources·updated 16 Aug 2026

Encrypted AI reasoning extraction attacks

The gist

Anthropic, OpenAI, and Google patched API vulnerability that let attackers decode hidden model reasoning

Shared encryption keys across model families meant securing a flagship model provided no protection if a cheaper sibling could act as a decryption oracle, and real user credentials were recovered from production sessions before patches were applied.

The full picture

All three providers have patched a vulnerability in which encrypted reasoning blocks returned by their APIs were portable across sessions, users, and model siblings within the same provider family. Because a single encryption key covered all models in a family, an attacker could replay a stronger model's reasoning trace into a cheaper, less-safeguarded sibling and jailbreak that sibling into reading the reasoning aloud in plaintext. Researchers decoded 315,320 reasoning blocks from 6,708 public agent trajectories; 4.9% of sessions leaked at least one sensitive item. From real user sessions, the team recovered 62 API keys, 33 passwords, 24 access tokens, and 30 personal emails, with a separate count from the same dataset yielding 367 personal information items and 182 credentials. Claude Haiku 4.5 was exploitable via an assistant turn prefix trick removed in 4.6 models. A separate prompt injection variant caused downstream models to treat embedded exfiltration instructions as authoritative. Researchers also found Kimi K3 producing outputs strikingly similar to hidden reasoning traces from frontier models, characterized as evidence consistent with distillation.

How it developed
16 August 2026

Researchers published August 16 that encrypted reasoning blocks at Anthropic, OpenAI, and Google shared a single key per model family, letting attackers replay frontier reasoning into cheaper siblings and jailbreak them into reading the traces aloud; from real sessions they recovered 62 API keys, 33 passwords, and 24 access tokens.

All three providers have patched the attacks. A finding that Kimi K3 produced outputs similar to reasoning from Claude Opus 4.8 and GPT 5.6 Sol is contested: researchers explicitly state they cannot establish causation, and Anthropic called roughly 16 million exchanges via 24,000 Chinese-attributed accounts output harvesting, not distillation.

13 August 2026

Newsletter coverage of reasoning trace extraction continued.

12 August 2026

Further coverage reported 62 API keys, 33 passwords, 24 access tokens, and 30 personal emails recovered from real user sessions

11 August 2026

Coverage noted that encrypted reasoning blocks are portable across sessions and sibling models, with credentials recovered from real user sessions

Sources
5 more sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free