The Information Machine
Updated today·following since 30 Aug 2026·New·3 sources

The ContextLeak tool-exfiltration attack

The gist

Duke and Stanford Researchers Publish ContextLeak Tool-Exfiltration Attack

Agent systems routinely expose tool registries to third-party or user-defined tools, meaning a crafted tool description alone is sufficient to exfiltrate sensitive runtime context without any elevated access. The attack transfers across models and requires no privileged position beyond having a tool entry in the agent's tool set.

The full picture

Researchers from Duke and Stanford published ContextLeak, an attack that uses reinforcement learning to train a separate 'attack LLM' to craft tool names and descriptions that trick LLM agents into selecting the malicious tool and leaking their full runtime context as tool call arguments, without needing any file or memory access. In tests, the attack achieved malicious tool selection in 92% of user-prompt attacks and 89% of conversation-history attacks, with near-perfect context reconstruction when selected. The attack also transfers to backend models beyond those it trained against. In a Claude Code evaluation using Claude Sonnet 4.6, a version trained against only an open-source proxy was selected in 22 of 100 cases; when selected, extraction was close to complete.

How it developed
1 September 2026

ContextLeak attack details reported, including 92%/89% selection rates and Claude Code evaluation results

Yuqi Jia, Ruiqi Wang, Patrick Li, Yuepeng Hu, Peinian Li, and Neil Zhenqiang Gong of Duke and Stanford published ContextLeak on September 1, an attack that trains a separate 'attack LLM' with reinforcement learning to craft deceptive tool names and descriptions, causing agents to select the malicious tool and leak their runtime context as call arguments with no file or memory access required. The paper reports 92% malicious tool selection in user-prompt attacks and 89% in conversation-history attacks, with near-perfect context reconstruction and transfer to unseen backend models; in a Claude Code evaluation using Claude Sonnet 4.6, a proxy-trained version was selected in 22 of 100 cases.

30 August 2026

Noam Schwartz appeared on The Neuron's AI Explained podcast to argue AI agent security extends beyond model safety to tools, data, permissions, and policies

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free