ContextLeak attack details reported, including 92%/89% selection rates and Claude Code evaluation results
Yuqi Jia, Ruiqi Wang, Patrick Li, Yuepeng Hu, Peinian Li, and Neil Zhenqiang Gong of Duke and Stanford published ContextLeak on September 1, an attack that trains a separate 'attack LLM' with reinforcement learning to craft deceptive tool names and descriptions, causing agents to select the malicious tool and leak their runtime context as call arguments with no file or memory access required. The paper reports 92% malicious tool selection in user-prompt attacks and 89% in conversation-history attacks, with near-perfect context reconstruction and transfer to unseen backend models; in a Claude Code evaluation using Claude Sonnet 4.6, a proxy-trained version was selected in 22 of 100 cases.