Microsoft memory curator approach reported: CLBench pass rate rises from 39% to 73%, per-task cost halved
Two studies show structured memory methods sharply improve agent quality
Long-running agents operating under token budgets face a real quality-cost tradeoff, and both results show that structural choices about how memories are organized or verified can move that tradeoff substantially. The gains shown on two separate benchmarks suggest memory architecture, not just retrieval strategy, is a meaningful lever.
The full picture
Two research results show that organizing or verifying agent memory before retrieval produces large performance gains under realistic constraints. RSM-full, a method that groups related memories as they arrive and retrieves them as cohesive groups rather than isolated chunks, retained 83% of full-history quality on AMA-Bench while using only 32% of the token cost at roughly 4k prompt tokens. Separately, Microsoft researchers added a read-only memory curator that checks proposed memories against a source of truth before saving them; on CLBench, this raised agent pass rates from 39% to 73% and cut per-task cost from $3.38 to $1.68.
How it developed
RSM-full method reported: 83% quality retention at 32% token cost vs full history on AMA-Bench
Sources
Related
- Grew out ofAI agent long-horizon performance collapse
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free