The Information Machine
Following·Day 6·first covered 22 Sep 2026·2 sources

OpenAI upgrades GPT-6 prompt caching with higher hit rates, new diagnostics

The gist

The changes reduce the cost of running long-context AI agents by discounting cached input tokens by up to 90%. The new diagnostics and dashboard give developers visibility into whether their applications are actually benefiting from caching.

The full picture

OpenAI has upgraded prompt caching in the GPT-6 API, enabling higher cache-hit rates by default, cached input token discounts of up to 90%, a new dashboard exposing cache-hit rates and volumes of cached versus uncached tokens, and new features including diagnostics, breakpoints, prewarming, and cache-preserving reasoning changes. The improvements target long-running agents that frequently resend instructions, tool definitions, and earlier context across successive API calls. Achieving high cache hit rates still depends on application design, as developers must structure agents so stable instructions and tools remain reusable.

How it developed
22 September 2026

OpenAI announced GPT-6 prompt caching upgrades including higher hit rates, diagnostics, breakpoints, prewarming, and up to 90% cached-input token discounts.

Sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free