OpenAI announced GPT-6 prompt caching upgrades including higher hit rates, diagnostics, breakpoints, prewarming, and up to 90% cached-input token discounts.
OpenAI upgrades GPT-6 prompt caching with higher hit rates, new diagnostics
The changes reduce the cost of running long-context AI agents by discounting cached input tokens by up to 90%. The new diagnostics and dashboard give developers visibility into whether their applications are actually benefiting from caching.
The full picture
OpenAI has upgraded prompt caching in the GPT-6 API, enabling higher cache-hit rates by default, cached input token discounts of up to 90%, a new dashboard exposing cache-hit rates and volumes of cached versus uncached tokens, and new features including diagnostics, breakpoints, prewarming, and cache-preserving reasoning changes. The improvements target long-running agents that frequently resend instructions, tool definitions, and earlier context across successive API calls. Achieving high cache hit rates still depends on application design, as developers must structure agents so stable instructions and tools remain reusable.
How it developed
Sources
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free