GPT-6 Prompt Caching Makes Latency and Cost an Engineering Choice
OpenAI's GPT-6 prompt caching brings dashboards, diagnostics, breakpoints, and prewarming. Netics on why caching is an economic layer agents must be architected around, not a vendor convenie
TL;DR
- OpenAI announced improved prompt caching for GPT-6 on September 22, 2026: higher cache hit rates by default, plus a dashboard, a diagnostics tool, explicit breakpoints, and prewarming.
- Cached input tokens are discounted up to 90% for eligible shared prefixes reused within a 30-minute window.
- The Prompt Caching Dashboard tracks hit rates; the diagnostics tool explains exactly why a request missed, with reasons like
tools_changedand token counts. - Cache-friendly practice is now documented as a design discipline: stable tool ordering, append-only instructions, breakpoints, and prewarming.
- Netics' take: prompt caching is an economic layer of the API, and the diagnostics are what turn it from a vendor black box into an auditable engineering choice.
- For teams: put hit rate on the spend-limit review cadence, keep tool definitions stable, and use diagnostics when a miss surprises you.

Caching was always the hidden economics of agents
Persistent agents make a series of API requests that build on one another, carrying forward the same instructions, tool definitions, and context across turns. Every one of those repeated input tokens is work the model has to redo — unless the provider caches the shared prefix. OpenAI has always kept some of that; what changed on September 22 is that the economics became visible and controllable. GPT-6 applications now get higher cache hit rates by default, a dashboard to monitor them, a diagnostics tool to explain misses, and explicit breakpoints to choose what is reused.
The 90% discount on cached input tokens is the part buyers will feel first. For a long-running agent whose context is mostly stable instructions and tool schemas, the difference between a cached and uncached run is the difference between a production budget and a fire. The 30-minute eligibility window matters too: it favors conversational workloads where turns arrive within a session, and it punishes applications that build a prompt from scratch on every call.

Diagnostics turn a black box into an audit trail
The most interesting tool in the announcement is diagnostics. You compare a request with a recent response, and the API tells you why the expected prefix was not reused: model_changed, tools_changed, input_changed, service_tier_changed, context_compacted, or a custom cache key change — each with estimated token counts so you can size the impact. That is the difference between guessing why your bill jumped and knowing the exact request attribute that invalidated the cache.
The guidance that follows is a documented design discipline, not a footnote. Keep tool definitions, schemas, and ordering stable, because any change before a breakpoint silences the discount. Use allowed_tools to restrict which tools run, or tool_choice: none, instead of removing definitions. Append new developer messages to the end of the context rather than rewriting earlier ones. These are the operational habits that determine whether caching is a discount or a myth.
The prefix is the contract
OpenAI's docs are blunt about the mechanism: the cache stores key-value tensors, and reuse requires the entire rendered prefix to match. Change the model, the tools, the service tier, or the input before the breakpoint, and the prefix no longer matches an eligible cache entry. The minimum cacheable length is 1,024 tokens for GPT-5.6 and later — a threshold worth knowing before you design for caching, because below it nothing is reused at all.
Breakpoints are the place teams will either win or lose. An implicit-mode design reuses whatever prefix boundary the engine chooses; explicit breakpoints let you pin the exact boundaries you reuse — after developer instructions and tool definitions, before user conversation. The guidance is specific: to reuse an implicit prefix, place an explicit breakpoint at that boundary in later requests. Extending a user message instead of appending a new one can silently kill reuse, because the old boundary stops being a message end.

Prewarming and reasoning-effort changes
Two features make caching friendlier for interactive agents. Prewarming lets an application prepare known context ahead of time — shared instructions, tool definitions, or reference material during startup — so the model starts responding sooner when the user asks the first question. The processing moves out of the user's wait time, which for a chat agent is the difference between "instant" and "spinner."
The second is subtler and genuinely useful: on GPT-6 models you can change reasoning effort between responses without breaking cache, by appending a configuration_update while leaving request-level reasoning effort unchanged. Raise effort for a hard task, lower it for a routine follow-up, and the reusable context survives the change. That removes a real tension in agent design, where adaptively spending more reasoning on hard steps used to cost you the entire cache.

What teams should actually do
First, put the Prompt Caching Dashboard on the same review cadence as your spend limits: hit rate is now a product metric, and a drop in cache hits usually means a deploy changed tool ordering or instructions. Second, treat stable tool definitions as a code review rule — agents that regenerate tool schemas per request are paying the uncached rate by design. Third, use diagnostics when a miss surprises you; the reason it returns is the difference between "we changed the model" and "our prompt template embeds a timestamp."
The honest caveat is that caching also touches data handling. Diagnostics records contain configuration metadata, token-count estimates, and hashes — and OpenAI states they are compatible with Zero Data Retention, which matters for teams that turned on ZDR as an architecture constraint. Caching itself stores your context server-side; that was already true, but the new dashboards make the retention question more visible, not less. If you already treat API rate limits as part of agent architecture, prompt caching is the same story on the cost axis: the boundary is now measurable, and what is measurable can be governed. For a pass over your own agent cost structure, the Netics homepage is where we start.
Sources
- Better prompt caching for GPT-6 — OpenAI, September 22, 2026.
- Prompt caching — OpenAI API Docs, accessed September 23, 2026.
- Prompt cache diagnostics — OpenAI API Docs, accessed September 23, 2026.
Source: "Better prompt caching for GPT-6" — openai.com, September 22, 2026.