GPT-6 Prompt Caching Makes Latency and Cost an Engineering Choice

OpenAI's GPT-6 prompt caching brings dashboards, diagnostics, breakpoints, and prewarming. Netics on why caching is an economic layer agents must be architected around, not a vendor convenie

Netics feature card for the OpenAI GPT-6 prompt caching article with the official OpenAI identity
Netics editorial feature card using the official OpenAI identity

TL;DR

  • OpenAI announced improved prompt caching for GPT-6 on September 22, 2026: higher cache hit rates by default, plus a dashboard, a diagnostics tool, explicit breakpoints, and prewarming.
  • Cached input tokens are discounted up to 90% for eligible shared prefixes reused within a 30-minute window.
  • The Prompt Caching Dashboard tracks hit rates; the diagnostics tool explains exactly why a request missed, with reasons like tools_changed and token counts.
  • Cache-friendly practice is now documented as a design discipline: stable tool ordering, append-only instructions, breakpoints, and prewarming.
  • Netics' take: prompt caching is an economic layer of the API, and the diagnostics are what turn it from a vendor black box into an auditable engineering choice.
  • For teams: put hit rate on the spend-limit review cadence, keep tool definitions stable, and use diagnostics when a miss surprises you.
Official OpenAI announcement screenshot: the Prompt Caching Dashboard showing cache hit rate, cache performance over time, and input token composition
Official OpenAI screenshot from "Better prompt caching for GPT-6" (openai.com, September 22, 2026): Prompt Caching Dashboard with cache hit rate, cache performance over time, and input token composition.

Caching was always the hidden economics of agents

Persistent agents make a series of API requests that build on one another, carrying forward the same instructions, tool definitions, and context across turns. Every one of those repeated input tokens is work the model has to redo — unless the provider caches the shared prefix. OpenAI has always kept some of that; what changed on September 22 is that the economics became visible and controllable. GPT-6 applications now get higher cache hit rates by default, a dashboard to monitor them, a diagnostics tool to explain misses, and explicit breakpoints to choose what is reused.

The 90% discount on cached input tokens is the part buyers will feel first. For a long-running agent whose context is mostly stable instructions and tool schemas, the difference between a cached and uncached run is the difference between a production budget and a fire. The 30-minute eligibility window matters too: it favors conversational workloads where turns arrive within a session, and it punishes applications that build a prompt from scratch on every call.

Original Netics diagram: up to 90% discount on cached input tokens for shared prefixes reused within a 30-minute window
Original Netics diagram: the cached-input discount up to 90% and the 30-minute eligibility window — the two numbers that define GPT-6 prompt-caching economics; source: OpenAI prompt caching announcement.

Diagnostics turn a black box into an audit trail

The most interesting tool in the announcement is diagnostics. You compare a request with a recent response, and the API tells you why the expected prefix was not reused: model_changed, tools_changed, input_changed, service_tier_changed, context_compacted, or a custom cache key change — each with estimated token counts so you can size the impact. That is the difference between guessing why your bill jumped and knowing the exact request attribute that invalidated the cache.

The guidance that follows is a documented design discipline, not a footnote. Keep tool definitions, schemas, and ordering stable, because any change before a breakpoint silences the discount. Use allowed_tools to restrict which tools run, or tool_choice: none, instead of removing definitions. Append new developer messages to the end of the context rather than rewriting earlier ones. These are the operational habits that determine whether caching is a discount or a myth.

The prefix is the contract

OpenAI's docs are blunt about the mechanism: the cache stores key-value tensors, and reuse requires the entire rendered prefix to match. Change the model, the tools, the service tier, or the input before the breakpoint, and the prefix no longer matches an eligible cache entry. The minimum cacheable length is 1,024 tokens for GPT-5.6 and later — a threshold worth knowing before you design for caching, because below it nothing is reused at all.

Breakpoints are the place teams will either win or lose. An implicit-mode design reuses whatever prefix boundary the engine chooses; explicit breakpoints let you pin the exact boundaries you reuse — after developer instructions and tool definitions, before user conversation. The guidance is specific: to reuse an implicit prefix, place an explicit breakpoint at that boundary in later requests. Extending a user message instead of appending a new one can silently kill reuse, because the old boundary stops being a message end.

Original Netics diagram: the four prompt-caching facts — 30-minute window, 1,024-token minimum, KV tensors, diagnostics at no extra cost
Original Netics diagram: four GPT-6 prompt-caching facts that define the operating boundary — 30-minute window, 1,024-token minimum cacheable prefix, KV tensors as the stored state, diagnostics at no additional cost; source: OpenAI docs.

Prewarming and reasoning-effort changes

Two features make caching friendlier for interactive agents. Prewarming lets an application prepare known context ahead of time — shared instructions, tool definitions, or reference material during startup — so the model starts responding sooner when the user asks the first question. The processing moves out of the user's wait time, which for a chat agent is the difference between "instant" and "spinner."

The second is subtler and genuinely useful: on GPT-6 models you can change reasoning effort between responses without breaking cache, by appending a configuration_update while leaving request-level reasoning effort unchanged. Raise effort for a hard task, lower it for a routine follow-up, and the reusable context survives the change. That removes a real tension in agent design, where adaptively spending more reasoning on hard steps used to cost you the entire cache.

Original Netics diagram: the cache-friendly checklist — stable tools, allowed_tools, append-only developer messages, explicit breakpoints, prewarming
Original Netics diagram: the prompt-caching discipline — stable tool definitions and ordering, allowed_tools instead of removal, append-only developer messages, explicit breakpoints, and prewarming; source: OpenAI prompt caching guide.

What teams should actually do

First, put the Prompt Caching Dashboard on the same review cadence as your spend limits: hit rate is now a product metric, and a drop in cache hits usually means a deploy changed tool ordering or instructions. Second, treat stable tool definitions as a code review rule — agents that regenerate tool schemas per request are paying the uncached rate by design. Third, use diagnostics when a miss surprises you; the reason it returns is the difference between "we changed the model" and "our prompt template embeds a timestamp."

The honest caveat is that caching also touches data handling. Diagnostics records contain configuration metadata, token-count estimates, and hashes — and OpenAI states they are compatible with Zero Data Retention, which matters for teams that turned on ZDR as an architecture constraint. Caching itself stores your context server-side; that was already true, but the new dashboards make the retention question more visible, not less. If you already treat API rate limits as part of agent architecture, prompt caching is the same story on the cost axis: the boundary is now measurable, and what is measurable can be governed. For a pass over your own agent cost structure, the Netics homepage is where we start.

Sources

Source: "Better prompt caching for GPT-6" — openai.com, September 22, 2026.