Exploring Mid-Session Changes · Part 2 — Context, caching and compaction: how AI costs add up
How conversation state, prompt caching and compaction shape the cost of a running AI session—and how to weigh a smaller context against rebuilding a reusable cache.
Reviewed
I What each turn carries forward
After the model-switch example in Part 1, we look more closely at the context carried between calls. Conversation state, prompt caching and compaction each affect a different part of that cost.
With OpenAI’s Responses API, previous_response_id lets the next call continue from an earlier response while the client sends only the new turn. The earlier input remains part of the model’s context and is still billed as input. Source 1.
Two quantities matter: what the client sends over the network and what the model processes for the next answer. State handling simplifies the first; caching and compaction change the cost of the second.
II Caching changes the price of repetition
Prompt caching attacks a different cost. OpenAI describes the cache as reusable KV state for an unchanged prompt prefix. For GPT-5.6 and later, cache writes are billed at 1.25× the ordinary input rate and cache reads at 0.1×. The write costs more once; later exact-prefix reuse can be much cheaper. Source 2.
The important word is prefix. Stable developer instructions, tool definitions and shared reference material belong early. Dynamic timestamps, IDs and other changing content belong later where possible. OpenAI’s current guidance explicitly recommends preserving earlier messages and appending new turns rather than rewriting history. Source 2.
III Compaction changes the cache boundary
Compaction reduces the context carried into later turns by replacing older conversation material with a smaller representation. That is valuable for long-running agents because fewer tokens can mean lower input cost and lower latency. Source 3.
But compaction also changes the rendered prompt. OpenAI’s prompt-caching documentation now lists context_management as a cache-affecting setting, and its cache diagnostics expose context_compacted as a specific reason why less of a previous prefix was reused. The docs are explicit: compaction can reduce cache reuse because the earlier conversation is replaced. Source 2 Source 5.
Cache reuse can survive in the unchanged part of the prompt. Stable instructions and tool definitions that remain before the changed section can still form a reusable prefix. After compaction, subsequent turns can also build a new reusable prefix around the compacted state. The engineering problem is therefore a trade-off: shrink the expensive growing history without needlessly changing the stable material that should remain cacheable.
IV The break-even moves after compaction
This is the same kind of break-even question as in Exploring Mid-Session Changes: Part 3, but inside one running conversation: a change can have an immediate transition cost while reducing the cost of the remaining work.
Consider an illustrative 100,000-token reusable conversation prefix across five future requests. Using GPT-6 Sol short-context Standard pricing checked on 28 September 2026—$2.00 per million ordinary input tokens, $0.20 cached input and $2.50 cache writes—one cache write plus four reads costs $0.33 for that prefix. Source 4.
Now assume compaction reduces the reusable carried context to 25,000 tokens. Further assume the first post-compaction request writes that new prefix and the next four requests reuse it fully. The comparable prefix cost is $0.0825: 25,000 × $2.50 / 1M + 100,000 × $0.20 / 1M. That is 75% below the 100,000-token cached case.
| Five future requests | Illustrative prefix cost |
|---|---|
| 100K prefix: one cache write + four reads | $0.3300 |
| 25K compacted prefix: one new write + four reads | $0.0825 |
Compare with the cache you already have. The table compares two initially written prefixes. It is not the saving from discarding an already-warm cache. If the 100K prefix is already cached at this checkpoint, five further reads cost $0.1000; the compacted path costs $0.0825 before producing the compaction. Its remaining advantage is only $0.0175 (17.5%), not 75%. Under these assumptions, the additional cost of producing the compaction must stay below $0.0175 to save money across those five requests.
This is a calculated example, not a benchmark. It excludes the cost of producing the compaction itself, new user turns, outputs, tools, partial cache hits and any long-context pricing threshold. Its point is narrower: a temporary cache reset can still be economically worthwhile when compaction removes enough future input.
V What good compaction looks like
The practical target is the lowest total cost for the required quality and latency. Cache-hit percentage is one measure within that comparison.
What I would measure around every compaction event
- Rendered input tokens immediately before and after compaction.
- Cached, cache-write and ordinary input tokens for the turns before and after it.
- The longest reusable prefix reported by cache diagnostics and any
context_compactedmiss reason. - How many later turns reuse the new compacted prefix before the next structural change.
- Output, tool and compaction-related usage separately from input-context cost.
OpenAI’s current guidance recommends comparing total input cost before and after compaction because fewer input tokens may save money even when cache reuse falls. Source 2. For long-running agents, that makes compaction part of cache design rather than a separate cleanup step.
Sources
- OpenAI — Conversation state
- OpenAI — Prompt caching
- OpenAI — Compaction
- OpenAI — GPT-6 Sol model and pricing
- OpenAI — Prompt cache diagnostics
Sources checked 2026-09-28.
