Agent engineering · Engineering

Agent Context Compaction, KV Cache, and Prompt Caching Explained

Understand three different memory mechanisms, calculate when compaction pays, and preserve agent state without mistaking cached tokens for durable memory.

An agent can have a million-token context window and still forget the instruction that matters. It can receive a large cache discount and still be expensive. It can summarize its conversation successfully and then lose the evidence needed to finish the task. These failures arise because developers use the word memory for several mechanisms with different guarantees.

The useful engineering distinction is between preserving computation, reducing the material presented to the model, and preserving authoritative application state. KV caching and prompt caching primarily help with the first. Compaction helps with the second. A durable task ledger, file store, or database helps with the third. A reliable agent usually needs all three, but each should have a separate owner and a separate test.

This guide explains that separation, then develops an explicit economic model for deciding when to compact. Its calculations are hypothetical engineering examples, not measurements of a commercial model. The companion Python program runs offline and checks the arithmetic; no provider API or model benchmark was run for this article.

Three mechanisms, three different questions

Mechanism Question it answers What it preserves What it does not guarantee
KV cache during inference Can we avoid recalculating earlier attention state? Model-specific intermediate tensors Durable facts, a correct answer, or a cross-provider format
Provider prompt caching Can later requests reuse a matching prefix? Eligible previously computed prefix state A hit on every request or indefinite retention
Context compaction Can we continue with less active context? A smaller representation of prior task state Lossless recall of every source detail
Application task ledger Which facts and actions are authoritative? Explicit records chosen by the application The model will use every record correctly

This taxonomy is an engineering framing. The underlying mechanisms are documented independently: Hugging Face explains inference caching, vLLM describes prefix caching, and OpenAI documents compaction.

Consider an agent reconciling invoices. It reads vendor records, discovers that invoice 804 was already paid, and proposes a correction. Reusing the invoice prompt prefix saves processing. Summarizing older turns reduces the next request. Neither mechanism should become the authoritative payment record. The database transaction and its confirmation are the authority; the agent's recollection is a convenient representation of that record.

This distinction changes implementation priorities. You first decide what must survive process termination. You then decide what the next model call needs. Only after those decisions do you optimize how much computation can be reused. Reversing the order encourages impressive cache metrics attached to fragile task behavior.

What a KV cache actually contains

In a causal attention model, a new token attends to representations of preceding tokens. The keys and values for earlier positions can be retained rather than recomputed for every decoding step. The cache belongs to a particular model execution and its representation; it is not a document summary or a collection of facts. The Transformers caching explanation provides the implementation background.

For a simplified full-attention model with a conventional KV layout, a useful planning equation is:

KV bytes = 2 × layers × KV heads × head dimension × tokens × bytes per element × concurrent sequences

The factor of two represents keys and values. This is a simplified storage estimate, not a sizing formula for every architecture. Compressed attention representations, sliding windows, quantization, sharing, allocator overhead, and serving implementation can change the result substantially.

Take an explicitly fictional deployment: 32 layers, 8 KV heads, a head dimension of 128, two-byte elements, and 100,000 retained tokens. The estimate is 13,107,200,000 bytes, or about 12.21 GiB for one sequence. Four separate sequences would require roughly 48.83 GiB for this component alone if no state is shared. Model weights, temporary activations, workspaces, and fragmentation are additional.

The example explains why a large supported context is not a promise of cheap concurrency. A server that accommodates one long request may struggle with a queue of them. When an operator publishes tokens per second, ask whether the measurement used short or long contexts, how many requests were active, and whether repeated prefixes were warmed beforehand.

There are several ways to trade memory for speed. Transformers exposes dynamic, static, and quantized cache strategies and discusses offloading; their supported features vary. A static allocation can help compilation but reserve space that short requests never use. Offloading can relieve device memory pressure while introducing transfers. These are deployment tradeoffs, not evidence that one strategy universally wins. Consult the current cache-strategy documentation for the exact implementation you intend to run.

Prefix caching is reuse, not understanding

Prompt caching extends the reuse idea across eligible requests. A matching beginning of a prompt can reuse prior work even though the new question at the end differs. It does not cache a universally reusable answer: different questions can still produce different generations.

vLLM's documentation describes automatic prefix caching as reusing KV state for shared prefixes and emphasizes that it helps prefill rather than the generation of new tokens. That distinction is useful even when a hosted provider hides its serving implementation. A long shared handbook with a short answer is a promising workload. A short question followed by a very long generated report has less input work available to save. vLLM automatic prefix caching.

For a concrete hosted example, current OpenAI documentation describes matching prefixes, model-specific breakpoints, and retention settings. GPT-5.6 and later have different cache controls from older generations, so old implementation advice must be checked against the selected model. The next section uses hypothetical prices to avoid presenting one provider's tariff as a universal rule. For provider-specific configuration and pricing, see our prompt caching economics guide. OpenAI prompt caching.

The key design lesson is to distinguish potentially shared tokens from billed cache-read tokens. A team may put a stable policy first in every prompt and still receive few hits because the configured breakpoint is beyond dynamic text, entries have expired, or a relevant request setting changed. Instrument actual usage fields rather than estimating savings from the string layout alone.

Cache reuse should also respect the application's isolation boundary. A prefix hash is useful for observability, but it must not become a way to disclose another customer's document or infer whether a sensitive prompt has been used. Keep accounting identifiers separate from raw content, and follow the provider's supported mechanisms for tenant separation.

Compaction changes the information budget

Compaction creates a smaller context for subsequent work. That is a semantic transformation. A good compacted state retains the goal, constraints, important observations, and unresolved work. It can omit repetitions and details that are no longer needed. A bad one removes the exception that controls the next action.

OpenAI's standalone compaction endpoint returns a context window containing an opaque encrypted compaction item and potentially other retained items. Its documentation tells developers to pass that returned window forward as-is. Server-side compaction offers a different integration path. These are provider API contracts, not ordinary prose summaries that an application should freely rewrite. OpenAI compaction.

Claude's documentation distinguishes on-demand compaction, compaction at a threshold, and a client-side summarizer. It also distinguishes summarization from clearing older tool results. Model support and beta requirements need checking before implementation. An integration written for one mechanism should not assume the output structure or continuity rules of another. Claude compaction overview.

The provider operation and the application's task ledger therefore play complementary roles. Preserve provider output according to its contract. Separately preserve action receipts, exact source references, permissions, and records required for audit or recovery. If an opaque continuation fails, the application should still know what it has done and what it is authorized to do next.

Why a huge context window does not settle quality

The original Lost in the Middle study found position-sensitive performance on its tested long-context retrieval tasks and models. It is historical evidence that supported context length and effective use of context are separate questions. It is not a benchmark of GPT-6, Claude 5.5, or a current Qwen release. Liu and colleagues, Lost in the Middle.

An application's evaluation should reproduce its own information topology. If the decisive exception appears halfway through a log, keep it halfway through some test cases. If tasks require connecting two sources far apart, evaluate that relationship rather than a single fact lookup. If compaction is used, test the answer before and after compaction with the same acceptance rule.

Suppose an order-processing task says, near the beginning, that customers may change an address. Later a tool says shipment has entered the carrier network. The correct next action now depends on the later state transition. A summary that retains the general policy but drops the shipping status can produce a confident and incorrect update. Longer context may preserve both facts, but it does not guarantee that the model resolves them correctly either.

The acceptance test should ask whether the agent observed the current shipment state, followed the relevant policy, and recorded the correct outcome. Counting summary tokens or judging prose fluency does not answer those questions.

The economic decision: compare future paths

Compaction should be compared with the path the agent would otherwise take. That path may include cheap cache reads, new tool input, generated reasoning, and a later context-price threshold. A simplistic comparison between old tokens and new tokens misses the opportunity cost of breaking a warm prefix.

Here is an intentionally simplified scenario. All rates are hypothetical USD per million tokens. The calculation concerns only the changed history; output and unchanged new input cancel between the alternatives.

Assumption Value
Original history 120,000 tokens
Compacted history 20,000 tokens
Remaining calls that reuse this history 8
Ordinary input rate $2.00
Cache-read rate $0.10
Cache-write rate $2.50
One-time compaction operation $0.25

If all eight future reads of the original history are hits, keeping it costs 8 × 0.12 × $0.10 = $0.096. Compacting costs $0.25 + 0.02 × $2.50 + 7 × 0.02 × $0.10 = $0.314. Under these assumptions, compaction increases this portion of spend by $0.218.

Now consider the other extreme: the original history is uncached at the ordinary input rate on every call. Keeping it costs 8 × 0.12 × $2 = $1.92. Compacting still costs $0.314 if the compacted prefix is written once and reused thereafter. The savings are $1.606. The same reduction in token count produces opposite economic conclusions because the counterfactual cache behavior differs.

The fictional tariff is chosen for arithmetic, not to represent a provider's missed-cache billing policy. Some systems bill newly cached eligible tokens at a write rate rather than an ordinary input rate; adapt the formula to actual usage categories. A realistic forecast also includes growing tool results and updates to the compacted summary, which can alter prefix reuse again.

For a stationary illustration with cache hit fraction h, the kept-history cost is:

calls × old_tokens / 1,000,000 × [h × read_rate + (1 − h) × fresh_rate]

The compacted path costs:

compaction_cost + new_tokens / 1,000,000 × write_rate + (calls − 1) × new_tokens / 1,000,000 × read_rate

At 90% hits in the example, kept history costs $0.2784, which is still slightly below $0.314. At 80% hits it costs $0.4608, so compaction wins on the modeled input component. The threshold is approximately an 88.05% original-history hit rate. Real decisions should use measured traffic distributions and cost categories, not this illustrative threshold.

The missing term: recovery and accepted outcomes

Input savings are only one line of the decision. Compaction may reduce latency or make the task more reliable. It may also trigger expensive recovery work. Measure accepted results rather than successful API responses.

Imagine that compaction adds a 2% probability of losing a necessary source reference, and each such loss creates a $5 recovery cost. The expected extra cost is $0.10 per task. This is another fictional scenario. If modeled token savings are $0.05, compaction is economically worse even before considering user frustration. If compaction prevents a costly context overflow, the tradeoff can go the other way.

The practical model is token costs + compaction costs + expected recovery costs + review costs, divided by accepted tasks. Keep these quantities observable independently. An input optimization that increases failed runs should be visible even if it improves cache hit rate or average token count.

A related distinction is latency. Record the time to a first useful action and the time to a verified completed task. A compaction pause may worsen the first and improve the second by avoiding repeated reading. Mean latency can conceal that some tasks stall catastrophically; inspect the tail and the fraction needing human rescue.

A state contract that survives compaction

For an agent changing a repository, a useful state contract includes the current objective, explicit constraints, repository revision, modified files, checks performed, checks still required, unresolved hypotheses, and authorization boundaries. Each action receipt should identify what happened and how the result was verified.

For a research agent, replace file revisions with source identifiers, publication dates, retrieved versions, exact figures, disagreements, and open questions. For a support agent, preserve customer identity, the relevant account state, policy version, promised actions, and the status of each external operation.

These records should remain short, but shortness is subordinate to precision. “Tests passed” is too vague when only a narrow unit suite ran. “All prices verified” is too vague when one promotional price expires next quarter. “Refund approved” is too vague when the model proposed a refund but the transaction never executed.

Distinguish observations from proposals and proposals from committed actions. An append-only receipt log is often easier to audit than a repeatedly overwritten paragraph. Build compacted summaries from that log and the required source material; do not allow a summary to retroactively convert an intention into a completed action.

When to compact in a running agent

A useful policy combines a hard context budget with task-aware triggers. Reserve room for the largest expected tool result and final output. Compact before a predictable large retrieval, before a price threshold becomes expensive, or at a coherent task boundary where completed work can be summarized safely.

Avoid summarizing between issuing a tool call and receiving its result unless the API explicitly supports that lifecycle. Treat incomplete tool pairs and pending external operations as state that needs precise continuity. A compacted conversation that loses an outstanding call identifier can fail even if its prose captures the high-level plan.

After a large source-reading phase, preserve a claim ledger and source pointers rather than every scraped paragraph. After a coding phase, preserve the exact changes and check receipts rather than every terminal log. Retain original documents in accessible storage so the agent can reopen them when a question requires more detail.

Use hysteresis to avoid repeated compaction around one threshold: compact down to a meaningfully smaller target and leave room for growth. The target should come from workload measurements. A fixed “summarize every ten turns” rule ignores that one turn may contain a sentence and another a 30,000-token tool result.

Offline calculator and reproducible checks

The companion context_budget.py implements the simplified model above, estimates conventional KV storage, and validates inputs. Run it with Python 3.11 or later:

sh
python examples/context_budget.py --self-test
python examples/context_budget.py

Its output includes the 12.21 GiB KV estimate, the $0.096 fully cached kept-history path, and the $0.314 compacted path. Those checks test arithmetic and boundary cases. They do not test whether a real summary preserves task information, whether a provider gives a cache hit, or whether any model can finish the invoice task.

For that separate quality test, create task fixtures with deliberately placed facts, expected end states, and source references. Run the same fixtures without compaction and with the proposed compaction policy. Include cancellation, contradictory tool responses, stale source versions, and user corrections. Count actions that should never occur as failures even when the final answer sounds plausible.

Checkpoint recovery with a versioned ledger

Make recovery a small application protocol. Before compaction, persist a checkpoint identifier, ledger revision, source-artifact versions, policy version, outstanding operation identifiers, and the last acknowledged action receipt. Commit these records durably before marking the checkpoint usable. A summary can refer to that checkpoint; it should not become the only copy of its evidence.

On resume, compare the checkpoint revision with the current ledger. Another worker or a user correction may have changed the task while compaction ran. If the revisions differ, reconcile the new receipts before planning another action. Use an expected-revision check when appending decisions so two workers cannot both advance the same task from an obsolete state. This is an application design proposal, not a guarantee offered by a provider’s context mechanism.

Version the checkpoint schema too. A reader should reject an unsupported schema rather than interpret a newly renamed status as a familiar older one. Migration belongs in explicit application code.

A lost citation should trigger evidence recovery, not a request to confidently recreate the quotation. Fetch the retained artifact by identifier and verify its recorded version. If only a newer source is available, mark the old observation stale and recheck the claim. If the original cannot be recovered, preserve the gap explicitly and suspend actions that depend on it. A current price page cannot silently replace the evidence for yesterday’s billed rate.

For an external operation with an uncertain outcome, query its authoritative status before repeating it. A compacted note saying “refund requested” does not reveal whether the payment service committed the refund. Store operation identity separately from the natural-language plan, and distinguish requested, acknowledged, failed, and unknown outcomes. Recovery should resolve unknown state rather than treating it as failure.

Test the protocol with injected interruptions: termination after checkpoint storage but before summary completion; a user correction during compaction; a missing source artifact; and a tool timeout after an action may have committed. The acceptance rule is that durable receipts survive, stale observations remain labeled, and no uncertain action is duplicated. These are proposed failure fixtures; no live recovery experiment was performed for this supplement.

Conclusion

The right memory design starts with the task's authority and evidence, then chooses a smaller active context, then optimizes reusable computation. This ordering makes it possible to reason clearly about failures: the durable record may be wrong, the compacted state may be incomplete, or the cache may simply have missed.

A long context is capacity. A warm prefix is an opportunity to reuse work. A compacted continuation is a representation of prior state. None is a substitute for an acceptance rule that checks what the agent actually accomplished. Instrument the three mechanisms separately, and judge changes by verified task quality and total cost.

All sources accessed 8 October 2026. Provider API behavior is version-dependent; historical research is labeled as such.

Related: Prompt caching economics, Benchmark and harness methodology.

Prepared with AI assistance. Source facts, original engineering recommendations, and hypothetical calculations are distinguished above. No live API calls or original model evaluations were performed.

RUN THE EXAMPLE

Keep the code close.

Python standard library. See the article for offline checks and live-integration limits.

Download context_budget.py
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology