Caching economics · Pricing

GPT-6 Prompt Caching Economics: Engineer for Reuse, Measure Accepted Work

A rigorous guide to GPT-6.1 Sol prompt caching, Claude and Gemini cache economics, break-even calculations, prefix design, and tested cost accounting.

The most useful question about prompt caching is not how large the advertised discount is. It is how much reusable context survives until the next useful request, and whether that reuse reduces the cost of completed work. A model with an excellent cached-input price can still be expensive when every request writes a different prefix, produces long answers, or needs repeated repairs. A model with a higher ordinary-input price can be economical when a stable reference pack serves many accepted tasks.

This article develops a practical accounting model for GPT-6 and GPT-6.1 Sol, compares the relevant Claude and Gemini mechanisms, and turns the arithmetic into engineering decisions. All workloads and hit rates below are explicitly hypothetical. They are calculations, not measured provider performance or independent model-quality benchmarks. The accompanying Python calculator runs offline and was tested locally; no paid API request was made.

The current pricing facts, separated from the analysis

The direct OpenAI model pages list the following standard text-token rates. For GPT-6.1 Sol, requests above 272,000 input tokens have doubled input and cache rates and 1.5 times output rates across the full request. The page also describes processing-tier and regional modifiers. Our numerical examples remain below that threshold and use standard processing. GPT-6.1 Sol model documentation, GPT-6 Sol model documentation.

Direct API model Ordinary input / million Cache write / million Cache read / million Output / million
GPT-6 Sol $2.00 $2.50 $0.20 $10.00
GPT-6.1 Sol $2.00 $2.50 $0.10 $10.00
Claude Sonnet 5.5 $2.00 $2.50 for 5 minutes; $4.00 for 1 hour $0.10 $10.00
Claude Opus 5.5 $4.00 $5.00 for 5 minutes; $8.00 for 1 hour $0.20 $20.00
Gemini 3.8 Flash, current promotional standard tier $0.75 Separate explicit-cache lifecycle; see storage discussion $0.075 $3.75

Claude figures are direct-platform prices. Gemini's listed 3.8 Flash promotion ends December 31, 2026: the corresponding input, cached-input, and output rates become $1.50, $0.15, and $7.50 on January 1, 2027. Its pricing table separately lists cache storage at $0.50 per million tokens per hour during the promotion, then $1.00. This table establishes token prices, not equivalent quality or equivalent tokenization. Claude pricing, Gemini Developer API pricing.

For current OpenAI generations, the documented minimum cacheable visible prefix is 1,024 tokens. Explicit mode uses prompt_cache_options.mode and content-block prompt_cache_breakpoint; prompt_cache_options.ttl supports "30m". Cache reads and writes are reported separately. Explicit mode without breakpoints does not cache, and a changing suffix beyond the chosen boundary receives ordinary input pricing. This is materially different from older OpenAI cache guidance. OpenAI prompt caching.

Claude Sonnet 5.5 and Opus 5.5 have a documented 512-token minimum, with five-minute and one-hour cache options. The five-minute lifetime starts at the request start, which matters for long streaming responses. Cache availability for parallel requests begins after the first response starts. Gemini's current guide lists a 4,096-token minimum for 3.8 Flash implicit caching, enabled by default; the Interactions API supports implicit caching only, while explicit caching requires generateContent. Claude prompt caching, Gemini context caching.

A bill has four token categories, not one discount

Represent each request with four disjoint counts: ordinary input, newly written cache input, cache reads, and billable output. If the corresponding per-million rates are p, w, r, and q, its modeled token cost is:

text
cost = (ordinary × p + written × w + read × r + output × q) / 1,000,000

Add storage, external tools, infrastructure, and other applicable charges separately. The partition is essential. Cache-write tokens do not also belong in ordinary input. If your adapter charges all input at the base rate and then adds the full cache-write rate, it double-counts that part of the request. Conversely, subtracting cached reads but ignoring newly written tokens understates the invoice whenever writes have a premium.

The adapter should retain original provider usage records. A normalized table makes cross-provider analysis possible, but the raw record lets finance or engineering reconstruct a disputed charge. Record model identifier, endpoint, processing tier, region, request timestamp, usage fields, and price-card version. Use a stable workload identifier rather than logging confidential prompt contents merely to group expenses.

Also distinguish two definitions of hit rate. Request hit rate asks whether a request reused any eligible context. Token hit rate asks what fraction of input tokens were actually read from cache. A request that reuses 1,024 tokens out of 100,000 has a request hit, yet almost all of its context remains outside that read. For budgeting, token counts and the write count are much more informative than a single green cache-hit indicator.

Neither definition is an acceptance metric. A fast, cheap response that fails validation still consumes money. The economic outcome needs both a trustworthy usage ledger and a trustworthy acceptance rule.

Worked example: a shared reference pack for 100 requests

Consider an application that sends a 20,000-token stable reference pack, a 2,000-token changing question, and receives 1,000 billable output tokens per request. Assume exactly 100 requests, identical token counts, GPT-6.1 Sol standard rates, no tools, no tax, and no regional modifiers. These are deliberately simplified assumptions so the comparison can be audited.

Without caching, input costs 100 × 22,000 × $2 / 1,000,000 = $4.40. Output costs 100 × 1,000 × $10 / 1,000,000 = $1.00. Total: $5.40.

With one successful prefix write and 99 full reads, the reference pack costs $0.05 + $0.198 = $0.248. Changing questions cost $0.40, and output still costs $1.00. Total: $1.648, a modeled saving of 69.48% of the complete token bill.

The stable-prefix portion alone falls from $4.00 to $0.248, a 93.8% reduction. Reporting that as the reduction in total application cost would be misleading. It excludes the changing question, output, and every non-token expense. Even under these favorable reuse assumptions, the output bill is over four times the cached-prefix bill.

Hypothetical cache behavior Prefix writes Full prefix reads Ordinary prefix requests Total token cost
No caching 0 0 100 $5.400
One warm epoch 1 99 0 $1.648
Five epochs 5 95 0 $1.840
Every prefix rewritten 100 0 0 $6.400

An epoch here means a modeled sequence with one write followed by reuse. It is our analytical unit, not a provider billing object. It could end because the prefix changes, retention expires, or a different deployment path no longer finds the previous entry. The five-epoch row assumes five such writes and 95 reads; it does not predict retention behavior.

The final row is the operational trap. Paying a premium to write content that never repeats is worse than processing it normally. Explicitly isolating the reusable material can avoid writing one-off questions, but a cache configuration is not a substitute for a reuse opportunity.

Break-even is a workload property

Suppose a fixed prefix has P tokens, appears in N requests, and is written once. Ignore every other cost because those parts are unchanged between the two alternatives. With w = 1.25p and r = 0.05p, caching is cheaper when:

text
P × [1.25p + (N − 1) × 0.05p] < P × Np
N > (1.25 − 0.05) / (1 − 0.05)
N > 1.263157...

Two actual appearances are enough in this simplified full-reuse case. This does not mean every two-request application saves money. The same prefix must be eligible, available, and successfully reused. The output, latency, and acceptance behavior must remain comparable. If the second request rewrites its prefix, the assumed read never happened.

For an arbitrary set of requests, let W be prefix writes, H full reads, and M ordinary prefix processing events. Against a baseline that processes all W + H + M appearances ordinarily, savings arise when:

text
H × (p − r) > W × (w − p)

At the Sol 6.1 rates, one write incurs an extra $0.50 per million prefix tokens and each read saves $1.90. Thus the full-read-to-write ratio must exceed approximately 0.2632. Misses processed ordinarily cancel out in this comparison. Misses that cause writes must instead be counted in W; confusing the two creates false confidence.

This expression also explains why write churn matters. A workload with numerous one-off requests and one highly reused prefix might still save money overall, but averages can hide costly tenants or document versions. Break the ledger down by reuse group. The useful optimization target is repeated content that generates actual reads, not a universal cache toggle.

Why the GPT-6.1 read discount is smaller than it looks

Return to the 100-request example and change only the GPT-6 Sol read price to $0.20. Its 99 reads cost $0.396, compared with $0.198 on Sol 6.1. All other modeled charges remain the same, so the totals become $1.846 and $1.648.

The read unit price halves; the complete token bill falls by approximately 10.73% in this workload. At 10,000 such campaigns, that difference is $1,980, provided the assumptions continue to hold. At low volume, migration engineering or regression risk could easily dominate the dollar benefit. At high volume, modest savings per task can justify serious optimization work.

Do not infer from matching price-card columns that Sol 6.1 and Sonnet 5.5 are interchangeable. This example holds token counts constant for algebraic clarity. Real adapters can have different tokenization, outputs, reasoning consumption, validation success, and tool-call patterns. Run the same tasks through each candidate, record actual usage, and compare the cost of accepted results. A price table is an input to that experiment, not its result.

Prefix design is an information architecture decision

A reusable prefix should contain material that is both stable and relevant: policy, output conventions, tool definitions, or a versioned reference pack. The changing task belongs after the boundary. This ordering supports a comprehensible prompt as well as reuse: the model receives general rules before the particular case.

Start with a versioned assembly contract. Define which components are rendered, their order, and how versions change. For example, a support assistant might assemble product policy version 17, tool schema version 4, and a locale-specific style guide before the user's account question. If every worker renders those components differently, apparently identical reference packs can become distinct sequences.

Canonicalize application-generated data when its order has no semantic meaning. Sort schema properties consistently; avoid random ordering of retrieved metadata; remove irrelevant request IDs from the stable block. Preserve meaningful order where it expresses precedence or chronology. Canonicalization should never silently reorder a legal policy or tool instruction whose ordering affects interpretation.

Put timestamps, user IDs, experiment assignments, and one-off permissions where they belong semantically, then consider the cache boundary. A per-request timestamp at the beginning of a giant reference pack is an expensive design when the rest is stable. However, authorization information must remain correct for the user and task. Never move or omit essential controls merely to increase the hit rate.

Treat content changes as versions rather than trying to conceal them. A new product price or policy revision ought to invalidate the old reference pack for that use. Cache efficiency should reward stable truth, not stale answers. Keep the source-document version in the assembly metadata and test that the model sees the current policy when a version advances.

The most important segmentation is often task-specific. An invoice extractor needs a different stable context from a code reviewer. Giving both a universal 80,000-token company handbook may generate impressive cache reads while making each request harder to evaluate. Measure whether smaller, relevant prefixes improve accepted-task cost even when the nominal hit rate falls.

Padding to the minimum: a calculation, not a default tactic

Adding useful reference material to reach a caching threshold can sometimes be rational. Adding meaningless text is a weak default: it increases context, may complicate reasoning, and consumes engineering attention. The comparison must include every added token across every request.

For a hypothetical 900-token relevant prefix, compare ordinary processing with an expanded 1,024-token prefix. Use ten total requests, full reuse, and Sol 6.1 rates. Ordinary prefix processing costs 900 × 10 × $2 / 1,000,000 = $0.018. Expanded caching costs 1,024 × ($2.50 + 9 × $0.10) / 1,000,000 = $0.0034816. The modeled token saving is real under these assumptions, but the absolute difference is under two cents.

At one million equivalent campaigns the same arithmetic becomes material. Before scaling, evaluate whether the added context changes accuracy, output size, or latency. A small deterioration in acceptance can overwhelm the prefix saving. Prefer useful examples or genuinely required instructions over filler, and never present this hypothetical calculation as a recommendation to pad every short prompt.

For very short instructions, minimum thresholds can make another model's caching mechanism relevant. Yet the threshold should remain one term in the decision. If changing providers requires different tooling, produces longer outputs, or increases retries, a smaller cache minimum does not automatically win.

Gemini storage adds a clock to the decision

Implicit reuse and explicitly managed cache storage are different operational commitments. An application choosing explicit storage needs to account for how much context is retained and for how long, as well as read charges and any applicable creation or ingestion charges. The current Gemini pricing page lists both cached-input and storage rates, while the current caching guide directs explicit users to generateContent. Those facts do not justify transplanting an OpenAI write/read formula unchanged.

For a storage-only sensitivity calculation, retaining 20,000 tokens for one hour at $0.50 per million token-hours costs $0.01. One cached read rather than ordinary input at the promotional Flash rates saves 20,000 × ($0.75 − $0.075) / 1,000,000 = $0.0135. Thus one such read covers that modeled hour's storage charge alone. It does not establish complete explicit-cache break-even: creation accounting, endpoint behavior, minimum eligibility, and actual lifetime must also be included.

For 24 retained hours, the storage-only charge is $0.24; 18 of those reads save $0.243, just enough to cover it. A document used once per day and a document used once per minute have very different economics even if their size is identical. Storage turns an engineering retention setting into a budget assumption.

Make retention a function of observed demand. A frequently queried reference manual can justify keeping a cache alive. A customer-specific document that is read once and abandoned should have a different lifecycle. Delete or expire cache objects according to the actual API and data policy, and reconcile their billed storage with the application inventory. An object that nobody remembers can still be a cost center.

The January 2027 change should be modeled explicitly in forecasts rather than discovered on an invoice. In this particular rate card, both read savings and storage rates double together, so the storage-only read count in the simplified example stays unchanged. The dollar exposure still doubles. Promotional prices affect forecasts even when some relative ratios remain constant.

One-hour Claude writes require a different reuse assumption

At the Sonnet 5.5 table rates, a one-hour write costs $4.00 per million prefix tokens, compared with $2.50 for five minutes. The extra $1.50 is paid to choose the longer retention option. For a 20,000-token prefix that difference is $0.03 per write.

Whether it is worthwhile depends on when readers arrive. Suppose a five-minute configuration causes three distinct writes during a sparse series of questions, while a one-hour configuration needs only one. The prefix-write expense is $0.15 versus $0.08. Under that hypothetical event trace, the longer option saves write expense; the associated reads must still be counted correctly, and the inference does not predict either provider's real traffic behavior.

If every request arrives within one five-minute episode, the longer write buys no modeled benefit. If requests arrive days apart, one hour may still be too short. Retention should be chosen from request timing, not the apparent prestige of the option. Analyze inter-arrival times, including weekends and delayed follow-up.

Streaming duration is part of this operational analysis. A user may think a follow-up arrived quickly after seeing the answer, while the lifetime already includes time spent producing that answer. Log request start times and response completion times separately. Then calculate eligibility from the documented lifecycle rather than from the UI impression of how recently the conversation was active.

A tested calculator you can adapt

The downloadable companion, examples/cache_economics.py, has no third-party dependencies and never contacts a provider. It uses decimal arithmetic, disjoint categories, nonnegative-count checks, and a campaign function that distinguishes writes, reads, and ordinary misses. It does not estimate tokens from characters, emulate cache eviction, or infer success rates.

python
from decimal import Decimal as D
from cache_economics import Rates, campaign

rates = Rates(D("2"), D("2.50"), D("0.10"), D("10"))
total = campaign(
    requests=100, prefix=20_000, suffix=2_000, output=1_000,
    writes=1, hits=99, rates=rates,
)
print(total)  # 1.648

Run the included file with Python to execute its tests and print the four campaign totals. Run the import example from the directory containing cache_economics.py. Tests cover the uncached baseline, warm reuse, five epochs, write churn, storage addition, invalid counts, and nonfinite monetary inputs. A test failure returns a failing process exit status. This validates the calculator's arithmetic and input checks. It does not validate provider invoicing, API request syntax, cache behavior, or model quality.

For production accounting, feed the calculator normalized usage counts instead of assuming a fixed prefix length. Attach the rate card selected for the actual model, processing tier, context bracket, and billing date. If a request falls into a higher-priced context bracket, apply that bracket according to the provider's documented rule; do not price only the tokens above the threshold unless the contract explicitly says to do so.

Evaluate caching on accepted tasks

Define acceptance before optimization. For extraction, acceptance might mean every required field matches a validated schema and a reference label. For coding, it could require tests passing plus reviewer approval. For a support draft, it could require factual correctness against current policy and successful human review. A superficial format check alone is rarely enough.

Then calculate:

text
cost per accepted task =
  (all attempts + repairs + tools + infrastructure + review cost)
  / accepted tasks

Suppose configuration A costs $100 for 100 tasks and accepts 90. Its cost per accepted task is approximately $1.111. Configuration B reduces the bill to $85 but accepts only 70: approximately $1.214. B is cheaper per attempted task and worse per accepted task. These invented numbers demonstrate why a caching experiment must preserve task quality.

Use paired task sets wherever feasible. Keep task inputs, validators, and time windows comparable; record output and reasoning consumption; include failed requests and repairs. Measure latency percentiles as well as averages, because a cold or rewritten prefix can affect the users who experience the slowest requests. Cache optimization should improve the product's useful throughput, not merely beautify a token dashboard.

Roll out with a causal measurement plan

Start with a ledger on the existing application. Establish how much cost belongs to stable input, changing input, writes, reads, output, and tools. Identify the largest reuse groups. A workload dominated by output needs a different optimization priority from one dominated by repeated reference input.

Choose one coherent change: for example, isolate a versioned reference pack behind an explicit boundary, or remove a request-specific marker from the stable assembly. Avoid simultaneously changing provider, model, prompt instructions, retrieval, and cache settings. If five things change, a lower bill cannot tell you which one helped or whether quality was traded away.

Compare a representative period or controlled cohort. Include quiet traffic, bursts, document updates, long answers, and human follow-ups. A benchmark that prewarms a cache and immediately fires identical requests measures a favorable reuse path. It does not describe the full production workload unless production actually behaves that way.

Review both successful reuse and expensive writes. If many prefixes are written only once, narrow the cached region or reconsider eligibility. If token hit rate is high but complete cost barely moves, examine output consumption. If billable cost improves while accepted-task cost worsens, investigate quality before shipping the change widely.

Governance belongs in this rollout because caches retain computation associated with context. Check the endpoint's data controls, residency, retention, access boundaries, and contractual requirements rather than equating an application-level cache key with an authorization mechanism. The OpenAI data-controls guide is the relevant official source for its platform policies. OpenAI data controls.

Conclusion

GPT-6.1 Sol's low cached-input rate creates a strong economic opportunity for stable, repeatedly used context. Separate reusable material from tasks, record reads and writes, and measure completed work. Claude's longer-write option adds a timing decision; Gemini's explicit-storage economics add token-hours and endpoint-specific lifecycle accounting.

Choose a configuration using a reconciled ledger, a quality gate, and a repeatable workload experiment.

All sources below were opened and accessed on 2026-10-08. These are provider documentation sources; this article makes no claim to independently measured model quality or production cache performance.

  1. OpenAI GPT-6.1 Sol model page — rates and model-specific modifiers.
  2. OpenAI GPT-6 Sol model page — predecessor's rates.
  3. OpenAI prompt caching guide — current eligibility, configuration, and accounting fields.
  4. Claude pricing — Sonnet/Opus write and read rates.
  5. Claude prompt caching — minimum prefixes and lifetime behavior.
  6. Gemini Developer API pricing — promotional rates and explicit storage prices.
  7. Gemini context caching — implicit eligibility and endpoint distinction.
  8. OpenAI data controls — platform governance reference.

Related reading in the LLM Scorebook editorial package: Prompt compaction, KV cache, and prompt caching, Model price wars: batch versus realtime economics, and Claude agent routing and cost per accepted task.

RUN THE EXAMPLE

Keep the code close.

Python standard library. See the article for offline checks and live-integration limits.

Download cache_economics.py
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology