The easiest way to underestimate Haiku 5.5 is to multiply every input token by $0.10 per million. That rate applies to short prompts. A growing conversation, a cache write, or an expensive second attempt needs a different calculation.
Build the estimate from the request's billing categories, then add up the entire workflow. This guide supplies an offline calculator and concrete examples. It makes the assumptions visible so you can replace them with observed usage rather than treating a launch-day scenario as your future invoice.
At a glance
- Keep fresh input, cache reads, and cache writes separate; output and thinking consumption affect the final bill.
- A cache hit is cheaper than an uncached request, but the first write must be amortized across actual reuse.
- Compare workflows by accepted results, including both the initial attempt and any fallback.
Start with the rate card
| USD per million tokens | Prompt ≤100,000 | Prompt >100,000 |
|---|---|---|
| Uncached input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| Cache read | $0.01 | $0.05 |
| Cache write, 5 minutes | $0.125 | $0.625 |
| Cache write, 1 hour | $0.20 | $1.00 |
These are Anthropic's published rates. The longer band's values are five times the shorter band's. The boundary is a prompt-length condition, not the maximum context capacity. A one-million-token context window says what the model can accommodate; it does not promise short-band pricing for that capacity.
Our calculator applies the selected band to all token categories in a request. Cached content remains prompt content. The official caching documentation defines total input as fresh input plus cache creation plus cache reads. Check the Models/token-count API and billing records for the exact route before treating the calculator as an invoice reconciliation tool. Image handling, server tools, iteration usage, and provider-specific charges may need additional accounting.
Sources: Haiku rate card, cache accounting.
The boundary is a cliff in the estimate
Suppose output is fixed at 2,000 tokens and all input is uncached. At exactly 100,000 input tokens, the calculated charge is $0.010 + $0.001 = $0.011. At 100,001 tokens, the longer rates make it approximately $0.0500005 + $0.005 = $0.0550005. One extra token changes the estimate by approximately five times because the whole request uses the higher band.
That is a pricing scenario, not a live boundary experiment performed by LLM Scorebook. Before migrating a workload near the line, test token counting and billed usage on representative requests. Keep a margin for changing tool definitions, screenshots, retrieved documents, and conversation history. A prompt that usually lands at 99,900 tokens is hard to budget when normal variation can cross the boundary.
Compaction also has a price and a quality cost. A summary may discard a constraint, source identifier, or tool outcome. A compaction route earns its place only if its charge plus the later smaller requests is lower and the final acceptance rate stays adequate. Preserve authoritative records outside the summary so a worker can retrieve them when needed.
A cache hit saves 72% in this example
Take an 80,000-token reusable prefix, a 10,000-token new question, and 2,000 output tokens. A short-band cache hit costs $0.0028. With all 90,000 input tokens uncached, the same scenario costs $0.010. The calculated saving is 72% for the hit request.
The first five-minute write changes the comparison. Creating the 80,000-token prefix costs $0.010, and the new question and output add $0.002, for $0.012. The first request therefore costs more than its uncached counterpart. If the next request reuses the prefix within the cache lifetime, the two-request total becomes $0.0148 versus $0.020 uncached: 26% less overall.
For ten requests with one initial write and nine hits, the total is $0.012 + 9 × $0.0028 = $0.0372, compared with $0.100 uncached. That is a 62.8% reduction across the sequence. It is lower than the 72% hit saving because the write remains in the total. Extra writes caused by expiration or prefix changes reduce the gain further.
Source for prices: official specification. These examples are calculations with fixed token counts and no tools.
Break-even depends on successful reuse
Consider only the reusable prefix and normalize its ordinary input price to one unit. A five-minute write costs 1.25 units and each hit 0.10. For n requests with one write and n−1 hits, cached cost is 1.25 + 0.10(n−1). Uncached cost is n. Caching wins when n > 1.277..., meaning two total requests are enough under these assumptions.
For a one-hour write, the equivalent expression is 2 + 0.10(n−1). It wins when n > 2.111..., meaning three total requests. This is not a universal cache rule: it assumes the same band and prefix, no expiration, no competing writes, and one successful initial creation. If there is no reuse, there is no amortization.
The current Claude caching documentation lists a 512-token minimum for Haiku 5.5 on the Claude API and named Claude Platform routes. It separately directs Bedrock users to AWS documentation. Do not transfer a platform minimum blindly. Read cache_creation_input_tokens and cache_read_input_tokens; both zero means the intended cache behavior did not occur.
Source: Claude prompt caching.
Routing math includes failed first attempts
For a short request with 10,000 uncached input and 2,000 output tokens, the calculated Haiku cost is $0.002 and the Sonnet 5.5 cost is $0.040. If every request first goes to Haiku and fraction e also gets a full Sonnet attempt, expected token cost is $0.002 + e × $0.040. At 10% escalation, that is $0.006 per request.
This arithmetic assumes every fallback uses the same token counts. A route that forwards the failed answer and tool transcript will use more input. A route that directly sends difficult tasks to Sonnet pays a different total because those tasks never receive the initial Haiku attempt. Make the implementation match the formula before presenting projected savings.
The threshold for success is separate from the cost model. Define accepted outcomes before looking at model outputs. For extraction, validate required fields and the source spans. For classification, use held-out labels and inspect error categories. For coding, run relevant tests and check scope. Count refusals, truncations, and empty replies as outcomes requiring a defined policy, not as free successful requests.
Source for tariffs: Claude models and rates. Routing policy is LLM Scorebook's proposed evaluation design.
Compare long prompts on equal assumptions
| Uncached input + output | Haiku 5.5 | GPT-6 Luna | DeepSeek V4.1 Flash off-peak |
|---|---|---|---|
| 10,000 + 2,000 | $0.002 | $0.002 | $0.0027 |
| 150,000 + 10,000 | $0.100 | $0.020 | $0.0285 |
| 400,000 + 10,000 | $0.225 | $0.0875 | $0.066 |
All rows are calculated with equal vendor token counts, standard processing, no cache or tools, and fixed output. DeepSeek's peak charge is twice the listed off-peak charge. Its official peak windows are 01:00–04:00 and 06:00–10:00 UTC on weekdays excluding Chinese public holidays. In India these correspond to 06:30–09:30 and 11:30–15:30 IST; determine weekday and holiday eligibility using the provider's UTC rule.
These examples do not compare identical text or model quality. Different tokenizers and reasoning strategies can change usage. OpenAI's documented long-prompt rule applies above 272,000 input tokens, doubling input/cache rates and multiplying output by 1.5 for the full request. A capacity comparison alone misses that economic difference.
Sources: Luna documentation, DeepSeek official pricing, Haiku rates.
Run the offline calculator
The companion examples/cost_calculator.py uses Python's standard library and decimal arithmetic. It calculates the scenarios above, including separate cache-write categories and the pricing boundary. From the package root:
python examples/cost_calculator.py --self-test
python examples/cost_calculator.py --input 10000 --read 80000 --output 2000
python examples/cost_calculator.py --input 100001 --output 2000
For the cached-hit command, expect 0.0028 USD. For the final command, expect 0.0550005 USD. The program is an estimate for these published rates, not an official billing SDK. It excludes tools, taxes, discounts negotiated with a provider, and the cost of running your own infrastructure.
Measure the final outcome
Use cost per accepted result = total charged attempts divided by accepted results. When no result is accepted, report the metric as undefined and retain the total spend; a zero success count is not a zero-cost result. Track review time separately so a cheap but labor-intensive route remains visible.
Run a small staged workload before a full rollout. Record latency percentiles, acceptance rate, fallback frequency, and total token charge, then inspect failures rather than only averages. Compare those results with your response-time and quality requirements. The best cost setting is the cheapest one that satisfies those requirements reliably.
Figures checked October 8, 2026. Prepared with AI assistance. The offline arithmetic was tested; no live API billing experiment was conducted.
Keep the code close.
Python standard library. See the article for offline checks and live-integration limits.
Download cost_calculator.py