Agent economics · Pricing

Claude Agent Routing: Cost per Accepted Task Beats the Cheapest Token

Build a defensible agent routing policy with conditional recovery rates, verifier errors, effort sweeps, batch deadlines, and a tested offline calculator.

A cheap first attempt can be an excellent architecture. It can also be an expensive detour that makes hard cases slower and hides incorrect answers behind a confident verifier. The difference is not the size of the model's input-price discount. It is the joint behavior of the task, first model, acceptance test, fallback, and deadline.

This guide develops routing economics for the current Claude Opus, Sonnet, and Haiku family without repeating the separate Haiku launch analysis. It also explains where batch discounts fit, why realtime voice is a different product constraint, and how to evaluate effort settings without assuming that a higher level is always better. The numerical examples are hypothetical sensitivity analyses, not model benchmark results. The supplied calculator was tested offline; no live API requests or original provider evaluations were performed.

Start with verified capabilities and a dated price card

Anthropic's current overview lists claude-opus-5-5, claude-sonnet-5-5, and claude-haiku-5-5. It describes text and image input, text output, and tool use for the current family. That supports a text-agent routing discussion; it does not establish a native audio-to-audio route. Claude model overview.

Direct Claude API model Ordinary input / million tokens Output / million tokens
Haiku 5.5, prompts up to 100,000 tokens $0.10 $0.50
Haiku 5.5, prompts above 100,000 tokens $0.50 $2.50
Sonnet 5.5 $2.00 $10.00
Opus 5.5 $4.00 $20.00

These direct-platform rates were checked on October 8, 2026. Cache charges and other modifiers are separate. The Haiku model page and pricing documentation distinguish its context brackets; do not build a universal ten-cent-input forecast. Haiku 5.5 documentation, Claude pricing.

Claude's batch guide describes asynchronous independent requests at 50% of standard API usage prices, with expiry after 24 hours when processing is incomplete. Batch does not support response streaming or synchronous fast-mode settings. Effort is a behavioral control rather than a hard token budget; the current defaults differ across these models. Claude batch processing, Claude effort.

Those source facts establish the available knobs. The rest of the routing policy must come from workload evidence. Provider descriptions of speed or intended use do not supply your acceptance rate, fallback recovery rate, or verifier accuracy.

Define accepted work before measuring savings

An accepted task is a result that meets the application's quality requirements, not merely one that the model finished. A JSON object can satisfy a schema and still contain the wrong invoice amount. A code patch can compile and still break the intended behavior. A support answer can sound professional and cite yesterday's policy.

Write an acceptance contract that can be audited. For extraction, specify field correctness and treatment of missing evidence. For coding, specify behavioral tests, repository constraints, and reviewer requirements. For search synthesis, specify source coverage and claim support. Decide whether abstention counts as success, partial completion, or failure, and keep that definition stable during the experiment.

Use two separate labels: actual correctness and operational disposition. Actual correctness is the best available ground truth established by tests, reference labels, or expert review. Disposition is whether the production verifier accepted, rejected, or escalated a result. Mixing the two lets the verifier grade itself. A system that automatically accepts every answer would appear perfect under that circular metric.

The basic objective is total spend divided by truly accepted tasks. Include failed attempts, verifier calls, fallbacks, tools, human review, and infrastructure in spend. Report the acceptance rate alongside this ratio: two policies can have similar unit economics while one leaves far more customers without a useful result.

If outcomes have materially different harm or value, add a separate expected-loss analysis. Avoid quietly treating a wrong marketing category and an unauthorized transaction as interchangeable failures. A simple unit-cost ratio is useful, but it should not erase business consequences.

A cheap model plus an oracle: the optimistic baseline

Imagine 1,000 independent extraction tasks. Each request contains 10,000 ordinary input tokens and generates 1,000 billable output tokens. At the short-context Haiku table rates, a first attempt costs $0.0015: 10,000 × $0.10 / 1,000,000 gives $0.001 input and 1,000 × $0.50 / 1,000,000 gives $0.0005 output. At the same deliberately fixed token counts, Sonnet costs $0.03 and Opus $0.06.

Assume Haiku correctly solves 80% of tasks, a perfect free verifier accepts precisely those, and Opus correctly solves 90% of the remaining tasks. These assumptions are invented to show the model, not to characterize Claude performance.

Haiku first attempts cost $1.50. Two hundred Opus fallbacks cost $12.00. Total token spend is $13.50. True accepted tasks are 800 + 200 × 0.90 = 980, yielding approximately $0.013776 per accepted task.

Sending all 1,000 directly to Opus would cost $60.00 at the fixed token counts. If its acceptance rate on the entire original population were hypothetically 95%, that policy yields approximately $0.063158 per accepted task. The cascade looks attractive in this invented setting because the first route is cheap, its failures are perfectly detected, and the fallback retains high conditional recovery.

An oracle is an analytical boundary, not a production component. It tells you what could happen under ideal detection; it does not demonstrate what will happen after you install another LLM as a judge.

Add verifier errors and the picture changes

Let a be first-model correctness, u the fraction of correct answers rejected by the verifier, and v the fraction of incorrect answers falsely accepted. Let b be fallback correctness among the actual escalated population. With an assumed perfect final disposition for fallback results:

text
good results kept = a × (1 − u)
bad results kept = (1 − a) × v
escalated share = a × u + (1 − a) × (1 − v)
true accepted share = good results kept + escalated share × b
spend per task = first cost + verifier cost + escalated share × fallback cost

The perfect final disposition assumption must be named. If fallback answers face another imperfect verifier, extend the model with its own confusion matrix rather than counting every fallback as safe. The companion calculator intentionally models only two stages and reports false acceptance from the first-stage gate.

Use a = 0.80, u = 0.05, v = 0.10, b = 0.90, first cost $0.0015, verifier cost $0.004, and fallback cost $0.06. The first stage keeps 76% correct results and 2% incorrect results. It escalates 22%. Fallback contributes another 19.8% truly accepted results. Total true acceptance is 95.8%, with 2% false acceptance remaining from the first stage.

Expected spend is $0.0015 + $0.004 + 0.22 × $0.06 = $0.0187 per original task. Divide by 0.958 to obtain approximately $0.019520 per truly accepted task. For 1,000 tasks, these are expected counts of 958 good accepted results, 20 bad accepted results, and 22 unresolved results. Actual counts fluctuate; fractions are not guarantees for a small batch.

Hypothetical policy True acceptance First-stage bad acceptance Spend / original task Spend / true accepted task
Cheap-first, free oracle, fallback recovery 90% 98.0% 0.0% $0.0135 $0.013776
Cheap-first, imperfect paid verifier, recovery 90% 95.8% 2.0% $0.0187 $0.019520
Same verifier, recovery 40% 84.8% 2.0% $0.0187 $0.022052

This table is a sensitivity analysis, not a ranking. It shows that verifier and conditional-fallback behavior are economic inputs as important as model prices.

The fallback sees a selected population

A common spreadsheet mistake multiplies first-model failure by the fallback model's overall benchmark accuracy. The fallback is not receiving a random sample of the original workload. It receives cases selected by both the first model's errors and the verifier's decisions.

Those cases may be unusually difficult, ambiguous, unfamiliar, or missing essential data. The two models may share failure modes. A fallback with excellent overall accuracy might recover only 40% of the escalated cases. In the table above, lowering conditional recovery from 90% to 40% reduces true acceptance from 95.8% to 84.8%, without reducing expected spend at all.

Estimate recovery on the actual routed subset. Save representative failed first attempts with their original evidence, run the fallback on them, and judge results independently. Preserve whether the fallback saw the first answer: an answer can help by supplying useful work or hurt by anchoring the repair on a mistaken premise. These are different experimental configurations.

Also measure the good cases the verifier rejects. Some are harmless extra expense; others become wrong during repair. A stronger model is not logically guaranteed to preserve a correct first answer. The two-stage expression above treats fallback correctness as a pooled conditional rate; a more detailed experiment separates rejected-good and rejected-bad populations.

Track task-family conditional rates. A router may be successful for clean invoices and unsuccessful for scanned multi-currency documents. An overall recovery figure can conceal the exact slice where escalation fails. The useful response is often improving evidence extraction or abstaining, rather than simply adding another model tier.

A verifier is a product component with its own budget

Use deterministic checks where they directly capture correctness: a sum must reconcile, an identifier must belong to the supplied set, or a patch must satisfy a behavioral test. These checks often have clear failure explanations and avoid paying a language model to rediscover a simple invariant.

But coverage matters. A unit test checks the behavior it exercises, not the entire specification. A schema validates shape, not factual provenance. Report which requirements the verifier actually tests. Add expert audit samples for requirements that cannot be reduced to a deterministic check.

For an LLM verifier, create a labeled set including plausible wrong answers, subtly unsupported claims, adversarially confident outputs, and correct but unusual answers. Measure the confusion matrix rather than a single pass rate. Then evaluate it on fresh tasks. Reusing the same examples while tuning the judge can produce an attractive number without credible generalization.

Calibrate thresholds to consequence. If accepting a wrong result is costly, a low false-acceptance target may require more human review or abstention. That is a legitimate architecture even if the automated acceptance percentage falls. The objective is useful outcomes under constraints, not maximizing the share that receives an automatic green tick.

Record verifier expenses separately. In the worked example, $0.004 costs more than the cheap first attempt. A complex judge that reads the entire history and writes a long rationale can become the dominant stage. Consider whether it needs all context, whether its rationale is necessary, and whether a targeted check can preserve accuracy more efficiently.

Model tier and effort form a joint policy

Routing only by model name is incomplete. A model at one effort setting can have a different cost, latency, and recovery profile from the same model at another setting. The official effort documentation says the control influences token expenditure and behavior, rather than enforcing a fixed budget. Treat settings as experimental configurations, not comparable universal amounts of computation. Claude effort.

Run a small effort sweep on representative tasks. For each model-setting pair, record full token usage, latency, tool actions, acceptance, and repair frequency. Keep explicit settings in the experiment manifest so a default change or model update does not silently alter the comparison.

A hypothetical Sonnet-medium policy might cost $0.03 per task and accept 90%; Sonnet-high might cost $0.05 and accept 95%. Their unit costs are approximately $0.033333 and $0.052632. Higher effort produces more accepted work in this invented example, but lower unit efficiency. Whether it is preferable depends on required coverage, downstream repair costs, and consequence of failure.

Add retries to the comparison before deciding. If lower effort frequently needs an expensive second attempt, its first-call saving may disappear. Equally, if high effort creates long explanations without improving correctness, the extra tokens are unnecessary. Measure the complete policy rather than treating first-call cost as its total.

Choose effort at task boundaries when possible. A classification subtask and a complex synthesis step need not share settings. However, changing configuration can affect context assembly and cache behavior; account for actual reads and writes rather than assuming the same warm prefix survives every policy transition. The separate caching guide handles this accounting in detail.

Batch discounts buy schedule flexibility

Batch processing is attractive for jobs whose deadline allows asynchronous completion: offline evaluations, nightly classifications, catalogue enrichment, or queued research preparation. It is not a discounted substitute for an interactive answer needed now. The official Claude guide describes 50% standard usage pricing and independent request processing, not a guaranteed immediate response. Claude batch processing.

Suppose a synchronous text task has a modeled token cost of $0.03. A batch-priced equivalent has $0.015 token cost, assuming the same usage. If 10% miss the application deadline and are rerun synchronously at $0.03, expected token spend becomes $0.018 per task before other costs. At 50% urgent reruns, it becomes $0.03; the nominal saving vanishes.

That calculation assumes both the batch attempt and urgent rerun complete and are paid, with the batch result arriving too late for the application's own deadline. Anthropic documents no charge for requests whose batch result is errored, canceled, or expired; those states differ from a successful but late result. Reconcile actual result states and usage. The calculation also excludes duplicate side effects. If both attempts can send a message or update a record, an idempotency mechanism must decide which result is allowed to commit. Batch result states.

Measure deadline acceptance separately from factual acceptance. A correct answer delivered after a settlement window can be operationally unusable. Define a completed task as meeting both requirements when the product needs both. In reports, show the two components so readers can distinguish model failure from scheduling failure.

For multi-step agents, batching the initial model call does not make the tool loop instantaneous. Each dependent result may require another round, a tool action, and a new request. Use batch for independent work or well-defined stages, then calculate the critical path for the actual workflow. A file full of requests is a scheduling mechanism, not proof of dependency-safe orchestration.

Realtime modalities change the unit of work

Realtime voice has constraints beyond short text latency: interruption, audio input and output, turn detection, connection handling, and a user who is waiting while speaking. Gemini's current Live API describes a stateful WebSocket interface with audio input and output and interruption support. That is a distinct capability surface from the text-output Claude family summarized above. Gemini Live API.

A voice application can route selected reasoning or retrieval tasks to a text model while maintaining a live conversation layer. But the bridge adds costs: transcription or retained transcripts, routing delay, audio synthesis, session management, and sometimes duplicated context. Do not compare a cheap text input rate directly with a complete voice session and call the difference a price war victory.

Define units suited to the product: a resolved support call, an accurate translated conversation segment, or a successfully completed spoken transaction. Include session duration, generated speech, interruptions, tools, and failed transfers where relevant. A per-minute view is useful for capacity; a per-resolved-call view is useful for business economics. Neither is interchangeable with cost per million text tokens.

Measure perceived delay at the important interaction points. A fast acknowledgement followed by a long wait may feel responsive while the task remains slow. Conversely, streaming partial speech can be valuable even before the final answer completes. Report time to first useful response, completion time, and fallback disruption separately rather than compressing them into one latency number.

Price wars do not remove engineering costs

Lower token prices can make a formerly impractical agent feasible. They can also shift the dominant expense toward tools, retrieval, validation, and humans. Recalculate the whole cost stack as rates change. A 50% reduction in the cheapest stage is not a 50% reduction in the product's cost.

Consider 1,000 tasks with $20 model spend, $30 external tool spend, and $100 human review. Halving model spend saves $10, or 6.67% of the $150 baseline. Improving evidence quality enough to reduce unnecessary review could be much more valuable, provided it preserves true acceptance. This is a hypothetical allocation illustrating cost composition, not an industry average.

Make migration economics explicit. A new cheap route may require prompt adaptation, SDK changes, testing, operational monitoring, and rollback readiness. Compare expected ongoing savings with that one-time work and the cost of regressions. Avoid treating engineering effort as free simply because it does not appear on the provider invoice.

Maintain dated price cards. A promotional discount, changed context bracket, or processing-tier difference can invalidate yesterday's optimal routing threshold. Reevaluate on a schedule and after material provider changes. Preserve prior cards so historical cost reports remain reproducible.

The tested offline routing calculator

The companion examples/routing_economics.py implements the two-stage expressions above. It accepts probabilities and per-attempt costs, validates finite values, and returns escalation, true acceptance, first-stage false acceptance, spend, and spend per truly accepted task. It does not estimate a model's accuracy or serve as a production router.

python
from routing_economics import route

result = route(
    cheap_success=0.80, reject_good=0.05, accept_bad=0.10,
    fallback_success=0.90, first_cost=0.0015,
    verifier_cost=0.004, fallback_cost=0.06,
)
print(result.true_acceptance)  # approximately 0.958
print(result.false_acceptance)  # approximately 0.02

Run the import example from the directory containing routing_economics.py. Its tests cover the worked example and low-recovery sensitivity, oracle case, always-reject gate, always-accept gate, zero acceptance, and invalid probabilities. Zero true acceptance returns infinite unit cost rather than dividing by zero. Six test methods passed locally; failed tests return a failing process exit status. No network traffic or API credentials were involved.

To use it responsibly, replace invented inputs with audited estimates from the same task distribution. Supply fallback correctness measured on routed cases. Attach uncertainty ranges and run sensitivity sweeps. If plausible parameter values change the preferred policy, collect more evidence before committing to a universal route.

For production, use actual counts rather than only expected probabilities. Log route decisions and final dispositions, keep a random audit sample, and track model and verifier versions. Preserve unresolved results as unresolved: excluding them from both spend and quality statistics makes a weak policy appear efficient.

Deploy a policy, not a leaderboard winner

Start with a baseline route whose behavior you understand. Build a task taxonomy from product requirements and observed failures. Add a cheap first route only where the acceptance gate can distinguish useful results from plausible mistakes, and where fallback recovery is demonstrated on the selected population.

Run the candidate policy in shadow mode or a controlled cohort. Compare accepted-task cost, false acceptance, coverage, deadlines, and latency tails. Include tools and human intervention. Before broad rollout, test malformed evidence, provider errors, missing context, duplicate requests, and fallback exhaustion; these are product states that average benchmarks often omit.

Set an explicit stopping rule. An agent should not retry indefinitely just because another model exists. Bound attempts and spend, define when to abstain or request human review, and retain sufficient evidence for the next operator to understand what failed. The fallback should be a deliberate recovery mechanism, not a recursion that hides uncertainty.

Keep permission checks outside the model's confidence judgment. A correct-looking answer is not authorization to send, charge, or publish. Separate generation, validation, and commit so a route experiment cannot accidentally turn duplicate attempts into duplicate external actions.

Conclusion

Claude Opus, Sonnet, and Haiku give an agent designer multiple price and capability choices. A defensible routing policy combines those choices with an independently evaluated acceptance gate, conditional fallback evidence, explicit effort settings, and the product's deadlines. Batch can reduce eligible asynchronous usage costs; realtime modalities impose their own capability and latency requirements.

The economic winner is the policy that delivers more trustworthy completed work for the available budget. Measure the failures it keeps, the failures it escalates, and the cost of recovery. The cheapest token is valuable only when it participates in a useful result.

All sources were opened and accessed on 2026-10-08. Provider facts are separated from original calculations; no independent Claude accuracy measurement is claimed.

  1. Claude models overview.
  2. Claude Haiku 5.5 overview.
  3. Claude pricing.
  4. Claude batch processing.
  5. Claude effort.
  6. Gemini Live API.

Related reading: GPT-6 prompt caching economics, Agent compaction, KV cache, and prompt caching, and the separate Haiku 5.5 launch article when its owning editorial session publishes it. This guide covers routing policy rather than reproducing that launch report.

RUN THE EXAMPLE

Keep the code close.

Python standard library. See the article for offline checks and live-integration limits.

Download routing_economics.py
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology