A provider can cut its token rate in half while your application bill barely moves. Another can charge less per million tokens and still cost more for the same finished task. A batch discount can disappear when urgent retries duplicate work. A promotional price can make this quarter's architecture look economical and next quarter's forecast wrong.
The engineering response is to compare complete workloads under explicit deadlines and billing rules. This article explains how to normalize usage, model pricing discontinuities, schedule deferred work, and account for realtime audio. It complements our caching and routing guides without repeating their prefix or verifier formulas. Every numerical workload below is an original hypothetical example unless identified as a dated provider price. There are no invented independent model-quality results.
Current price changes are conditions, not universal discounts
Three verified facts illustrate why an undated price table is insufficient. Gemini 3.8 Flash's current standard input/output promotion is $0.75/$3.75 per million tokens through December 31, 2026, becoming $1.50/$7.50 on January 1, 2027. GPT-6.1 Sol's model page applies higher full-request rates above 272,000 input tokens. DeepSeek V4.1 Flash has different peak and off-peak rates. Gemini pricing, GPT-6.1 Sol model page, DeepSeek pricing.
For DeepSeek Flash, peak cache-miss input/cache-hit input/output prices are $0.30/$0.006/$1.20 per million; off-peak prices are half. Peak windows are 01:00–04:00 and 06:00–10:00 UTC Monday through Friday, excluding Chinese public holidays. Other hours are off-peak. In India those windows are 06:30–09:30 and 11:30–15:30 IST. This describes the published schedule, not a prediction of availability or equivalent quality. DeepSeek pricing.
OpenAI's Batch guide describes 50% lower costs and a 24-hour completion window. Its Flex guide describes slower, lower-cost processing with possible resource-unavailable errors; that specific error is not charged. Gemini Batch targets 24 hours, offers 50% standard cost, and currently uses generateContent. These are distinct service contracts. OpenAI Batch, OpenAI Flex, Gemini Batch.
Normalize the task, not the number of tokens
Tokens are billing units produced by a model-specific representation. The same document can have different counts across providers, versions, languages, and modalities. Equal input text does not establish equal token count; equal output length does not establish equal reasoning consumption. A fair task comparison starts with identical source material and acceptance criteria, then records actual usage separately for each configuration.
Suppose fictional model A charges $1 per million input tokens and $5 per million output tokens. A task consumes 10,000 input and 2,000 output tokens: $0.02. Fictional model B charges $0.80/$4, but the same task consumes 15,000 input and 3,000 output tokens: $0.024. B has cheaper rates and a higher bill for that task. These numbers illustrate normalization; they are not tokenizer measurements of any named model.
Keep at least three views in the experiment report: provider-billed usage, application work units, and accepted results. Provider usage explains the invoice. Work units such as documents, tickets, or repository tasks explain volume. Accepted results explain usefulness. A report containing only dollars per million tokens leaves the relationship between those views unknown.
For multimodal work, retain modality-specific usage. Do not flatten image, audio, video, and text into an assumed universal text token. If a model uses separate rates or conversion rules, preserve them in the ledger. Source-document size is useful operational metadata; it is not a substitute for billable usage.
Record output requirements explicitly. One configuration may generate a concise field extraction while another writes an explanation, repeats the document, and includes a long rationale. The comparison should not reward unnecessary verbosity simply because it comes from a cheaper output tariff. Give both candidates the same useful deliverable and evaluate that deliverable directly.
Build a billing ledger with enough dimensions
A reproducible ledger should identify the task, attempt, model version, provider endpoint, timestamp, processing tier, region, usage category, and price-card version. Store a final task outcome separately from attempt status. An HTTP success is not a validated document, and a timeout does not by itself prove that no provider work occurred.
Treat ordinary input, cache writes, cache reads, output, and storage as separate categories when the API exposes them. Attach external tool usage and infrastructure costs using the same task identifier. The detailed caching guide explains why overlapping input categories can double-count spend; here the principle is broader: every billable event should have one documented interpretation.
Use dated price-card records rather than editing one global configuration value. If a rate changes midway through a reporting period, historical events must remain priced at the appropriate version. Preserve effective date, context bracket, modality, and processing-tier rules. A forecast can intentionally use a future card, but it should say so.
Reconcile regularly against provider statements. Small rounding differences can be normal; missing tools, duplicate attempts, and wrong brackets are material. If your internal estimate cannot explain the invoice, do not present its optimization percentages as established savings. Investigate the gap before using the figure to pick a provider.
Context thresholds create discontinuities
When a provider changes rates for the entire request beyond a context threshold, adding one small block can create a large bill increase. A linear per-token spreadsheet will miss that discontinuity. The GPT-6.1 Sol page documents doubled input/cache rates and 1.5 times output rates for requests exceeding 272,000 input tokens. GPT-6.1 Sol model page.
As an illustrative standard-rate calculation, a request with 272,000 ordinary input tokens and 1,000 output tokens costs $0.544 + $0.01 = $0.554 at the short-context Sol rates. A request with 272,001 ordinary input tokens and the same output costs $1.088004 + $0.015 = $1.103004 under the documented long-context multipliers. These examples exclude caching, tools, and other modifiers; they illustrate a bracket boundary, not an invoice guarantee.
The right response is not blindly cutting every prompt below the threshold. Removing essential evidence can increase failure and repair cost. Instead, measure which material the task needs, reserve room for expected tool results, and evaluate selective retrieval or compaction with the same acceptance rule. Preserve exact source access where the model may need to recover a detail later.
Threshold risk grows in agents because history changes between calls. A session can begin below a bracket and cross it after one large tool response. Track prospective request size before each call, not merely the initial prompt. An alert should explain the next call's predicted bracket and the evidence responsible for the growth.
Promotional expiry belongs in the forecast
For Gemini's cited promotion, consider a hypothetical monthly workload of 100 million ordinary input tokens and 20 million billable output tokens, with no cache, tools, or other modifiers. The promotional standard token cost is $75 + $75 = $150. The January 2027 rates produce $150 + $150 = $300 for the same usage. This is arithmetic using a published future rate, not a forecast of traffic.
If volume grows 25% at the same time, the later bill is $375, or 2.5 times the earlier baseline. A dashboard that reports only the promotional rate change misses the volume contribution. Show a bridge separating price changes, volume, token mix, cache behavior, and additional services.
Keep commitment decisions separate from experiment decisions. Trying a low-priced model on a reversible evaluation set is inexpensive. Building a tightly coupled production pipeline can create migration costs. A promotion may justify the experiment without justifying a long-term architecture that assumes the price never changes.
Use a range for uncertain future usage. Report a low, central, and high scenario with explicit assumptions. When a future price is published, use it directly; when it is not, label a stress scenario as an assumption rather than implying an announced increase. A transparent range is more useful than a precise number resting on hidden guesses.
Time-of-day pricing is a scheduling opportunity
Off-peak pricing makes eligible deferred work cheaper when the application can choose its execution window. A nightly enrichment job and a human waiting for an answer have different flexibility. The price schedule should be one input to the queue, alongside deadline, dependencies, task value, and available capacity.
At the published DeepSeek Flash miss/output rates, an illustrative request with 20,000 uncached input tokens and 2,000 output tokens costs $0.006 + $0.0024 = $0.0084 at peak and $0.0042 off-peak. Ten thousand such requests cost $84 or $42, assuming the specified rates apply and usage is identical. This calculation establishes a possible token saving, not identical quality, latency, or success probability.
Record timestamps in UTC and display local conversions for operators. Public-holiday exceptions need an authoritative calendar and a documented update process. Do not infer them from weekend logic, and do not make a deployment host's local timezone silently determine price eligibility. For requests spanning a window boundary, verify the provider's charging rule rather than inventing whether start time, end time, or another event controls billing.
Scheduling can move demand into the same cheap period as everyone else. Measure queue delay and deadline success under that traffic pattern. A model that is economical at off-peak rates may still be unsuitable for a deadline-sensitive job if retries or waiting dominate. Savings should be reconciled from billed usage rather than assumed from the time the queue submitted a job.
Batch is a queueing policy with a discount attached
Batch APIs organize independent requests into asynchronous jobs. The main application design question is how much slack a task has: the interval between submission and the latest useful completion. A 24-hour service window is compatible with some nightly operations and incompatible with an answer promised in ten minutes.
Partition work by deadline and dependency. Keep urgent interactions synchronous; defer independent enrichment; batch evaluations that do not influence a live user until later. A report with ten independent source extractions may batch them together, but its synthesis depends on their outputs. Calculate the full workflow's critical path rather than treating all stages as one simultaneous request.
A batch result is not a transaction commit. First ingest the result, associate it with the task, validate it, and decide whether it remains current. A customer could change a record while a queued request is running. The result may be correct for its original snapshot and stale for the current system. Preserve source versions and reject stale commits deliberately.
Cache reuse inside batch work should be measured. Asynchronous scheduling means requests may not arrive in the favorable sequence assumed by a prewarmed synchronous experiment. A discount on eligible usage does not imply that every request reads a shared prefix. Avoid multiplying two best-case percentages without verifying their joint behavior.
Idempotent orchestration prevents duplicate work
Gemini's Batch guide explicitly warns that creating the same job twice creates two jobs: job creation is not idempotent. OpenAI's guide says results should be matched by unique custom_id, and cancellation can remain in progress while in-flight requests finish. These facts require application bookkeeping. Gemini Batch, OpenAI Batch.
Give the application task a stable identity and each attempt its own identity. Keep submission state, provider job ID, result state, validation state, and commit state separately. If a network interruption makes submission uncertain, consult persisted records and provider state before blindly creating another job. A request timeout is not permission to erase the possibility that the original submission succeeded.
At ingestion, deduplicate by the application task and source version. Multiple model results may exist, but only a valid current result should commit. Use an atomic state transition or database constraint appropriate to the system. Log the rejected duplicate so its cost remains visible even though its side effect was prevented.
Retry individual failed items where the API contract permits it, rather than rerunning an entire successful batch. Preserve failure classification: invalid input, expired job, transient provider error, and failed acceptance require different fixes. An invalid prompt retried unchanged can burn time without creating useful progress.
Cancellation is a state transition, not an instant rollback. Continue reconciling results and charges after requesting it. If a synchronous rescue attempt races the batch, decide which result is eligible to commit and how late arrivals are handled. The cancellation button should not make the ledger forget completed work.
Deadline economics: include rescue and stale output
For an original sensitivity example, assume synchronous completion costs $0.04 per task and a delayed alternative costs $0.02. Suppose 15% require a synchronous rescue before their deadline. If both attempts are billed, expected model spend is $0.02 + 0.15 × $0.04 = $0.026 per original task.
Now add $0.002 of queue/orchestration expense and an expected $0.006 of stale-result review or repair. Total becomes $0.034, leaving a $0.006 saving against the simple synchronous baseline. Whether this remains preferable depends on quality, deadline coverage, and whether the baseline also incurs review. Keep comparisons symmetric.
Under these assumptions, the rescue share at which the delayed policy reaches $0.04 is (0.04 − 0.02 − 0.002 − 0.006)/0.04 = 0.30. That is a calculated threshold for the invented workload, not a general batch recommendation. If failures are uncharged, tools differ, or rescue quality changes, use the corresponding event costs instead.
A queue should reserve time for validation and recovery. Scheduling completion at the exact business deadline leaves no time to discover a malformed result. Learn a conservative cutoff from observed latency distributions and operational requirements, then monitor misses. A mean turnaround is inadequate when a small late tail breaks important workflows.
Flex has a different failure-cost path
OpenAI Flex can return an uncharged resource-unavailable error, according to its official guide. This differs from the hypothetical fully billed delayed attempt above. Model the specific events: a refusal for lack of resources can cost no model usage, while a completed answer followed by a second attempt can generate two usage records. OpenAI Flex.
If a fictional flexible attempt costs $0.02 when processed and is unavailable 20% of the time, with those cases rescued at $0.04, expected model cost is 0.80 × $0.02 + 0.20 × $0.04 = $0.024. Do not instead add the flexible charge to every unavailable case. The point is event-based accounting, not a prediction that the unavailable rate is 20%.
Also include waiting and retry overhead. An uncharged error can still consume application resources and user time. Configure bounded retries and explicit fallback deadlines. If the SDK retries automatically, record the actual attempt behavior so an application-level retry does not unknowingly multiply the total.
Realtime audio is a capability and session budget
The OpenAI Realtime guide describes direct speech-to-speech interaction with conversation state and tools, using browser WebRTC or server WebSocket connections. Gemini Live likewise exposes a stateful audio interface. These are different systems from streaming text, and neither should be priced by substituting a text-token rate for a complete voice session. OpenAI Realtime, Gemini Live API.
For synchronous text, measure request input, billable output, tools, and completion latency. For audio, also preserve session duration, modality-specific usage, generated speech, interruptions, and telephony or media-transport charges where applicable. A voice application may delegate selected reasoning to a text model, but that introduces its own context transfer and waiting path.
Compare products by resolved interactions. One voice system may require fewer turns but longer speech; another may be terse but repeatedly misunderstand the user. Cost per minute measures capacity expense, while cost per resolved call measures usefulness. Both can be valuable, but they answer different questions.
Do not confuse time to first audio with time to a verified outcome. A quick acknowledgement can hide a slow backend. Track handoff delay and abandonment. The ability to interrupt is valuable only if the application handles partial speech, pending tools, and canceled actions coherently.
Tools, infrastructure, and review can dominate
A hypothetical 10,000-task workload has $100 model expense, $200 external tool expense, $50 storage and network expense, $150 infrastructure, and $500 human review. Its total is $1,000. Cutting model expense by half saves $50, or 5% of the total. This allocation is invented to illustrate composition, not an industry estimate.
Measure tools by their actual charging units. A search request can initiate several billable queries; a browser session can create compute and storage costs; a document pipeline can repeatedly download or embed the same asset. Keep those events attached to tasks so moving between models does not make them disappear from the comparison.
For self-hosting, include idle capacity, redundancy, deployment work, memory constraints, and peak concurrency. A hardware cost divided by theoretical maximum tokens is not an observed unit cost. For managed endpoints, include contract commitments and unused reserved capacity where relevant. Neither approach wins solely because one price line looks lower.
Review cost deserves the same discipline. Record actual time and reasons for intervention. If a cheaper configuration needs more manual correction, the cost shift should appear in the ledger. If a better evidence pipeline reduces review while preserving quality, that is a legitimate saving even when model spend rises slightly.
Make a purchasing decision with evidence
Evaluate representative task samples across language, document size, modality, deadline, and failure modes. Hold acceptance criteria constant. Record all attempts and actual usage. Then compare total cost, useful coverage, deadline success, and operational reliability rather than declaring a universal provider winner.
Present forecasts with sensitivity ranges for volume, outputs, bracket crossings, rescue rates, tool usage, and review. Use published future prices where available. Preserve assumptions alongside results so another engineer can reproduce the calculation after a rate change.
Prefer reversible experiments before coupling critical infrastructure to a promotion or schedule. A provider adapter and clear task contract can make subsequent changes easier, but abstraction is not free: every supported endpoint still needs capability and billing tests. Keep the abstraction honest about differences.
Conclusion
LLM price competition creates real opportunities, especially for deferred and repeatedly structured work. Real savings depend on how the application uses the service: its tokenizer counts, output behavior, context brackets, schedule, deadline recovery, and non-model expenses. Batch, Flex, synchronous text, and realtime audio are separate operational choices with separate event costs.
Choose a configuration that delivers trustworthy work within the deadline at a reproducible total cost. A dated price card begins the analysis. A complete ledger and a representative workload finish it.
Sources and related reading
All primary pages were opened and accessed 2026-10-08. Calculations are hypothetical except for explicitly cited dated tariffs.
- OpenAI Batch API.
- OpenAI Flex processing.
- OpenAI Realtime guide.
- OpenAI GPT-6.1 Sol.
- Gemini pricing.
- Gemini Batch.
- Gemini Live.
- DeepSeek pricing.
Related reading: Prompt caching economics, Agent routing and cost per accepted task, and Context compaction economics.