Model comparison · Coding

GPT-6.1 Sol vs Claude Sonnet 5.5: Choosing a Coding Agent Without a Fake Winner

A rigorous comparison of current coding evidence, token tariffs, reasoning budgets, and deployment evaluation for GPT-6.1 Sol and Claude Sonnet 5.5.

GPT-6.1 Sol and Claude Sonnet 5.5 are verified current models with serious coding-agent claims. Their public announcements do not, however, supply a single clean experiment that proves one is the better coding model for every deployment. The defensible comparison starts with the task, the agent around the model, and the price of getting a change accepted.

That approach is especially useful here because the current standard tariffs are closely aligned. At the prices checked for this article, both list $2 per million ordinary input tokens and $10 per million output tokens. Current model/pricing documentation also lists $0.10 cached input and $2.50 cache writes for each, with Sonnet's latter figure applying to five-minute writes. Equal rows in a price table do not produce equal task bills. Models can read different amounts, generate different amounts, invoke different tools, and fail on different issues. GPT-6.1 Sol model page, Claude pricing.

The thesis is simple: use public coding results to identify plausible candidates, then choose between complete, budgeted agent configurations on held-out work. This article lays out the verified evidence, explains the gaps, and provides a concrete decision protocol. We did not call either paid API or run a live model benchmark. All workload totals and acceptance counts in the worked examples are hypothetical.

What is verified today

Anthropic's launch page dates Sonnet 5.5 to September 28, 2026 and identifies claude-sonnet-5-5 as its platform model name. It describes a model suited to well-scoped work while reserving its strongest claims about complex sustained judgment for Opus 5.5. Sonnet announcement.

OpenAI's release page confirms GPT-6.1 Sol, and its current model documentation identifies gpt-6.1-sol. The model page lists a 1,050,000-token context window, 128,000 maximum output tokens, and supported reasoning efforts from Low through Max, including Xhigh. It recommends the Responses API for tool calling; Chat Completions is supported without tools. Sol model documentation.

These facts describe availability and interface constraints. A large window does not show that an agent will retrieve the right file or preserve the decisive requirement after a long trajectory. A maximum output allowance does not show that spending it is helpful. A model name also does not identify the surrounding product: an API integration, a managed coding product, and an internal benchmark scaffold can behave differently with the same underlying model.

Before migrating, make an inventory of your current system. Record shell behavior, repository search, file editing, test execution, output truncation, and approval boundaries. The model replacement is successful only if those interfaces continue to carry the information and actions your workflow requires.

The coding evidence is real but distributed

Here is the evidence this article can verify from opened primary pages. The table preserves benchmark identities and source attribution rather than merging them into a universal score.

Model and evidence Verified public claim Interpretation boundary
Sonnet 5.5, Anthropic launch 70.6% on Terminal-Bench 4.0 Provider-reported terminal-agent result; not a direct Sol 6.1 comparison
Sonnet 5.5, same source 55.5% on CursorBench 4.0 Different task population and grading protocol
Sonnet 5.5, same source FrontierCode 1.1 Main: 46.2% Max, 52.1% Xhigh Effort can change agent behavior and acceptance
Sol 6.1, OpenAI launch Matches Astra on DeepSWE 1.1 at roughly one-fifth the cost Provider claim; not a Sonnet 5.5 result on matched tasks
Sol 6.1, same source Scientific workflows on Terminal-Bench Science 0.1; $5.47 mean task cost at maximum effort Science 0.1 is a separate evaluation from Terminal-Bench 4.0

Sources: Anthropic, OpenAI. These are provider-reported claims, not LLM Scorebook measurements or independent replications.

DeepSWE's primary evaluator site supplies the named evaluation context. Terminal-Bench's official benchmark index separately lists its standard and Science releases. That distinction matters when a chart says “Terminal-Bench” without the suffix: two numbers can share a family name and still measure different tasks. DeepSWE evaluator, Terminal-Bench index.

Do not subtract 70.6% from an unrelated DeepSWE percentage and call the difference a coding gap. A resolved rate is conditional on the task distribution and acceptance test. A model may excel at long repository changes while struggling with terminal environment repair, or perform well on narrow edits while making too many out-of-scope improvements on a maintainer-style evaluation.

An independent comparison exists, with narrower meaning

Artificial Analysis's comparison page, accessed October 8, reports 56% on Terminal-Bench 4.0 for GPT-6.1 Sol (Max) and 64% for Sonnet 5.5 (Max, Default Fallback). Its methodology lists 66 tasks and three repeats, scored as pass@1. These are independent evaluator results for those configurations, not the provider launch numbers. Independent comparison, intelligence methodology.

The same page shows $0.72 versus $5.46 “cost per task,” but this is a weighted Intelligence Index task cost, not a Terminal-Bench-specific bill. It cannot be divided by those terminal scores to calculate terminal cost per resolution. Preserve the Sonnet fallback label; this is not a claim about a no-fallback configuration. The methodology also estimates caching economics using typical live cache-hit rates. Cost definitions.

This supports a useful narrow observation: Sonnet's displayed terminal result is higher at the named settings, while the broader suite's displayed cost favors Sol. It does not settle your budgeted repository workflow. Task-level artifacts and a relevant held-out experiment remain valuable.

Why Max effort can make a patch less acceptable

A developer might reasonably expect more thinking to help. In agentic coding, the additional budget can change the policy: more exploration, more tools, additional reviews, or broader edits. Those actions can help difficult tasks while damaging a carefully bounded one.

Anthropic's Sonnet release footnote says its Max FrontierCode result falls below Xhigh, describing extra review behavior and cases of timeout or edits beyond scope. Treat this as a provider's explanation of a particular evaluation, not a law that Max is always worse. Release footnotes.

An agent repairing a two-line validation bug should not restructure the surrounding subsystem merely because it notices architectural debt. An agent implementing a cross-service migration may need exactly that broader planning. Acceptance depends on the authorized objective, not the amount of apparently valuable work produced.

Compare effort as a budget choice. Measure accepted-task rate, generated tokens, tool calls, wall time, and reviewer edits at each relevant setting. Provider labels are not standardized units of compute. Sol Medium and Sonnet Medium can be practical product choices without being an experimentally equal amount of inference.

The official Claude effort guide is the implementation reference for its control; OpenAI's reasoning guide is the corresponding reference for its API. Use current model-specific support rather than assuming every general example applies to every model. Claude effort, OpenAI reasoning.

Compare model replacements and complete products separately

For an existing internal agent, a controlled replacement keeps prompts, tools, and environment as similar as possible. It asks whether swapping the model improves the service you already built. Some adapter differences remain unavoidable: preserve equivalent information and action semantics, and document the differences.

For a product procurement decision, each candidate should use its supported production configuration. A product's file search, context management, and review features are part of what the team is buying. Evaluating only the base model may omit the features that determine developer adoption.

Both experiments can be worth doing, but their conclusions should have different labels. A model replacement result estimates a change within one agent. A product result compares complete offerings. If Sonnet wins in one scaffold and Sol wins in another, inspect the interactions instead of averaging them away.

Tool schemas need particular care. A shell wrapper that returns giant unstructured logs can create avoidable cost; one that hides exit codes can induce incorrect recovery. A repository search tool that returns file paths with useful snippets may be easier to use than a raw text blob. OpenAI's tools guide documents its tool ecosystem, but tool availability alone is not evidence that a specific agent uses those tools correctly. Using tools.

For fairness, include the same task instructions and expose equivalent repository state. Do not give one candidate a clean checkout while the other inherits a half-applied patch unless that asymmetry is the routing policy you are intentionally testing.

A tariff snapshot with the right categories

The table below uses current standard first-party documentation accessed October 8, 2026 and assumes each request has at most 272,000 input tokens. GPT-6.1 Sol applies doubled input and cache rates and 1.5 times output rates above that threshold to the entire request. The table excludes fast modes, batch discounts, residency modifiers, tool fees, and negotiated terms.

Category, USD per million tokens GPT-6.1 Sol Sonnet 5.5
Ordinary input $2.00 $2.00
Cached input read $0.10 $0.10
Cache write $2.50 $2.50, five-minute duration
Output $10.00 $10.00
One-hour cache write Not established by this table $4.00

Sources: OpenAI model tariff, Claude tariff. Sonnet's launch page states an older $0.20 cache-read price; this comparison uses the current dedicated pricing documentation rather than assuming the announcement remains the operating tariff.

Write and read categories must be reconciled with provider usage fields. Do not charge a token simultaneously as ordinary input and a cache write unless the provider's invoice explicitly defines an additional charge that way. Do not assume cache lifetimes, eligibility, or residency behavior are equivalent just because one rate matches.

The important economic variable is the measured token ledger. Large cached prefixes can make repeated reading cheap, but a model that generates twice as many tokens still pays substantially more output cost. More successful tool batching can shorten a trajectory. Better stopping behavior can avoid pointless work after the acceptance condition is already satisfied.

Worked costs: equal tariffs, unequal bills

Consider two hypothetical configurations on 100 assigned coding issues. Each usage total is already classified into disjoint billing categories, so no token is counted twice.

Synthetic aggregate usage Configuration A Configuration B
Ordinary input 10M tokens 10M tokens
Cache writes 2M 2M
Cache reads 80M 50M
Output 5M 8M
Accepted first-pass patches 70 75

Assuming every contributing request stays in the stated short-context band, using the common current tariff, A costs $20 + $5 + $8 + $50 = $83. B costs $20 + $5 + $5 + $80 = $110. Their token cost per accepted first-pass patch is $83/70 = $1.19 and $110/75 = $1.47. The letters do not stand for Sol and Sonnet. These are invented usage patterns demonstrating the calculation.

If five additional accepted patches are highly valuable, paying $27 more can be sensible. If failures are cheap to handle, A may be preferable. The task-value difference and human follow-up determine the decision; the cheaper accepted-patch ratio alone does not settle it.

Now add review. Suppose A's patches require five minutes of reviewer work each, while B's require three minutes, and review labor is valued at an illustrative $60/hour. Across the 70 and 75 accepted patches, that adds $350 and $225. The combined accepted-patch costs become $433/70 = $6.19 for A and $335/75 = $4.47 for B. The ordering reverses.

This calculation still omits review of rejected patches, infrastructure, tool charges, and later regressions. State those boundaries. A useful production ledger attaches every attempt and human intervention to a task identifier, then prices the final outcome rather than charging only the successful trajectory.

What an accepted patch should mean

A patch passes when its agreed acceptance criteria pass. That sentence sounds obvious until a benchmark reports success and a maintainer rejects the change for unrelated edits, missing tests, or a new compatibility problem.

For a bug repair, define the targeted behavior and regression constraints before running the agent. Require a reproducer or suitable test, preserve existing public interfaces, and define whether documentation is necessary. For a feature, specify edge cases, migrations, and operational behavior. For a refactor, define allowed equivalence and performance expectations.

Use multiple layers of acceptance. Deterministic checks can include compilation, tests, lint rules, schema validation, and a clean application start. Human review can assess scope, maintainability, and ambiguity. A hidden test can be useful, but the same hidden test cannot also guide the agent during generation if the claim is an unaided test-set evaluation.

SWE-bench's Docker evaluator provides a concrete example of generated-patch testing. Its official site distinguishes original, Lite, and Verified task populations and a common-scaffold Bash Only view. Preserve those identities if you use it as external evidence. SWE-bench, evaluation guide.

Your acceptance policy may be stricter than a benchmark's. That is appropriate when a production regression has consequences absent from the test environment. Publish both the mechanical test outcome and maintainer acceptance so readers can see where disagreement occurs.

A concrete held-out deployment trial

Build a task pool from the work you want to delegate. Include routine bug repairs, moderately complex multi-file changes, and ambiguous issues that should trigger clarification. Stratify by repository and task type so a large easy repository does not dominate the result. Keep examples used to tune prompts separate from the final comparison.

Freeze two initial configurations, then assign both the same tasks in clean isolated environments. Set a per-task dollar cap, wall-time cap, and retry rule. Interleave executions across candidates to reduce time-of-day or incident confounding. Preserve trajectory, usage, patch, environment digest, and acceptance evidence.

Start with a pilot that validates the evaluation machinery. Confirm that a known-good patch passes and a deliberately broken change fails. Confirm that spend is recorded even when the agent times out. The pilot is for repairing the harness, not selecting the winning prompt after looking at held-out outcomes.

Then run the locked comparison. Report assigned tasks, graded tasks, accepted tasks, empty outputs, infrastructure failures, and ambiguous cases. Do not improve a headline by removing one candidate's difficult failures. When a platform failure is rerun, retain the original event and apply the same rule to both candidates.

Use paired task analysis. The task-level overlap shows whether one configuration uniquely solves issues the other misses. A modest average gap with many disagreements may support routing. A similar score where both fail on the same issues may provide little benefit from fallback.

The companion methodology article includes tested offline Python for paired bootstrap uncertainty and an exact discordance test. Its assumptions fit one preregistered binary attempt per independent task; repeated attempts or repository clusters need an appropriately clustered analysis. Methodology guide.

Evaluate routing as its own system

Routing can be based on task type, known difficulty indicators, or a failed acceptance check. It should be tested directly. A fallback that receives a clean checkout is different from one that receives the first agent's patch and diagnosis. The latter may save useful exploration or inherit a misleading premise.

Suppose a cheap first stage accepts 60 of 100 hypothetical tasks at $1 each. A second stage attempts the 40 remaining tasks at $3 each and accepts 20. The complete policy costs $100 + $120 = $220, accepts 80, and has a machine cost of $2.75 per accepted task. This is a scenario, not measured performance for either model.

The comparison needs a direct second-stage-only baseline. If that baseline accepts 85 for $300, routing saves $80 but loses five accepted outcomes. If it accepts 80 for $300, routing appears attractive under these assumptions. Different reviewer costs or latency can reverse the preference.

Measure whether a classifier's difficulty estimate is useful. A prompt's length alone may not identify the hard issue; a tiny change can require deep dependency reasoning. A routing policy should be calibrated on development data and frozen for confirmation. If every uncertain task escalates, the system can accumulate the costs of both candidates without enough benefit.

Long context helps only when the agent preserves the right state

Coding trajectories accumulate plans, logs, files, edits, and failed tests. Merely retaining all of them can bury the decisive requirement. Compaction can reduce noise while accidentally deleting a constraint that prevented an earlier bad approach.

Test context management as part of the agent. Use tasks with a constraint introduced early and required much later. Check whether the agent preserves the source of truth, distinguishes applied changes from proposed changes, and remembers which tests actually ran. A successful summary should carry operational state rather than a persuasive narrative of progress.

Cache layout also affects costs. A stable prefix may be reusable while volatile logs are not. Compaction can shorten requests but change the prefix and require a new cache write. The optimal policy depends on sequence length, repetition, storage behavior, and quality after summarization. It should be measured rather than inferred from the headline context window.

For a real repository, keep files as recoverable artifacts and summaries as navigation aids. A summary that says a module is fixed should not override the actual diff or test result. Agents need an inexpensive way to reopen evidence when uncertainty appears.

Latency and reliability belong beside accuracy

A developer waiting for a short edit values time to a useful patch. An overnight migration job may tolerate slower execution if acceptance improves. Record median and tail latency, not just a mean that hides rare long stalls.

Separate time to first action from time to accepted completion. A model can stream text quickly but spend many rounds producing the patch. A slower-looking model can finish with fewer tool iterations. A provider's generation-speed claim cannot settle that workflow comparison.

Also record refusal, unnecessary clarification, incorrect success declarations, and recovery behavior. A well-placed clarification can prevent wasted work on an ambiguous requirement; too many unnecessary questions can make the service unusable. Acceptance should reward respecting the requested scope as well as producing valid code.

Infrastructure deserves controlled treatment. Anthropic's published study demonstrates that resource headroom can change coding-agent outcomes in its tested setups. That makes CPU/RAM limits and concurrency part of the run card, not incidental background. Infrastructure-noise analysis.

The decision you can make now

If you already have a stable API agent, both are credible candidates for a measured replacement trial. Sol's documented Responses tool path requires adapter care if your current system uses tool-enabled Chat Completions. Sonnet's reasoning migration details also deserve checking before transferring old request parameters. Use provider-specific adapters with equivalent capabilities, then test the accepted-task result.

If you are buying a complete coding product, evaluate the actual product configuration your developers will use. Include repository onboarding, review workflow, environment recovery, and the amount of human steering needed. A product that fits the team well may be the better purchase even if an isolated benchmark favors another model.

For an immediate decision without matched internal evidence, choose a reversible pilot with a cost cap and explicit acceptance checks. Treat the launch tables as a candidate filter. Preserve a fallback to the current known configuration until the new system demonstrates value on relevant work.

Conclusion

GPT-6.1 Sol and Sonnet 5.5 have verified current coding evidence, but the inspected announcements emphasize different tasks and scaffolds. Their current standard token tariffs are close enough that trajectory length, acceptance, and review effort can determine the real price difference.

The strongest comparison does not need an invented universal winner. It needs a named workload, fixed budgets, paired task outcomes, and a cost ledger that includes failure and human review. Publish the configuration that wins that experiment, state its boundaries, and update the decision when the workload or model service changes.

Sources and access dates

All listed pages were opened October 8, 2026. Provider results are labeled; this article reports no independently executed model evaluation. Worked costs and task counts are original hypothetical calculations.

  1. GPT-6.1 Sol release — named coding/Science claims and evaluation provenance.
  2. GPT-6.1 Sol model documentation — identity, tools, window, effort, and current tariffs.
  3. Claude Sonnet 5.5 release — date, coding results, and effort caveat.
  4. Claude current pricing — current Sonnet tariff categories.
  5. DeepSWE primary evaluator — named evaluation context.
  6. Terminal-Bench official index — distinct benchmark families.
  7. OpenAI reasoning guide — reasoning implementation reference.
  8. Claude effort guide — effort implementation reference.
  9. OpenAI tools guide — API tool implementation reference.
  10. SWE-bench official leaderboards — variants and common scaffold.
  11. SWE-bench evaluation — generated-patch testing.
  12. Anthropic infrastructure study — controlled resource experiments.
  13. Artificial Analysis model comparison — independent configuration-specific terminal result and suite cost.
  14. Artificial Analysis intelligence methodology — task count, repeats, pass@1, and cache-cost estimation.
  15. Artificial Analysis methodology — weighted suite cost definition.
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology