Benchmark methodology · Coding

Terminal-Bench 4.0 and SWE-bench: How to Compare Coding Agents Honestly

Understand Terminal-Bench 4.0, SWE-bench, harness effects, paired statistics, and cost per resolved task—with tested Python and current source evidence.

A coding-agent score describes an entire experiment. It combines a model with a task set, an agent program, tools, resource limits, and a grader. Treating the resulting percentage as a permanent property of the underlying model hides the variables that often decide whether a deployment succeeds.

That distinction matters especially during a release cycle. Claude Sonnet 5.5 and GPT-6.1 Sol are verified current releases, but their announcement pages emphasize different coding evaluations. A number from one cannot fill a blank cell in the other's benchmark table. The useful question is more specific: which configuration completes your kind of task, within your budget, under your acceptance policy?

This guide explains what Terminal-Bench 4.0 and SWE-bench can establish, how to read recent claims without manufacturing a ranking, and how to build a local comparison with useful uncertainty and cost measures. The worked dataset and all dollar amounts in the statistical example are synthetic. We did not run either frontier model, Terminal-Bench, or SWE-bench for this article.

1. The score's unit is a system

Imagine two agents using the same model. One gets a shell, a repository search tool, and enough memory to run tests. Another sees a truncated directory listing, repeatedly loses command output, and gets killed while installing dependencies. Their different scores would tell you something important about the agents, but would not isolate a change in model intelligence.

For practical evaluation, define a configuration as:

model × agent policy × tools × environment × task distribution × budget × grader

The expression is a reminder of interacting factors, not a mathematical multiplication. A stronger model may use a tool more effectively. A tool optimized for one model's calling conventions may disadvantage another. A long output allowance may let an agent recover from mistakes, while a poor stopping policy spends that allowance on repetitive exploration.

There are consequently two legitimate comparisons. A controlled-model comparison keeps the surrounding system as similar as the APIs allow and asks what replacing the model changes. A product comparison evaluates each product's complete preferred setup and asks what a user can buy. Mixing the two produces misleading conclusions: an excellent product result does not automatically show which base model is best in a common scaffold.

State your estimand before collecting results. If the intended purchase is an integrated coding product, stripping away its useful orchestration may answer the wrong question. If you are choosing an API model for an existing agent, allowing every competitor a completely different agent can also answer the wrong question.

2. What Terminal-Bench 4.0 changed

Terminal-Bench's official benchmark index dates 4.0 to August 28, 2026. The release note describes resource calibration, task repairs, and removal of saturated or problematic tasks. It sets a flat eight-hour agent timeout and identifies environment and task-set changes as breaking changes that require fresh trials. Its published dataset selector is terminal-bench/terminal-bench@4.0.0, using Harbor. Terminal-Bench release, benchmark index.

Those facts change how the score should be used. An eight-hour allowance measures what can be completed with considerable elapsed-time headroom; it does not establish a service-level guarantee for an interactive developer. Removing tasks changes the denominator and difficulty mix. Fixing instructions can change what information an agent receives. Resource calibration changes which strategies are feasible.

The version number therefore belongs in every headline, chart label, and stored result. A 3.0 score and a 4.0 score cannot be interpreted as a clean before-and-after model improvement. Even if the model were unchanged, the measurement changed. To measure improvement, rerun the old and new configurations on the same pinned release.

For a new evaluation, the official selector is useful:

bash
harbor run -d terminal-bench/terminal-bench@4.0.0

This is a documentation reference, not a command we executed. It omits the agent and model configuration needed for a substantive experiment; follow Harbor's current installation and configuration documentation and store the resolved versions. Harbor describes itself as a framework for evaluating and optimizing agents in container environments. Harbor.

The broader lesson is that benchmark maintenance is necessary. A broken task is not a useful permanent anchor. But maintenance requires versioned comparisons. A leaderboard that silently combines incompatible task releases creates apparent progress that nobody can attribute correctly.

3. SWE-bench answers a different question

The official SWE-bench site lists 2,294 instances for the original benchmark, 500 for Verified, and 300 for Lite. It describes Verified as a human-filtered subset and its default Bash Only view as models evaluated in the same mini-SWE-agent environment. These views must be named when quoting a score; “SWE-bench” alone is insufficient. Official SWE-bench leaderboards.

SWE-bench is particularly useful when the application is repository issue repair. The evaluation guide describes applying generated patches to repositories and executing tests in Docker. Its prediction format contains instance_id, model_name_or_path, and model_patch. Evaluation guide.

The grading API exposes fail-to-pass and pass-to-pass concepts. The first checks whether the targeted broken behavior is fixed; the second checks whether previously passing behavior remains intact. These categories help explain why a plausible-looking patch can fail evaluation. Harness API.

Passing the selected tests remains narrower than proving that a patch is production-ready. A patch might introduce a performance regression, an undocumented interface change, or a new failure outside the selected checks. That is an analytical limitation of test-based acceptance, not a claim that benchmark maintainers endorse those changes. Your release policy may need security review, broader tests, migration checks, and a maintainer's assessment.

Terminal tasks and repository issue tasks overlap in tool use but differ in distribution. An operations agent that repairs an environment may benefit from Terminal-Bench evidence. A bug-fixing assistant for an established Python repository may find SWE-bench evidence closer to its work. Neither automatically estimates performance on greenfield frontend creation, proprietary monorepos, or tasks that need stakeholder clarification.

4. A current score provenance table, not a fabricated duel

The following records verified claims from opened primary pages. “Provider-reported” describes the source of the claim, even when the underlying benchmark is maintained elsewhere. We have not independently reproduced these results.

Configuration or claim Benchmark and reported value Provenance What a comparison still needs
Claude Sonnet 5.5 Terminal-Bench 4.0: 70.6% Anthropic release page, September 28, 2026 Exact run configuration, trial artifacts, effort, budget, and uncertainty
Claude Opus 5.5 Terminal-Bench 4.0: 66.4% Same page; footnote identifies Xhigh as its highest score Matched budgets and acceptance of the provider's setup
GPT-6.1 Sol DeepSWE 1.1: OpenAI says it matches Astra at roughly one-fifth the cost OpenAI release page; no absolute score extracted here Primary evaluator artifacts and a matched Sonnet run
GPT-6.1 Sol Terminal-Bench Science 0.1: $5.47 average cost per task at maximum effort OpenAI release page Resolution rate and matched Science configuration; this is not TB 4.0
SWE-bench Verified Bash Only 500-task common mini-SWE-agent environment Official benchmark site Specific model entries, dated result export, model/agent versions

Sources: Sonnet 5.5 announcement, GPT-6.1 Sol announcement, SWE-bench.

The correct reading of the first two rows is that Anthropic's reported Sonnet result exceeds its reported Opus result on that named evaluation. It does not imply that Sonnet dominates Opus across all coding tasks. It certainly does not establish a Sonnet-versus-Sol winner. OpenAI's announcement also explicitly warns that its research environment or API evaluation can differ from production ChatGPT and that competitor results come from public reports.

Be equally careful with reasoning effort. “High,” “Xhigh,” and “Max” are provider controls, not calibrated units of computation across vendors. A matched effort label does not prove matched latency, token use, or dollars. Anthropic's release footnote also describes a case in which Sonnet's Max setting scored below Xhigh on FrontierCode because additional review behavior caused timeout or scope problems. More computation is not guaranteed to improve a task's acceptance rate.

5. Pin a run card before you run

A reproducible comparison starts with a compact run card. It should travel with the score rather than living in a forgotten notebook.

Field Minimum record Why it matters
Task identity Dataset release, split, exact instance IDs Detects changed denominators and selective exclusions
Model identity API ID, snapshot when available, provider, request dates Aliases and serving behavior can change
Agent identity Repository commit, prompts, tool schema, policy Separates model changes from orchestration changes
Generation controls Effort, temperature if supported, output limit Determines both search behavior and spend
Budgets Wall time, tokens, tool calls, retry allowance Prevents an unlimited system winning a constrained claim
Environment Image digest, CPU/RAM requests and limits, network policy Captures feasibility and infrastructure differences
Grading Evaluator commit, test definition, judge model if any Makes acceptance semantics recoverable
Outcomes Per-instance status, trajectory, patch, logs Supports paired analysis and error audit
Accounting Usage categories, tariffs, tools, infrastructure, fallback costs Distinguishes cheap tokens from cheap accepted work

Record both resource requests and resource limits when your scheduler distinguishes them. Guaranteed capacity and a kill threshold are different controls. Also record concurrency. Eight simultaneous agents competing for a host can see a different environment from one isolated agent, even if their nominal container specifications match.

Randomize or interleave execution order. Running all A tasks overnight and all B tasks during a provider incident confounds provider condition with configuration. If you cannot interleave, report time blocks and repeat an overlapping subset to assess whether the time difference matters.

Do not quietly revise prompts after inspecting test-set failures. Use development tasks to tune the agent, freeze the policy, then evaluate on held-out tasks. If you examine and tune against a public benchmark repeatedly, disclose that process. It is valuable engineering, but the resulting score is a weaker estimate of performance on fresh tasks.

6. Infrastructure is a confounder and sometimes a capability constraint

Anthropic's infrastructure study held a model, harness, and task set constant while changing resource headroom on Terminal-Bench 2.0. It reports that uncapped resources increased success by six percentage points over strict enforcement in its setup, and distinguishes reliability gains from additional resources enabling different solution strategies. Its SWE-bench experiment also found a smaller RAM-related score change. These are provider-run experiments on specified configurations, not a universal correction factor. Infrastructure noise study.

The practical inference is to distinguish two scenarios. In the first, a transient environment failure prevents a valid strategy from completing: for example, a container never starts. In the second, the agent chooses a strategy that exceeds your intended resource budget: for example, installing a large stack when a lightweight implementation would fit. Calling both “infrastructure errors” erases the capability question your application may care about.

Predefine handling. A confirmed platform failure can receive a bounded rerun under a symmetric policy. An agent that uses its whole token allowance is generally a budgeted task failure, not a reason to give that particular configuration free extra attempts. Ambiguous failures should remain visible and receive sensitivity analysis.

Report the total assigned denominator as well as the completed denominator. If A resolves 70 of 100 assigned tasks but you omit its 10 failed environments, reporting 70/90 makes its headline 77.8%. B's 75/100 could then appear inferior despite completing more assigned work. A conditional completion statistic can be useful; label it and show the full count beside it.

One counterintuitive implementation trap deserves its own warning. SWE-bench's evaluation guide says result caching uses run_id and instance_id and can reuse the first result even if the prediction diff changes. Use a fresh run ID when evaluating changed patches. Docker image caching and evaluation-result caching solve different problems. A fast rerun can otherwise look like a new measurement while reproducing a stored old one. SWE-bench evaluation guide.

7. Measure uncertainty with paired tasks

Suppose two systems each attempt the same 100 tasks once. A resolves 70 and B resolves 60. The ten-percentage-point difference is easy to calculate. Its evidential strength depends on which tasks they solved.

Consider this explicitly synthetic outcome table:

Outcome Tasks
Both resolve 50
Only A resolves 20
Only B resolves 10
Neither resolves 20
Total 100

The paired comparison uses the 30 discordant outcomes. An exact two-sided McNemar test conditions on those discordant cases and asks how unusual a 20-versus-10 split would be under equal directional probabilities. Here its p-value is approximately 0.0987. That is not a probability that A and B are equal. It describes how compatible this observed discordance is with that specific null model.

A paired task bootstrap resamples task rows while preserving A/B pairing. For this example, the accompanying script's fixed seed and 10,000 draws produce a percentile interval of 0 to 21 percentage points for A minus B. The bootstrap interval and exact test are different procedures; they need not agree at a sharp threshold. Neither converts a small convenience sample into a representative estimate of all software work.

The important interpretation is practical: the point estimate favors A, but the evidence remains uncertain and the sampled task mix is narrow. Decide whether the likely benefit is large enough to justify its expense, and collect more relevant observations if the decision is sensitive to the gap.

8. Runnable offline analysis

The companion paired_benchmark_analysis.py uses Python's standard library, accepts a CSV, rejects duplicate tasks and invalid pass/cost values, and reports paired uncertainty and cost per resolution. It was executed with Python 3.14 in this production workflow. Its self-test checks the synthetic totals, the exact p-value, reproducibility, and malformed inputs.

bash
python paired_benchmark_analysis.py --self-test
python paired_benchmark_analysis.py
python paired_benchmark_analysis.py results.csv --bootstrap 10000

A real input file has this shape:

csv
task_id,a_pass,b_pass,a_cost,b_cost
private-001,1,0,2.10,1.35
private-002,1,1,1.85,1.10
private-003,0,1,2.45,1.70

Those rows are illustrative. Insert audited trial outcomes and measured dollars, not the example values. The statistical core is compact enough to inspect:

python
discordant = a_only + b_only
tail = sum(math.comb(discordant, k)
           for k in range(min(a_only, b_only) + 1))
p = min(1.0, 2 * tail / (2 ** discordant)) if discordant else 1.0

The full file also bootstraps paired task differences. Its inference assumes independent sampled task rows with one preregistered trial each. Multiple issues from the same repository may be correlated. Repeated stochastic attempts on the same issue are also correlated. For those designs, preserve clusters and use a hierarchical or cluster bootstrap rather than feeding every attempt into this file as an independent task.

The code intentionally analyzes binary acceptance. Do not convert an arbitrary partial-reward score into pass/fail after seeing results. Predefine the acceptance threshold. For graded outcomes, preserve the reward and choose an analysis appropriate to its scale. The example also does not account for multiple comparisons: testing twenty candidates and highlighting the best p-value requires a separate selection and confirmation strategy.

9. Price the accepted outcome

Using the same synthetic sample, let A cost $2 per assigned task and B cost $1.20. Then:

A: $200 / 70 resolved = $2.86 per resolved task

B: $120 / 60 resolved = $2.00 per resolved task

A resolves more tasks, but B is cheaper per resolution. The two observations can both be true. If the value of an additional resolved issue is high, A may be the rational choice. If a human will cheaply handle remaining failures, B may be better.

“Cost per resolved task” is total spend divided by accepted outcomes. It must include failed trials. Charging only successful trajectories understates the spend needed to obtain those successes. Distinguish this retrospective ratio from the mean cost of successful trajectories; the latter answers a different question.

A routing policy requires paired evidence. In the synthetic table, sending B's 40 failures to A would recover the 20 A-only tasks if A's isolated performance carries over. Assuming the same $2 cost for each fallback attempt, total spend would be $120 + $80 = $200 and acceptance would be 80/100, giving $2.50 per accepted task. This is an illustrative estimate, not a tested routing result.

The assumption is substantial. A may see B's misleading patch, a modified repository, or a summarized diagnosis. That different starting state can improve or damage its success. A proper routing experiment must specify whether the fallback gets a clean environment or the failed trajectory and must measure the combined policy directly.

For a business decision, add review and rework. A $3 machine resolution that needs 40 minutes of maintainer repair can be more expensive than a $5 resolution accepted immediately. Store reviewer time, reopened issues, regression incidents, and abandonment. These measures take longer to collect but align the benchmark with the outcome you are buying.

10. Retries and pass@k change the product

One attempt, five independent attempts, and five attempts with an oracle choosing the passing patch are different services. A production user typically lacks the benchmark's hidden tests and may be unable to identify the successful candidate.

For a simplified model with independent attempt success probability p, the probability of at least one success in k tries is 1 - (1-p)^k. With p = 0.4 and three attempts, that is 78.4%. This is a mathematical illustration, not a forecast: real attempts can repeat the same failure, share context, and vary in cost.

Even if three attempts yield one correct patch, an automated selector may choose the wrong one. A deployable best-of-k policy should therefore evaluate candidate generation and candidate selection. Report selection tools, access to tests, total attempted spend, and elapsed time. An oracle-selected benchmark result is an upper-bound style capability observation under that oracle, not directly a production acceptance rate.

Retry infrastructure failures according to the frozen policy, and retain the original failure record. Do not keep rerunning difficult tasks until they pass and then call the run pass@1. Also distinguish agent-internal repair—one trajectory correcting its own code—from restarting the entire task with a fresh trajectory.

11. Build an evaluation that matches your deployment

For a repository maintenance assistant, start with issues resembling your actual backlog: the same languages, dependency policies, repository sizes, and review constraints. Include fixes that require tests, documentation, and interface care. Decide whether the issue's acceptance definition includes a maintainer approval, not just a selected test suite.

For an operations agent, include diagnosis under your network and permission policy. A benchmark result obtained with unrestricted downloads does not estimate success in an isolated production environment. Capture whether the agent escalates appropriately when a required credential is unavailable; blindly continuing can be worse than a clear blocked result.

For a frontend agent, add visual and interaction checks. A patch that builds can still have an inaccessible form, broken mobile layout, or a missing loading state. Use a fixed browser environment and preserve screenshots or recordings as review artifacts. Separate subjective design ratings from deterministic functional acceptance.

Budget a pilot before a broad run. Validate that gold or known-good outcomes pass, malformed outputs fail, logs are retained, and spend is captured. Then freeze the setup. Keep pilot tasks separate from the final test set.

Publish the preferred configuration, task population, acceptance policy, budget, and observed limitations as a decision record.

12. Editorial rules for a trustworthy scorebook

LLM Scorebook should store score provenance as structured data. Keep provider-reported results separate from benchmark-maintainer runs, independent evaluator runs, and our own measurements. “Independent” describes who ran the evaluation, not automatically how well it was controlled; methodology still needs inspection.

Preserve immutable evidence snapshots where licensing permits and timestamp every retrieval. A live leaderboard can change after a provider bug fix. Keep the original result date and the latest verification date separate. Annotate corrections rather than silently updating a historical article to imply the new evidence existed at publication.

Avoid a universal aggregate when benchmark versions and task definitions conflict. If publishing a composite, expose components, weights, missing-data policy, and sensitivity to those weights. A model missing an evaluation should be marked missing, not assigned zero or given an invented interpolation.

For this release cycle, the supportable editorial conclusion is deliberately specific: Sonnet 5.5 has a verified provider-reported Terminal-Bench 4.0 result; GPT-6.1 Sol has verified current coding and scientific-workflow claims on other named evaluations. These provider announcements do not themselves supply a matched comparison. Independent configuration-specific evidence is discussed in our Sol versus Sonnet coding article; it must retain its harness, effort, and fallback labels. Neither source should be expanded into a universal model ranking.

Conclusion

Terminal-Bench 4.0 and SWE-bench offer useful evidence when their task sets and run conditions match the question being asked. Their scores become misleading when version changes, agent behavior, budgets, or acceptance rules disappear from the comparison.

Choose the deployment outcome first. Pin the system, preserve task-level artifacts, analyze paired outcomes, and include failures in the cost denominator. Use public results to form hypotheses, then confirm the decision on a relevant held-out workload. The strongest benchmark article tells readers both what the evidence supports and exactly what experiment would settle the remaining question.

Sources and verification record

Research cutoff: October 8, 2026. The GitHub release reference awaits confirmation. Source facts are summarized briefly; the statistical examples, comparison protocol, and business calculations are LLM Scorebook's original analysis. No vendor models were benchmarked for this article.

  1. Terminal-Bench 4.0 release — release changes, timeout, version selector, and rerun requirement.
  2. Terminal-Bench benchmark index — release date and distinct benchmark family entries.
  3. Terminal-Bench v4.0.0 GitHub release — unconfirmed reference.
  4. Harbor framework — benchmark execution framework scope.
  5. Official SWE-bench leaderboards — dataset variants, counts, and common Bash Only scaffold.
  6. SWE-bench evaluation guide — prediction format, outputs, evaluation-result caching.
  7. SWE-bench harness reference — patch application, containers, grading workflow.
  8. SWE-bench harness API — grading and failure categories.
  9. Introducing SWE-bench Verified — primary subset background.
  10. Anthropic infrastructure-noise study — resource-enforcement experiments and their limitations.
  11. Introducing Claude Sonnet 5.5 — dated release, provider score claims, effort caveats.
  12. Introducing GPT-6.1 Sol — current release existence, named evaluations, provider provenance warning.
  13. GPT-6.1 Sol model documentation — current model identity corroboration.
RUN THE EXAMPLE

Keep the code close.

Python standard library. See the article for offline checks and live-integration limits.

Download paired_benchmark_analysis.py
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology