MODEL COMPARISON

See the tradeoffs.
Side by side.

Compare configurations under the same units. The interactive controls can build your own shortlist.

Latest updates

This default comparison is an attributed historical selection. Different publisher harnesses remain separate; evaluation cost is not API token pricing.

EvidenceClaude Opus 5.5 (max with fallback)GPT-6 Astra (max)MiMo-V2.6-Pro
AA Intelligence Index57.622369810296352.67366939551346.3242065310383
Coding · FrontierCode54.4%R——
Coding · DeepSWE———
Coding · LiveCodeBench———
Reasoning · HLE61.4%54.7%49.4%
Reasoning · GPQA—96.1%—
Agents · BrowseComp———
Agents · Toolathlon———
Agents · MCP Atlas———
Price per task · Per task$5.98$3.26$0.13
Price per task · Per success———

Select configurations from the leaderboard

Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology