MODEL COMPARISONSee the tradeoffs.
Latest updatesSee the tradeoffs.
Side by side.
Compare configurations under the same units. The interactive controls can build your own shortlist.
This default comparison is an attributed historical selection. Different publisher harnesses remain separate; evaluation cost is not API token pricing.
| Evidence | Claude Opus 5.5 (max with fallback) | GPT-6 Astra (max) | MiMo-V2.6-Pro |
|---|---|---|---|
| AA Intelligence Index | 57.6223698102963 | 52.673669395513 | 46.3242065310383 |
| Coding · FrontierCode | 54.4%R | — | — |
| Coding · DeepSWE | — | — | — |
| Coding · LiveCodeBench | — | — | — |
| Reasoning · HLE | 61.4% | 54.7% | 49.4% |
| Reasoning · GPQA | — | 96.1% | — |
| Agents · BrowseComp | — | — | — |
| Agents · Toolathlon | — | — | — |
| Agents · MCP Atlas | — | — | — |
| Price per task · Per task | $5.98 | $3.26 | $0.13 |
| Price per task · Per success | — | — | — |
Select configurations from the leaderboard
ADVERTISEMENT