THE INTELLIGENCE DESK
Latest updatesA clearer view of AI.
Four benchmark groups, original sources, and the conditions behind each result.
Selected snapshot: 8 October 2026. Reference rank uses the AA Intelligence Index within these 14 configurations. It is not a complete live market ranking. R marks vendor-reported setups; a dash means unavailable.
| Rank | Model configuration | Coding | Reasoning | Agents | Price per task | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| FrontierCode | DeepSWE | LiveCodeBench | HLE | GPQA | BrowseComp | Toolathlon | MCP Atlas | Per task | Per success | ||
| 1 | Claude Opus 5.5Anthropic · max with fallback | 54.4%R | — | — | 61.4% | — | — | — | — | $5.98 | — |
| 2 | Claude Sonnet 5.5Anthropic · max with fallback | 46.2%R | — | — | 55.0% | — | — | — | — | $5.46 | — |
| 3 | GPT-6 AstraOpenAI · max | — | — | — | 54.7% | 96.1% | — | — | — | $3.26 | — |
| 4 | Gemini 4 ArgonGoogle · high | — | — | — | 57.1% | — | — | — | — | $1.99 | — |
| 5 | GPT-6 AstraOpenAI · xhigh | — | — | — | 54.6% | 96.3% | — | — | — | $2.31 | — |
| 6 | Muse Spark 1.3Meta · max | — | — | — | 48.7% | 93.5% | — | — | — | $1.60 | — |
| 7 | Grok 4.7SpaceXAI · xhigh | — | — | — | 43.1% | — | — | — | — | $3.74 | — |
| 8 | MiMo-V2.6-ProXiaomi · as listed | — | — | — | 49.4% | — | — | — | — | $0.13 | — |
| 9 | Qwen3.8 MaxAlibaba · 0902 | — | — | — | 43.1% | 92.8% | — | — | — | $5.41 | — |
| 10 | GLM-5.3Z AI · max | — | — | — | 42.3% | 91.7% | — | — | — | $2.01 | — |
| 11 | Step 5 PreviewStepFun · as listed | — | — | — | 46.5% | — | — | — | — | $0.72 | — |
| 12 | Kimi K3Kimi · max | — | 67.5%R | — | 46.9% | 93.5% | 91.2%R | 76.5%R | 84.2%R | $2.00 | — |
| 13 | Ling 3.1 FlashInclusionAI · as listed | — | — | — | 39.4% | — | — | — | — | $0.99 | — |
| 14 | Gemini 3.8 FlashGoogle · high | — | — | — | 47.8% | 95.3% | — | — | — | $1.24 | — |
Read the ranking and benchmark methodology · Browse model references
ADVERTISEMENT