THE INTELLIGENCE DESK

A clearer view of AI.

Four benchmark groups, original sources, and the conditions behind each result.

Latest updates

Selected snapshot: 8 October 2026. Reference rank uses the AA Intelligence Index within these 14 configurations. It is not a complete live market ranking. R marks vendor-reported setups; a dash means unavailable.

RankModel configurationCodingReasoningAgentsPrice per task
FrontierCodeDeepSWELiveCodeBenchHLEGPQABrowseCompToolathlonMCP AtlasPer taskPer success
1
Claude Opus 5.5Anthropic · max with fallback
54.4%R——61.4%————$5.98—
2
Claude Sonnet 5.5Anthropic · max with fallback
46.2%R——55.0%————$5.46—
3
GPT-6 AstraOpenAI · max
———54.7%96.1%———$3.26—
4
Gemini 4 ArgonGoogle · high
———57.1%————$1.99—
5
GPT-6 AstraOpenAI · xhigh
———54.6%96.3%———$2.31—
6
Muse Spark 1.3Meta · max
———48.7%93.5%———$1.60—
7
Grok 4.7SpaceXAI · xhigh
———43.1%————$3.74—
8
MiMo-V2.6-ProXiaomi · as listed
———49.4%————$0.13—
9
Qwen3.8 MaxAlibaba · 0902
———43.1%92.8%———$5.41—
10
GLM-5.3Z AI · max
———42.3%91.7%———$2.01—
11
Step 5 PreviewStepFun · as listed
———46.5%————$0.72—
12
Kimi K3Kimi · max
—67.5%R—46.9%93.5%91.2%R76.5%R84.2%R$2.00—
13
Ling 3.1 FlashInclusionAI · as listed
———39.4%————$0.99—
14
Gemini 3.8 FlashGoogle · high
———47.8%95.3%———$1.24—

Read the ranking and benchmark methodology · Browse model references

Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology