OUR METHODOLOGY

Evidence before rankings.

The numbers are useful when you understand what they measure.

Latest updates

Name the measurement

Coding, Reasoning, Agents and Price per task contain individual subcolumns. No category average is invented. Versions and conditions travel with every observation.

Understand the reference rank

Reference rank uses raw AA Intelligence Index precision within 14 selected configurations, captured 8 October 2026. It remains stable on sorting and filtering. It is not a global or continuously live rank.

Separate producers and setups

R identifies vendor release observations. They can use different effort, tools, context and harnesses from independent evaluations. Do not average incompatible setups.

Keep cost units visible

Per task refers to the evaluator’s workload. Cost per successful task remains unknown without a matching success denominator and total charge. Per-million-token API prices answer another question.

Preserve unknowns

A dash is unavailable evidence. HLE and GPQA retain the captured AA protocol label; exact revisions and tool details should be checked in the cited source.

Read the editorial evidence

Articles distinguish vendor claims, independent observations, our own arithmetic and executed local tests. We do not infer a workload winner from a launch score.

Original AA leaderboard source · Editorial and corrections policy

Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology