New to this topic? Which AI Is Best? Reading the Scoreboard Like a Pro is a one-hour workshop deck on this site that walks through everything on this page: what benchmarks measure, why they saturate, whose numbers to trust, and how to spot a wrapper app. It's the plain-language companion to the catalogue below.
There is no single "best AI". Capability is jagged; the same model can win a gold-medal score on Olympiad mathematics and still misread an analogue clock, so choose by task and treat every score as directional rather than final. Benchmarks also saturate in months, not years, and vendor-reported scores routinely run higher than independent ones; prefer independent leaderboards, note the date on every number, and distrust any single headline figure. One useful side effect: apps that quietly repackage a cheaper model behind a polished interface never publish independent benchmark results, so asking "where does this sit on Arena or the Epoch hub?" quickly separates frontier models from wrappers.
Reasoning & Knowledge
Humanity's Last Exam (HLE)
2,500 questions written by subject experts across roughly 100 fields, built to be the last closed-ended academic exam models can't pass; top models sat between about 44% and 53% in July 2026 and are climbing fast. A February 2026 audit (HLE-Verified) found a material error rate in the questions themselves, so read small score gaps sceptically. Release paper.
GPQA Diamond
198 graduate-level science questions designed to be "Google-proof"; even PhD holders outside the sub-field score poorly. Frontier models now sit at 93% to 95%, and much of what remains is label noise, yet it is still quoted in nearly every model launch. Release paper.
ARC-AGI (versions 2 and 3)
Fluid-intelligence puzzles that are easy for people and hard for models, testing novel-problem solving rather than memorised knowledge; version 3 moves to interactive game environments where frontier models scored under 1% at launch against near-perfect human play. The best official verified ARC-AGI-2 score was 54% in July 2026, so both versions still separate the field. ARC-AGI-2 paper; ARC-AGI-3 documentation.
MMLU and MMLU-Pro
The multiple-choice knowledge exam that defined "smart" for the ChatGPT era, and its harder 2024 successor; both now sit around 90% or above, and MMLU itself carries an estimated 9% label-error rate. Kept here because thousands of older claims cite them; treat any MMLU number as a history lesson. MMLU paper; MMLU-Pro paper.
Mathematics
FrontierMath
Around 340 unpublished research-level mathematics problems in four difficulty tiers, written by professional mathematicians; the version 2 release of June 2026 corrected errors an audit found in roughly a third of the problems. Also a benchmark-integrity teaching story: OpenAI funded its creation with dataset access under an undisclosed agreement, and Epoch now keeps a 50-problem holdout no lab ever sees. Release paper.
MathArena live competitions
Runs models on real mathematics competitions (IMO, USAMO, HMMT, Putnam) within hours of the problems being released, which makes it the cleanest contamination-free signal in mathematics; in July 2025 it certified a gold-medal IMO performance by Gemini Deep Think. Methodology paper.
AIME and MATH
The two staples of maths benchmarking for years: the American Invitational Mathematics Examination and the MATH competition-problem dataset. Both are now saturated at the frontier; five models scored 100% on AIME 2025, and MATH sits at 97% to 99%, so neither can tell strong models apart any more. MATH paper.
Coding
SWE-bench Verified
500 human-validated real GitHub issues the model must fix in actual Python repositories; the most-quoted coding number in the industry. Top agents now resolve roughly 85% to 95%, so its power to separate frontier models is fading. Release paper.
SWE-bench Pro
1,865 enterprise-grade software tasks across 41 repositories and four languages, with contamination-controlled splits. Standardised public scores sit around 59% while vendor-run claims land 10 to 30 points higher; a tidy demonstration of why the test harness matters as much as the model. Release paper.
Terminal-Bench 2.0
Real tasks completed inside a live terminal and container environment, from build failures to system administration; closer to actual developer work than issue-fixing alone. Top agents reached 83% to 85% by May 2026, so watch for saturation at the top. Release paper.
HumanEval
164 short Python function-writing problems; the benchmark that launched code generation as a research field. Solved years ago (96% to 98% and above) and kept here purely to show how far the goalposts have moved: from writing one function to fixing real enterprise codebases. Release paper.
Agents & Computer Use
OSWorld-Verified
369 real tasks in a live operating system: spreadsheets, browsers, file managers, multi-app workflows. Top agents now score 82% to 85%, above the roughly 72% human baseline, which says as much about the benchmark's ceiling as about the agents. Release paper.
tau2-bench
Customer-service agents that must follow business policy while using tools mid-conversation, scored on whether they succeed reliably across repeated runs rather than once. Headline splits are close to solved (97% to 99% on telecom), and aggregate indices have already moved to a harder successor. Release paper.
BrowseComp
1,266 questions built to require persistent, creative web research; agents scored 0.6% at launch and roughly 90% by July 2026, one of the fastest saturation arcs on record. A stricter successor (BrowseComp-Plus) restores the difficulty. Release paper.
Vending-Bench 2
A long-horizon business simulation: the model runs a vending machine company for a simulated year from $500 of capital, scored on closing net worth; it punishes the drift and forgetfulness that short benchmarks never see. In July 2026 the leading model finished with $10,937. Release paper.
METR time horizons
Rather than a pass rate, METR measures how long a task a model can complete at 50% reliability, expressed in human working time; the January 2026 update put the frontier at roughly a 320-minute horizon, doubling about every 89 days since 2024. The single best chart for showing a general audience how fast agentic capability is moving. Methodology paper.
Cybersecurity
Cybench
40 professional capture-the-flag security challenges with graded subtasks, drawn from real CTF competitions. Frontier models are closing in on the ceiling, but it remains a fixture of the safety sections in model system cards. Release paper.
CyberGym
Over 1,500 tasks built from real vulnerabilities in open-source projects: agents must reproduce known flaws and sometimes discover new ones. Cited directly in lab safety evaluations, including the Claude Opus 4.5 system card (50.63% pass rate, November 2025). Release paper.
ExploitBench
A capability ladder against 41 real Chrome browser vulnerabilities, scoring how far an agent climbs from reaching the flawed code to a full working exploit; at release the best public model achieved the top rung on just one bug. Only months old, so treat it as promising rather than established. Release paper.
Multimodal & Long Context
MMMU-Pro
College-level problems that require reading images, charts and diagrams together with text; the Pro version strips shortcuts that let models answer without truly looking at the picture. The standard citation for multimodal reasoning. Release paper.
RULER
Measures a model's effective context length: how much of the advertised window it can actually use before quality collapses, which is routinely far less than the marketing number. Discrimination between frontier models now happens at 256k tokens and beyond. Release paper.
MRCR
Multi-round co-reference: the model must track many similar items scattered through a very long conversation and retrieve exactly the right one. Cited by the major labs, and its hardest splits remain unsolved, which makes it a better long-context test than simple retrieval.
Needle in a Haystack
The original long-context test: hide one sentence in a mass of text and ask the model to find it. Every serious model now passes, so a perfect score means little; it survives as a sanity check and a reminder that retrieval is not comprehension. No vetted public link yet; the title is searchable.
Leaderboards & Aggregate Indices
Arena (formerly LMArena)
Millions of blind side-by-side votes from real users, ranked by Elo rating; the largest human-preference leaderboard, with the top labs now separated by only a handful of points. Read it alongside "The Leaderboard Illusion" (April 2025), which documented how private variant testing can flatter the rankings. The Leaderboard Illusion paper.
Artificial Analysis Intelligence Index
An independent composite score built from agentic, coding, scientific-reasoning and general benchmarks, re-run in-house on every major model with cost and speed alongside; it drops components as they saturate. A single convenient number, with all the caveats a single number deserves.
Epoch AI Benchmarking Hub
Independent trend charts across the major benchmarks, with costs and a note on who actually ran each evaluation; the most transparent place to see saturation curves and vendor-versus-independent gaps at a glance.
