ntworld.ink
Circuitry, a keyboard and a mouse holding a lit bulb
Resources · Scoreboards

AI Benchmarks

The tests used to answer "which AI is best?", grouped by what they measure. Each entry links the benchmark itself and its release paper, says what it tests, and carries a status flag: Watch means the benchmark still separates today's frontier models; nearing saturation means top scores are bunching near the ceiling and the number is losing meaning; Retired, historical means the test is solved and is listed here to show how the field got here. Scores quoted were current at the linked leaderboards in July 2026; they age quickly.

Jump to: Reasoning & Knowledge · Mathematics · Coding · Agents & Computer Use · Cybersecurity · Multimodal & Long Context · Leaderboards & Indices

Start here

New to this topic? Which AI Is Best? Reading the Scoreboard Like a Pro is a one-hour workshop deck on this site that walks through everything on this page: what benchmarks measure, why they saturate, whose numbers to trust, and how to spot a wrapper app. It's the plain-language companion to the catalogue below.

How to read this page

There is no single "best AI". Capability is jagged; the same model can win a gold-medal score on Olympiad mathematics and still misread an analogue clock, so choose by task and treat every score as directional rather than final. Benchmarks also saturate in months, not years, and vendor-reported scores routinely run higher than independent ones; prefer independent leaderboards, note the date on every number, and distrust any single headline figure. One useful side effect: apps that quietly repackage a cheaper model behind a polished interface never publish independent benchmark results, so asking "where does this sit on Arena or the Epoch hub?" quickly separates frontier models from wrappers.

Reasoning & Knowledge

Scale AI & Center for AI Safety · January 2025 · Watch

Humanity's Last Exam (HLE)

2,500 questions written by subject experts across roughly 100 fields, built to be the last closed-ended academic exam models can't pass; top models sat between about 44% and 53% in July 2026 and are climbing fast. A February 2026 audit (HLE-Verified) found a material error rate in the questions themselves, so read small score gaps sceptically. Release paper.

Rein et al., NYU · November 2023 · Watch, nearing saturation

GPQA Diamond

198 graduate-level science questions designed to be "Google-proof"; even PhD holders outside the sub-field score poorly. Frontier models now sit at 93% to 95%, and much of what remains is label noise, yet it is still quoted in nearly every model launch. Release paper.

ARC Prize Foundation · v2 March 2025; v3 March 2026 · Watch

ARC-AGI (versions 2 and 3)

Fluid-intelligence puzzles that are easy for people and hard for models, testing novel-problem solving rather than memorised knowledge; version 3 moves to interactive game environments where frontier models scored under 1% at launch against near-perfect human play. The best official verified ARC-AGI-2 score was 54% in July 2026, so both versions still separate the field. ARC-AGI-2 paper; ARC-AGI-3 documentation.

Hendrycks et al.; TIGER-Lab · 2020 and 2024 · Retired, historical

MMLU and MMLU-Pro

The multiple-choice knowledge exam that defined "smart" for the ChatGPT era, and its harder 2024 successor; both now sit around 90% or above, and MMLU itself carries an estimated 9% label-error rate. Kept here because thousands of older claims cite them; treat any MMLU number as a history lesson. MMLU paper; MMLU-Pro paper.

Mathematics

Epoch AI · November 2024 · Watch

FrontierMath

Around 340 unpublished research-level mathematics problems in four difficulty tiers, written by professional mathematicians; the version 2 release of June 2026 corrected errors an audit found in roughly a third of the problems. Also a benchmark-integrity teaching story: OpenAI funded its creation with dataset access under an undisclosed agreement, and Epoch now keeps a 50-problem holdout no lab ever sees. Release paper.

MathArena · Ongoing, since 2025 · Watch

MathArena live competitions

Runs models on real mathematics competitions (IMO, USAMO, HMMT, Putnam) within hours of the problems being released, which makes it the cleanest contamination-free signal in mathematics; in July 2025 it certified a gold-medal IMO performance by Gemini Deep Think. Methodology paper.

MAA competition, tracked by MathArena; Hendrycks et al. · Yearly; 2021 · Retired, historical

AIME and MATH

The two staples of maths benchmarking for years: the American Invitational Mathematics Examination and the MATH competition-problem dataset. Both are now saturated at the frontier; five models scored 100% on AIME 2025, and MATH sits at 97% to 99%, so neither can tell strong models apart any more. MATH paper.

Coding

Princeton; Verified split by OpenAI · 2023; August 2024 · Watch, nearing saturation

SWE-bench Verified

500 human-validated real GitHub issues the model must fix in actual Python repositories; the most-quoted coding number in the industry. Top agents now resolve roughly 85% to 95%, so its power to separate frontier models is fading. Release paper.

Scale AI · September 2025 · Watch

SWE-bench Pro

1,865 enterprise-grade software tasks across 41 repositories and four languages, with contamination-controlled splits. Standardised public scores sit around 59% while vendor-run claims land 10 to 30 points higher; a tidy demonstration of why the test harness matters as much as the model. Release paper.

Stanford & Laude Institute · May 2025; v2.0 November 2025 · Watch

Terminal-Bench 2.0

Real tasks completed inside a live terminal and container environment, from build failures to system administration; closer to actual developer work than issue-fixing alone. Top agents reached 83% to 85% by May 2026, so watch for saturation at the top. Release paper.

OpenAI · July 2021 · Retired, historical

HumanEval

164 short Python function-writing problems; the benchmark that launched code generation as a research field. Solved years ago (96% to 98% and above) and kept here purely to show how far the goalposts have moved: from writing one function to fixing real enterprise codebases. Release paper.

Agents & Computer Use

OSWorld team · 2024; Verified July 2025 · Watch, nearing saturation

OSWorld-Verified

369 real tasks in a live operating system: spreadsheets, browsers, file managers, multi-app workflows. Top agents now score 82% to 85%, above the roughly 72% human baseline, which says as much about the benchmark's ceiling as about the agents. Release paper.

Sierra · June 2025 · Watch, nearing saturation

tau2-bench

Customer-service agents that must follow business policy while using tools mid-conversation, scored on whether they succeed reliably across repeated runs rather than once. Headline splits are close to solved (97% to 99% on telecom), and aggregate indices have already moved to a harder successor. Release paper.

OpenAI · April 2025 · Watch, nearing saturation

BrowseComp

1,266 questions built to require persistent, creative web research; agents scored 0.6% at launch and roughly 90% by July 2026, one of the fastest saturation arcs on record. A stricter successor (BrowseComp-Plus) restores the difficulty. Release paper.

Andon Labs · February 2025; v2 November 2025 · Watch

Vending-Bench 2

A long-horizon business simulation: the model runs a vending machine company for a simulated year from $500 of capital, scored on closing net worth; it punishes the drift and forgetfulness that short benchmarks never see. In July 2026 the leading model finished with $10,937. Release paper.

METR · March 2025, updated January 2026 · Watch

METR time horizons

Rather than a pass rate, METR measures how long a task a model can complete at 50% reliability, expressed in human working time; the January 2026 update put the frontier at roughly a 320-minute horizon, doubling about every 89 days since 2024. The single best chart for showing a general audience how fast agentic capability is moving. Methodology paper.

Cybersecurity

Stanford · August 2024 · Watch, nearing saturation

Cybench

40 professional capture-the-flag security challenges with graded subtasks, drawn from real CTF competitions. Frontier models are closing in on the ceiling, but it remains a fixture of the safety sections in model system cards. Release paper.

UC Berkeley · June 2025 · Watch

CyberGym

Over 1,500 tasks built from real vulnerabilities in open-source projects: agents must reproduce known flaws and sometimes discover new ones. Cited directly in lab safety evaluations, including the Claude Opus 4.5 system card (50.63% pass rate, November 2025). Release paper.

Carnegie Mellon · May 2026 · Watch, very new

ExploitBench

A capability ladder against 41 real Chrome browser vulnerabilities, scoring how far an agent climbs from reaching the flawed code to a full working exploit; at release the best public model achieved the top rung on just one bug. Only months old, so treat it as promising rather than established. Release paper.

Multimodal & Long Context

MMMU team · September 2024 · Watch

MMMU-Pro

College-level problems that require reading images, charts and diagrams together with text; the Pro version strips shortcuts that let models answer without truly looking at the picture. The standard citation for multimodal reasoning. Release paper.

NVIDIA · April 2024 · Watch

RULER

Measures a model's effective context length: how much of the advertised window it can actually use before quality collapses, which is routinely far less than the marketing number. Discrimination between frontier models now happens at 256k tokens and beyond. Release paper.

OpenAI · 2025 · Watch

MRCR

Multi-round co-reference: the model must track many similar items scattered through a very long conversation and retrieve exactly the right one. Cited by the major labs, and its hardest splits remain unsolved, which makes it a better long-context test than simple retrieval.

Community test · 2023 · Retired, historical

Needle in a Haystack

The original long-context test: hide one sentence in a mass of text and ask the model to find it. Every serious model now passes, so a perfect score means little; it survives as a sanity check and a reminder that retrieval is not comprehension. No vetted public link yet; the title is searchable.

Leaderboards & Aggregate Indices

Arena · Ongoing · Watch

Arena (formerly LMArena)

Millions of blind side-by-side votes from real users, ranked by Elo rating; the largest human-preference leaderboard, with the top labs now separated by only a handful of points. Read it alongside "The Leaderboard Illusion" (April 2025), which documented how private variant testing can flatter the rankings. The Leaderboard Illusion paper.

Artificial Analysis · Ongoing; Index v4.1 June 2026 · Watch

Artificial Analysis Intelligence Index

An independent composite score built from agentic, coding, scientific-reasoning and general benchmarks, re-run in-house on every major model with cost and speed alongside; it drops components as they saturate. A single convenient number, with all the caveats a single number deserves.

Epoch AI · Ongoing, since February 2025 · Watch

Epoch AI Benchmarking Hub

Independent trend charts across the major benchmarks, with costs and a note on who actually ran each evaluation; the most transparent place to see saturation curves and vendor-versus-independent gaps at a glance.

Last updated: 28 July 2026