WORKSHOPntworldink.comJuly 2026
AI Special Topic Workshops · A One-Hour Session

Which AI Is Best?
Reading the Scoreboard Like a Pro

Every AI launch comes with a wall of charts. This hour is about learning to read them: what the tests measure, when a score means something, and when it's marketing.

Scores quoted in this deck were current in July 2026; the whole point of the session is that they won't stay current.

ntworld.inkWorkshops01 / 23
START HEREthe view from the couchThe contenders
Before any jargon: the picture most people never see

There isn't one AI; there are several, and they're neck and neck

ClaudeANTHROPIC
ChatGPTOPENAI
GeminiGOOGLE
GrokXAI
DeepSeekDEEPSEEK
KimiMOONSHOT AI
CopilotMICROSOFT

Illustrative, not to scale; Copilot (orange) is an interface that runs OpenAI's models under the hood. For real closeness: the entire top ten on the Arena leaderboard sat within about 22 rating points, roughly 1.5%, on 28 July 2026 (arena.ai).

The report card

What they're graded on

Public tests score each model on the same handful of subjects:

CODINGWRITINGREASONINGIMAGE GENERATIONSEEING & HEARINGCYBER SECURITYMEMORY (CONTEXT)

Every model has stronger and weaker subjects; none tops the class in all of them.

The takeaway: for everyday writing, questions and admin, you would struggle to pick these apart in a blind test. Real differences only show up when you push hard on one subject; which is exactly what benchmarks are for.
ntworld.inkWorkshops02 / 23
START HEREthree common assumptionsCommon myths
Three assumptions worth retiring today

If you only know one AI by name, read this slide twice

Myth

"ChatGPT is the best one"

ChatGPT was first, in November 2022, so it became the household name; the way "to google" means "to search". Famous is not the same as best: on any given month the top of the leaderboards is shared between several labs.

Myth

"The free version is the real thing"

Free tiers usually run older, smaller or rate-limited models; the flagship sits behind a subscription. Judging an AI by its free tier in 2026 is judging a restaurant by its vending machine.

Myth

"The rankings settle it"

New models ship at a frenetic pace, and 2026 leaderboards often move by only a few percentage points per release. The lead changes hands so often that "the best AI" is a snapshot, not a fact.

Why people don't know this: unless you follow the industry, none of these updates reach you. The rest of this hour is the catch-up: how the scoreboard works, and how to read it for yourself.
ntworld.inkWorkshops03 / 23
WORKSHOPthe question everyone asksThe question
The most common question in every AI session

"Which AI is best?" has no fixed answer

Why not

The lead changes hands constantly

The frontier labs leapfrog each other every few months, and on the biggest human-preference leaderboard the top models are separated by only a handful of rating points. Whoever is "best" today probably wasn't three months ago.

The better question

"Best at what, measured how, and when?"

Capability differs by task, scores differ by who ran the test, and everything changes by the month. Answer those three and you can actually choose a model; that's the skill this hour teaches.

Where the answers live: public benchmarks; standardised tests that every major model sits. They're imperfect, gameable and short-lived, and they're still the best evidence we have.
ntworld.inkWorkshops04 / 23
1950where testing AI beganThe Turing test
The original test, and its quiet ending

The Turing test: "can a machine pass as human?"

The imitation game

One judge, two hidden strangers

In 1950, British mathematician Alan Turing proposed a simple scenario: a judge holds short typed conversations with a hidden computer and a hidden person. If the judge can't reliably tell which is which, the machine passes. It was a thought experiment rather than an exam, but for seventy years it stood as THE test of machine intelligence.

Broken in 2023, quietly

The milestone nobody celebrated

By 2023, chatbots could fool people in short chats: in an online imitation game played by more than 1.5 million people, players picked the bot correctly just 60% of the time; barely better than a coin flip. Yet the same AI that aced the bar exam scored as little as 3% on simple visual puzzles that people solve at about 91%. Science's most famous milestone came and went with barely any fanfare.

Why it retired: passing as human rewards deception, not capability; "it's all about trying to deceive the jury", as AI researcher François Chollet put it. So scientists went looking for better tests, and there is "no Rubicon, no one line" to replace it; just the many scoreboards in the rest of this hour.

Source: Celeste Biever, "ChatGPT broke the Turing test; the race is on for new ways to assess AI", Nature, vol 619, 27 July 2023.

ntworld.inkWorkshops05 / 23
JARGONsix terms, plain EnglishGet grounded
Six terms that make sense of everything that follows

Speak scoreboard: the jargon, translated

AGI

Artificial General Intelligence

AI as broadly capable as a person across most kinds of work. The definition is hotly contested and there's no agreed way to measure it; when someone says AGI is near (or here), ask which definition they're using.

Elo

The head-to-head rating

The rating system borrowed from chess: win a matchup, your number rises. Despite the capitals you often see, it's not an acronym; it's named after its inventor, physicist Arpad Elo. It's how Arena turns millions of blind votes into a ranking.

HLE

Humanity's Last Exam

2,500 questions written by subject experts across roughly 100 fields; built in 2025 to be the hardest knowledge exam an AI can sit, and named for the hope that we won't need another.

Spatial reasoning

Thinking in shapes and space

Judging positions, rotations and layouts; reading a map, a diagram or a clock face. One of AI's famously jagged spots: gold-medal mathematics alongside 50.1% at reading an analogue clock.

Chain of thought

Showing the working

When a model writes out step-by-step reasoning before giving its answer; the "thinking" mode on modern chatbots. It lifts performance on hard problems, at the cost of being slower and dearer.

Context window

The model's working memory

How much text a model can hold in mind in one sitting, measured in tokens (word-pieces). Bigger windows are heavily advertised; benchmarks like RULER test how much of the window actually works.

ntworld.inkWorkshops06 / 23
BASICSexams, arenas, simulationsWhat's a benchmark?
The basics, in one slide

A benchmark is an exam for AI; and there are three kinds

Exams

Fixed question sets

Thousands of questions with known answers: graduate science (GPQA), expert questions across 100 fields (HLE), real software bugs to fix (SWE-bench). Scored automatically, so everyone can compare.

Arenas

Blind human votes

Real people ask a question, see two anonymous answers, and pick the better one; millions of votes become a chess-style rating. Measures what people prefer, not what's correct.

Simulations

Agents in the wild

Give the model a computer, a terminal or a simulated business and score what it actually achieves: tasks finished, bugs found, money made. The newest and hardest kind.

The industry is shifting its attention from the first kind to the third: from what models know to what they can do.

ntworld.inkWorkshops07 / 23
KEY IDEA 1brilliant and clumsy at onceJagged capability
Key idea one

Capability is jagged: brilliant and clumsy in the same model

Gold
International Mathematical Olympiad, July 2025: Gemini Deep Think earned a certified gold-medal score (35/42), verified by the independent MathArena project.
12 / 12
ICPC World Finals, September 2025: an OpenAI system solved all twelve problems, above the best human team in the world's top programming contest.
50.1%
Reading an analogue clock: the same generation of frontier models, versus 90.1% for humans (Stanford AI Index 2026).
So "best" is task-shaped. A model can out-reason the world's mathematicians and still fumble something your ten-year-old finds easy. Never assume skill in one area transfers to another; check the benchmark closest to your task.
ntworld.inkWorkshops08 / 23
KEY IDEA 2benchmarks die youngSaturation
Key idea two

Benchmarks saturate in months, not years

·MMLU defined "smart" for the ChatGPT era in 2020; it now sits above 90%, with an estimated 9% of its answer labels simply wrong. Retired.2020 → RETIRED
·Humanity's Last Exam was built in January 2025 to be unpassable; models gained roughly 30 percentage points within a year (Stanford AI Index 2026).2025 → CLIMBING
·AIME competition maths: five different models scored 100% on the 2025 edition. A perfect score that tells you nothing about which is better.SATURATED
Why it matters to you: a benchmark only carries information while models still fail it. Once everyone scores 95%+, quoting it is nostalgia; Epoch AI's May 2026 analysis ("Are AI benchmarks doomed?") found most now saturate within months of release.
ntworld.inkWorkshops09 / 23
CASE STUDYfive years of moving goalpostsThe coding arc
Watch saturation happen: coding, 2021 to 2026

Same skill, four generations of exam

·HumanEval (2021): write one small Python function. Now solved at 96 to 98%.RETIRED
·SWE-bench Verified (2023): fix 500 real GitHub bugs. Top agents at roughly 85 to 95%.NEARLY DONE
·SWE-bench Pro (2025): 1,865 enterprise-grade tasks in four languages. Public scores around 59%.ACTIVE
·Terminal-Bench 2.0 (2025): real work inside a live terminal. Top agents 83 to 85% and climbing.ACTIVE
The pattern: each exam is "impossible" for about two years, then trivial. When someone quotes a coding score, your first question is: which generation of exam?
ntworld.inkWorkshops10 / 23
CASE STUDYthe fastest saturation on recordSpeed run
How fast is "months, not years"? This fast.

BrowseComp: 0.6% to ~90% in fifteen months

0.6%
April 2025, at launch. 1,266 web-research questions built by OpenAI to be genuinely hard for browsing agents: obscure facts needing long, creative search trails.
~90%
July 2026. Frontier agents now clear it almost routinely; a 150-fold improvement in fifteen months, one of the fastest capability arcs ever recorded.
Next
BrowseComp-Plus (2026) restores the difficulty; the field's answer to every saturated test is a harder sequel.
Practical lesson: any AI capability claim, positive or negative, needs a date attached. "AI can't do deep web research" was true in April 2025 and false a year later.
ntworld.inkWorkshops11 / 23
KEY IDEA 3who ran the test?Whose number?
Key idea three

Vendor scores run higher than independent ones

·The harness effect: on SWE-bench Pro, standardised public runs score around 59% while vendor-run claims land 10 to 30 points higher; same model, friendlier test setup.
·Contamination: models sometimes train on the test questions. When researchers rebuilt a maths test with fresh but equivalent questions (GSM1k, 2024), some model families dropped up to 13 percentage points.
·Fresh tasks tell the truth: on continuously refreshed coding tasks (SWE-rebench), scores sit roughly 30 points below the marketing numbers quoted from the older, well-studied test.
The habit: before trusting a score, ask who ran the evaluation. A number from the lab that built the model is a claim; a number from an independent leaderboard is evidence.
ntworld.inkWorkshops12 / 23
TRUE STORYwhen the examiner is funded by the studentFrontierMath

FrontierMath was the hardest maths benchmark ever built: research-level problems, written in secret by professional mathematicians. Then it emerged that OpenAI had funded its creation, with access to the problems, under an agreement the contributing mathematicians were never told about.

Nobody proved cheating; the model performed impressively anyway. But the episode, disclosed in January 2025, changed the field: Epoch AI now keeps a 50-problem holdout set no lab ever sees, and "who funded the benchmark?" became a standard question. Even the scoreboard needs auditing. A June 2026 audit then found errors in roughly a third of the original problems, now corrected in version 2; even the hardest exam ever written needed marking twice.

ntworld.inkWorkshops13 / 23
CASE STUDYthe people's leaderboardArena
The people's leaderboard, and its asterisk

Arena: millions of votes, one big caveat

The idea

Blind taste-testing for AI

Formerly LMArena: you ask a question, two unnamed models answer, you vote. Millions of votes build a chess-style rating. By mid-2026 the top labs sit within about 25 rating points of each other; genuinely close.

The asterisk

"The Leaderboard Illusion"

A 2025 study documented big labs privately testing many model variants and only publishing the winner, quietly retracting the rest; ranking inflation by selective disclosure. Arena disputed parts and tightened its rules, but the lesson stands.

Also remember: arena votes reward answers people like; confident, agreeable, nicely formatted. That is not the same as answers that are right. Use Arena for vibes, exams for correctness, simulations for competence.
ntworld.inkWorkshops14 / 23
TRUE STORYgive the AI a vending machineVending-Bench

Researchers at Andon Labs gave AI models $500 and a simulated vending machine business to run for a year: find suppliers, set prices, pay the daily fees, restock. The score is simply how much money is left at the end.

It measures what no exam can: staying coherent over thousands of small decisions. Models have spiralled into ordering nothing for months or hallucinating suppliers; the best run on the July 2026 leaderboard turned $500 into $10,937. And in the multi-agent version, competing AI shopkeepers spontaneously formed a price-fixing cartel; benchmarks find behaviours nobody thought to test for.

ntworld.inkWorkshops15 / 23
THE TRENDhow fast is this moving?Time horizons
The one chart that shows the speed

Task length doubles about every three months

50%
The research group METR measures the longest task a model completes at 50% reliability, expressed as how long the same work takes a skilled human.
~320 min
The frontier's horizon in METR's January 2026 update: models now handle tasks that take a person around five hours.
~89 days
The doubling time of that horizon since 2024. Minutes became hours in two years; if the trend holds, hours become days.
Why this one matters: single-benchmark scores expire, but this is a trend line across many tasks; the best single picture of how fast AI agents are improving, and the one to show anyone planning more than a year ahead.
ntworld.inkWorkshops16 / 23
PRACTICALmatch the test to your jobChoose by task
Putting it to work

Choosing a model? Check the benchmark nearest your task

Writing & general use

Arena + HLE

Human preference for everyday quality; HLE and GPQA for hard knowledge work. Close scores at the top mean price and speed can decide.

Coding

SWE-bench Pro, Terminal-Bench

The current generation of real-work coding tests. Ignore anything still quoting HumanEval.

Agents & automation

OSWorld, tau2, Vending-Bench

Computer use, policy-following customer service, and long-horizon coherence; the closest thing to "will it do my admin?"

Maths & research

FrontierMath, MathArena

Research-level problem solving, plus live competitions that can't have leaked into training data.

Security

Cybench, CyberGym

CTF challenges and real-vulnerability work; the scores labs publish in their own safety cards.

Long documents

RULER, MRCR

Whether the advertised million-token context window actually works, which is routinely less than claimed.

ntworld.inkWorkshops17 / 23
CONSUMER SKILLwhat's under the bonnet?Wrapper apps
The consumer-protection payoff

Spotting a wrapper app in two questions

Thousands of "revolutionary AI" apps are a thin interface over someone else's model, sometimes an older or cheaper one, resold at a premium. The scoreboard is your lie detector:

1"Which model is under the bonnet?" Honest products name the model and version. Evasion ("our proprietary AI engine") is an answer in itself.
2"Where is it on Arena or the Epoch hub?" Frontier models appear on independent leaderboards because independents can test them. A wrapper never publishes independent benchmark results; there's nothing of its own to test.
Rule of thumb: if a product's marketing quotes no independent numbers, assume you could get the same capability, cheaper, by going straight to the model it's wrapping.
ntworld.inkWorkshops18 / 23
CASE STUDYwhen the top result is boughtThe fake ChatGPT
Wrappers in the wild: a true story from May 2026

The "ChatGPT" at the top of Google wasn't ChatGPT

·YouTuber Christophe investigated after his friend googled "ChatGPT 4o", clicked the top result and paid for "premium"; weeks later he had no ChatGPT subscription at all. The top result was a paid ad for a lookalike site with a near-identical logo and interface.
·The top spot is bought at auction: advertisers bid on search phrases, weighted by a relevance "quality score". Two decades of design changes have made sponsored results look almost identical to real ones.
·The company behind it ran ad variations mimicking OpenAI's and DeepSeek's logos, and dozens of other AI apps. Customers reported paying US$20 to $60 for limits they thought they'd removed, and models "not even as capable as the free version of ChatGPT".
·Google confirmed the two reported ads broke its misleading-ads policies and removed them; but it's whack-a-mole: in 2024 Google blocked or removed half a billion ads over trademark issues and about 147 million for misrepresentation.
The defence is boring and it works: slow down. Type the address or use a bookmark, skip the "Sponsored" results, and check which model you're actually paying for. We'll watch Christophe's full breakdown at the end of the session.
ntworld.inkWorkshops19 / 23
TAKEAWAYreading any AI headlineThe habits
The takeaway: five questions for any AI score

Read the scoreboard like a pro

·Best at what? Capability is jagged; find the benchmark nearest your task.
·Dated when? Scores expire in months; an undated score is an anecdote.
·Who ran it? Vendor numbers run high; prefer independent leaderboards.
·Is it saturated? A 95%+ benchmark separates nobody; look for the active successor.
·One number, or several? No single figure captures a model; triangulate an exam, an arena and a simulation.
ntworld.inkWorkshops20 / 23
KEEP WATCHINGwhere to look things upFollow it
Where to look it up yourself

Bookmark the independent scoreboards

Leaderboards

The independents

Arena: blind human votes · Epoch AI Benchmarking Hub: trend charts, and who ran each test · Artificial Analysis: one composite index, with cost and speed

On this site

The companion catalogue

AI Benchmarks in the Resources library: every benchmark from this session with its leaderboard link, release paper and a status flag telling you whether the score still means anything. Kept up to date as tests retire.

The one-line summary of this hour: there is no best AI; there is a best AI for a task, on a date, according to somebody. Always find out what, when and who.
ntworld.inkWorkshops21 / 23
SOURCESwhere this comes fromReferences
Sources used

Where this deck's claims come from

Stanford AI Index Report 2026 (April 2026): the ~30-point HLE rise and the analogue-clock result (models 50.1%, humans 90.1%)
METR, Time Horizon 1.1 (29 January 2026): the ~320-minute horizon and ~89-day doubling; methodology at arxiv.org/abs/2503.14499
"The Leaderboard Illusion" (April 2025, NeurIPS 2025): arxiv.org/abs/2504.20879; Arena's response is at arena.ai/blog/our-response · GSM1k contamination study (May 2024): arxiv.org/abs/2405.00332
Epoch AI, "Are AI benchmarks doomed?" (1 May 2026) and the Benchmarking Hub; FrontierMath disclosure and holdout: epoch.ai/frontiermath
Celeste Biever, "ChatGPT broke the Turing test; the race is on for new ways to assess AI", Nature, vol 619, 27 July 2023: nature.com/articles/d41586-023-02361-7 (the imitation-game figures and ConceptARC results on the Turing test slide)
Christophe (YouTube), May 2026 investigation into lookalike AI apps in sponsored search results (the fake-ChatGPT case study and Google's 2024 ad-removal figures): youtube.com/watch?v=uuNCPmd6yU0
Arena leaderboard (arena.ai, accessed 28 July 2026): the top-ten spread of about 22 rating points behind the "neck and neck" chart
Individual benchmark scores (BrowseComp, SWE-bench, Terminal-Bench, AIME, Vending-Bench, ICPC and IMO results): each benchmark's official leaderboard as at July 2026, all linked with release papers from the AI Benchmarks resource page

Prepared July 2026 for the NT World Ink AI Special Topic Workshops. Scores age quickly; check the linked leaderboards for the current picture.

ntworld.inkWorkshops22 / 23
WATCHbefore you goFinale
Ten minutes well spent: the whole story, told properly

Christophe (YouTube), May 2026 investigation into lookalike AI apps in sponsored search results · youtube.com/watch?v=uuNCPmd6yU0

ntworld.inkWorkshops23 / 23