The complete, honest map of AI model benchmarks in September 2026: which of the classic 50 evals are dead, which still separate frontier models, and how to read the scores without being fooled.
Computer-use agents crossed the human baseline this spring, so the field threw the test away. On the repaired OSWorld-Verified benchmark, frontier agents pushed past the 72.36% human baseline the original study measured - OSWorld. The response was not celebration; it was a harder instrument. OSWorld 2.0 shipped on June 26, 2026 with 108 long-horizon workflows whose median task takes a skilled human about 1.6 hours, and at launch the best model on it, Claude Opus 4.8, completed just 20.6% of tasks - OSWorld 2.0. In eight weeks, "computer use is superhuman" became "the best agent finishes one task in five." Nothing about the models changed. The ruler did. Then the models moved too, and faster: on September 3, 2026 ARC Prize published its measurement of GPT-6 Astra on ARC-AGI-3, the interactive benchmark where frontier systems scored 0.51% at its March launch, and Astra returned 62.7% through ARC Prize's standard harness and 99.9% through OpenAI's own provider adapter - ARC Prize. Same weights, same benchmark, 37 points of difference produced by the software wrapped around the model.
That is the problem with every benchmark list you will find online, including the previous edition of this one. The August 5, 2026 version of this guide called ARC-AGI-3 "wide open" at 0.51% and printed an Intelligence Index leader of 61, four weeks before an OpenAI model scored 62.7% on the first and the index rebased the second down to 53. Its thesis was "benchmarks are instruments, and instruments wear out," and it failed to apply that thesis to its own most operationally relevant number. This edition corrects that, and everything else that broke in the weeks since: Claude Fable 5.1 and Mythos 5.1 (September 1), Gemini 3.8 Flash and Meta's Muse Spark 1.3 (September 2), and GPT-6 Astra (September 3) all landed inside a single three-day window - llm-stats, and nearly every leaderboard fact from early August is now one generation stale.
This guide is the full September 2026 ledger: all 50 benchmarks that have mattered across the modern LLM era, each with an honest dead-or-alive verdict (saturated, contaminated, repaired, superseded, or live), then deep profiles of the instruments that still discriminate, the trust problem behind self-reported scores, and what it costs to run evals at verified September 2026 prices. If you want a raw interactive database, llm-stats and lmcouncil do that better than any article can. What a database cannot give you is verdicts, history, and an operator's view of which numbers predicted real production behavior. That is what this page is for.
Contents
- What Changed in the Five Weeks Since the August Edition
- The Full 50: Benchmark Status Ledger
- Knowledge and Reasoning: HLE, GPQA, and the ARC-AGI Family
- Math: FrontierMath's Saturation and the Open Problems Era
- Coding: SWE-bench Pro, Terminal-Bench 2.0, and a Stranger on the Leaderboard
- Computer Use, Web Research, and Tool Calling
- Economically Grounded Evaluation: GDPval, METR, and Vending-Bench
- Meta-Benchmarks: Arena Elo, ECI, and the Intelligence Index
- The Lab-by-Lab Scorecard (September 2026)
- The Trust Problem: Self-Reported Scores and Harness Effects
- Versioning, Error Correction, and Contamination Defenses
- What It Costs to Run Evals in September 2026
- How to Pick Benchmarks for Your Use Case
- Conclusion: The Evaluation Playbook, Rewritten Again
The Scoreboard: Which Benchmarks Still Discriminate in September 2026
Before the detailed profiles, here is the master assessment of the benchmarks that still matter. Each is scored on four criteria: frontier discrimination (30%: does it separate the top models from each other), contamination resistance (25%: can training-data leakage or targeted optimization game it), real-work signal (25%: does the score predict usefulness on economically valuable tasks), and verifiability (20%: are the published numbers independently reproduced or self-reported). Scores run 0-10, and the final column is the weighted average.
The table is sorted by final score, highest first. Notice what tops it: instruments whose task pools are private or live, where the leaders still leave a third of the work on the table. That is not a paradox; it is the definition of a good instrument. A test with headroom, live environments, and hour-scale tasks carries more decision-relevant information than any exam the frontier has already solved. The famous names from every 2024 model card sit at the bottom for the same reason, inverted.
| # | Benchmark | What It Measures | Discrimination (30%) | Contamination Resistance (25%) | Real-Work Signal (25%) | Verifiability (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | OSWorld 2.0 | Long-horizon desktop workflows | 10 - 20.6% at launch, headroom still enormous | 9 - live OS, 1.6h median tasks, no answer key | 9 - real multi-application office workflows | 8 - XLang harness, public leaderboard | 9.1 |
| 2 | GDPval / GDPval-AA v2 | Real deliverables across 44 occupations | 9 - Elo spread 1764 to 1703 across frontier | 8 - anonymized pairwise judging of fresh outputs | 10 - tasks from the 9 largest US GDP industries | 9 - independent Artificial Analysis rerun | 9.0 |
| 3 | SWE-bench Pro | Real GitHub issue resolution at scale | 8 - public top 61.5%, wide spread below | 9 - private commercial + held-out splits | 8 - production repos, real diffs | 8 - Scale AI runs the splits | 8.3 |
| 4 | METR time horizons | Length of human task done at 50% reliability | 7 - ~14.5h top point, but 16h instrument ceiling | 8 - curated task suite, not scraped | 9 - denominated in hours of human labor | 9 - METR runs all models itself | 8.2 |
| 5 | ARC-AGI-3 | Novel-skill acquisition in interactive games | 9 - 0.51% in March, 62.7% in September, still unsaturated | 10 - interactive, no instructions, unseen environments | 4 - abstract games, not job tasks | 5 - same weights read 62.7% or 99.9% by harness | 7.6 |
| 6 | Terminal-Bench 4.0 | Agentic terminal and CLI work | 8 - re-versioned twice past 2.0, spread intact | 7 - curated live-shell tasks, 5 attempts each | 8 - matches real DevOps and build work | 8 - public leaderboard names every harness | 7.5 |
| 7 | Epoch Capabilities Index | Cross-benchmark composite capability | 7 - live dashboard, leader within noise of #2 | 7 - aggregates many sources, hard to game one | 7 - correlates with broad usefulness | 9 - Epoch methodology public, anchored scale | 7.4 |
| 8 | Vending-Bench 2 | Long-horizon business operation P&L | 7 - large dollar-outcome spread between models | 8 - simulation with emergent states | 8 - literally measured in money | 6 - Andon Labs runs it, less third-party rerun | 7.3 |
| 9 | Humanity's Last Exam | Expert-level academic knowledge | 8 - 59.1% top, real spread below | 6 - static public questions, memorizable | 5 - academic recall, not deliverables | 9 - Artificial Analysis independent run | 7.0 |
| 10 | tau2-bench | Policy-compliant conversational agents | 6 - frontier clusters but failures differ | 7 - dual-control simulation, versioned domains | 9 - direct customer-service relevance | 6 - mostly lab self-runs | 7.0 |
| 11 | BFCL V4 | Holistic function calling and agentic tool use | 6 - top models cluster in V4 | 6 - public dataset, versioned to stay ahead | 7 - tool calling is core agent plumbing | 8 - Berkeley runs the leaderboard | 6.7 |
| 12 | FrontierMath + Open Problems | Research-level mathematics | 7 - base set near-solved at 89%, Open Problems at 3/50 | 8 - private problems, expert-written | 4 - research math, niche economic value | 7 - Epoch runs it; top scores vendor-reported | 6.7 |
| 13 | Arena Elo (arena.ai) | Human preference at scale | 6 - top 5 within 8 Elo points | 5 - style and formatting gameable | 6 - preference, not task completion | 10 - 8.00M community votes | 6.6 |
| 14 | BrowseComp | Hard web research retrieval | 5 - top five within 1.4 points of each other | 6 - static questions, answers findable | 7 - deep research is a real workload | 3 - 63 entries, all self-reported, zero verified | 5.4 |
| 15 | GPQA Diamond | Graduate-level science QA | 4 - mid-90s at top, ceiling reached | 5 - public set, likely trained on | 4 - exam knowledge only | 7 - widely rerun by third parties | 4.9 |
| 16 | SWE-bench Verified | Curated GitHub issue resolution | 4 - top claims near-saturation | 3 - public since 2024, heavily optimized | 7 - still real repos underneath | 4 - overwhelmingly self-reported | 4.5 |
| 17 | MMLU-Pro | Harder multiple-choice knowledge | 2 - frontier clusters at 88-94% | 3 - fully public, in training corpora | 3 - multiple choice, no work product | 7 - trivially reproducible | 3.5 |
Four things changed in this table since August, and each is a story the rest of the guide tells in full. OSWorld 2.0 keeps the top slot as the instrument with the most room left to measure: computer use went from "solved" to 20.6% overnight because the test got honest about task length. ARC-AGI-3 fell from third to fifth, and not on difficulty: ARC Prize's September run of GPT-6 Astra scored the identical weights at 62.7% on its own harness and 99.9% on OpenAI's provider adapter - ARC Prize, and an instrument whose reading moves 37 points with the wrapper is measuring two things at once. BrowseComp collapsed from mid-table to fourteenth, not because the benchmark changed but because its tracker now shows 63 self-reported results and zero verified ones - llm-stats, which under this table's verifiability axis is disqualifying for decision use. And METR's time horizons dropped a rank for the most on-thesis reason imaginable: METR itself now posts that measurements above 16 hours are unreliable with the current task suite, meaning the field's favorite yardstick has hit the end of its own tape measure.
1. What Changed in the Five Weeks Since the August Edition
The previous edition of this guide went out on August 5, 2026 with a scoreboard, a status ledger, and verified prices. Within five weeks, most of its headline facts were stale. That is not an apology; it is the single most important datum in this article, because the decay rate of benchmark knowledge is itself the thing you need to plan around. Walk through what one month did.
The frontier turned over. Claude Fable 5.1 and Claude Mythos 5.1 shipped September 1 at $10/$50 per million tokens, with cache reads cut 75% to $0.25 per million, a 0.025x multiplier on a page where every other Claude model charges 0.1x - Anthropic. Fable 5.1 immediately took the top of the independent HLE and GDPval-AA leaderboards and tied for the top of the Intelligence Index, as Sections 3, 7 and 8 detail, while our Opus 5 vs Opus 4.8 breakdown still covers the tier directly beneath it. Two days later GPT-6 Astra arrived at $10/$50 as OpenAI's computer-use-first flagship, and GPT-5.6 Sol was repriced down to $4/$20 - OpenAI, the ladder our GPT-5.6 benchmark and pricing guide walks through in full. Gemini 3.8 Flash and Meta's Muse Spark 1.3 both landed on September 2 - llm-stats, on a tracker that now counts 389+ model releases overall.
The instruments turned over too, which matters more. OSWorld re-versioned (Section 6). FrontierMath's base set effectively saturated and Epoch answered with FrontierMath: Open Problems, a set of 50 genuinely unsolved research questions of which AI has so far solved three - Epoch AI. The Arena top five, an all-Anthropic sweep in July, now has a Meta model in it. Artificial Analysis moved its Intelligence Index from v4.1 to v4.3, dropping GPQA Diamond, adding AutomationBench-AA and jumping Terminal-Bench from v2.1 to v4.0, which moved every model's index number without any model moving. And METR still posts the 16-hour reliability ceiling on its own chart. A benchmark guide that updates monthly is still a snapshot of a moving object; the honest response is not to pretend otherwise but to date every number, link every live leaderboard, and tell you which claims will rot fastest. That is how this edition is written.
The deeper point deserves first-principles framing, because it explains why this keeps happening. A benchmark is an information instrument: it exists to reduce your uncertainty about which model to deploy. An instrument only carries information while its measurements spread across the range you care about. Static, public, finite test sets have two fatal properties: they leak into training corpora, and capability growth pins every model to their ceiling. Between 2023 and 2025 both failure modes hit the entire classic suite at once, a saturation wave Stanford's AI Index documented across the academic benchmarks - Stanford HAI. The 2026 replacements are live environments, private task pools, and scores denominated in hours, dollars, and deliverables. OSWorld 2.0 is simply the newest turn of that same wheel, and it will not be the last: the correct prior is that at least one instrument in this guide's top ten will be re-versioned or superseded before the next refresh.
2. The Full 50: Benchmark Status Ledger
This is the heart of the guide: the complete list of 50 benchmarks that have defined LLM evaluation, each with a September 2026 status verdict. Interactive databases enumerate hundreds of benchmarks; what they do not tell you is which ones are dead. Knowing that a benchmark died, when, and what replaced it is worth more than knowing it exists, because dead benchmarks still circulate in marketing material and stale listicles, pointing readers at scores that no longer discriminate between any models you would actually consider.
A note on how to read the status column. Live means the benchmark still separates frontier models and its results should influence decisions. Saturated means top models cluster at or near the ceiling, so it can only tell you a model is not frontier, never which frontier model is better. Contaminated means test data is believed to be in training corpora, so scores are unreliable in both directions. Superseded means a newer version or instrument replaced it (use the successor, never the old scores). Retired means the field has formally or effectively stopped reporting it. Historical model names appear only as history; none of the pre-2026 models here belong in a current comparison.
| # | Benchmark | Category | Status (Sep 2026) | The 2026 verdict |
|---|---|---|---|---|
| 1 | MMLU | Knowledge | Retired (saturated) | Above 90% for all frontier models; dropped from frontier comparisons |
| 2 | MMLU-Pro | Knowledge | Saturated | Clusters at 88-94%; no longer discriminates |
| 3 | GPQA Diamond | Science QA | Retired (saturated) | Mid-90s self-reports at the top; dropped from the Intelligence Index basket at v4.3 - Artificial Analysis |
| 4 | ARC (AI2 science) | Science QA | Retired | The 2018 science-question set; solved years ago, not to be confused with ARC-AGI |
| 5 | AGIEval | Exams | Retired | Human-exam aggregate, saturated alongside MMLU |
| 6 | BIG-Bench Hard (BBH) | Reasoning | Retired | Chain-of-thought made it trivial; historical only |
| 7 | HellaSwag | Commonsense | Retired | Mid-90s since 2023; zero frontier signal |
| 8 | WinoGrande | Commonsense | Retired | Same story as HellaSwag |
| 9 | PIQA | Commonsense | Retired | Physical commonsense, solved |
| 10 | BoolQ | Reading | Retired | Yes/no comprehension, solved |
| 11 | DROP | Reading + math | Retired | Discrete reasoning over paragraphs, saturated |
| 12 | TruthfulQA | Factuality | Retired | Methodology criticized, replaced by live factuality suites |
| 13 | CommonsenseQA | Commonsense | Retired | Solved, historical only |
| 14 | TriviaQA | Knowledge | Retired | Closed-book recall, saturated |
| 15 | Natural Questions | Knowledge | Retired | Superseded by browsing-based research evals |
| 16 | SQuAD | Reading | Retired | The 2016-2018 era; museum piece |
| 17 | GLUE | NLU | Retired | Pre-LLM instrument, historical only |
| 18 | SuperGLUE | NLU | Retired | Saturated in 2021; historical only |
| 19 | Humanity's Last Exam | Frontier knowledge | Live | Top score 59.1% (Claude Fable 5.1, independent run) - Artificial Analysis |
| 20 | ARC-AGI-1 | Abstract reasoning | Effectively solved | Superseded by ARC-AGI-2 |
| 21 | ARC-AGI-2 | Abstract reasoning | Live (closing) | GPT-6 Astra at 95.0% tops the tracker - llm-stats |
| 22 | ARC-AGI-3 | Interactive reasoning | Live (harness-split) | "Humans score 100%. Frontier AI scores 0.51%" at March launch - ARC Prize; GPT-6 Astra now scores 62.7% on the standard harness and 99.9% on OpenAI's adapter - ARC Prize |
| 23 | GSM8K | Grade-school math | Retired (solved) | Every frontier model at ceiling; contaminated besides |
| 24 | MATH | Competition math | Retired (solved) | The "hard" 2024 benchmark, now trivial |
| 25 | AIME | Competition math | Saturated | Essentially solved by reasoning models |
| 26 | FrontierMath | Research math | Saturating (v2) | GPT-5.6 Sol at 89.0% on the tracker; Open Problems (50 unsolved, 3 AI-solved) is the new frontier - llm-stats |
| 27 | HumanEval | Coding | Contaminated | Everything frontier above 88%; historical footnote only |
| 28 | MBPP | Coding | Contaminated | Same era and fate as HumanEval |
| 29 | SWE-bench (original) | Coding agents | Superseded | Replaced by Verified, then Pro |
| 30 | SWE-bench Verified | Coding agents | Near-saturated | Leaderboard dominated by vendor self-reports; use Pro instead |
| 31 | SWE-bench Pro | Coding agents | Live | Public split (731 tasks, 41 repos) topped at 61.5% - Scale AI |
| 32 | LiveCodeBench | Coding | Live | Post-cutoff problem generation keeps it clean - LiveCodeBench |
| 33 | CodeContests / competitive programming | Coding | Retired | The AlphaCode-era instrument; superseded by agentic coding evals |
| 34 | Terminal-Bench 2.0 | Agentic coding | Superseded (4.0 live) | The 2.0 board topped out at 84.7%; Artificial Analysis now runs v4.0 in its basket - tbench.ai |
| 35 | WebArena | Web agents | Superseded | The 812-task pioneer; replaced by BrowseComp and live-web evals |
| 36 | Mind2Web | Web agents | Superseded | Static snapshots aged badly against live-web tests |
| 37 | AgentBench | Agents | Retired | First-gen agent suite, superseded by task-specific evals |
| 38 | GAIA | Assistant tasks | Saturated | The 2023 assistant benchmark is largely solved in 2026 |
| 39 | OSWorld / OSWorld-Verified | Computer use | Superseded (June 26, 2026) | Agents beat the 72.36% human baseline, so the instrument was replaced - OSWorld |
| 40 | OSWorld 2.0 | Computer use | Live (wide open) | 108 long-horizon tasks; Opus 4.8 tops it at 20.6% binary completion - OSWorld 2.0 |
| 41 | BrowseComp | Web research | Live (trust caveat) | 1,266 hard questions; top five within 1.4 points, all self-reported - llm-stats |
| 42 | BFCL (V1-V3) | Function calling | Superseded | Old single-call scores not comparable to V4 |
| 43 | BFCL V4 | Agentic tool use | Live | Holistic multi-turn + memory eval, updated April 12, 2026 - Berkeley |
| 44 | tau-bench / tau2-bench | Policy-bound agents | Live (re-versioned) | v1.0.1 (July 2026) added voice and banking domains; pre-1.0.1 results not comparable - Sierra Research |
| 45 | Vending-Bench 2 | Long-horizon business | Live | Business P&L as the score - Andon Labs |
| 46 | GDPval / GDPval-AA v2 | Real knowledge work | Live | Fable 5.1 (max) at 1764 Elo vs human baseline 1000 - Artificial Analysis |
| 47 | METR time horizons | Autonomous work length | Live (instrument ceiling) | ~14.5h top point; METR warns measurements above 16 hrs are unreliable - METR |
| 48 | MT-Bench | Chat quality | Retired | LLM-judge chat scoring, abandoned |
| 49 | AlpacaEval | Instruction following | Retired | Gameable LLM-judge format, abandoned |
| 50 | Arena Elo (was Chatbot Arena / LMArena) | Human preference | Live | 7,999,020 votes across 400 models as of September 2 - arena.ai |
Count the statuses and the story tells itself: of the 50, roughly 15 are live instruments in September 2026, and even among the live ones, three carry explicit caveats this edition added (OSWorld's supersession, BrowseComp's all-self-reported tracker, METR's 16-hour ceiling). Around 30 are retired, saturated, or contaminated, including almost everything a 2024-era reader would recognize from model cards. The most important row is the pair at 39-40: it is the only place in the table where you can watch a benchmark die of success and its replacement arrive inside a single quarter.
It is also worth naming the pattern in what survived, because it predicts next year's table. The survivors share three design properties: private or regenerating task pools (FrontierMath's unpublished problems, LiveCodeBench's post-cutoff generation, SWE-bench Pro's held-out split), interaction instead of recall (ARC-AGI-3's games, OSWorld 2.0's live desktop, Terminal-Bench's real shell), and grading against reality (GDPval's human-anchored Elo, Vending-Bench's simulated bank balance, METR's human-hours yardstick). Any benchmark lacking all three properties, whatever its citation count, belongs on the left side of this ledger within a year. The rest of the guide walks through the live instruments one by one.
3. Knowledge and Reasoning: HLE, GPQA, and the ARC-AGI Family
The knowledge-and-reasoning tier is where benchmark lists age most visibly, because it is where the frontier moves fastest and where the marketing pressure is highest. A year ago the story was "MMLU is saturated, and the replacement exams are unsolvable." Today the replacement exams are half-solved, and the only tests that still humble frontier models are interactive ones that cannot be memorized at all. Understanding this progression matters for anyone reading model cards, because labs still quote whichever exam makes their release look best, and the exams differ wildly in how much signal they carry.
Start with Humanity's Last Exam, the benchmark built specifically to outlive MMLU: 2,500 questions across more than 100 subjects, assembled by the Center for AI Safety and Scale AI from expert submissions that frontier models of the time could not answer - arXiv. In late 2025, the best public claim was around 25%, and that number was treated as evidence the exam would hold for years. On Artificial Analysis' independent run (the 2,158 text-only questions of the full set, across 28 evaluated models), Claude Fable 5.1 now scores 59.1% at max effort, 58.7% at xhigh and 55.9% at high - Artificial Analysis. The exam designed to be the last one standing went from unsolvable to nearly six-tenths solved in under two years, and the August-to-September update is its own micro-lesson: the top three scores on the board are one model at three effort settings, spread 3.2 points apart, so an HLE figure quoted without its effort dial is barely a figure at all.
HLE also illustrates the reporting problem this guide keeps returning to. The official leaderboard at lastexam.ai still shows Gemini 3 Pro on top at 38.3%, with a results table whose displayed dataset date is April 3rd, 2025 - lastexam.ai. The independent tracker and the official page now disagree by twenty-one points and more than a model generation. Official benchmark pages lagging third-party trackers by months is normal in 2026, and it means the date and the runner of a score matter as much as the score itself. If a comparison you are reading cites the official HLE leaderboard, it is describing a frontier that no longer exists.
GPQA Diamond, the graduate-level science exam that replaced MMLU in most 2025 model cards, has now met the ceiling that killed MMLU: top self-reports sit in the mid-90s, and Artificial Analysis dropped it outright when its Intelligence Index moved to v4.3, a basket of ten evaluations that contains no multiple-choice science exam at all - Artificial Analysis. When the top of a multiple-choice leaderboard reaches the mid-90s, the remaining gap is mostly question ambiguity and grading noise, not capability. The previous edition of this guide predicted that removal "within a few quarters" and it took one; the same pattern is documented in our Gemini 3.1 Pro guide, where the static exams cluster and the agentic evals spread.
The reasoning benchmarks that still discriminate are the ARC-AGI family, which tests novel-skill acquisition: abstract puzzles where the solver must infer a rule from a few examples, with no training data to lean on. The family's three generations have sharply different September 2026 statuses. ARC-AGI-1 is effectively solved. ARC-AGI-2 is live but closing: the tracker's top five are GPT-6 Astra at 95.0%, GPT-5.5 at 85.0%, Gemini 3.1 Pro at 77.1%, GPT-5.4 at 73.3%, and Gemini 3.5 Flash at 72.1% - llm-stats. ARC-AGI-3, launched March 25, 2026 as hundreds of handcrafted interactive game environments with no instructions and no stated goals, opened at "Humans score 100%. Frontier AI scores 0.51%," with over $2 million attached to ARC Prize 2026 - ARC Prize.
Two updates since sharpen the ARC-AGI-3 picture. Anthropic claims Claude Opus 5's ARC-AGI-3 score is "three times as high as the next-best model" - Anthropic, a vendor claim worth flagging as exactly that until the official leaderboard reflects it; and the ARC Prize leaderboard now discloses a $10,000 compute cap for displayed systems - ARC Prize, an underrated design choice, because unlimited-compute submissions were becoming a way to buy leaderboard positions rather than demonstrate efficient learning. That single-digit gap, the largest human-AI spread measured anywhere in the field, lasted five months. On September 3, 2026 ARC Prize published its own run of GPT-6 Astra: 62.7% on ARC-AGI-3 Semi-Private at max reasoning through ARC Prize's standard harness for $26,098, and 99.9% at high reasoning through OpenAI's provider adapter for $18,817, with the same weights and the same tasks in both runs - ARC Prize. Read the second number carefully before you repeat it. The adapter preserves the model's opaque reasoning state between requests and compacts long context, so it is scoring a system, not a model; ARC Prize records that the adapter runs were also 3.66x faster and used 49% fewer tokens - ARC Prize. Note also what the compute cap does to both figures: at $26,098 and $18,817 these runs sit well above the $10,000 ceiling for systems the public leaderboard displays, so the headline result and the leaderboard are answering different questions. Interaction still defeats memorization, which is why the static exams died. What it does not defeat is scaffolding.
4. Math: FrontierMath's Saturation and the Open Problems Era
Mathematics is the cleanest case study in benchmark lifecycle, because math problems have unambiguous answers and therefore unambiguous saturation points. The classic ladder (GSM8K as baseline, MATH as the serious test, olympiad problems as the ceiling) is entirely saturated: GSM8K is solved and contaminated, MATH is solved, and AIME differences are within run-to-run variance for frontier reasoning models. What replaced the ladder was FrontierMath, Epoch AI's research-level benchmark, and the biggest math story of this refresh is that FrontierMath itself is now saturating, one version and two years after it was designed to be unsolvable.
The current state, on the live tracker: GPT-5.6 Sol scores 89.0%, GPT-5.6 Terra 84.9%, and GPT-5.6 Luna 78.6%, with the leaderboard refreshed on September 8, 2026 and still carrying no GPT-6 Astra entry - llm-stats. Hold that against the benchmark's own history: at launch in November 2024, no model solved even 2% of the set. The July edition of this guide recorded GPT-5.5 Pro leading at 52.4%; four weeks and one model family later, the top score jumped 36 points. A benchmark of unpublished, expert-written research mathematics went from "nobody can touch it" to "the frontier solves nine problems in ten" in twenty-one months. Note the caveat that belongs on this number: OpenAI has exclusive access to a subset of the benchmark - Epoch AI, and the tracker's own status line reads 17 self-reported results, zero verified, so treat the 89% as directionally solid but not independently reproduced.
The structural details still matter for reading any FrontierMath number. The current version is v2, released June 12, 2026, after a correction review rebuilt the set to 338 problems (295 across Tiers 1-3, 43 in the Tier 4 expansion), of which only twelve are public - Epoch AI. Scores from before the v2 release were measured against a partially different answer key and are not comparable; any FrontierMath figure you cite needs a version attached, a discipline Section 11 generalizes to every re-versioned benchmark.
Epoch's answer to its own benchmark's saturation is the most interesting new instrument of the summer: FrontierMath: Open Problems, launched July 31, 2026, containing 50 significant, unsolved problems from research mathematics, of which "AI has solved three so far" - Epoch AI. Read that design carefully, because it is qualitatively new. Every previous benchmark tested models against answers somebody already knew; held-out sets were secret, but they were solved. An open-problems benchmark has no answer key at all: a solve is a contribution to mathematics, verifiable by mathematicians, impossible to leak because nobody on Earth has the answer to leak. Three solves out of fifty is simultaneously a tiny score and a historically strange sentence to be able to write. If the pattern of this guide holds (instruments saturate, replacements get closer to reality), open-problem benchmarks are the logical endpoint for math evaluation, and worth watching even if your workload is nowhere near research mathematics, because they are the first evals where "beating the benchmark" and "doing new valuable work" are literally the same event.
For model selection, the practical read is unchanged from July: unless your workload is actual research mathematics, FrontierMath standing is a proxy for deep multi-step reasoning stamina rather than a direct predictor, and the economically grounded evals in Section 7 will tell you more about business work. What changed is the label: FrontierMath's base set now belongs with AIME in the "capability confirmed, discrimination fading" bucket, and the live math signal is the Open Problems count and Tier 4.
5. Coding: SWE-bench Pro, Terminal-Bench 2.0, and a Stranger on the Leaderboard
Coding is the category where benchmark scores most directly drive purchasing decisions, because coding agents are the largest real workload for frontier models. It is also the category with the clearest instrument hierarchy in September 2026: HumanEval is a contaminated historical footnote, SWE-bench Verified is near-saturated and dominated by vendor self-reports, and the live signal comes from two instruments with structurally different defenses, plus one newcomer that this guide flags with an honest shrug.
SWE-bench Pro, built by Scale AI, scales the real-GitHub-issue format with a contamination defense built into its splits: a public split of 731 tasks across 41 repositories, plus private commercial and held-out splits drawn from codebases that cannot be in any training corpus - Scale AI. The September 2026 public leaderboard still reads: Muse Spark 1.1 at 61.5% (±3.1), gpt-5.4 (xHigh) at 59.1%, Muse Spark at 55.0%, claude-opus-4-6 (thinking) at 51.9%, and gemini-3.1-pro (thinking) at 46.1%. Two honest observations about that list. First, the top of it moved five points since July, when the best public score was around 59%. Second, and more interesting: the leaderboard does not say who makes Muse Spark. A previously unknown entrant holds the top of the most contamination-resistant public coding benchmark with no lab attribution on the leaderboard itself, and the attribution turned up on other people's pages instead: Arena lists muse-spark-1.2 xHigh under Meta - arena.ai, llm-stats logs Meta shipping Muse Spark 1.3 on September 2 - llm-stats, and that release now sits fifth on GDPval-AA at 1703 Elo - Artificial Analysis. That a frontier coding leaderboard can be topped by a system whose provenance is not stated on the page is itself a 2026 data point about how fast this space moves and how thin the metadata around "scores" still is.
Terminal-Bench 2.0, a Stanford and Laude collaboration, attacks contamination from the environment side: live shell tasks (builds, debugging, sysadmin work) attempted five times each, with the harness named alongside the model on every leaderboard row - tbench.ai. The 2.0 board's final top five: NexAU-AHE + GPT-5.5 at 84.7% (May 14, 2026), LemonHarness at 84.5% (multiple models), Capy + GPT-5.5 at 83.1%, Codex CLI + GPT-5.5 at 82.2%, and Polaris at 82.2% - Terminal-Bench. Look at what that table actually shows: the same GPT-5.5 model spans 2.5 points across three harnesses in the top five alone. Section 10 treats this as a core trust problem, but for buyers the practical version is simple: on agentic benchmarks, you are never buying a model score, you are buying a model-plus-scaffolding score, and the scaffolding is a product choice you control. Our coding agent frameworks benchmark ranks that layer specifically. One versioning footnote, and it resolved faster than expected: tbench.ai's front page now serves Terminal-Bench 4.0 - tbench.ai, and Artificial Analysis' Intelligence Index runs v4.0 in its v4.3 basket - Artificial Analysis, so every 2.0 figure above is a snapshot of a retired instrument and belongs in a comparison only with other 2.0 figures.
Two smaller instruments round out the live coding stack. LiveCodeBench stays clean by generating problems from contests published after model training cutoffs, making contamination chronologically impossible rather than merely discouraged - LiveCodeBench. And Epoch's benchmarking hub posted a new coding leaderboard note on August 3, 2026: on MirrorCode, "Claude Fable 5 leads with a score of 64%, followed by GPT-5.6 Sol at 20%" - Epoch AI. A 44-point gap between the top two frontier models is the largest spread on any coding eval this guide tracks, which makes it exactly the kind of number to handle carefully: Epoch had not published a standalone methodology page we could load at write time, so treat MirrorCode as a promising signal awaiting documentation rather than a decision-grade instrument.
The practical reading order for a coding decision in September 2026 is therefore: SWE-bench Pro standing first (private-split numbers if published), Terminal-Bench with attention to the harness column second, LiveCodeBench for contamination-proof freshness third, and old Verified or HumanEval numbers not at all. Weight long-horizon behavior separately if your agents run unattended: pass rates say nothing about whether an agent catches its own mistakes at hour six, a distinction we explored in our guide to long-running coding agents and one that OSWorld 2.0's results, next section, just made unavoidable.
6. Computer Use, Web Research, and Tool Calling
Agentic benchmarks beyond coding split into three workloads: operating a computer, researching the open web, and calling tools under policy constraints. All three changed materially since the July edition, and computer use changed most, in the way this guide's whole thesis predicts: the instrument that got solved got replaced.
Here is the before-and-after in full, because no competitor page tells it as one story. The original OSWorld (369 desktop tasks) was the pioneering computer-use benchmark; at publication, its authors measured that humans could accomplish over 72.36% of the tasks while the best model of that era managed about 12% - OSWorld. Audits later found broken tasks, and the repaired OSWorld-Verified became the standard. Through the spring of 2026, frontier agents climbed past the human baseline on that repaired set (the July edition of this guide recorded Claude Opus 4.8 at 83.5%). Then, on June 26, 2026, the XLang team's own site posted the supersession banner: "OSWorld 2.0 is now available" - OSWorld. The new instrument is a different class of test: 108 long-horizon tasks with a median of about 1.6 hours of skilled-human time each, requiring an average of 318 tool calls per task (measured with Claude Opus 4.7) against roughly 30 on the prior version - OSWorld 2.0. On it, Claude Opus 4.8 with maximum thinking and batched tool calls scored best at launch with 20.6% binary completion and a 54.8% partial score, while GPT-5.5 completes about 14%, notably more token-efficiently. The maintainers' own summary: current agents "are still far from professional-level computer use," struggling with hidden state recovery and constraint maintenance across extended workflows.
Read the flip correctly, because both wrong readings are circulating. "Computer use is superhuman" was true only of short, well-specified desktop tasks: the old instrument's range. "Computer use barely works" is equally wrong: a 54.8% partial score on 1.6-hour workflows means agents complete most of the steps of most professional-length tasks and fail on endurance, state-tracking, and constraint-keeping. For anyone deploying agents that operate real software (the entire premise of platforms like O-mega, whose agent workforces run browsers and desktops for business users), OSWorld 2.0's failure taxonomy is the single most useful document published this year: it says the binding constraint is no longer perception or clicking but long-horizon coherence, which is also exactly what METR's data shows from a different angle in the next section. Our computer-use benchmark guide and the deeper agentic computer-use guide track this space between editions. And note Anthropic's launch claim that Opus 5 surpasses "Fable 5's best result at just over a third of the cost" on OSWorld 2.0 - Anthropic: a vendor claim not yet on the public leaderboard at write time, flagged accordingly.
Web research consolidated around BrowseComp, OpenAI's benchmark of 1,266 questions built on an asymmetry: answers easy to verify but genuinely hard to find, requiring long multi-hop browsing across obscure sources - arXiv. The September 2026 tracker shows how far the instrument decayed: the top five (GPT-6 Astra at 91.5%, Kimi K3 at 91.2%, Claude Opus 5 at 90.8%, GPT-5.6 Sol at 90.4%, GPT-5.5 Pro at 90.1%) sit within 1.4 points of each other, and all 63 tracked results are self-reported, with zero independently verified - llm-stats. A clustered top plus purely vendor-reported numbers is the exact configuration in which a benchmark stops being decision-useful, which is why BrowseComp fell hard in this edition's scoreboard. The one genuinely notable fact it still carries: an open-weights model sits second by three tenths of a point, which is not close to an accident, as Section 9 shows.
Tool calling and policy compliance both re-versioned. BFCL V4, updated April 12, 2026, remains the live function-calling instrument, holistic across multi-turn chains and memory, with scores incomparable to the V1-V3 numbers still circulating in older articles - Berkeley. tau2-bench, the policy-bound customer-service benchmark from Sierra, hit v1.0.1 in July 2026: it now spans five domains (airline, retail, telecom, banking knowledge, and a mock testing domain) plus a full-duplex voice mode for audio-native evaluation, and its changelog states plainly that "results produced with tau2-bench < 1.0.1 are not comparable with >= 1.0.1" - Sierra Research. That sentence is the version-discipline rule of Section 11 written by the maintainers themselves. tau2's framing remains the correct one for customer-facing deployment: an agent that completes the task while violating policy fails, because a refund issued against policy is worse than no refund at all. If you are deploying conversational agents with real authority over money or commitments, tau2 standing should outweigh every academic exam a vendor shows you.
The category also gained a genuinely new instrument this month, and it is the most operationally relevant addition of the year so far. AutomationBench-AA, built by Zapier and run independently by Artificial Analysis, scores an agent on the share of each task's objectives completed without guardrail violations across 657 tasks in six business domains (finance, HR, marketing, operations, sales and support) inside simulated Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira and HubSpot environments, on a private test set - Artificial Analysis. The current top three are GPT-6 Astra (max) at 68.5%, Astra (xhigh) at 67.2%, and Grok 4.6 (xhigh) at 67.0%, and it is now one of the ten components of the Intelligence Index. Read it beside tau2 rather than instead of it: both grade policy compliance rather than raw completion, and both have the three properties that decide which instruments survive, because the task pool is private, the environment is live business software, and a violated guardrail scores as a failure no matter how good the output looked. That a third of the objectives still go uncompleted at the top is the number to carry into any automation business case.
7. Economically Grounded Evaluation: GDPval, METR, and Vending-Bench
The deepest shift in evaluation since 2025 is not any single benchmark; it is the unit of measurement. The classic suite measured models in percent correct on questions. The instruments in this section measure them in units the economy already uses: deliverables judged against professionals' work, hours of human labor completed autonomously, and dollars of simulated profit and loss. The logic is first-principles: a business does not buy "88% on a multiple-choice exam," it buys completed work, and an eval denominated in work plugs directly into a build-versus-buy spreadsheet while an exam score always needs a translation layer nobody has validated.
GDPval-AA v2 is the flagship. The underlying GDPval task set, built by OpenAI, draws real knowledge-work deliverables from 44 occupations across the nine largest US GDP industries; Artificial Analysis runs an independent agentic version on 220 tasks with shell and web access, where outputs are anonymized, judged pairwise, and aggregated into an Elo rating anchored to a human baseline of 1,000 - Artificial Analysis. The September 2026 standings: Claude Fable 5.1 (max effort) at 1764, Fable 5.1 (xhigh) at 1745, Claude Opus 5 (max) at 1735, Opus 5 (xhigh) at 1708, and Muse Spark 1.3 (max) at 1703. Two readings matter. The headline one: in blind pairwise judging, frontier model deliverables now beat the human-anchored baseline by enormous margins on these task types, and the August-to-September turnover at the top is the fastest lead change this eval has seen, with a point release taking the two highest slots inside a day of shipping and pushing the model that led in August down to third. The honest caveat: the pairwise judge is an LLM, not a human panel, which makes GDPval-AA an instrument with a known systematic-bias risk that its Elo spread cannot express. It is still the best real-work signal available; it is not a court ruling.
METR's time horizons measure the orthogonal quantity: not how good the output is, but how long a task a model can carry autonomously, defined as the human-task length at which the model succeeds 50% of the time. The last stable published point put Claude Opus 4.6 at roughly 14.5 hours of human-equivalent work, and METR's chart was last updated May 8, 2026, adding an early Claude Mythos Preview measurement - METR. The new fact this edition adds is the one METR itself now posts on the page: "Measurements above 16 hrs are unreliable with our current task suite." Sit with that. The field's favorite capability yardstick, the chart in every forecasting debate, has formally announced that the frontier is running off the end of its tape measure, roughly one model generation after crossing the 14-hour mark. The instrument that measures how long agents can work now needs longer tasks than its designers built, which is the METR chart and the OSWorld 2.0 reset saying the same thing in two different units: long-horizon capability is outrunning long-horizon measurement.
The third instrument, Vending-Bench 2 from Andon Labs, is the most literal: the model runs a simulated vending-machine business over a long horizon (ordering stock, setting prices, negotiating with suppliers, handling events), and the score is the final bank balance. It sounds whimsical and is anything but: long-horizon coherence failures that no exam detects (forgetting inventory, hallucinating supplier agreements, death-spiraling on pricing) show up as bankruptcies, denominated in the one unit executives never misread. For a business evaluating agent platforms, this economic tier is the one to read first, because it answers the questions a deployment actually poses: will the deliverable be as good as my analyst's (GDPval), how much of a workday can it carry unsupervised (METR), and does it stay coherent over weeks of operation (Vending-Bench). It is also why agent-workforce platforms like O-mega orient around task outcomes rather than exam scores: at the deployment layer, the operative benchmark is whether this week's real work got done.
8. Meta-Benchmarks: Arena Elo, ECI, and the Intelligence Index
When individual benchmarks saturate quickly and disagree with each other, the field's response is aggregation, and meta-benchmarks are now the default way to talk about overall capability. Three aggregate views dominate: community-preference Elo, Epoch's cross-benchmark index, and Artificial Analysis' Intelligence Index. Each compresses a different kind of evidence, and each has failure modes worth knowing before you quote it.
The largest human-preference signal is Arena Elo at arena.ai (the platform formerly known as Chatbot Arena and LMArena), which as of September 2, 2026 holds 7,999,020 votes across 400 models - arena.ai. The top five: claude-fable-5 at 1507, claude-opus-4-6-high at 1505, claude-fable-5.1-max at 1504, claude-opus-4-7-high at 1502, and muse-spark-1.2 xHigh at 1499. That fifth row is the update that matters, and it is the same name that has been sitting unattributed at the top of SWE-bench Pro since July: Arena labels Muse Spark as a Meta model - llm-stats, three Elo points off fourth place and well inside the confidence intervals. The top five now span 8 Elo points, which translates to near-coin-flip win rates between adjacent models. Arena tells you which model people prefer in open-ended chat; it does not tell you which completes a refund workflow without violating policy, and its known biases (verbosity, formatting polish, sycophancy) are style dimensions agentic work does not reward. Use it as a popularity prior, never a deployment verdict.
Epoch AI's Capabilities Index (ECI) aggregates scores across many benchmarks onto a single anchored scale (calibrated so Claude 3.5 Sonnet = 130 and GPT-5 = 150, with no maximum achievable score) - Epoch AI. This edition handles ECI differently than a database would, and deliberately: the dashboard is interactive and its current leader changes with each ingested result, so any static number printed here would be expired on arrival. The July edition's snapshot (Fable 5 setting a then-record 161) predates Opus 5, Fable 5.1 and GPT-6 Astra alike, so treat it as history and read the live dashboard for the current leader. ECI's real value was never the leader anyway: it is the anchored scale, which lets you say "the frontier gained roughly two Claude-generations of capability in a year" with a straight face, and the slope of that line matters more than which lab holds the top pixel this week.
The Artificial Analysis Intelligence Index (now v4.3, a basket of ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1) is the widest independently run net, with every score produced by Artificial Analysis itself - Artificial Analysis. The September 2026 top is a four-way tie: Claude Fable 5.1 (max) and (xhigh) at 53, GPT-6 Astra (max) and (xhigh) at 53, then Fable 5.1 (high) at 51. Hold that against the 61 the previous edition printed for Opus 5 and you have Section 11's version discipline in a single line: no model regressed, the basket changed, and an index number without its index version attached means nothing. The open-weights headline rewrote itself again: the highest open model is now GLM-5.3 (max) at 45, eight points off the global lead, in the family whose previous generation is covered in our GLM-5.2 guide. That eight-point gap, Kimi K3's three-tenths-of-a-point miss on BrowseComp from Section 6, and Meta's arrival in the Arena top five, taken together, mark September 2026 as the month the frontier stopped being a two-lab story; our Kimi K3 breakdown and open-source LLM guide go deeper.
The correct way to use meta-benchmarks is as a screening filter, not a verdict. An aggregate index tells you which four or five models belong on your shortlist; the task-specific evals from Sections 5 through 7 tell you which of those to deploy. Inverting that order (picking by Arena rank, then rationalizing with task scores) remains the most common model-selection mistake, because an 8-point Elo band contains models with materially different coding, computer-use, and policy-compliance profiles.
9. The Lab-by-Lab Scorecard (September 2026)
An honest scorecard has to be given category by category, with dates, because leads now change hands within weeks, not quarters. What follows is the verified state as of this writing, using independent runs wherever they exist and flagging vendor claims where they do not.
Anthropic leads the aggregate and real-work signals, and did something economically notable to get there: it cut the price of the lead model without changing its list rate. Claude Fable 5.1 (September 1, $10/$50, with cache reads at $0.25, a quarter of what Fable 5 charges) - Anthropic tops HLE (59.1%) and GDPval-AA (1764) and ties GPT-6 Astra at the head of the Intelligence Index (53), while Claude Opus 5 at $5/$25 - Anthropic holds the tier below it on both; backgrounds on both generations are in our Fable 5 and Mythos 5 guide and Opus 4.8 guide. The Fable generation's brief suspension under US export-control requirements earlier this summer, covered in our export-controls analysis, stands as the reminder that model availability is now a regulatory variable as well as a technical one.
OpenAI owns the hard-reasoning column and the widest price ladder. GPT-6 Astra (September 3, $10/$50, no none effort setting) sits at the top of ARC-AGI-2 at 95.0% - llm-stats, BrowseComp at 91.5%, and AutomationBench-AA at 68.5%, and ties for the Intelligence Index lead; beneath it the GPT-5.6 family was repriced to Sol $4/$20, Terra $2/$12, Luna $0.20/$1.20, with a separate gpt-5.6-cyber tier at $12.50/$75 - OpenAI, and Sol still holds FrontierMath at 89.0% - llm-stats. The pattern worth naming is the shape of the ladder: OpenAI now sells its computer-use flagship at exactly Fable's list price and its previous flagship at 40% of it; the head-to-head deployment question is worked through in our GPT-5.6 vs Claude Opus 5 comparison.
Google remains the price-performance and context play, now with the fastest refresh cadence at the low end of any lab: Gemini 3.8 Flash shipped September 2 at $0.75/$3.75 through December 31, 2026, the same rate the 3.6 and 3.7 Flash models were repriced to and half the output cost of the 3.5 Flash they replaced - Google, while Gemini 3.1 Pro Preview holds at $2/$12 under 200k context and stays within a few points of the frontier on most agentic evals, per Section 5's SWE-bench Pro table where it posts 46.1%. For high-volume workloads where cost per task dominates, Google's column wins even where its peak scores do not; the down-market economics are in our Gemini 3.5 Flash guide.
xAI is this edition's biggest single correction, and the correction is the information. The July edition reported Grok 4.5 in private beta with no public access and zero independently verified scores, and framed xAI as the lab that had opted out of measurement. That narrative lasted eighteen days: Grok 4.5 was publicly released July 16, 2026, available via API at $2/$6 per million tokens with a 500K context window - llm-stats. That gap has now closed from the other end as well: Grok 4.6 (xhigh) sits third on AutomationBench-AA at 67.0%, a private-test-set score run by a third party rather than by xAI - Artificial Analysis, which retires the "benchmark-sparse" label the last two editions carried. Our Grok 4.5 guide tracks the coverage as it fills in.
The open-weights tier stopped being a footnote and became a per-benchmark contender. GLM-5.3 (max) is now the highest open model on the independent Intelligence Index at 45, eight points off the tied global lead. Kimi K3 (Moonshot, July 16) sits second on BrowseComp at 91.2%, three tenths of a point behind GPT-6 Astra. DeepSeek-V4-Flash-0731 (July 31) debuted at 50 on the Intelligence Index as an explicitly fast, open release - llm-stats, with its cost math against closed flagships worked in our DeepSeek V4-Flash comparison. The 2025 framing of open weights as trailing baselines is dead; the operative question in September 2026 is deployment preference, not capability class.
10. The Trust Problem: Self-Reported Scores and Harness Effects
Everything in this guide so far assumes the published numbers mean something, and that assumption needs its own section, because the biggest evaluation story of 2026 is not any model's score: it is who measured it. As benchmark results became the primary marketing surface for billion-dollar launches, the incentive to optimize the reporting, not just the model, grew accordingly, and the verification infrastructure is only partially catching up.
The cleanest exhibit this month is BrowseComp: 63 tracked results, every one self-reported, zero independently verified - llm-stats, on a leaderboard whose top five sit within 1.4 points of each other. When the spread between "winning" and fifth place is smaller than the plausible variance between labs' private harness configurations, the ranking is not information, it is formatting. SWE-bench Verified decayed the same way before it (its leaderboard has long been dominated by vendor self-reports, which is why this guide routes coding decisions to Scale's independently run Pro splits instead), and the general lesson compounds: a benchmark's trust profile decays faster than its difficulty, because self-reporting arrives the moment a leaderboard becomes marketing-relevant, which is well before the tasks get easy.
The harness effect is the second, less appreciated distortion, and Terminal-Bench 2.0's top five stated it with unusual precision: the same GPT-5.5 model scores 84.7% under NexAU-AHE, 83.1% under Capy, and 82.2% under Codex CLI - tbench.ai. Two and a half points of pure scaffolding, inside the top five of one leaderboard, before counting the older harnesses further down. September produced the extreme version of the same effect, roughly fifteen times larger, and it is the single most important number in this edition: ARC Prize ran identical GPT-6 Astra weights on ARC-AGI-3 twice and got 62.7% through its own standard harness and 99.9% through OpenAI's provider adapter, which preserves opaque reasoning state between requests and compacts long context - ARC Prize. The better run was also 3.66x faster on aggregate elapsed time and used 49% fewer tokens - ARC Prize. Thirty-seven points of a headline benchmark score belonged to the wrapper rather than the weights, on the benchmark specifically designed to be immune to everything except genuine novel-skill acquisition. Any agentic benchmark comparison that does not name the harness is incomplete to the point of meaninglessness, and for buyers the harness is part of the product: a well-engineered agent platform wrapped around a second-tier model routinely outperforms a frontier model in a naive loop. The third distortion is newer and subtler: the best real-work eval we have, GDPval-AA, uses an LLM judge for its pairwise grading - Artificial Analysis, which independent execution mitigates but does not eliminate, since a systematic judge preference (for length, structure, or a family's house style) would move every Elo in the table and be invisible in the error bars.
Three structural checks separate trustworthy numbers from marketing numbers in 2026:
- Independent execution: was the score produced by a third party (Artificial Analysis, METR, Epoch, Scale's splits) or by the vendor
- Named harness and settings: does the report specify scaffolding, retry budget, tool access, and effort configuration
- Version and date: is the benchmark version stated, and is the score current against a leaderboard that actually updates
Apply those checks to any model card and most of its benchmark table fails at least one; apply them to this article and you will find the vendor claims explicitly flagged as such (Opus 5's ARC-AGI-3 multiple, its OSWorld 2.0 cost claim, FrontierMath's tracker-reported 89% on a board that logs zero verified results). That is not nihilism, it is a reading protocol. The verified subset of the evidence is smaller than the reported subset, but it is entirely sufficient to make good decisions, and it is the subset this guide quotes.
11. Versioning, Error Correction, and Contamination Defenses
If 2024-2025 was the era of new benchmarks, 2026 is the era of repaired and re-versioned benchmarks, and the repair wave is better news than any new leaderboard: it means the field audits its own instruments, finds them wanting, and fixes them in public. But it also means every score now needs metadata (version, date, runner) attached before it can be compared to anything, and this section is the maintenance manual.
The re-versioning events now span every major category, and the maintainers have started saying the quiet part in their changelogs. FrontierMath v2 (June 12, 2026) rebuilt the set to 338 problems after a correction review, with only twelve public - Epoch AI; pre-v2 scores were measured against a partially different answer key. OSWorld went original, then Verified, then 2.0 within its lifetime, with the 2.0 supersession dated June 26, 2026 on the project's own banner. tau2-bench v1.0.1 ships the bluntest version of the rule ever written by a benchmark team: "results produced with tau2-bench < 1.0.1 are not comparable with >= 1.0.1" - Sierra Research. BFCL rolled through four versions, with V4's holistic agentic scoring incompatible with the single-call accuracy of V1-V3 - Berkeley. And the aggregators re-version too, which almost nobody accounts for: Artificial Analysis' Intelligence Index went from v4.1 to v4.3 in five weeks, dropping GPQA Diamond, adding AutomationBench-AA and moving Terminal-Bench from v2.1 to v4.0 - Artificial Analysis, which is why the top score on that index reads 53 in this edition and 61 in the last one with no model having gotten worse. Each event rearranged rankings without any model changing, which is the point: some fraction of every pre-repair leaderboard was measurement error, not capability difference.
The complementary story is contamination defense, because repair fixes broken tasks but not leaked ones. The defenses form a clear hierarchy of strength. Static-but-private task pools (FrontierMath's unpublished problems, SWE-bench Pro's commercial and held-out splits) defeat direct memorization but depend on the maintainer's operational security. Post-cutoff generation (LiveCodeBench) is stronger because leakage is chronologically impossible. Interactive environments (ARC-AGI-3's games, OSWorld 2.0's live desktop, Terminal-Bench's real shell) are stronger still, because there is no answer key to leak: the agent must produce behavior, not recall. And July 2026 added the endpoint of the hierarchy: unsolved-problem benchmarks (FrontierMath: Open Problems), where no answer exists anywhere, and a solve is simultaneously a score and a contribution to human knowledge.
For practitioners, the discipline reduces to a habit: never cite a score without its version, date, and runner, and never compare across versions. A "SWE-bench" number without a suffix is uninterpretable. An HLE score without a runner might describe the official page's years-old snapshot or this month's independent run, a twenty-one-point difference. An OSWorld number without a version is off by a factor of four. The July edition of this guide is itself the cautionary example twice over: every headline number it printed was correct on July 8, most were stale by August 5, and one (OSWorld-Verified as the live frontier) was already superseded the day it was published. Treat this edition identically: the linked leaderboards outrank this text the moment you read it.
12. What It Costs to Run Evals in September 2026
Benchmark scores are produced by spending tokens, and the price list moved again since July, in one direction at the top and with one deadline at the mid-tier. All prices below were verified against the official pricing pages this run, per million input/output tokens.
Anthropic - pricing: Claude Fable 5.1 and Mythos 5.1 at $10/$50 (Mythos flagged limited availability), the Fable 5 and Mythos 5 generation still listed at the same $10/$50, Claude Opus 5 at $5/$25 (the frontier-adjacent tier at half Fable's price), the Opus 4.5-4.8 line also at $5/$25, and Claude Haiku 4.5 at $1/$5. The 5.1 generation's real price move is not in the headline rate at all: its cache reads bill at 0.025x base input, or $0.25 per million tokens, against the 0.1x every other Claude model charges, so a workload that reuses a large prompt pays a quarter of what Fable 5 charged for the same context. And the price step-up this guide told you to model has been cancelled: Claude Sonnet 5's $2/$10 was introductory through August 31, 2026, and Anthropic's pricing page now states that it is the standard price and that "the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur," so any cost model still carrying $3/$15 for that tier is 50% too high; our Sonnet 5 cost breakdown covers it. The research-preview fast mode for Opus 5 and Opus 4.8 holds at $10/$50, trading money for latency at exactly Fable list price. Batch is 50% off, cache reads are 0.1x outside that 5.1 exception, and the 1M-token context window is standard-priced on Claude 4.6 and later.
OpenAI - pricing: GPT-6 Astra sits on top at $10/$50, and the GPT-5.6 family below it was repriced into a cleaner ladder: Sol $4/$20 (down from $5/$30), Terra $2/$12, Luna $0.20/$1.20, with long-context at 2x base rates and a separate gpt-5.6-cyber tier at $12.50/$75. GPT-5.4 holds at $2.50/$15, mini at $0.75/$4.50, nano at $0.20/$1.25. Cached input is discounted 90% and batch is 50% off. Google - pricing: Gemini 3.1 Pro Preview at $2/$12 under 200k context ($4/$18 above), the new Gemini 3.8 Flash at $0.75/$3.75 through December 31, 2026 (the 3.6 and 3.7 Flash models now match it, while the older 3.5 Flash still lists at $1.50/$9.00), Gemini 3.5 Flash-Lite at $0.30/$2.50, and batch at 50% off across the line. xAI: Grok 4.5 at $2/$6 - llm-stats, the most aggressive flagship list price on the board, and the successor Grok 4.6 now has an independently run agentic score to go with it.
List prices understate real eval costs in three specific ways that anyone budgeting a benchmark run must model. First, tokenizer inflation: per Anthropic's own pricing doc, Claude 4.7 and later models (and Claude Mythos Preview) use a newer tokenizer that produces approximately 30% more tokens for the same text, while Sonnet 4.6 and earlier use the previous one - Anthropic, which silently inflates every cross-generation cost comparison. Second, reasoning-token blowup: effort dials and agentic loops generate thinking and tool transcripts that dwarf the visible answer; OSWorld 2.0's 318 tool calls per task is what "one task" costs now, and a full run of a 2,158-question exam or a 731-task coding split lands in the hundreds to thousands of dollars per model at flagship rates. That range is no longer an estimate, because the independent evaluators now publish the bill: Artificial Analysis reports spending $5,324.10 to run GPT-6 Astra through the ten evaluations of Intelligence Index v4.3 - Artificial Analysis, which is what one model, one basket, one effort setting actually costs at $10/$50 with the reasoning tokens counted. Third, infrastructure add-ons: on Anthropic's meter, web search runs $10 per 1,000 searches, managed agent sessions bill $0.08 per session-hour, and code execution is free alongside web tools (otherwise $0.05 per container-hour after 1,550 free monthly hours).
The mitigation stack is the same one production workloads use: batch APIs (50% off at all three major labs) for anything not latency-sensitive, cache reads at a tenth of input price for shared benchmark prompts, and mid-tier models for pre-screening before spending flagship tokens. The full cost-modeling treatment lives in our AI model benchmarks and pricing guide. The strategic point stands independent of any single price: capability-per-dollar improved again this summer (Opus 5 delivering near-Fable intelligence at half the price is the clearest single instance), which means running your own evaluations, once a luxury reserved for labs, is now affordable for any team automating a meaningful workload. Section 13 says exactly how.
13. How to Pick Benchmarks for Your Use Case
Fifty benchmarks is a research syllabus, not a decision tool. A business choosing a model or agent platform in September 2026 needs five to ten numbers, selected by workload, all from the live column of the Section 2 ledger, all passing the Section 10 trust checks. This section is that selection guide, organized by the four workloads that cover most real deployments.
For coding agents, read SWE-bench Pro first (with the caveat that its public leader is a Meta model the leaderboard itself does not name), Terminal-Bench 4.0 second with attention to the harness column, and LiveCodeBench for contamination-proof freshness; ignore Verified and HumanEval entirely. For computer-use automation, OSWorld 2.0 is now the primary signal, and read it two-layered: the binary completion rate (20.6% for the best model at the benchmark's launch) tells you what runs unattended, the partial score (54.8%) tells you what runs with a human checkpoint, and the gap between them prices the supervision you still need to budget. Pair it with AutomationBench-AA, which asks the same question about business software rather than a desktop and grades guardrail violations as failures. For research and retrieval workloads, treat BrowseComp as a self-reported directional signal only, and lean on HLE's independent run for depth of expert knowledge. For customer-facing conversational agents and back-office automation, tau2-bench v1.0.1 (including its new voice mode, if you are deploying audio), AutomationBench-AA and BFCL V4 matter more than every academic exam combined; our broader agent evals guide covers this tier in depth.
Two cross-cutting rules complete the method. First, always convert to cost per completed task, not cost per token: a cheaper model with a lower completion rate can be more expensive per merged pull request or per finished workflow, and Section 12's price list only becomes decision-grade after multiplying by success rates, retry counts, and (for computer use) OSWorld-2.0-scale tool-call volumes. Second, run your own eval on 20-50 tasks drawn from your actual workload before committing, because public benchmarks are population averages and your task distribution is not the population. The public numbers narrow the shortlist to two or three candidates; your private eval picks the winner. This is the pattern mature agent deployments follow, whether built in-house or run on platforms like O-mega, where the operative benchmark is the completion rate on your own recurring business tasks, tracked week over week.
The final heuristic is a time discipline, and this edition upgrades it: quarterly re-checks are no longer enough at the top of the market. Between August 5 and September 8, the GDPval, HLE and Intelligence Index leads all changed hands, ARC-AGI-3 went from 0.51% to 62.7%, the Intelligence Index rebased its entire scale, and a scheduled price increase this guide told you to model was cancelled outright. Re-check the leaderboards for your two or three decision-critical benchmarks monthly, and re-run your private eval whenever a model one price tier below your incumbent posts scores within a few points of it, because that is the event that actually moves your unit economics.
14. Conclusion: The Evaluation Playbook, Rewritten Again
The through-line of this guide is that benchmarks are instruments, and instruments wear out, and the five weeks since the last edition provided the cleanest demonstration in the field's history: the interactive benchmark built to be unmemorizable went from 0.51% to 62.7% and then to 99.9% on a change of harness alone, the widest independent index rebased every score it publishes without a single model moving, and the yardstick for autonomous work still posts a notice that it cannot reliably measure past 16 hours. None of those events were model failures. They were measurements catching up to reality, which is what good instruments do right before they get replaced.
The decision playbook compresses to five moves. Screen with meta-benchmarks (Arena, ECI, the Intelligence Index) to build a shortlist, remembering the top five sit within a coin flip of each other. Deep-dive with workload-matched live evals from Section 13. Apply the trust checks: independent runner, named harness, stated version and date. Convert to economics: cost per completed task at verified September 2026 prices, with tokenizer inflation and effort-dial blowup modeled in, and with the Sonnet 5 step-up to $3/$15 taken back out, because Anthropic cancelled it. Verify on your own tasks, and re-verify monthly at the frontier. A team that follows those five steps will outperform any team choosing by headline score, whatever the headlines say by the time you read this.
And the honest closing note is this edition's own correction record. Five weeks ago this guide reported an Intelligence Index leader at 61 (the index rebased and the number is now 53), an ARC-AGI-3 human-AI gap it called the largest in the field (it closed), an unattributed system at the top of SWE-bench Pro that it declined to guess at (it is Meta's), and a September price increase it told you to budget for (Anthropic cancelled it). Every one of those corrections made the article more useful, which is the entire argument for reading a dated, versioned, verdict-carrying ledger instead of a static list: the value is not that it is always right, it is that it tells you exactly when and how it went stale. The evaluation landscape finally mirrors the market it measures: competitive, versioned, audited, and moving too fast for any snapshot, this one included, to stay true for long.
Written by Yuma Heymans (@yumahey), founder and CEO of O-mega. Running an agent-workforce platform means every benchmark claim in this guide eventually gets tested the hard way: against paying customers' real tasks.
This guide reflects the AI benchmark landscape as of September 8, 2026. Scores, leaderboards, and prices in this space now change weekly at the frontier: verify current numbers against the linked leaderboards before making decisions.