The complete, honest map of AI model benchmarks in August 2026: which of the classic 50 evals are dead, which still separate frontier models, and how to read the scores without being fooled.
Computer-use agents crossed the human baseline this spring, so the field threw the test away. On the repaired OSWorld-Verified benchmark, frontier agents pushed past the 72.36% human baseline the original study measured - OSWorld. The response was not celebration; it was a harder instrument. OSWorld 2.0 shipped on June 26, 2026 with 108 long-horizon workflows whose median task takes a skilled human about 1.6 hours, and on it the best model in the world, Claude Opus 4.8, completes just 20.6% of tasks - OSWorld 2.0. In eight weeks, "computer use is superhuman" became "the best agent finishes one task in five." Nothing about the models changed. The ruler did.
That is the problem with every benchmark list you will find online, including the previous edition of this one. The July 8, 2026 version of this guide still presented OSWorld-Verified as the live computer-use frontier, twelve days after OSWorld 2.0 had already replaced it. Its thesis was "benchmarks are instruments, and instruments wear out," and it failed to apply that thesis to its own most operationally relevant number. This edition corrects that, and everything else that broke in the four weeks since: Claude Opus 5 (July 24), the GPT-5.6 family going GA, Grok 4.5 going public, Kimi K3, Gemini 3.6 Flash, and Qwen3.8 Max all landed within a single month, and nearly every leaderboard fact from early July is now one generation stale.
This guide is the full August 2026 ledger: all 50 benchmarks that have mattered across the modern LLM era, each with an honest dead-or-alive verdict (saturated, contaminated, repaired, superseded, or live), then deep profiles of the instruments that still discriminate, the trust problem behind self-reported scores, and what it costs to run evals at verified August 2026 prices. If you want a raw interactive database, llm-stats and lmcouncil do that better than any article can. What a database cannot give you is verdicts, history, and an operator's view of which numbers predicted real production behavior. That is what this page is for.
Contents
- What Changed in the Four Weeks After the July Edition
- The Full 50: Benchmark Status Ledger
- Knowledge and Reasoning: HLE, GPQA, and the ARC-AGI Family
- Math: FrontierMath's Saturation and the Open Problems Era
- Coding: SWE-bench Pro, Terminal-Bench 2.0, and a Stranger on the Leaderboard
- Computer Use, Web Research, and Tool Calling
- Economically Grounded Evaluation: GDPval, METR, and Vending-Bench
- Meta-Benchmarks: Arena Elo, ECI, and the Intelligence Index
- The Lab-by-Lab Scorecard (August 2026)
- The Trust Problem: Self-Reported Scores and Harness Effects
- Versioning, Error Correction, and Contamination Defenses
- What It Costs to Run Evals in August 2026
- How to Pick Benchmarks for Your Use Case
- Conclusion: The Evaluation Playbook, Rewritten Again
The Scoreboard: Which Benchmarks Still Discriminate in August 2026
Before the detailed profiles, here is the master assessment of the benchmarks that still matter. Each is scored on four criteria: frontier discrimination (30%: does it separate the top models from each other), contamination resistance (25%: can training-data leakage or targeted optimization game it), real-work signal (25%: does the score predict usefulness on economically valuable tasks), and verifiability (20%: are the published numbers independently reproduced or self-reported). Scores run 0-10, and the final column is the weighted average.
The table is sorted by final score, highest first. Notice what tops it: a benchmark that is six weeks old, where the best model completes one task in five. That is not a paradox; it is the definition of a good instrument. A test with headroom, live environments, and hour-scale tasks carries more decision-relevant information than any exam the frontier has already solved. The famous names from every 2024 model card sit at the bottom for the same reason, inverted.
| # | Benchmark | What It Measures | Discrimination (30%) | Contamination Resistance (25%) | Real-Work Signal (25%) | Verifiability (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | OSWorld 2.0 | Long-horizon desktop workflows | 10 - top score 20.6%, massive headroom | 9 - live OS, 1.6h median tasks, no answer key | 9 - real multi-application office workflows | 8 - XLang harness, public leaderboard | 9.1 |
| 2 | GDPval / GDPval-AA v2 | Real deliverables across 44 occupations | 9 - Elo spread 1852 to 1730 across frontier | 8 - anonymized pairwise judging of fresh outputs | 10 - tasks from the 9 largest US GDP industries | 9 - independent Artificial Analysis rerun | 9.0 |
| 3 | ARC-AGI-3 | Novel-skill acquisition in interactive games | 10 - humans 100%, frontier AI 0.51% at launch | 10 - interactive, no instructions, unseen environments | 4 - abstract games, not job tasks | 9 - ARC Prize verified, $10k compute cap disclosed | 8.3 |
| 4 | SWE-bench Pro | Real GitHub issue resolution at scale | 8 - public top 61.5%, wide spread below | 9 - private commercial + held-out splits | 8 - production repos, real diffs | 8 - Scale AI runs the splits | 8.3 |
| 5 | METR time horizons | Length of human task done at 50% reliability | 7 - ~14.5h top point, but 16h instrument ceiling | 8 - curated task suite, not scraped | 9 - denominated in hours of human labor | 9 - METR runs all models itself | 8.2 |
| 6 | Terminal-Bench 2.0 | Agentic terminal and CLI work | 8 - 82-85% top band with real spread | 7 - curated live-shell tasks, 5 attempts each | 8 - matches real DevOps and build work | 8 - public leaderboard names every harness | 7.8 |
| 7 | Epoch Capabilities Index | Cross-benchmark composite capability | 7 - live dashboard, leader within noise of #2 | 7 - aggregates many sources, hard to game one | 7 - correlates with broad usefulness | 9 - Epoch methodology public, anchored scale | 7.4 |
| 8 | Vending-Bench 2 | Long-horizon business operation P&L | 7 - large dollar-outcome spread between models | 8 - simulation with emergent states | 8 - literally measured in money | 6 - Andon Labs runs it, less third-party rerun | 7.3 |
| 9 | Humanity's Last Exam | Expert-level academic knowledge | 8 - 53.3% top, real spread below | 6 - static public questions, memorizable | 5 - academic recall, not deliverables | 9 - Artificial Analysis independent run | 7.0 |
| 10 | tau2-bench | Policy-compliant conversational agents | 6 - frontier clusters but failures differ | 7 - dual-control simulation, versioned domains | 9 - direct customer-service relevance | 6 - mostly lab self-runs | 7.0 |
| 11 | BFCL V4 | Holistic function calling and agentic tool use | 6 - top models cluster in V4 | 6 - public dataset, versioned to stay ahead | 7 - tool calling is core agent plumbing | 8 - Berkeley runs the leaderboard | 6.7 |
| 12 | FrontierMath + Open Problems | Research-level mathematics | 7 - base set near-solved at 89%, Open Problems at 3/50 | 8 - private problems, expert-written | 4 - research math, niche economic value | 7 - Epoch runs it; top scores vendor-reported | 6.7 |
| 13 | Arena Elo (arena.ai) | Human preference at scale | 6 - top 5 within 13 Elo points | 5 - style and formatting gameable | 6 - preference, not task completion | 10 - 7.57M community votes | 6.6 |
| 14 | BrowseComp | Hard web research retrieval | 5 - top five within 4 points of each other | 6 - static questions, answers findable | 7 - deep research is a real workload | 3 - 58 entries, all self-reported, zero verified | 5.4 |
| 15 | GPQA Diamond | Graduate-level science QA | 4 - mid-90s at top, ceiling reached | 5 - public set, likely trained on | 4 - exam knowledge only | 7 - widely rerun by third parties | 4.9 |
| 16 | SWE-bench Verified | Curated GitHub issue resolution | 4 - top claims near-saturation | 3 - public since 2024, heavily optimized | 7 - still real repos underneath | 4 - overwhelmingly self-reported | 4.5 |
| 17 | MMLU-Pro | Harder multiple-choice knowledge | 2 - frontier clusters at 88-94% | 3 - fully public, in training corpora | 3 - multiple choice, no work product | 7 - trivially reproducible | 3.5 |
Three things changed in this table since July, and each is a story the rest of the guide tells in full. OSWorld 2.0 took the top slot from GDPval by being the instrument with the most room left to measure: computer use went from "solved" to 20.6% overnight because the test got honest about task length. BrowseComp collapsed from mid-table to fourteenth, not because the benchmark changed but because its tracker now shows 58 self-reported results and zero verified ones - llm-stats, which under this table's verifiability axis is disqualifying for decision use. And METR's time horizons dropped a rank for the most on-thesis reason imaginable: METR itself now posts that measurements above 16 hours are unreliable with the current task suite, meaning the field's favorite yardstick has hit the end of its own tape measure.
1. What Changed in the Four Weeks After the July Edition
The previous edition of this guide went out on July 8, 2026 with a scoreboard, a status ledger, and verified prices. Within four weeks, most of its headline facts were stale. That is not an apology; it is the single most important datum in this article, because the decay rate of benchmark knowledge is itself the thing you need to plan around. Walk through what one month did.
The frontier turned over. Claude Opus 5 shipped July 24 at $5/$25 per million tokens with a four-position effort dial (low, high, xhigh, max), positioned by Anthropic as coming "close to the frontier intelligence of Claude Fable 5 at half the price" - Anthropic. It immediately took the top of the independent GDPval-AA and Intelligence Index leaderboards, as Section 7 and Section 8 detail, and our Opus 5 vs Opus 4.8 breakdown covers the upgrade math. The GPT-5.6 family (Sol, Terra, Luna) went GA in early July, covered in our GPT-5.6 benchmark and pricing guide, and promptly rewrote the FrontierMath leaderboard. Grok 4.5 went publicly available on July 16 - llm-stats, killing the "xAI opted out of measurement" narrative eighteen days after the July edition printed it. The same one-month window added Kimi K3 (July 16), Gemini 3.6 Flash (July 21), DeepSeek-V4-Flash-0731 (July 31), and Qwen3.8 Max (August 2), on a tracker that now counts 337+ model releases overall.
The instruments turned over too, which matters more. OSWorld re-versioned (Section 6). FrontierMath's base set effectively saturated and Epoch answered with FrontierMath: Open Problems, a set of 50 genuinely unsolved research questions of which AI has so far solved three - Epoch AI. The Arena top five, an all-Anthropic sweep in July, was broken by an open-weights model. And METR posted the 16-hour reliability ceiling on its own chart. A benchmark guide that updates monthly is still a snapshot of a moving object; the honest response is not to pretend otherwise but to date every number, link every live leaderboard, and tell you which claims will rot fastest. That is how this edition is written.
The deeper point deserves first-principles framing, because it explains why this keeps happening. A benchmark is an information instrument: it exists to reduce your uncertainty about which model to deploy. An instrument only carries information while its measurements spread across the range you care about. Static, public, finite test sets have two fatal properties: they leak into training corpora, and capability growth pins every model to their ceiling. Between 2023 and 2025 both failure modes hit the entire classic suite at once, a saturation wave Stanford's AI Index documented across the academic benchmarks - Stanford HAI. The 2026 replacements are live environments, private task pools, and scores denominated in hours, dollars, and deliverables. OSWorld 2.0 is simply the newest turn of that same wheel, and it will not be the last: the correct prior is that at least one instrument in this guide's top ten will be re-versioned or superseded before the next refresh.
2. The Full 50: Benchmark Status Ledger
This is the heart of the guide: the complete list of 50 benchmarks that have defined LLM evaluation, each with an August 2026 status verdict. Interactive databases enumerate hundreds of benchmarks; what they do not tell you is which ones are dead. Knowing that a benchmark died, when, and what replaced it is worth more than knowing it exists, because dead benchmarks still circulate in marketing material and stale listicles, pointing readers at scores that no longer discriminate between any models you would actually consider.
A note on how to read the status column. Live means the benchmark still separates frontier models and its results should influence decisions. Saturated means top models cluster at or near the ceiling, so it can only tell you a model is not frontier, never which frontier model is better. Contaminated means test data is believed to be in training corpora, so scores are unreliable in both directions. Superseded means a newer version or instrument replaced it (use the successor, never the old scores). Retired means the field has formally or effectively stopped reporting it. Historical model names appear only as history; none of the pre-2026 models here belong in a current comparison.
| # | Benchmark | Category | Status (Aug 2026) | The 2026 verdict |
|---|---|---|---|---|
| 1 | MMLU | Knowledge | Retired (saturated) | Above 90% for all frontier models; dropped from frontier comparisons |
| 2 | MMLU-Pro | Knowledge | Saturated | Clusters at 88-94%; no longer discriminates |
| 3 | GPQA Diamond | Science QA | Saturating | Mid-90s self-reports at the top; still in the Intelligence Index basket but fading - Artificial Analysis |
| 4 | ARC (AI2 science) | Science QA | Retired | The 2018 science-question set; solved years ago, not to be confused with ARC-AGI |
| 5 | AGIEval | Exams | Retired | Human-exam aggregate, saturated alongside MMLU |
| 6 | BIG-Bench Hard (BBH) | Reasoning | Retired | Chain-of-thought made it trivial; historical only |
| 7 | HellaSwag | Commonsense | Retired | Mid-90s since 2023; zero frontier signal |
| 8 | WinoGrande | Commonsense | Retired | Same story as HellaSwag |
| 9 | PIQA | Commonsense | Retired | Physical commonsense, solved |
| 10 | BoolQ | Reading | Retired | Yes/no comprehension, solved |
| 11 | DROP | Reading + math | Retired | Discrete reasoning over paragraphs, saturated |
| 12 | TruthfulQA | Factuality | Retired | Methodology criticized, replaced by live factuality suites |
| 13 | CommonsenseQA | Commonsense | Retired | Solved, historical only |
| 14 | TriviaQA | Knowledge | Retired | Closed-book recall, saturated |
| 15 | Natural Questions | Knowledge | Retired | Superseded by browsing-based research evals |
| 16 | SQuAD | Reading | Retired | The 2016-2018 era; museum piece |
| 17 | GLUE | NLU | Retired | Pre-LLM instrument, historical only |
| 18 | SuperGLUE | NLU | Retired | Saturated in 2021; historical only |
| 19 | Humanity's Last Exam | Frontier knowledge | Live | Top score 53.3% (Claude Fable 5, independent run) - Artificial Analysis |
| 20 | ARC-AGI-1 | Abstract reasoning | Effectively solved | Superseded by ARC-AGI-2 |
| 21 | ARC-AGI-2 | Abstract reasoning | Live (closing) | GPT-5.5 at 85.0% tops the tracker - llm-stats |
| 22 | ARC-AGI-3 | Interactive reasoning | Live (wide open) | "Humans score 100%. Frontier AI scores 0.51%" at March launch - ARC Prize |
| 23 | GSM8K | Grade-school math | Retired (solved) | Every frontier model at ceiling; contaminated besides |
| 24 | MATH | Competition math | Retired (solved) | The "hard" 2024 benchmark, now trivial |
| 25 | AIME | Competition math | Saturated | Essentially solved by reasoning models |
| 26 | FrontierMath | Research math | Saturating (v2) | GPT-5.6 Sol at 89.0% on the tracker; Open Problems (50 unsolved, 3 AI-solved) is the new frontier - llm-stats |
| 27 | HumanEval | Coding | Contaminated | Everything frontier above 88%; historical footnote only |
| 28 | MBPP | Coding | Contaminated | Same era and fate as HumanEval |
| 29 | SWE-bench (original) | Coding agents | Superseded | Replaced by Verified, then Pro |
| 30 | SWE-bench Verified | Coding agents | Near-saturated | Leaderboard dominated by vendor self-reports; use Pro instead |
| 31 | SWE-bench Pro | Coding agents | Live | Public split (731 tasks, 41 repos) topped at 61.5% - Scale AI |
| 32 | LiveCodeBench | Coding | Live | Post-cutoff problem generation keeps it clean - LiveCodeBench |
| 33 | CodeContests / competitive programming | Coding | Retired | The AlphaCode-era instrument; superseded by agentic coding evals |
| 34 | Terminal-Bench 2.0 | Agentic coding | Live | Top harness+model combo at 84.7% - tbench.ai |
| 35 | WebArena | Web agents | Superseded | The 812-task pioneer; replaced by BrowseComp and live-web evals |
| 36 | Mind2Web | Web agents | Superseded | Static snapshots aged badly against live-web tests |
| 37 | AgentBench | Agents | Retired | First-gen agent suite, superseded by task-specific evals |
| 38 | GAIA | Assistant tasks | Saturated | The 2023 assistant benchmark is largely solved in 2026 |
| 39 | OSWorld / OSWorld-Verified | Computer use | Superseded (June 26, 2026) | Agents beat the 72.36% human baseline, so the instrument was replaced - OSWorld |
| 40 | OSWorld 2.0 | Computer use | Live (wide open) | 108 long-horizon tasks; Opus 4.8 tops it at 20.6% binary completion - OSWorld 2.0 |
| 41 | BrowseComp | Web research | Live (trust caveat) | 1,266 hard questions; top five within 4 points, all self-reported - llm-stats |
| 42 | BFCL (V1-V3) | Function calling | Superseded | Old single-call scores not comparable to V4 |
| 43 | BFCL V4 | Agentic tool use | Live | Holistic multi-turn + memory eval, updated April 12, 2026 - Berkeley |
| 44 | tau-bench / tau2-bench | Policy-bound agents | Live (re-versioned) | v1.0.1 (July 2026) added voice and banking domains; pre-1.0.1 results not comparable - Sierra Research |
| 45 | Vending-Bench 2 | Long-horizon business | Live | Business P&L as the score - Andon Labs |
| 46 | GDPval / GDPval-AA v2 | Real knowledge work | Live | Opus 5 (max) at 1852 Elo vs human baseline 1000 - Artificial Analysis |
| 47 | METR time horizons | Autonomous work length | Live (instrument ceiling) | ~14.5h top point; METR warns measurements above 16 hrs are unreliable - METR |
| 48 | MT-Bench | Chat quality | Retired | LLM-judge chat scoring, abandoned |
| 49 | AlpacaEval | Instruction following | Retired | Gameable LLM-judge format, abandoned |
| 50 | Arena Elo (was Chatbot Arena / LMArena) | Human preference | Live | 7,571,037 votes across 386 models as of August 1 - arena.ai |
Count the statuses and the story tells itself: of the 50, roughly 15 are live instruments in August 2026, and even among the live ones, three carry explicit caveats this edition added (OSWorld's supersession, BrowseComp's all-self-reported tracker, METR's 16-hour ceiling). Around 30 are retired, saturated, or contaminated, including almost everything a 2024-era reader would recognize from model cards. The most important row is the pair at 39-40: it is the only place in the table where you can watch a benchmark die of success and its replacement arrive inside a single quarter.
It is also worth naming the pattern in what survived, because it predicts next year's table. The survivors share three design properties: private or regenerating task pools (FrontierMath's unpublished problems, LiveCodeBench's post-cutoff generation, SWE-bench Pro's held-out split), interaction instead of recall (ARC-AGI-3's games, OSWorld 2.0's live desktop, Terminal-Bench's real shell), and grading against reality (GDPval's human-anchored Elo, Vending-Bench's simulated bank balance, METR's human-hours yardstick). Any benchmark lacking all three properties, whatever its citation count, belongs on the left side of this ledger within a year. The rest of the guide walks through the live instruments one by one.
3. Knowledge and Reasoning: HLE, GPQA, and the ARC-AGI Family
The knowledge-and-reasoning tier is where benchmark lists age most visibly, because it is where the frontier moves fastest and where the marketing pressure is highest. A year ago the story was "MMLU is saturated, and the replacement exams are unsolvable." Today the replacement exams are half-solved, and the only tests that still humble frontier models are interactive ones that cannot be memorized at all. Understanding this progression matters for anyone reading model cards, because labs still quote whichever exam makes their release look best, and the exams differ wildly in how much signal they carry.
Start with Humanity's Last Exam, the benchmark built specifically to outlive MMLU: 2,500 questions across more than 100 subjects, assembled by the Center for AI Safety and Scale AI from expert submissions that frontier models of the time could not answer - arXiv. In late 2025, the best public claim was around 25%, and that number was treated as evidence the exam would hold for years. On Artificial Analysis' independent run (the 2,158 text-only questions of the full set, across 27 evaluated models), Claude Fable 5 now scores 53.3%, with Claude Opus 5 immediately behind at 52.6% on max effort and 52.5% on xhigh - Artificial Analysis. The exam designed to be the last one standing went from unsolvable to half-solved in under a year, and the July-to-August update is its own micro-lesson: a model at half Fable 5's price now sits within 0.7 points of it, which changes the cost-per-point calculus for anyone actually paying for this capability.
HLE also illustrates the reporting problem this guide keeps returning to. The official leaderboard at lastexam.ai still shows Gemini 3 Pro on top at 38.3%, with a results table whose displayed dataset date is April 3rd, 2025 - lastexam.ai. The independent tracker and the official page now disagree by fifteen points and more than a model generation. Official benchmark pages lagging third-party trackers by months is normal in 2026, and it means the date and the runner of a score matter as much as the score itself. If a comparison you are reading cites the official HLE leaderboard, it is describing a frontier that no longer exists.
GPQA Diamond, the graduate-level science exam that replaced MMLU in most 2025 model cards, is meanwhile approaching the same ceiling that killed MMLU: top self-reports sit in the mid-90s, and it survives mainly as one of nine components in Artificial Analysis' Intelligence Index basket - Artificial Analysis. When the top of a multiple-choice leaderboard reaches the mid-90s, the remaining gap is mostly question ambiguity and grading noise, not capability. Expect GPQA to be formally dropped from frontier comparisons within a few quarters, exactly as MMLU was; the pattern is documented in our Gemini 3.1 Pro guide, where the static exams cluster and the agentic evals spread.
The reasoning benchmarks that still discriminate are the ARC-AGI family, which tests novel-skill acquisition: abstract puzzles where the solver must infer a rule from a few examples, with no training data to lean on. The family's three generations have sharply different August 2026 statuses. ARC-AGI-1 is effectively solved. ARC-AGI-2 is live but closing: the tracker's top five are GPT-5.5 at 85.0%, Gemini 3.1 Pro at 77.1%, GPT-5.4 at 73.3%, Gemini 3.5 Flash at 72.1%, and Claude Opus 4.6 at 68.8% - llm-stats. ARC-AGI-3, launched March 25, 2026 as hundreds of handcrafted interactive game environments with no instructions and no stated goals, remains wide open: at launch, "Humans score 100%. Frontier AI scores 0.51%," with over $2 million attached to ARC Prize 2026 - ARC Prize.
Two August-2026 updates sharpen the ARC-AGI-3 picture. Anthropic claims Claude Opus 5's ARC-AGI-3 score is "three times as high as the next-best model" - Anthropic, a vendor claim worth flagging as exactly that until the official leaderboard reflects it; and the ARC Prize leaderboard now discloses a $10,000 compute cap for displayed systems - ARC Prize, an underrated design choice, because unlimited-compute submissions were becoming a way to buy leaderboard positions rather than demonstrate efficient learning. A benchmark where humans score 100% and the best AI systems score in low single digits is the largest human-AI gap measured anywhere in the field, and it exists precisely because interaction defeats every memorization strategy that killed the static exams.
4. Math: FrontierMath's Saturation and the Open Problems Era
Mathematics is the cleanest case study in benchmark lifecycle, because math problems have unambiguous answers and therefore unambiguous saturation points. The classic ladder (GSM8K as baseline, MATH as the serious test, olympiad problems as the ceiling) is entirely saturated: GSM8K is solved and contaminated, MATH is solved, and AIME differences are within run-to-run variance for frontier reasoning models. What replaced the ladder was FrontierMath, Epoch AI's research-level benchmark, and the biggest math story of this refresh is that FrontierMath itself is now saturating, one version and two years after it was designed to be unsolvable.
The current state, on the live tracker: GPT-5.6 Sol scores 89.0%, GPT-5.6 Terra 84.9%, and GPT-5.6 Luna 78.6%, with the leaderboard updated in August 2026 - llm-stats. Hold that against the benchmark's own history: at launch in November 2024, no model solved even 2% of the set. The July edition of this guide recorded GPT-5.5 Pro leading at 52.4%; four weeks and one model family later, the top score jumped 36 points. A benchmark of unpublished, expert-written research mathematics went from "nobody can touch it" to "the frontier solves nine problems in ten" in twenty-one months. Note the caveat that belongs on this number: OpenAI has exclusive access to a subset of the benchmark - Epoch AI, and the tracker's top entries are vendor-reported, so treat the 89% as directionally solid but not independently reproduced.
The structural details still matter for reading any FrontierMath number. The current version is v2, released June 12, 2026, after a correction review rebuilt the set to 338 problems (295 across Tiers 1-3, 43 in the Tier 4 expansion), of which only twelve are public - Epoch AI. Scores from before the v2 release were measured against a partially different answer key and are not comparable; any FrontierMath figure you cite needs a version attached, a discipline Section 11 generalizes to every re-versioned benchmark.
Epoch's answer to its own benchmark's saturation is the most interesting new instrument of the summer: FrontierMath: Open Problems, launched July 31, 2026, containing 50 significant, unsolved problems from research mathematics, of which "AI has solved three so far" - Epoch AI. Read that design carefully, because it is qualitatively new. Every previous benchmark tested models against answers somebody already knew; held-out sets were secret, but they were solved. An open-problems benchmark has no answer key at all: a solve is a contribution to mathematics, verifiable by mathematicians, impossible to leak because nobody on Earth has the answer to leak. Three solves out of fifty is simultaneously a tiny score and a historically strange sentence to be able to write. If the pattern of this guide holds (instruments saturate, replacements get closer to reality), open-problem benchmarks are the logical endpoint for math evaluation, and worth watching even if your workload is nowhere near research mathematics, because they are the first evals where "beating the benchmark" and "doing new valuable work" are literally the same event.
For model selection, the practical read is unchanged from July: unless your workload is actual research mathematics, FrontierMath standing is a proxy for deep multi-step reasoning stamina rather than a direct predictor, and the economically grounded evals in Section 7 will tell you more about business work. What changed is the label: FrontierMath's base set now belongs with AIME in the "capability confirmed, discrimination fading" bucket, and the live math signal is the Open Problems count and Tier 4.
5. Coding: SWE-bench Pro, Terminal-Bench 2.0, and a Stranger on the Leaderboard
Coding is the category where benchmark scores most directly drive purchasing decisions, because coding agents are the largest real workload for frontier models. It is also the category with the clearest instrument hierarchy in August 2026: HumanEval is a contaminated historical footnote, SWE-bench Verified is near-saturated and dominated by vendor self-reports, and the live signal comes from two instruments with structurally different defenses, plus one newcomer that this guide flags with an honest shrug.
SWE-bench Pro, built by Scale AI, scales the real-GitHub-issue format with a contamination defense built into its splits: a public split of 731 tasks across 41 repositories, plus private commercial and held-out splits drawn from codebases that cannot be in any training corpus - Scale AI. The August 2026 public leaderboard reads: Muse Spark 1.1 at 61.5% (±3.1), gpt-5.4 (xHigh) at 59.1%, Muse Spark at 55.0%, claude-opus-4-6 (thinking) at 51.9%, and gemini-3.1-pro (thinking) at 46.1%. Two honest observations about that list. First, the top of it moved five points since July, when the best public score was around 59%. Second, and more interesting: the leaderboard does not say who makes Muse Spark. A previously unknown entrant holds the top of the most contamination-resistant public coding benchmark, with no lab attribution on the leaderboard itself, and this guide will not guess at one. That a frontier coding leaderboard can be topped by a system whose provenance is not stated on the page is itself a 2026 data point about how fast this space moves and how thin the metadata around "scores" still is.
Terminal-Bench 2.0, a Stanford and Laude collaboration, attacks contamination from the environment side: live shell tasks (builds, debugging, sysadmin work) attempted five times each, with the harness named alongside the model on every leaderboard row - tbench.ai. The current top five: NexAU-AHE + GPT-5.5 at 84.7% (May 14, 2026), LemonHarness at 84.5% (multiple models), Capy + GPT-5.5 at 83.1%, Codex CLI + GPT-5.5 at 82.2%, and Polaris at 82.2% - Terminal-Bench. Look at what that table actually shows: the same GPT-5.5 model spans 2.5 points across three harnesses in the top five alone. Section 10 treats this as a core trust problem, but for buyers the practical version is simple: on agentic benchmarks, you are never buying a model score, you are buying a model-plus-scaffolding score, and the scaffolding is a product choice you control. Our coding agent frameworks benchmark ranks that layer specifically. One versioning footnote: Artificial Analysis' Intelligence Index has already moved to running Terminal-Bench v2.1 in its basket - Artificial Analysis, so expect the 2.0 leaderboard to be superseded on the usual schedule.
Two smaller instruments round out the live coding stack. LiveCodeBench stays clean by generating problems from contests published after model training cutoffs, making contamination chronologically impossible rather than merely discouraged - LiveCodeBench. And Epoch's benchmarking hub posted a new coding leaderboard note on August 3, 2026: on MirrorCode, "Claude Fable 5 leads with a score of 64%, followed by GPT-5.6 Sol at 20%" - Epoch AI. A 44-point gap between the top two frontier models is the largest spread on any coding eval this guide tracks, which makes it exactly the kind of number to handle carefully: Epoch had not published a standalone methodology page we could load at write time, so treat MirrorCode as a promising signal awaiting documentation rather than a decision-grade instrument.
The practical reading order for a coding decision in August 2026 is therefore: SWE-bench Pro standing first (private-split numbers if published), Terminal-Bench with attention to the harness column second, LiveCodeBench for contamination-proof freshness third, and old Verified or HumanEval numbers not at all. Weight long-horizon behavior separately if your agents run unattended: pass rates say nothing about whether an agent catches its own mistakes at hour six, a distinction we explored in our guide to long-running coding agents and one that OSWorld 2.0's results, next section, just made unavoidable.
6. Computer Use, Web Research, and Tool Calling
Agentic benchmarks beyond coding split into three workloads: operating a computer, researching the open web, and calling tools under policy constraints. All three changed materially since the July edition, and computer use changed most, in the way this guide's whole thesis predicts: the instrument that got solved got replaced.
Here is the before-and-after in full, because no competitor page tells it as one story. The original OSWorld (369 desktop tasks) was the pioneering computer-use benchmark; at publication, its authors measured that humans could accomplish over 72.36% of the tasks while the best model of that era managed about 12% - OSWorld. Audits later found broken tasks, and the repaired OSWorld-Verified became the standard. Through the spring of 2026, frontier agents climbed past the human baseline on that repaired set (the July edition of this guide recorded Claude Opus 4.8 at 83.5%). Then, on June 26, 2026, the XLang team's own site posted the supersession banner: "OSWorld 2.0 is now available" - OSWorld. The new instrument is a different class of test: 108 long-horizon tasks with a median of about 1.6 hours of skilled-human time each, requiring an average of 318 tool calls per task (measured with Claude Opus 4.7) against roughly 30 on the prior version - OSWorld 2.0. On it, Claude Opus 4.8 with maximum thinking and batched tool calls scores best at 20.6% binary completion and a 54.8% partial score, while GPT-5.5 completes about 14%, notably more token-efficiently. The maintainers' own summary: current agents "are still far from professional-level computer use," struggling with hidden state recovery and constraint maintenance across extended workflows.
Read the flip correctly, because both wrong readings are circulating. "Computer use is superhuman" was true only of short, well-specified desktop tasks: the old instrument's range. "Computer use barely works" is equally wrong: a 54.8% partial score on 1.6-hour workflows means agents complete most of the steps of most professional-length tasks and fail on endurance, state-tracking, and constraint-keeping. For anyone deploying agents that operate real software (the entire premise of platforms like O-mega, whose agent workforces run browsers and desktops for business users), OSWorld 2.0's failure taxonomy is the single most useful document published this year: it says the binding constraint is no longer perception or clicking but long-horizon coherence, which is also exactly what METR's data shows from a different angle in the next section. Our computer-use benchmark guide and the deeper agentic computer-use guide track this space between editions. And note Anthropic's launch claim that Opus 5 surpasses "Fable 5's best result at just over a third of the cost" on OSWorld 2.0 - Anthropic: a vendor claim not yet on the public leaderboard at write time, flagged accordingly.
Web research consolidated around BrowseComp, OpenAI's benchmark of 1,266 questions built on an asymmetry: answers easy to verify but genuinely hard to find, requiring long multi-hop browsing across obscure sources - arXiv. The August 2026 tracker shows how the instrument decayed: the top five (Kimi K3 at 91.2%, Claude Opus 5 at 90.8%, GPT-5.6 Sol at 90.4%, GPT-5.5 Pro at 90.1%, GPT-5.6 Terra at 87.5%) sit within four points of each other, and all 58 tracked results are self-reported, with zero independently verified - llm-stats. A clustered top plus purely vendor-reported numbers is the exact configuration in which a benchmark stops being decision-useful, which is why BrowseComp fell hard in this edition's scoreboard. The one genuinely notable fact it still carries: an open-weights model tops it, and it is not close to an accident, as Section 9 shows.
Tool calling and policy compliance both re-versioned. BFCL V4, updated April 12, 2026, remains the live function-calling instrument, holistic across multi-turn chains and memory, with scores incomparable to the V1-V3 numbers still circulating in older articles - Berkeley. tau2-bench, the policy-bound customer-service benchmark from Sierra, hit v1.0.1 in July 2026: it now spans five domains (airline, retail, telecom, banking knowledge, and a mock testing domain) plus a full-duplex voice mode for audio-native evaluation, and its changelog states plainly that "results produced with tau2-bench < 1.0.1 are not comparable with >= 1.0.1" - Sierra Research. That sentence is the version-discipline rule of Section 11 written by the maintainers themselves. tau2's framing remains the correct one for customer-facing deployment: an agent that completes the task while violating policy fails, because a refund issued against policy is worse than no refund at all. If you are deploying conversational agents with real authority over money or commitments, tau2 standing should outweigh every academic exam a vendor shows you.
7. Economically Grounded Evaluation: GDPval, METR, and Vending-Bench
The deepest shift in evaluation since 2025 is not any single benchmark; it is the unit of measurement. The classic suite measured models in percent correct on questions. The instruments in this section measure them in units the economy already uses: deliverables judged against professionals' work, hours of human labor completed autonomously, and dollars of simulated profit and loss. The logic is first-principles: a business does not buy "88% on a multiple-choice exam," it buys completed work, and an eval denominated in work plugs directly into a build-versus-buy spreadsheet while an exam score always needs a translation layer nobody has validated.
GDPval-AA v2 is the flagship. The underlying GDPval task set, built by OpenAI, draws real knowledge-work deliverables from 44 occupations across the nine largest US GDP industries; Artificial Analysis runs an independent agentic version on 220 tasks with shell and web access, where outputs are anonymized, judged pairwise, and aggregated into an Elo rating anchored to a human baseline of 1,000 - Artificial Analysis. The August 2026 standings: Claude Opus 5 (max effort) at 1852, Opus 5 (xhigh) at 1819, Claude Fable 5 at 1743, Opus 5 (high) at 1735, and GPT-5.6 Sol (max) at 1730. Two readings matter. The headline one: in blind pairwise judging, frontier model deliverables now beat the human-anchored baseline by enormous margins on these task types, and the July-to-August turnover at the top (Fable 5 led at 1760 a month ago; Opus 5 now clears it by 100+ Elo at half the list price) is the fastest lead change this eval has seen. The honest caveat: the pairwise judge is an LLM, not a human panel, which makes GDPval-AA an instrument with a known systematic-bias risk that its Elo spread cannot express. It is still the best real-work signal available; it is not a court ruling.
METR's time horizons measure the orthogonal quantity: not how good the output is, but how long a task a model can carry autonomously, defined as the human-task length at which the model succeeds 50% of the time. The last stable published point put Claude Opus 4.6 at roughly 14.5 hours of human-equivalent work, and METR's chart was last updated May 8, 2026, adding an early Claude Mythos Preview measurement - METR. The new fact this edition adds is the one METR itself now posts on the page: "Measurements above 16 hrs are unreliable with our current task suite." Sit with that. The field's favorite capability yardstick, the chart in every forecasting debate, has formally announced that the frontier is running off the end of its tape measure, roughly one model generation after crossing the 14-hour mark. The instrument that measures how long agents can work now needs longer tasks than its designers built, which is the METR chart and the OSWorld 2.0 reset saying the same thing in two different units: long-horizon capability is outrunning long-horizon measurement.
The third instrument, Vending-Bench 2 from Andon Labs, is the most literal: the model runs a simulated vending-machine business over a long horizon (ordering stock, setting prices, negotiating with suppliers, handling events), and the score is the final bank balance. It sounds whimsical and is anything but: long-horizon coherence failures that no exam detects (forgetting inventory, hallucinating supplier agreements, death-spiraling on pricing) show up as bankruptcies, denominated in the one unit executives never misread. For a business evaluating agent platforms, this economic tier is the one to read first, because it answers the questions a deployment actually poses: will the deliverable be as good as my analyst's (GDPval), how much of a workday can it carry unsupervised (METR), and does it stay coherent over weeks of operation (Vending-Bench). It is also why agent-workforce platforms like O-mega orient around task outcomes rather than exam scores: at the deployment layer, the operative benchmark is whether this week's real work got done.
8. Meta-Benchmarks: Arena Elo, ECI, and the Intelligence Index
When individual benchmarks saturate quickly and disagree with each other, the field's response is aggregation, and meta-benchmarks are now the default way to talk about overall capability. Three aggregate views dominate: community-preference Elo, Epoch's cross-benchmark index, and Artificial Analysis' Intelligence Index. Each compresses a different kind of evidence, and each has failure modes worth knowing before you quote it.
The largest human-preference signal is Arena Elo at arena.ai (the platform formerly known as Chatbot Arena and LMArena), which as of August 1, 2026 holds 7,571,037 votes across 386 models - arena.ai. The top five: claude-fable-5 at 1509, claude-opus-4-6-thinking at 1505, claude-opus-4-7-thinking at 1502, claude-opus-4-6 at 1497, and qwen3.8-max at 1496. That fifth row is the update that matters: a month ago Anthropic held all five top slots, and the model that broke the sweep is an open-lab Alibaba model released on August 2 - llm-stats, sitting one Elo point (well inside the confidence intervals) from fourth place. The top five now span 13 Elo points, which translates to near-coin-flip win rates between adjacent models. Arena tells you which model people prefer in open-ended chat; it does not tell you which completes a refund workflow without violating policy, and its known biases (verbosity, formatting polish, sycophancy) are style dimensions agentic work does not reward. Use it as a popularity prior, never a deployment verdict.
Epoch AI's Capabilities Index (ECI) aggregates scores across many benchmarks onto a single anchored scale (calibrated so Claude 3.5 Sonnet = 130 and GPT-5 = 150, with no maximum achievable score) - Epoch AI. This edition handles ECI differently than a database would, and deliberately: the dashboard is interactive and its current leader changes with each ingested result, so any static number printed here would be expired on arrival. The July edition's snapshot (Fable 5 setting a then-record 161) predates both Opus 5 and GPT-5.6, so treat it as history and read the live dashboard for the current leader. ECI's real value was never the leader anyway: it is the anchored scale, which lets you say "the frontier gained roughly two Claude-generations of capability in a year" with a straight face, and the slope of that line matters more than which lab holds the top pixel this week.
The Artificial Analysis Intelligence Index (now v4.1, a basket of nine evaluations including GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, and Humanity's Last Exam) is the widest independently run net: 175 models evaluated, 99 of them open weights, every score run by Artificial Analysis itself - Artificial Analysis. The August 2026 top five: Claude Opus 5 (max) at 61, Opus 5 (xhigh) at 60, Claude Fable 5 at 60, GPT-5.6 Sol (max) at 59, and Opus 5 (high) at 59. The open-weights headline rewrote itself since July: the highest open model is now Kimi K3 (max) at 57, four points off the global lead, displacing GLM-5.2 (still respectable at 51, its story told in our GLM-5.2 guide), with the brand-new DeepSeek V4 Flash 0731 already at 50. Kimi K3's 57, its BrowseComp lead from Section 6, and Qwen3.8 Max's Arena top-five entry, taken together, mark August 2026 as the month the open-weights tier stopped trailing by a generation and started trailing by a configuration; our Kimi K3 breakdown and open-source LLM guide go deeper.
The correct way to use meta-benchmarks is as a screening filter, not a verdict. An aggregate index tells you which four or five models belong on your shortlist; the task-specific evals from Sections 5 through 7 tell you which of those to deploy. Inverting that order (picking by Arena rank, then rationalizing with task scores) remains the most common model-selection mistake, because a 13-point Elo band contains models with materially different coding, computer-use, and policy-compliance profiles.
9. The Lab-by-Lab Scorecard (August 2026)
An honest scorecard has to be given category by category, with dates, because leads now change hands within weeks, not quarters. What follows is the verified state as of this writing, using independent runs wherever they exist and flagging vendor claims where they do not.
Anthropic leads the aggregate and real-work signals, and did something economically notable to get there: the lead model is no longer the expensive one. Claude Opus 5 (July 24, $5/$25, effort dial from low to max) - Anthropic tops the independent Intelligence Index (61) and GDPval-AA (1852), while the pricier Claude Fable 5 still holds Arena (1509) and HLE (53.3%); backgrounds on both generations are in our Fable 5 and Mythos 5 guide and Opus 4.8 guide. The Fable generation's brief suspension under US export-control requirements earlier this summer, covered in our export-controls analysis, stands as the reminder that model availability is now a regulatory variable as well as a technical one.
OpenAI owns the hard-reasoning column and the widest price ladder. The GPT-5.6 family (Sol $5/$30, Terra $2/$12, Luna $0.20/$1.20) - OpenAI took FrontierMath to 89.0% - llm-stats, GPT-5.5 still tops ARC-AGI-2 at 85.0%, and GPT-5.5-based harnesses hold the Terminal-Bench 2.0 summit at 84.7%. GPT-5.6 Sol also sits fourth on the independent Intelligence Index and fifth on GDPval-AA, making it the strongest non-Anthropic model on the real-work tier; the head-to-head deployment question is worked through in our GPT-5.6 vs Claude Opus 5 comparison.
Google remains the price-performance and context play, now with a faster refresh cadence at the low end: Gemini 3.6 Flash shipped July 21 at $1.50/$7.50 (cheaper on output than the 3.5 Flash it succeeds) - Google, while Gemini 3.1 Pro Preview holds at $2/$12 under 200k context and stays within a few points of the frontier on most agentic evals, per Section 5's SWE-bench Pro table where it posts 46.1%. For high-volume workloads where cost per task dominates, Google's column wins even where its peak scores do not; the down-market economics are in our Gemini 3.5 Flash guide.
xAI is this edition's biggest single correction, and the correction is the information. The July edition reported Grok 4.5 in private beta with no public access and zero independently verified scores, and framed xAI as the lab that had opted out of measurement. That narrative lasted eighteen days: Grok 4.5 was publicly released July 16, 2026, available via API at $2/$6 per million tokens with a 500K context window - llm-stats. The honest August status is "released, priced aggressively, still benchmark-sparse": the model's tracker pages show benchmark sections unpopulated and scores pending, so aggressive pricing is currently the most verifiable thing about it. Our Grok 4.5 guide tracks the coverage as it fills in.
The open-weights tier stopped being a footnote and became a per-benchmark contender. Kimi K3 (Moonshot, July 16) is the highest open model on the independent Intelligence Index at 57 and tops BrowseComp outright at 91.2%. Qwen3.8 Max (Alibaba, August 2) broke into Arena's top five within days of release. DeepSeek-V4-Flash-0731 (July 31) debuted at 50 on the Intelligence Index as an explicitly fast, open release - llm-stats, with its cost math against closed flagships worked in our DeepSeek V4-Flash comparison. The 2025 framing of open weights as trailing baselines is dead; the operative question in August 2026 is deployment preference, not capability class.
10. The Trust Problem: Self-Reported Scores and Harness Effects
Everything in this guide so far assumes the published numbers mean something, and that assumption needs its own section, because the biggest evaluation story of 2026 is not any model's score: it is who measured it. As benchmark results became the primary marketing surface for billion-dollar launches, the incentive to optimize the reporting, not just the model, grew accordingly, and the verification infrastructure is only partially catching up.
The cleanest exhibit this month is BrowseComp: 58 tracked results, every one self-reported, zero independently verified - llm-stats, on a leaderboard whose top five sit within four points of each other. When the spread between "winning" and fifth place is smaller than the plausible variance between labs' private harness configurations, the ranking is not information, it is formatting. SWE-bench Verified decayed the same way before it (its leaderboard has long been dominated by vendor self-reports, which is why this guide routes coding decisions to Scale's independently run Pro splits instead), and the general lesson compounds: a benchmark's trust profile decays faster than its difficulty, because self-reporting arrives the moment a leaderboard becomes marketing-relevant, which is well before the tasks get easy.
The harness effect is the second, less appreciated distortion, and this month's Terminal-Bench top five states it with unusual precision: the same GPT-5.5 model scores 84.7% under NexAU-AHE, 83.1% under Capy, and 82.2% under Codex CLI - tbench.ai. Two and a half points of pure scaffolding, inside the top five of one leaderboard, before counting the older harnesses further down. Any agentic benchmark comparison that does not name the harness is incomplete to the point of meaninglessness, and for buyers the harness is part of the product: a well-engineered agent platform wrapped around a second-tier model routinely outperforms a frontier model in a naive loop. The third distortion is newer and subtler: the best real-work eval we have, GDPval-AA, uses an LLM judge for its pairwise grading - Artificial Analysis, which independent execution mitigates but does not eliminate, since a systematic judge preference (for length, structure, or a family's house style) would move every Elo in the table and be invisible in the error bars.
Three structural checks separate trustworthy numbers from marketing numbers in 2026:
- Independent execution: was the score produced by a third party (Artificial Analysis, METR, Epoch, Scale's splits) or by the vendor
- Named harness and settings: does the report specify scaffolding, retry budget, tool access, and effort configuration
- Version and date: is the benchmark version stated, and is the score current against a leaderboard that actually updates
Apply those checks to any model card and most of its benchmark table fails at least one; apply them to this article and you will find the vendor claims explicitly flagged as such (Opus 5's ARC-AGI-3 multiple, its OSWorld 2.0 cost claim, FrontierMath's tracker-reported 89%). That is not nihilism, it is a reading protocol. The verified subset of the evidence is smaller than the reported subset, but it is entirely sufficient to make good decisions, and it is the subset this guide quotes.
11. Versioning, Error Correction, and Contamination Defenses
If 2024-2025 was the era of new benchmarks, 2026 is the era of repaired and re-versioned benchmarks, and the repair wave is better news than any new leaderboard: it means the field audits its own instruments, finds them wanting, and fixes them in public. But it also means every score now needs metadata (version, date, runner) attached before it can be compared to anything, and this section is the maintenance manual.
The re-versioning events now span every major category, and the maintainers have started saying the quiet part in their changelogs. FrontierMath v2 (June 12, 2026) rebuilt the set to 338 problems after a correction review, with only twelve public - Epoch AI; pre-v2 scores were measured against a partially different answer key. OSWorld went original, then Verified, then 2.0 within its lifetime, with the 2.0 supersession dated June 26, 2026 on the project's own banner. tau2-bench v1.0.1 ships the bluntest version of the rule ever written by a benchmark team: "results produced with tau2-bench < 1.0.1 are not comparable with >= 1.0.1" - Sierra Research. BFCL rolled through four versions, with V4's holistic agentic scoring incompatible with the single-call accuracy of V1-V3 - Berkeley. Each event rearranged rankings without any model changing, which is the point: some fraction of every pre-repair leaderboard was measurement error, not capability difference.
The complementary story is contamination defense, because repair fixes broken tasks but not leaked ones. The defenses form a clear hierarchy of strength. Static-but-private task pools (FrontierMath's unpublished problems, SWE-bench Pro's commercial and held-out splits) defeat direct memorization but depend on the maintainer's operational security. Post-cutoff generation (LiveCodeBench) is stronger because leakage is chronologically impossible. Interactive environments (ARC-AGI-3's games, OSWorld 2.0's live desktop, Terminal-Bench's real shell) are stronger still, because there is no answer key to leak: the agent must produce behavior, not recall. And July 2026 added the endpoint of the hierarchy: unsolved-problem benchmarks (FrontierMath: Open Problems), where no answer exists anywhere, and a solve is simultaneously a score and a contribution to human knowledge.
For practitioners, the discipline reduces to a habit: never cite a score without its version, date, and runner, and never compare across versions. A "SWE-bench" number without a suffix is uninterpretable. An HLE score without a runner might describe the official page's years-old snapshot or this month's independent run, a fifteen-point difference. An OSWorld number without a version is off by a factor of four. The July edition of this guide is itself the cautionary example twice over: every headline number it printed was correct on July 8, most were stale by August 5, and one (OSWorld-Verified as the live frontier) was already superseded the day it was published. Treat this edition identically: the linked leaderboards outrank this text the moment you read it.
12. What It Costs to Run Evals in August 2026
Benchmark scores are produced by spending tokens, and the price list moved again since July, in one direction at the top and with one deadline at the mid-tier. All prices below were verified against the official pricing pages this run, per million input/output tokens.
Anthropic - pricing: Claude Fable 5 and Mythos 5 at $10/$50 (Mythos 5 now flagged limited availability), Claude Opus 5 at $5/$25 (the frontier-adjacent tier at half Fable's price, which is the pricing event of the summer), the Opus 4.5-4.8 line also at $5/$25, and Claude Haiku 4.5 at $1/$5. Claude Sonnet 5's introductory $2/$10 ends August 31, 2026: from September 1 the standard price is $3/$15, so any cost model you build this month should already use the higher number; our Sonnet 5 cost breakdown covers that tier. New since July: a research-preview fast mode for Opus 5 and Opus 4.8 at $10/$50, trading money for latency at exactly Fable-5 list price. Batch is 50% off, cache reads are 0.1x, and the 1M-token context window is standard-priced on Claude 4.6 and later.
OpenAI - pricing: the GPT-5.6 family prices as a clean ladder: Sol $5/$30, Terra $2/$12, Luna $0.20/$1.20, with long-context at 2x base rates. GPT-5.5 holds at $5/$30, the Pro tiers (GPT-5.5 Pro, GPT-5.4 Pro) at $30/$180, GPT-5.4 at $2.50/$15, mini at $0.75/$4.50, nano at $0.20/$1.25. Cached input is discounted 90% and batch is 50% off. Google - pricing: Gemini 3.1 Pro Preview at $2/$12 under 200k context ($4/$18 above), the new Gemini 3.6 Flash at $1.50/$7.50, Gemini 3.5 Flash-Lite at $0.30/$2.50, and batch at 50% off across the line. xAI: Grok 4.5 at $2/$6 - llm-stats, the most aggressive flagship list price on the board, with the benchmark coverage to justify it still pending.
List prices understate real eval costs in three specific ways that anyone budgeting a benchmark run must model. First, tokenizer inflation: per Anthropic's own pricing doc, Claude 4.7 and later models (and Claude Mythos Preview) use a newer tokenizer that produces approximately 30% more tokens for the same text, while Sonnet 4.6 and earlier use the previous one - Anthropic, which silently inflates every cross-generation cost comparison. Second, reasoning-token blowup: effort dials and agentic loops generate thinking and tool transcripts that dwarf the visible answer; OSWorld 2.0's 318 tool calls per task is what "one task" costs now, and a full run of a 2,158-question exam or a 731-task coding split lands in the hundreds to thousands of dollars per model at flagship rates. Third, infrastructure add-ons: on Anthropic's meter, web search runs $10 per 1,000 searches, managed agent sessions bill $0.08 per session-hour, and code execution is free alongside web tools (otherwise $0.05 per container-hour after 1,550 free monthly hours).
The mitigation stack is the same one production workloads use: batch APIs (50% off at all three major labs) for anything not latency-sensitive, cache reads at a tenth of input price for shared benchmark prompts, and mid-tier models for pre-screening before spending flagship tokens. The full cost-modeling treatment lives in our AI model benchmarks and pricing guide. The strategic point stands independent of any single price: capability-per-dollar improved again this summer (Opus 5 delivering near-Fable intelligence at half the price is the clearest single instance), which means running your own evaluations, once a luxury reserved for labs, is now affordable for any team automating a meaningful workload. Section 13 says exactly how.
13. How to Pick Benchmarks for Your Use Case
Fifty benchmarks is a research syllabus, not a decision tool. A business choosing a model or agent platform in August 2026 needs five to ten numbers, selected by workload, all from the live column of the Section 2 ledger, all passing the Section 10 trust checks. This section is that selection guide, organized by the four workloads that cover most real deployments.
For coding agents, read SWE-bench Pro first (with the caveat that this month's public leader is unattributed), Terminal-Bench 2.0 second with attention to the harness column, and LiveCodeBench for contamination-proof freshness; ignore Verified and HumanEval entirely. For computer-use automation, OSWorld 2.0 is now the primary signal, and read it two-layered: the binary completion rate (20.6% at the top) tells you what runs unattended, the partial score (54.8%) tells you what runs with a human checkpoint, and the gap between them prices the supervision you still need to budget. For research and retrieval workloads, treat BrowseComp as a self-reported directional signal only, and lean on HLE's independent run for depth of expert knowledge. For customer-facing conversational agents, tau2-bench v1.0.1 (including its new voice mode, if you are deploying audio) and BFCL V4 matter more than every academic exam combined; our broader agent evals guide covers this tier in depth.
Two cross-cutting rules complete the method. First, always convert to cost per completed task, not cost per token: a cheaper model with a lower completion rate can be more expensive per merged pull request or per finished workflow, and Section 12's price list only becomes decision-grade after multiplying by success rates, retry counts, and (for computer use) OSWorld-2.0-scale tool-call volumes. Second, run your own eval on 20-50 tasks drawn from your actual workload before committing, because public benchmarks are population averages and your task distribution is not the population. The public numbers narrow the shortlist to two or three candidates; your private eval picks the winner. This is the pattern mature agent deployments follow, whether built in-house or run on platforms like O-mega, where the operative benchmark is the completion rate on your own recurring business tasks, tracked week over week.
The final heuristic is a time discipline, and this edition upgrades it: quarterly re-checks are no longer enough at the top of the market. Between July 8 and August 5, the GDPval and Intelligence Index leads changed, FrontierMath jumped 36 points, the Arena top five broke, and a benchmark this guide called the computer-use frontier was confirmed superseded. Re-check the leaderboards for your two or three decision-critical benchmarks monthly, and re-run your private eval whenever a model one price tier below your incumbent posts scores within a few points of it, because that is the event that actually moves your unit economics.
14. Conclusion: The Evaluation Playbook, Rewritten Again
The through-line of this guide is that benchmarks are instruments, and instruments wear out, and the four weeks since the last edition provided the cleanest demonstration in the field's history: computer use crossed the human baseline on one instrument and scored 20.6% on its replacement, math's "unsolvable" benchmark hit 89% and spawned a set of genuinely open problems, and the yardstick for autonomous work posted a notice that it cannot reliably measure past 16 hours. None of those events were model failures. They were measurements catching up to reality, which is what good instruments do right before they get replaced.
The decision playbook compresses to five moves. Screen with meta-benchmarks (Arena, ECI, the Intelligence Index) to build a shortlist, remembering the top five sit within a coin flip of each other. Deep-dive with workload-matched live evals from Section 13. Apply the trust checks: independent runner, named harness, stated version and date. Convert to economics: cost per completed task at verified August 2026 prices, with tokenizer inflation, effort-dial blowup, and the Sonnet 5 price step-up on September 1 modeled in. Verify on your own tasks, and re-verify monthly at the frontier. A team that follows those five steps will outperform any team choosing by headline score, whatever the headlines say by the time you read this.
And the honest closing note is this edition's own correction record. A month ago this guide reported an all-Anthropic Arena top five (an open Alibaba model broke it), a lab that had opted out of public measurement (it shipped a public API eighteen days later), and a live computer-use benchmark (it had been superseded before publication). Every one of those corrections made the article more useful, which is the entire argument for reading a dated, versioned, verdict-carrying ledger instead of a static list: the value is not that it is always right, it is that it tells you exactly when and how it went stale. The evaluation landscape finally mirrors the market it measures: competitive, versioned, audited, and moving too fast for any snapshot, this one included, to stay true for long.
Written by Yuma Heymans (@yumahey), founder and CEO of O-mega. Running an agent-workforce platform means every benchmark claim in this guide eventually gets tested the hard way: against paying customers' real tasks.
This guide reflects the AI benchmark landscape as of August 5, 2026. Scores, leaderboards, and prices in this space now change weekly at the frontier: verify current numbers against the linked leaderboards before making decisions.