The practical guide to what actually changed between Claude Opus 4.8 and Claude Opus 5, what it costs, and which one you should run.
Anthropic shipped Claude Opus 5 on July 24, 2026 at exactly the same price as Claude Opus 4.8: $5 per million input tokens and $25 per million output tokens - Fortune. That single fact is the whole story, and it is also the trap. When two models carry an identical price sheet, people assume the choice between them is a coin flip, or that the newer one is simply better in every way. Neither is true. The interesting question is not "which is cheaper per token" (they are the same), it is "which one spends fewer tokens to finish your job, and how much more capable is the extra spend actually buying."
Here is the problem the sticker price hides. Claude Opus 5 runs with extended thinking on by default, it is measurably more verbose, and on identical tasks at the same effort setting it can burn more output tokens than Opus 4.8 while charging the same rate for each one. So the real bill can go up even though the price per token did not move. In independent testing, running the full Artificial Analysis Intelligence Index cost $3,835.51 on Opus 5 versus $3,752.55 on Opus 4.8, and Opus 5 generated roughly 100 million output tokens during that evaluation against a model-median of about 63 million. Same price, more tokens, bigger invoice. Whether that trade is worth it depends entirely on your workload.
This guide breaks down every benchmark where the two models have been measured side by side, the exact pricing mechanics that decide your monthly bill, the behavioral changes that quietly move your token spend, and a decision framework for choosing between them (and against the wider frontier field of GPT-5.6, Gemini 3.1 Pro, Kimi K3, and Anthropic's own Claude Fable 5). It is written for people who are actually paying the invoice, not just reading the launch tweet.
Contents
- The one-sentence answer
- The frontier field at a glance (scored)
- What actually changed under the hood
- Benchmarks: the head-to-head that matters
- The independent read: where the marketing and the leaderboards diverge
- Cost part one: why identical rates do not mean identical bills
- Cost part two: effort, caching, and batch as the real levers
- Cost part three: cost in the wild (Claude Code, Max, Bedrock, Vertex)
- Migrating from 4.8 to Opus 5 without a surprise invoice
- Where Opus 5 sits in the wider field
- First principles: what cheaper intelligence at the same price changes
- The decision framework
- Sources and further reading
The frontier field at a glance (scored). Before the deep dive, here is the entire relevant field ranked on a single weighted scale so you can see where Opus 5 and Opus 4.8 actually land relative to each other and to their competitors. Each cell carries the score and the data point behind it, so the number is not a black box. The methodology and its caveats follow the table.
| # | Model | What It Does | Intelligence (25%) | Agentic Coding (30%) | Cost Efficiency (25%) | Speed (10%) | Ecosystem (10%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic's frontier agentic model, half of Fable 5 price | 10 - AA Index 61, #1 of 191 | 10 - Terminal-Bench 89.1%, SWE-Pro 79.2% | 7 - $5/$25 but ~$2.03/task (verbose) | 4 - 54.8 t/s, 65s to first token | 10 - API, AWS, Vertex, Foundry, Max default | 8.7 |
| 2 | GPT-5.6 Sol | OpenAI's flagship coding variant | 9 - AA Index 59 | 10 - SWE-Verified 96.2%, Terminal-Bench 89.5% | 6 - premium tier, pricing undisclosed | 6 - moderate | 8 - OpenAI API + Azure | 8.2 |
| 3 | Claude Opus 4.8 | The prior Opus, still fully supported | 8 - AA Index 56 | 8 - SWE-Verified 88.6%, Terminal-Bench 74.6% | 8 - $5/$25, ~$1.80/task, Priority Tier | 5 - 55.2 t/s, 29s to first token | 10 - everywhere, fast mode + Priority | 7.9 |
| 4 | Gemini 3.1 Pro | Google's flagship, long-context strength | 8 - strong reasoning, cheap tokens | 6 - SWE-Verified ~80.6% | 9 - $2/$12 per 1M | 8 - fast serving | 9 - Google Cloud scale | 7.8 |
| 5 | Kimi K3 | Strongest open-weights coder | 7 - AA Index 57 | 8 - SWE-Verified 93.4% | 9 - open weights, cheap hosting | 6 - provider-dependent | 6 - self-host or third-party | 7.6 |
| 6 | Claude Fable 5 | Anthropic's top-tier frontier model | 10 - AA Index 60, frontier | 9 - SWE-Verified ~95%, SWE-Pro ~80% | 4 - $10/$50, ~$2.75/task | 4 - very long turns | 8 - API only, 30-day retention required | 7.4 |
| 7 | GLM-5.2 | Deep-value open-weights model | 6 - mid-tier reasoning | 7 - SWE-Pro 62.1% | 10 - $1.40/$4.40 per 1M | 7 - fast | 6 - open/self-host | 7.4 |
Criteria and weights: Intelligence (25%) uses the Artificial Analysis Intelligence Index at max effort. Agentic Coding (30%) is weighted highest because it is what most people actually buy Opus for, drawing on SWE-bench Verified, SWE-bench Pro, and Terminal-Bench. Cost Efficiency (25%) blends the published per-token price with real cost-per-task, not just the rate. Speed (10%) uses measured output tokens per second and time-to-first-token. Ecosystem (10%) covers availability across clouds, tooling, and subscription access. Rows are ordered by final score descending, and where two models tie (Fable 5 and GLM-5.2 at 7.4) they are ordered alphabetically. Competitor pricing and per-variant scores carry more uncertainty than the Anthropic figures, which is exactly why the sections below separate what is firmly sourced from what is not.
1. The one-sentence answer
If you want the whole guide compressed to a sentence: Claude Opus 5 is a real generational step over Opus 4.8 on hard, agentic, long-horizon work at the same per-token price, but it spends more tokens to get there, so pick 4.8 when the task is routine and cost-sensitive, and pick Opus 5 when the task is difficult enough that raw capability, not token thrift, decides the outcome. Everything else in this guide is the evidence, the mechanics, and the edge cases behind that sentence.
The reason the one-sentence answer is not simply "use the new one" comes down to a structural fact about how these two models are priced and how they behave. Anthropic did not change the price. Claude Opus 5 arrived at $5 per million input tokens and $25 per million output tokens, the identical tier that Opus 4.8, Opus 4.7, Opus 4.6, and Opus 4.5 have all occupied - Claude pricing documentation. In an industry where every generation usually raised the flagship price, holding it flat while pushing capability up is genuinely aggressive. Anthropic's own framing is that Opus 5 delivers "greatly improved performance for the same cost as its predecessor" and comes "close to the frontier intelligence of Claude Fable 5 at half the price" - Anthropic.
But a flat per-token price is not a flat bill. The two models spend tokens differently, and Opus 5 spends more. That is the crux. A model that is 9 percent smarter on a benchmark index but 5 to 15 percent more expensive to run on a given task has not made your decision for you, it has handed you a genuine trade-off. For a chatbot answering support tickets, the extra capability is invisible and the extra cost is real, so 4.8 wins. For an autonomous agent refactoring a large codebase overnight, the extra capability is the difference between a merged pull request and a broken build, and the extra cost is a rounding error against an engineer's salary, so Opus 5 wins. The skill in 2026 is not picking "the best model," it is matching the model to the altitude of the task, which is a theme we have returned to repeatedly in our guide to building AI agents.
This is also why "benchmarks and cost" belong in the same guide rather than two separate ones. A benchmark number in isolation tells you which model is more capable. A price in isolation tells you which is cheaper. Neither answers the only question that matters operationally, which is capability per dollar on your actual workload. Anthropic's own launch charts understood this: nearly every one plots score against cost-per-task, not score alone, precisely because the frontier of interest is the trade-off curve, not the peak.
2. What actually changed under the hood
Before a single benchmark, it is worth understanding the mechanical differences between the two models, because those differences drive both the capability gains and the cost changes. Opus 5 is not just "Opus 4.8 with more training." Several concrete API behaviors changed, and each one has a downstream effect on what you pay and how the model acts. Understanding these first makes every later number make sense, and it prevents the single most common migration surprise, which is a bill that jumps for reasons the price sheet does not explain.
The headline mechanical change is that extended thinking is on by default on Opus 5. On Opus 4.8, if you sent a request without a thinking parameter, the model answered directly with no reasoning tokens. On Opus 5, that same parameter-free request runs adaptive thinking automatically - Claude platform docs. This is the root cause of most of the cost delta. Thinking tokens are billed as output tokens at the full $25 per million rate, and they count against your max output budget. A workload that was tuned on 4.8 to answer in 2,000 tokens can suddenly produce 2,000 tokens of answer plus several thousand tokens of thinking, and if your max_tokens ceiling was set tight, the response can even truncate mid-answer.
A second change tightens how you can turn that behavior off. On Opus 5, thinking cannot be disabled at the two highest effort levels. Sending a disabled-thinking request at xhigh or max effort returns a 400 error, whereas Opus 4.8 accepted that combination at any effort. The practical implication is that if you want a fast, cheap, non-reasoning response on Opus 5, you must also drop the effort to high or below. The two settings are now coupled in a way they were not before, and code that carried a disabled-thinking configuration forward from 4.8 will break the moment it also requests high-intensity effort.
The remaining structural changes cut in your favor. A short reference of the deltas that matter for cost and behavior:
- Prompt-cache minimum dropped to 512 tokens on Opus 5, down from 1,024 on Opus 4.8, so shorter prefixes now cache and earn the cheap read rate.
- Tool-use system-prompt overhead shrank to roughly 286 to 406 tokens on both Opus 5 and 4.8, versus 675 to 804 on Opus 4.7, cutting fixed cost per agentic request.
- Opus 5 has a separate rate-limit bucket from the combined Opus 4.x pool, so shifting traffic to it neither frees nor inherits your existing Opus headroom.
- Mid-conversation tool changes and a server-side fallbacks parameter are new on Opus 5, letting cyber-category refusals route automatically to Opus 4.8.
Each of these is small on its own, but together they reshape the cost profile of an agentic workload. The lower cache minimum matters enormously for agents that resend a large stable system prompt on every turn: with a 512-token floor, more of that prefix qualifies for the roughly 0.1x cache-read price. The reduced tool overhead matters for agents that carry dozens of tool definitions in context. And the separate rate bucket matters for anyone running at scale, because it means capacity planning for Opus 5 is a fresh exercise, not a migration of existing limits. If you are optimizing an agent's economics at this level of detail, the mechanics in our Claude Agent SDK deep dive go deeper on how caching and tool overhead compound across a long run.
The behavioral changes are just as consequential as the API ones, and they are the part most teams underestimate. Opus 5 verifies its own work without being told to, it writes longer responses and longer files by default, it delegates to subagents more readily than 4.8, and it can expand the scope of a task beyond what you literally asked. Anthropic's own migration guidance tells developers to delete the "double-check your answer" instructions they wrote for older models, because on Opus 5 those instructions now cause redundant over-verification. Every one of those behaviors spends tokens. None of them is a bug. They are the model being more autonomous and more thorough, which is exactly what you want for a hard task and exactly what you do not want to pay for on an easy one.
3. Benchmarks: the head-to-head that matters
Now the numbers. The important framing before any specific score is that Anthropic itself did not lead its Opus 5 announcement with a clean table of SWE-bench percentages. Instead it emphasized a set of harder, more agentic evaluations plotted against cost, which is a tell about where the real gains landed. The gains are largest on tasks that are difficult, multi-step, and open-ended, and smallest on the saturated single-turn benchmarks that older models already handled well. That pattern is the single most useful thing to understand about this generation, and it maps directly onto the one-sentence answer from section one.
Start with Frontier-Bench v0.1, Anthropic's own agentic-coding evaluation and the benchmark it chose to headline. On this 74-task successor to Terminal-Bench, Opus 5 scored 43.3% at max effort (and 44.4% at xhigh) versus 18.7% for Opus 4.8 - MarkTechPost. That is more than double, which matches Anthropic's official phrasing that Opus 5 "more than doubles" its predecessor's Frontier-Bench performance at a lower cost per task. The official chart makes the trade-off vivid, because it plots score against dollars spent per attempt.
Read that chart carefully, because it contains the entire cost argument in one picture. Opus 5's cheapest data point sits at roughly $5.50 per attempt and scores about 25 percent, which is already higher than Opus 4.8's most expensive data point at roughly $16 per attempt scoring 18.7 percent. In other words, on this agentic-coding task, the whole Opus 5 curve dominates the Opus 4.8 curve: at any given budget, Opus 5 delivers more, and to match Opus 5's floor, Opus 4.8 cannot spend enough. This is the honest version of "same price, more capability," and it is far more persuasive than any single headline percentage.
The pattern repeats across the agentic evaluations, though the magnitudes vary. A same-version comparison of the two models on the coding and computer-use benchmarks where both have published or attributed scores:
| Benchmark | Opus 4.8 | Opus 5 | What it measures |
|---|---|---|---|
| SWE-bench Pro | 69.2% | 79.2% | Harder, verified software engineering tasks |
| Terminal-Bench v2.1 | 74.6% | 89.1% | Autonomous terminal and shell task completion |
| Frontier-Bench v0.1 | 18.7% | 43.3% | Anthropic's hardest agentic-coding eval |
| OSWorld 2.0 | 55.7% | 70.6% | Computer use in a real desktop environment |
The chart below turns that table into the visual it deserves, because the gap is easier to feel than to read. Note that these are all same-version comparisons: OSWorld 2.0 is the newer, harder variant, not the older OSWorld-Verified where Opus 4.8 scores 83.4%, and mixing the two versions in one column would be a common and misleading error.
The most dramatic single number is on ARC-AGI-3, the novel-problem-solving benchmark from the ARC Prize foundation. Here Opus 5 reached roughly 30.2% at high effort against just 1.5% for Opus 4.8, which Anthropic characterizes as roughly three times the next-best published model - Anthropic. ARC-AGI is deliberately designed to resist memorization and reward genuine reasoning on unfamiliar puzzles, so a 20x jump on it is a strong signal that Opus 5's improvement is not just benchmark-gaming on familiar problem shapes. The official chart, again plotted against total evaluation cost, shows Opus 5 sitting alone at the top.
Independent verification from the ARC Prize team lines up with Anthropic's claims and adds detail. Their verified results put Opus 5 at 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2 at max effort - ARC Prize. That is important because ARC Prize runs the evaluation itself rather than accepting a vendor's self-report, so these are among the most trustworthy numbers in the entire comparison. The gains extend well beyond coding, too. Anthropic reports that Opus 5 scores 10.2 percentage points higher on organic-chemistry tasks and 7.7 points higher on protein tasks than Opus 4.8, which matters for anyone using these models in scientific or life-sciences workflows rather than pure software.
Where does that leave the most famous benchmark of all, SWE-bench Verified? In an unusually honest place. Third-party leaderboards such as Vals.ai list Opus 5 at roughly 96 to 97% against Opus 4.8's 88.6%, which would mean the benchmark is effectively saturated. But Anthropic's own announcement page pointedly did not publish an exact SWE-bench Verified figure for Opus 5, and the third-party numbers disagree with each other (96 versus 97) and in one case predate the launch. The responsible reading is that Opus 5 is at or near the top of SWE-bench Verified, that the benchmark is close to saturation and therefore losing its discriminating power, and that Frontier-Bench and the ARC-AGI suite are now the more informative signals. That nuance is exactly why we treat any single decimal-point SWE-bench claim with suspicion, a discipline we applied throughout our Opus 4.8 benchmark guide as well.
Anthropic's second headline coding evaluation, CursorBench 3.2, tells the same story in a different register. It measures agentic coding inside a realistic editor workflow rather than an isolated task harness, and Opus 5 reaches roughly 70% against Opus 4.8's low-to-mid-60s, again plotted against dollars spent per task so the trade-off stays visible. The point of showing a second coding chart is not redundancy, it is confirmation that the improvement generalizes across coding contexts rather than being an artifact of one benchmark's particular task distribution.
What both coding charts share is a subtle but important detail: Opus 5's curve is not just higher than Opus 4.8's, it starts at a lower cost point and rises more steeply, so the capability gain is not simply "spend more, get more." Opus 5 is genuinely on a better cost-performance frontier for agentic coding, which is the honest technical basis for Anthropic's "same price, more capability" claim. On these two evaluations at least, the marketing and the measured data agree.
Beyond coding, the two models are closer than the coding gap suggests on saturated knowledge benchmarks, and this is itself informative. Opus 4.8 already scored 93.6% on GPQA Diamond graduate-level science questions and 57.9% on Humanity's Last Exam with tools, so Opus 5's improvements on these established evaluations are incremental rather than dramatic, simply because there is little headroom left at the top. Where Opus 5's scientific gains are clean and specific is on the specialized, unsaturated tasks: Anthropic reports 10.2 percentage points on organic chemistry and 7.7 points on protein sequences over 4.8. That pattern (small gains where the benchmark is already near-solved, large gains where it is not) is the through-line of the entire generation, and it is why a buyer should care far more about Frontier-Bench and ARC-AGI than about another decimal point on a benchmark everyone already aces.
4. The independent read: where the marketing and the leaderboards diverge
Vendor benchmarks, even honest ones, are chosen by the vendor. The more valuable exercise is to look at independent evaluators who run the same harness across every model and publish the cost of doing so. The single most useful of these is Artificial Analysis, whose Intelligence Index aggregates a broad basket of evaluations into one comparable number and, critically, reports the dollar cost of running that basket on each model. Their read on Opus 5 versus 4.8 is both more measured and more useful than any launch chart, and it is where the "same price, bigger bill" story becomes concrete rather than theoretical.
On the Artificial Analysis Intelligence Index, Opus 5 at max effort scores 61, ranking first of the 191 models they track, while Opus 4.8 at max effort scores 56. A five-point gap on this composite is a real generational step, but it is a step, not a leap, and it puts the "modest on paper" framing in perspective. In their own words Opus 5 (61) is narrowly the most intelligent model available, effectively tied with Claude Fable 5 (60) and ahead of GPT-5.6 Sol (59) and Kimi K3 (57). The chart below shows the top of that leaderboard at max effort.
The cost side of that same evaluation is where the surprise lives, and it is the number every buyer should internalize. Running the full index cost $3,835.51 on Opus 5 versus $3,752.55 on Opus 4.8, and Artificial Analysis flagged Opus 5 as unusually verbose, generating around 100 million output tokens during evaluation against a model-median of roughly 63 million. At the per-task level, they measured Opus 5 at max effort at about $2.03 per task versus $1.80 for Opus 4.8, with Fable 5 at $2.75 - Artificial Analysis. This is the empirical confirmation of the mechanical argument from section two: identical per-token pricing, higher token consumption, therefore a higher bill. Opus 5 costs roughly 13 percent more per task than 4.8 despite an identical price sheet, purely because it thinks and writes more.
One number on that same evaluation is worth pulling out because it reframes the "modest step" reading. On GDPval-AA v2, a benchmark built to approximate economically valuable professional work rather than academic puzzles, Opus 5 set the highest score Artificial Analysis has measured to date at 1,861 Elo - Artificial Analysis. Composite indices compress a lot of variance into a single point, so a five-point index gain can hide the fact that on the specific dimension of real, revenue-bearing knowledge work, Opus 5 opened a clear lead. This is the tension a careful buyer has to hold: the gap looks small on the aggregate and large on the tasks that most resemble paid professional output, which is precisely the kind of task an autonomous agent is deployed to do.
The verbosity finding deserves one more beat, because it is easy to read "generated 100 million output tokens" as a mere curiosity rather than a design signal. Opus 5 is verbose because it thinks more, verifies more, and narrates more, and those are the same behaviors that lift its completion rate on hard tasks. So the extra tokens are not waste in the abstract, they are the mechanism of the capability gain. The operational lesson is that you cannot have Opus 5's thoroughness and Opus 4.8's token thrift at the same time from the same model. If you want thrift, you use 4.8 or a smaller model, and if you want thoroughness, you pay for the tokens that produce it. Pretending you can prompt away the cost while keeping the capability is the mistake that produces both a disappointing model and a confusing bill.
Speed is the other place where the independent numbers puncture a comfortable assumption. Both Opus models are slow relative to the reasoning-model field. Artificial Analysis measured Opus 5 at about 54.8 tokens per second with a 65-second time-to-first-token, and Opus 4.8 at 55.2 tokens per second with a much faster 29-second time-to-first-token, against a reasoning-model median around 72 to 76 tokens per second. Notice that Opus 5 is actually a touch slower than 4.8 on throughput and dramatically slower to first token, because more of its thinking happens before the first visible token appears. For an interactive product where a human is staring at a loading spinner, that 65-second first-token latency is a genuine user-experience cost that no benchmark score captures. This is one more reason the choice is a trade-off rather than a clean upgrade.
The most important thing the independent data reveals is that "best model" breaks down by benchmark. Artificial Analysis ranks Opus 5 first on its composite index, but on Terminal-Bench v2.1 the leader is GPT-5.6 Sol at 89.5% with Opus 5 just behind at 89.1%, and on ARC-AGI-2 some third-party trackers put GPT-5.6 Sol ahead of Opus 5. Meanwhile developer reception has been mixed in an instructive way: in one widely discussed review, Simon Willison described Opus 5 as "brilliant but annoying," with a personality that in one session refused to touch a merge conflict and kept working long past where a human would have stopped. That is the flip side of "verifies its own work and finishes the job." It is a model with strong opinions about how a task should be done, which is a feature for autonomous overnight runs and a friction for quick interactive edits. Anyone deciding between the two should weigh temperament, not just benchmarks, a point that also came up in our Claude Sonnet 5 benchmarks and cost breakdown.
Temperament is an underrated selection criterion precisely because no benchmark reports it, yet it shapes daily experience more than a two-point index gap. A model that keeps working until the job is genuinely finished is exactly what you want supervising an unattended agent at 3 a.m., and exactly what you do not want when you asked for a one-line change and got a thorough refactor plus a paragraph of unsolicited advice. Opus 4.8's more compliant, more literal disposition makes it the calmer partner for tight, well-scoped interactive work, while Opus 5's autonomy and persistence pay off when the task is ambiguous and long-running and you would rather the model push through than stop to ask. In practice, teams that run both often reserve Opus 5 for the asynchronous, high-stakes lane and keep Opus 4.8 in the interactive, human-in-the-loop lane, which is a temperament decision as much as a cost or capability one. Benchmarks tell you which model is smarter; only usage tells you which one you can stand to work with.
5. Cost part one: why identical rates do not mean identical bills
The pricing is genuinely identical between the two models, and it is worth stating the exact figures so there is no ambiguity, because the equality is the foundation of every cost decision that follows. Both Claude Opus 5 and Claude Opus 4.8 cost $5.00 per million input tokens and $25.00 per million output tokens on the standard first-party API - Claude pricing documentation. Prompt caching is also priced identically: a five-minute cache write costs 1.25x the base input rate at $6.25 per million, a one-hour cache write costs 2x at $10 per million, and a cache read (a hit) costs just 0.1x at $0.50 per million. The Batch API applies a flat 50% discount to both models, bringing them to $2.50 input and $12.50 output. Nothing on the price sheet distinguishes the two.
So if the rates are the same, the entire cost difference is a token-volume difference, and the token volume is driven by behavior, not billing. This is the single most important idea in the guide, and it is worth belaboring because it inverts most people's intuition. A newer, "better" model on the same price sheet feels like it should be a pure upgrade. But when the newer model thinks by default, verifies by default, delegates by default, and writes longer by default, it consumes more of the thing you pay for. The diagram below traces exactly how a byte-identical price sheet produces two different invoices.
The practical consequence is that your Opus 4.8 cost model does not transfer to Opus 5 unchanged. If you budgeted a workload at a certain dollar figure on 4.8 and swap the model string to Opus 5 without touching anything else, expect the bill to rise by roughly 10 to 15 percent on comparable tasks at the same effort level, purely from the extra tokens. Whether that is acceptable depends on what the extra tokens buy. On a hard agentic task where Opus 5 finishes a job that 4.8 would have failed or looped on, the higher per-task cost is dwarfed by the value of a completed task and the avoided cost of a human cleaning up a failure. On a simple classification or extraction task where both models would have succeeded, the extra tokens buy nothing and the 4.8 bill is strictly better.
There is one more subtlety that cuts across both models and inflates the effective cost of text versus older Opus generations. Opus 5 and Opus 4.8 both use the newer tokenizer introduced with Opus 4.7, which produces roughly 30% more tokens for the same text than Opus 4.6 and earlier. The per-token price did not change, but the number of tokens per page of English (or per file of code) went up, so the effective cost per unit of real content rose across the whole 4.7-and-later line. If you are migrating from an older Opus, do not reuse token counts measured on 4.6 to forecast Opus 5 spend, because you will underestimate by a meaningful margin. Re-baseline against the current tokenizer instead.
Finally, a genuinely favorable cost change worth not overlooking. Both models carry the full 1M-token context window with no long-context surcharge, so a 900,000-token request is billed at the same per-token rate as a 9,000-token one. Many competitors charge a premium above 200,000 tokens, so for retrieval-heavy or whole-repository workloads this is a real structural advantage of the Opus line that partly offsets the verbosity tax. We break down how that long-context economics compares across providers in our AI model benchmarks and pricing analysis, and it is one of the reasons whole-codebase agents have gravitated to Claude.
6. Cost part two: effort, caching, and batch as the real levers
If the price sheet is fixed and the model's default behavior drives token volume, then the way you actually control your bill is through the levers that change how many tokens get spent. There are three that matter most, and using them well is the difference between an Opus 5 deployment that costs a fortune and one that costs less than an Opus 4.8 deployment that was left on default settings. The levers are effort, prompt caching, and batching, and each deserves a proper explanation because each is routinely misunderstood.
The most powerful lever is the effort parameter. Both models support five levels: low, medium, high, xhigh, and max, with high as the default. Crucially, effort does not change the per-token price at all, it changes how many tokens the model spends thinking and acting - Claude effort documentation. Lower effort means fewer thinking tokens, fewer tool calls, and terser output, which directly lowers cost and latency. Higher effort means the opposite. On Opus 5 specifically, the low and medium settings are unusually strong, often matching what older models produced at their highest settings, which makes effort the single best cost dial available. The chart below shows how Opus 5's measured intelligence scales with effort, and the shape of that curve is the whole argument for tuning it.
Read that curve for what it tells you about waste. Opus 5 at medium effort scores 56, which is exactly Opus 4.8's max-effort score, and it climbs only to 61 at max. The jump from medium to high (56 to 59) buys three points, the jump from high to xhigh buys one, and the jump from xhigh to max buys one more, while token spend rises steeply the whole way. In plain terms, most of the intelligence is available at medium and high effort, and the top two settings are where you pay a premium for the last few points. The correct operational habit is to sweep effort on your own evaluation set and pick the lowest level that still passes, rather than defaulting to max because it is the ceiling. One important caveat on Opus 5: changing effort does not reliably shorten the visible response length, only the thinking volume, so if your problem is verbose user-facing output you control it with prompting, not with effort.
The second lever is prompt caching, and its economics are dramatic for the right workload. A cache read costs one-tenth of the standard input price, so any large, stable prefix that gets reused pays for its slightly-more-expensive write almost immediately. The break-even math is simple and worth memorizing: a five-minute cache write costs 1.25x and a read costs 0.1x, so a prefix that is read even once after being written already comes out ahead (1.25 + 0.1 versus 2.0 for two uncached passes), and a one-hour cache at 2x write pays off after two reads. For an agent that resends a 50,000-token system prompt on every turn of a long session, caching is not an optimization, it is the difference between a viable and an unviable deployment. Opus 5's lower 512-token cache minimum widens the set of prefixes that qualify, which is a quiet but real cost win over 4.8.
The third lever is batching, which is the bluntest and most reliable discount available. Any workload that is not latency-sensitive - overnight report generation, bulk classification, large-scale data extraction, evaluation runs - can go through the Batch API for a flat 50% off both input and output tokens, bringing Opus 5 to $2.50 input and $12.50 output. This stacks with the verbosity problem in a useful way: even if Opus 5 spends more tokens than 4.8, running those tokens at half price through a batch can leave you net cheaper than an interactive 4.8 call. The catch is that batching and fast mode are mutually exclusive, and batch results can take up to 24 hours, so it is strictly for work that can wait. Between these three levers, a disciplined team can run Opus 5 at a cost that undercuts a naive Opus 4.8 deployment, which is the counterintuitive but correct conclusion: the newer, more expensive-to-run model can end up cheaper in practice if you actually use the controls.
A concrete scenario makes the levers tangible, using the measured per-task figures as a starting point. Imagine an agent that runs 10,000 tasks per month. Sent entirely to Opus 4.8 at roughly $1.80 per task, the monthly bill is about $18,000. Swap the model string to Opus 5 with no other change and, at roughly $2.03 per task, the bill rises to about $20,300, the $2,300 verbosity tax that the price sheet does not warn you about. Now apply the framework instead of a blind swap: route only the hard 20 percent of tasks (2,000 of them) to Opus 5 and send the routine 80 percent to a cheaper model such as Sonnet 5 at perhaps $0.30 per task. The math becomes 2,000 tasks at $2.03 plus 8,000 tasks at $0.30, or about $4,060 plus $2,400, roughly $6,460 per month, a two-thirds cut against the all-4.8 baseline while the hard tasks now run on a more capable model. Push the asynchronous portion through the Batch API for its 50 percent discount and the number falls further. These figures are illustrative rather than a quote for your workload, but the shape is the lesson: the model choice is a smaller lever than the routing choice, and a team that only argues about "Opus 5 versus 4.8" while sending every task to one model is optimizing the wrong variable.
7. Cost part three: cost in the wild (Claude Code, Max, Bedrock, Vertex)
Raw API pricing is where most cost analysis stops, but very few people consume these models purely through raw metered API calls. Most access Opus 5 through a subscription, a coding tool, or a cloud marketplace, and each of those wrappers changes the effective economics in ways the per-token rate does not show. Understanding the wrappers is the difference between a theoretical cost model and the number that actually lands on your card, so this section walks through the real consumption paths in the order most people encounter them.
The most common path for individuals is a Claude subscription. Opus 5 is the default model on the Max plans and the strongest model available on Pro. Pro costs $20 per month, Max 5x costs $100 per month for roughly five times Pro's usage, and Max 20x costs $200 per month for roughly twenty times - Claude Max plan documentation. For a heavy individual user, a subscription is almost always cheaper than metered API access, because the plans are priced to be a flat all-you-can-reasonably-eat rather than a strict pass-through. The catch on the Max plans is a set of weekly usage caps plus a shared five-hour session limit, so extremely heavy automated use can still hit a ceiling. But for a developer using Opus 5 interactively through Claude Code all day, the subscription is the correct default, and we walk through the full plan math in our Claude Code pricing guide.
For teams building agents, the picture shifts to metered access, and here the deployment surface adds its own costs on top of tokens. If you run Claude through Anthropic's Managed Agents, the session runtime is billed at $0.08 per session-hour on top of standard token rates, accruing only while a session is actively running. Server tools stack too: web search is $10 per 1,000 searches, web fetch is free, and code execution has 1,550 free container-hours per month per organization before $0.05 per hour after. None of these are large individually, but for an agent fleet running thousands of sessions they add up, and they are invisible if you only model token cost. The full architecture of running these agents at scale is covered in our Claude Managed Agents guide.
The cloud marketplaces introduce a pricing wrinkle that catches teams off guard. On Amazon Bedrock and Google Cloud Vertex AI, regional and multi-region endpoints carry a 10% premium over global endpoints for Opus 4.5 and every later model, including both Opus 5 and 4.8, whereas the first-party Claude API is global by default with no such surcharge. There is a further residency multiplier on the first-party API itself: setting inference to US-only applies a 1.1x multiplier across all token categories. And on Claude Platform on AWS and Microsoft Foundry, usage is billed in Claude Consumption Units where 100 units equal one dollar. The point is not that any of these is expensive, it is that the "same $5/$25" you read on the pricing page is the floor, and the real number depends on which door you walk through. A short reference of the paths and their overheads:
- First-party Claude API - the baseline $5/$25, global by default, no surcharge.
- Bedrock or Vertex regional - add roughly 10% for a regional or multi-region endpoint.
- Managed Agents - add $0.08 per session-hour on top of tokens.
- US-only residency - a 1.1x multiplier across every token category.
The practical takeaway is that fast mode is the one wrapper that dramatically changes the token price itself, and it does so upward. Fast mode doubles the rate to $10 input and $50 output in exchange for up to 2.5x higher output throughput, and it is available only on the first-party Claude API for Opus 5 and Opus 4.8, not on the cloud marketplaces. It exists for latency-critical interactive products where the 65-second first-token wait on Opus 5 is a dealbreaker, and it is a poor choice for anything batchable. For most teams the right mental model is a ladder: subscription for individuals, standard metered API for production agents, batch for anything asynchronous, and fast mode reserved for the narrow case where a human is waiting and every second counts.
8. Migrating from 4.8 to Opus 5 without a surprise invoice
Swapping the model string from claude-opus-4-8 to claude-opus-5 is a one-character change that can produce three unpleasant surprises if you do nothing else: a truncated response, a rejected request, and a bill that climbed for no visible reason. All three are avoidable, and all three come directly from the behavioral changes in section two. This section is the practical checklist for making the move cleanly, and it is worth doing deliberately rather than discovering the issues in production.
The first thing to fix is max_tokens. Because Opus 5 now thinks by default, your max output budget must cover thinking plus the answer, not just the answer. A ceiling that was comfortable on 4.8 (where a parameter-free request produced no thinking) can now truncate the response mid-sentence, because the thinking tokens ate the budget before the answer finished. The fix is to raise max_tokens with real headroom, and Anthropic specifically recommends starting at 64,000 when running at xhigh or max effort. If you genuinely need the old non-thinking behavior for a cheap, fast path, you can pass a disabled-thinking configuration, but only at high effort or below, because pairing disabled thinking with xhigh or max returns a 400 error on Opus 5.
The second thing to fix is your prompt scaffolding, because instructions that helped older models now hurt Opus 5. The migration guidance is counterintuitive: delete your verification instructions. Phrases like "double-check your answer" or "include a final verification step," which were sound advice on prior models, now trigger redundant over-verification on Opus 5 because the model already verifies its own work. The same applies to "delegate to subagents" hints you may have added for Opus 4.8, which under-reached for subagents; Opus 5 reaches for them freely, so you likely want a cap rather than encouragement. And because Opus 5 is more verbose and can expand scope, adding an explicit conciseness instruction and a "deliver exactly what was asked" scope-discipline instruction typically reduces both token spend and unwanted extra work.
The behavioral re-tuning is genuinely where the value is, so it is worth being concrete about the highest-leverage prompt changes:
- Remove self-check scaffolding - "double-check" and "verify before responding" now cause over-verification, not accuracy.
- Add a conciseness instruction - Opus 5 writes longer by default, and effort will not shorten visible output; prompting will.
- Add scope discipline - a "do exactly what was asked, flag but do not expand" instruction cuts scope creep to near zero.
- Cap subagent delegation - Opus 5 delegates readily, and each subagent multiplies cost, so bound the count.
After the prompt work, run an effort sweep rather than assuming the old setting transfers. Because Opus 5's low and medium settings are so strong, a route that ran at xhigh on Opus 4.8 for adequate quality may hold quality at medium on Opus 5, which would more than pay back the verbosity tax. The correct migration is not "same effort, new model," it is "re-tune effort on the new model's curve." The diagram below lays out the migration decision as a flow, so the sequence is unambiguous.
Finally, remember the rate-limit reset. Opus 5 draws from a separate bucket than the combined Opus 4.x pool, so shifting production traffic to it neither frees your existing Opus headroom nor inherits it. Capacity planning for Opus 5 is a fresh exercise. If you are moving a high-volume workload, confirm your tier's Opus 5 limits before the cutover rather than discovering them under load. Teams already running on the prior generation will find the full behavioral background in our Opus 4.8 full benchmark and cost guide, which documents the baseline you are migrating from.
9. Where Opus 5 sits in the wider field
Choosing between Opus 5 and Opus 4.8 is the narrow question, but neither model exists in a vacuum, and a serious buyer is really choosing among the whole frontier. As of late July 2026 the field is unusually crowded, and the honest summary is that Opus 5 sits at or very near the top of the coding-and-agentic frontier while trading blows with a small number of genuine peers, rather than standing clearly alone. Getting the current model names right matters here, because this space turns over monthly and a stale comparison is worse than none.
The most direct rival is OpenAI's GPT-5.6, released July 9, 2026, which ships in three variants ranked least to most capable as Luna, Terra, and Sol, with Sol positioned as the flagship coding model - OpenAI. On the Vals.ai SWE-bench Verified leaderboard, GPT-5.6 Sol scores about 96.2%, essentially neck and neck with Opus 5, and on Terminal-Bench v2.1 Sol actually edges Opus 5 at 89.5% to 89.1%. These two models are the coding frontier, and which one leads depends on the specific benchmark and harness. Below them, Google's Gemini 3.1 Pro (released February 19, 2026) offers strong reasoning at a markedly lower token price of roughly $2 input and $12 output, trading some coding capability for cost and Google Cloud scale, a trade we examine in our Gemini 3.1 Pro guide.
The open-weights challengers are the most interesting part of the field for cost-conscious teams, because they attack the price dimension directly. Moonshot's Kimi K3, launched July 16 and open-weighted July 27, 2026, scores about 93.4% on SWE-bench Verified, which puts a self-hostable model within striking distance of the closed frontier. Z.ai's GLM-5.2 (June 13, 2026) is priced at roughly $1.40 input and $4.40 output, a fraction of Opus pricing, and lands respectably on the harder SWE-bench Pro variant, which we cover in our GLM-5.2 benchmarks and cost guide. And DeepSeek's V4-Pro continues the pattern of near-frontier scores at deep-value pricing. For workloads where a model does not need to be the absolute best, these open options can be five to ten times cheaper per output token than Opus 5, which reframes the entire decision.
The comparison table below places the field on a common footing, with the confidence caveat that competitor pricing and per-variant scores are less firmly sourced than the Anthropic figures, and should be verified against each provider's own page before you commit budget:
| Model | Provider | SWE-bench Verified | Output $/1M | Positioning |
|---|---|---|---|---|
| Claude Opus 5 | Anthropic | ~96% | $25 | Frontier agentic, same price as 4.8 |
| GPT-5.6 Sol | OpenAI | 96.2% | Premium | Co-leader on coding |
| Claude Fable 5 | Anthropic | ~95% | $50 | Top-tier, twice Opus price |
| Kimi K3 | Moonshot | 93.4% | Open weights | Best open-weights coder |
| Claude Opus 4.8 | Anthropic | 88.6% | $25 | Prior Opus, cheaper per task |
| Grok 4.5 | xAI | 86.6% | ~$6 | 500K context, cost play |
| Gemini 3.1 Pro | ~80.6% | ~$12 | Reasoning + Google scale |
The most important within-Anthropic comparison is Opus 5 against Claude Fable 5, the company's genuine top-tier model. Fable 5 scores a hair higher on some raw-intelligence measures but costs exactly twice as much at $10 input and $50 output, and Anthropic explicitly positions Opus 5 as delivering close to Fable 5's frontier intelligence at half the price - CNBC. For the vast majority of production workloads, that makes Opus 5, not Fable 5, the correct frontier choice, with Fable 5 reserved for the genuinely hardest long-horizon problems where the last few points of capability justify double the cost. We work through that specific trade-off in our Claude Fable 5 and Mythos 5 benchmarks guide, and the fact that Opus 5 undercuts Fable 5 so aggressively is arguably the most consequential pricing move of the generation.
10. First principles: what cheaper intelligence at the same price changes
It is easy to read a model comparison as a shopping decision and miss the structural shift underneath it. So step back from the benchmark tables and ask the fundamental question: what does it mean that the frontier of machine intelligence moved up meaningfully while the price held exactly flat? The answer is not "a slightly better chatbot." It is that the cost of a unit of competent cognitive work is falling on a curve, and every business built on applying that work is affected by the slope of that curve, not by which model wins this month's leaderboard.
Reason it out from the inputs. What a company actually sells is not software or intelligence, it is outcomes: a filed legal document, a shipped feature, a resolved support ticket, a closed sale, a reconciled ledger. Each outcome is produced by combining some amount of cognitive work with domain knowledge, tools, and accountability. When the cognitive-work input gets cheaper (and Opus 5 holding Fable 5-class capability at half of Fable's price is precisely that input getting cheaper), the economics of producing every one of those outcomes changes. The businesses that win are not the ones that own the cheapest intelligence, they are the ones that combine it most effectively with the other inputs to deliver an outcome a customer will pay for. This is the same structural logic we traced in our analysis of self-improving Founden companies and it is why a benchmark guide is really a strategy document in disguise.
That reframing changes how you should read the "same price, more tokens" cost story from earlier sections. On the surface, Opus 5 costing 13 percent more per task than 4.8 looks like a regression on the cost axis. But the right denominator is not tokens, it is completed outcomes. If Opus 5's extra tokens turn a 60-percent task-completion rate into an 85-percent rate, then the cost per completed outcome fell even though the cost per task rose, because you are no longer paying for the failed attempts and the human remediation they trigger. The verbosity that inflates the token bill is often the same behavior (thinking, self-verification, thoroughness) that raises the completion rate. Judging these models on cost per token rather than cost per outcome is the single most common analytical error in the space, and it consistently points people at the wrong model.
There is a genuine counter-argument worth taking seriously, because pressure-testing your own conclusion is the point of first-principles reasoning. The counter-argument is that for the large majority of real-world tasks, which are routine rather than frontier, the marginal capability of Opus 5 over 4.8 (or over a cheap open-weights model) is simply unused, and paying for thinking tokens on a task a smaller model would have nailed is pure waste. This is correct, and it is why the answer is not "always use the best model." The correct architecture for most real systems is a cascade: route the easy majority of tasks to a cheap model, escalate only the hard minority to Opus 5, and reserve Fable 5 for the rare cases where even Opus 5 struggles. The frontier model is a scalpel, not a hammer, and the teams getting the best economics in 2026 are the ones building the routing layer that decides when to reach for it.
This is the design philosophy behind platforms like O-mega, which lets you build an autonomous company through one conversation and runs the website, app, billing, content, and admin as a coordinated workforce rather than a single model call. The value is not in having access to Opus 5, which anyone can buy, it is in the orchestration layer that decides which model, at which effort level, does which piece of work, so that the frontier model is deployed exactly where its extra cost buys a better outcome and nowhere else. Yuma Heymans (@yumahey), O-mega's founder and previously a co-founder of the AI recruiting platform HeroHunt.ai, has argued that the autonomous-enterprise thesis depends less on which model tops the leaderboard and more on the routing intelligence around it, which is exactly the lesson the Opus 5 versus 4.8 cost data teaches: the model is a commodity input, and the leverage is in how you spend it.
11. The decision framework
All of the analysis reduces to a small number of decisions that you can make quickly once you know your workload. The mistake to avoid is treating "which Claude Opus model" as a single global choice. It is a per-workload choice, and a mature deployment often runs both models (and sometimes a cheaper one below them) for different jobs. The framework below is meant to be applied route by route, not once for your whole company.
Use Opus 5 when the task is genuinely hard: multi-file refactors, long-horizon autonomous agents, complex code review, scientific reasoning, or any job where a failed attempt is expensive to clean up. On these, Opus 5's higher completion rate makes it cheaper per outcome despite the higher per-task token cost, and its lead over 4.8 on Frontier-Bench, ARC-AGI, and the agentic evaluations is exactly the capability the task demands. Run it at high effort by default, sweep down to medium where quality holds, and reserve xhigh and max for the cases where correctness clearly matters more than cost.
Use Opus 4.8 when the task is routine and cost-sensitive, or when latency to first token matters. Its 29-second time-to-first-token beats Opus 5's 65 seconds by a wide margin, it costs less per task because it is less verbose, and it remains eligible for the Priority Tier that Opus 5 does not cover. For high-volume production workloads where both models would succeed and the extra capability is unused, 4.8 is the strictly better economic choice and there is no shame in staying on it. The decision tree below captures the routing logic in one view.
For the large volume of easy work that neither Opus model needs to touch, drop down the ladder entirely. Claude Sonnet 5 at roughly $3 input and $15 output (with introductory pricing lower through August 2026) handles most production tasks at a fraction of Opus cost, and the open-weights options are cheaper still, which is why we recommend a cascade rather than a single model in our practical Sonnet 5 guide. The correct spend curve for a real system is a pyramid: a cheap model handling the base of routine volume, Opus 5 handling the hard minority, and Fable 5 handling the rare frontier cases at the apex.
The last piece of the framework is discipline about measurement. Do not choose based on the launch chart or the leaderboard. Choose based on your own evaluation set, run at several effort levels, measured on cost per completed outcome rather than cost per token. The benchmarks in this guide tell you where to start; your own workload tells you where to land. The teams that win the cost game in 2026 are not the ones with a strong opinion about which model is best, they are the ones with a measurement harness that lets them route each task to the cheapest model that reliably completes it, and re-check that routing every time a new model ships. For computer-use and browser tasks specifically, the routing calculus differs again, which we break down in our computer-use benchmarks guide.
12. Sources and further reading
The primary sources for this guide are Anthropic's official Claude Opus 5 announcement and Opus 4.8 announcement, the Claude pricing documentation, and the Opus 5 system card. Independent evaluation data comes from Artificial Analysis and the ARC Prize verified results. Where third-party benchmark figures appear, they are labeled as such, because in a field this fast the gap between a vendor's chart and an independent leaderboard is exactly where the useful insight lives. For the models that shipped after this piece, apply the same discipline: read the system card, check the independent evaluators, and measure on your own workload before you trust any single number.
If you found this useful, our related deep dives on the Claude Opus 4.7 complete guide, the Claude Mythos preview, and the broader May 2026 model benchmarks and pricing round out the picture of how Anthropic's lineup evolved into the Opus 5 generation.
This guide reflects the AI model landscape as of late July 2026. Claude Opus 5 launched on July 24, 2026, and the competitive field moves monthly. Pricing, benchmark scores, and available models change frequently, so verify current details against the official provider pages before making a purchasing or architecture decision. Benchmark figures attributed to third parties carry more uncertainty than those from Anthropic's own documentation and independent evaluators, and are labeled accordingly throughout.