What OpenAI's and Anthropic's rival cost-per-task claims really mean once an agent pays for every token, retry and failure
On September 29, 2026, OpenAI said GPT-6.1 Sol beats Claude Opus 5.5 on business automation "at roughly a third of the cost." One week earlier, Anthropic said Claude Opus 5.5 matches OpenAI's flagship on Terminal-Bench 4.0 "for about 40% of the cost."
Both statements are accurate, and both were chosen with care. OpenAI's AutomationBench comparison is made at medium reasoning effort, where GPT-6.1 Sol scores 2.2 points higher than Opus 5.5 - OpenAI. The chart data on OpenAI's own page shows that at maximum effort the order flips: Opus 5.5 reaches 42.5% against Sol's 36.1%. Anthropic's cost claims, meanwhile, compare Claude Opus 5.5 with GPT-6 Astra and GPT-5.6 Sol, because GPT-6.1 Sol did not exist when Anthropic published on September 22 - Anthropic. Each lab now picks the benchmark, the effort level and the fallback setting on which its model wins, and prints the ratio.
That leaves buyers with a narrower and harder question. For the agent you actually run, which model finishes the work for less money once you count every cached and uncached token, every retry, every safety fallback, and every failed task that a person has to clean up? A price page cannot answer that. Neither can a single benchmark chart, because the same model at two effort levels behaves like two different products, and the two vendors meter caching, long contexts and refusals in different ways.
This guide answers it with data rather than launch copy. It rebuilds both vendors' cost charts from the numbers embedded in their own announcement pages, sets them against independent runs from Artificial Analysis, Vals AI, Mercor, Surge AI and Zapier, and breaks cost per task into its three real drivers: price per token, tokens per task and attempts per success. It then derives a break-even formula that tells you, workload by workload, when Opus 5.5's extra accuracy pays for its larger bill, and when GPT-6.1 Sol's lower bill is the better deal even after its extra failures.
The timing matters. Opus 5.5 shipped on September 22, Claude Sonnet 5.5 on September 28, and GPT-6.1 Sol at OpenAI's DevDay on September 29, alongside dots, always-on agents that run on GPT-6 Astra with "their own cloud computer" - OpenAI. Meta's Muse runs on a dedicated "Muse Secure VM" with a separate Sentinel agent that must approve anything Muse sends to the internet - Meta, and Microsoft moved its new Autopilot teammate onto usage-based billing - Microsoft. Every one of these products pays for long, tool-heavy trajectories, which is why cost per task, not price per token, became the battleground this month.
Contents
- What each vendor claimed, and what the footnotes say
- The price sheets: every line item differs by exactly 2x
- Cost per task from first principles: price, tokens, attempts
- The independent evidence, benchmark by benchmark
- Cost per solved task and the break-even price of a failure
- The hidden multipliers: caching, long context, fallbacks and speed
- Safety, hallucination and liability are costs too
- Verdicts by workload
- Subscriptions, managed runtimes and where the bill lands
- How to run your own cost-per-task bake-off
- The bottom line
The ten configurations at a glance
Comparing "GPT-6.1 Sol" with "Claude Opus 5.5" as if each were a single product hides the most important decision you will make. Both models expose five reasoning effort levels (low, medium, high, xhigh and max), and Artificial Analysis now benchmarks every one of them separately. The spread inside each model is enormous: Opus 5.5 costs $0.55 per index task at low effort and $5.98 at max, an 11x range, while GPT-6.1 Sol runs from $0.13 to $0.72 - Artificial Analysis. The useful comparison is therefore between configurations, not brand names.
The table below scores all ten configurations on four criteria that decide whether an agent configuration is worth running, using the same evidence for every row. Capability and cost come from the Artificial Analysis Intelligence Index v4.3.2, which now includes Terminal-Bench 4.0, GDPval-AA and AutomationBench-AA among its ten evaluations. Time per task comes from the same source. Operational risk draws on each vendor's system card and on third-party safety results, and it is scored per model because effort level barely changes it. Each cell shows the score and the data behind it, so you can disagree with a weight without reconstructing the evidence.
| # | Configuration | What It Does | Capability (40%) | Cost per task (30%) | Time per task (10%) | Operational risk (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6.1 Sol (high) | Best balance of capability and cost for most agent steps | 6.0 - AA index 50, Terminal-Bench 4.0 51.5% | 7.9 - $0.32 per AA task | 6.3 - 3.8 min per task | 6.5 - injection data is vendor-run only (98.98% robust); 54% AA hallucination rate | 6.7 |
| 2 | GPT-6.1 Sol (medium) | OpenAI's default setting, cheapest capable worker | 5.0 - AA index 48, Terminal-Bench 4.0 48.0% | 8.9 - $0.21 per AA task | 7.6 - 2.6 min per task | 6.5 - same model-level evidence as row 1 | 6.7 |
| 3 | GPT-6.1 Sol (xhigh) | Sol's sweet spot for long agent loops | 6.5 - AA index 51, AutomationBench-AA 66.6% | 7.4 - $0.39 per AA task | 5.0 - 5.5 min per task | 6.5 - same model-level evidence as row 1 | 6.6 |
| 4 | Claude Opus 5.5 (high) | Best Opus value: beats Sol max on capability | 8.0 - AA index 54, Terminal-Bench 4.0 56.6% | 3.8 - $1.82 per AA task | 5.4 - 5.0 min per task | 7.5 - 1.0% Gray Swan injection success at 15 tries; fallbacks on 3.99% of Vals tasks | 6.4 |
| 5 | Claude Opus 5.5 (xhigh) | Near-max quality for long-horizon work | 9.0 - AA index 56, GDPval-AA 1820 Elo | 2.3 - $3.46 per AA task | 3.6 - 8.5 min per task | 7.5 - same model-level evidence as row 4 | 6.2 |
| 6 | GPT-6.1 Sol (max) | Sol's ceiling; often not worth the extra time | 7.0 - AA index 52, Terminal-Bench 4.0 56.1% | 6.0 - $0.72 per AA task | 3.1 - 9.9 min per task | 6.5 - same model-level evidence as row 1 | 6.2 |
| 7 | Claude Opus 5.5 (medium) | Anthropic's default; matches Sol max on GDPval-AA | 6.5 - AA index 51, GDPval-AA 1576 Elo | 4.5 - $1.34 per AA task | 6.4 - 3.7 min per task | 7.5 - same model-level evidence as row 4 | 6.1 |
| 8 | GPT-6.1 Sol (low) | Fast triage and extraction, weak on long tasks | 2.0 - AA index 42, Terminal-Bench 4.0 30.8% | 10 - $0.13 per AA task | 10 - 1.2 min per task | 6.5 - same model-level evidence as row 1 | 6.1 |
| 9 | Claude Opus 5.5 (max) | Highest measured intelligence of any public model | 10 - AA index 58 (#1), GDPval-AA 1846 Elo | 1.0 - $5.98 per AA task | 2.0 - 13.7 min per task | 7.5 - same model-level evidence as row 4 | 6.0 |
| 10 | Claude Opus 5.5 (low) | Cheapest Opus, but no better than Sol low | 2.0 - AA index 42, Terminal-Bench 4.0 31.3% | 6.6 - $0.55 per AA task | 9.5 - 1.4 min per task | 7.5 - same model-level evidence as row 4 | 5.2 |
How the criteria work. Capability (40%) maps the Artificial Analysis index linearly, with 58 scoring 10 and each point below it costing half a point, because errors compound across agent steps and a few index points translate into large differences in finished tasks. Cost per task (30%) uses Artificial Analysis's weighted average cost per index task on a log scale from $0.13 (10) to $5.98 (1), because buyers think in multiples, not dollar differences. Time per task (10%) uses the same source's measured minutes per task on a log scale, since latency matters for interactive agents but not for overnight ones. Operational risk (20%) rewards third-party adversarial evidence and penalizes hallucination rates and unpredictable detours; it is explained in section 7. Final scores are weighted averages rounded half up, and ties are ordered alphabetically.
Read the table as a default for a generic agent workload, not as a verdict for yours. The ranking says that GPT-6.1 Sol at high or medium effort gives the most capability per dollar for typical steps, and that Opus 5.5 at high effort is the best way to buy Anthropic's extra intelligence. It also says the two maximum-effort settings are the worst value on this weighting. Section 5 shows the important exception: when a failed task is expensive enough, the ordering inverts, and Opus 5.5 at max becomes the cheapest way to get work done.
1. What each vendor claimed, and what the footnotes say
Launch posts are marketing documents, but this month they are unusually precise marketing documents. Both OpenAI and Anthropic now publish accuracy-versus-cost charts for every effort level, and both embed the underlying numbers in their pages. That makes it possible to reconstruct exactly which point each vendor chose to quote, and which points it chose not to. The pattern that emerges is not dishonesty. It is selection: each claim is true at the configuration where it was measured, and silent about the others.
The selection matters because readers tend to carry a headline ratio into a budget. If you take "a third of the cost" from OpenAI's post and "40% of the cost" from Anthropic's, you can construct almost any conclusion you like about which model is cheaper. The way through is to look at the whole curve for each benchmark, note which effort level and which safety setting each quoted point used, and then compare like with like.
OpenAI's claims for GPT-6.1 Sol
OpenAI's post names Opus 5.5 three times, and each comparison is framed in cost per task. On AutomationBench, Zapier's benchmark of multi-step business workflows across 47 simulated apps, "GPT-6.1 Sol scores 2.2 percentage points above Opus 5.5 at medium reasoning effort, at roughly a third of the cost" - OpenAI. On GDP.pdf, Surge AI's test of answering professional questions from complex PDFs, it "scores higher than Opus 5.5 with fallbacks at less than half the cost per task across the tested reasoning settings." On Terminal-Bench Science 0.1, it costs "$5.47 per task on average, compared with $23.21 for Opus 5.5 and $23.80 for Astra."
The chart data behind those sentences is more revealing than the sentences. On AutomationBench, Sol at medium scores 31.7% for $0.19 per task, and Opus 5.5 at medium scores 29.5% for $0.65, which is where the 2.2 points and the one-third ratio come from. But the same chart shows Opus 5.5 at max effort at 42.5% for $1.44, against Sol's best of 36.1% at max for $0.30. Zapier's public leaderboard confirms the Opus figure and lists it fourth overall, behind two Gemini 4 Argon settings and Claude Sonnet 5.5 - Zapier. OpenAI quoted the effort level where its model leads; at the top of the curve, Opus leads by 6.4 points.
The two lines are almost indistinguishable from low to xhigh effort, which is the honest summary of this benchmark: on structured business workflows, the models are peers until Opus is allowed its maximum budget, at which point it pulls away. The cost lines, which the chart does not show, never converge: Sol costs between $0.16 and $0.30 per task across its five settings, and Opus between $0.51 and $1.44. Whether the extra 6.4 points at max are worth an extra $1.14 per task is exactly the question section 5 answers.
The GDP.pdf numbers come with a different footnote. OpenAI compares Sol with "Opus 5.5 with fallbacks", meaning runs where Anthropic's safety classifiers sent some steps to an older Claude model. The scores in OpenAI's chart (Sol 32.0% at high effort against Opus 28.8% at high) match Artificial Analysis's own GDP.pdf results exactly, so they are independently measured rather than OpenAI's private runs - Artificial Analysis. But Surge AI's own GDP.pdf leaderboard, which uses its own scoring, puts Opus 5.5 at 30.6% and does not yet list GPT-6.1 Sol at all - Surge AI. The direction of OpenAI's claim holds on Artificial Analysis's runs; its size depends on whose harness you trust.
OpenAI presented GPT-6.1 Sol on stage at DevDay, where it was one of more than 20 announcements. The keynote is the primary source for how OpenAI positions Sol against Astra, and the official recap describes Sol as delivering "near-Astra intelligence to everyone at a fifth of its standard input and output token prices" - OpenAI DevDay recap. In other words, OpenAI now sells Sol as the default worker for agents rather than as a budget tier.
What the keynote framing leaves out is the long-context surcharge, covered in section 2, and the fact that OpenAI's computer-use comparison (OSWorld 2.0, where Sol lands within 2.1 points of Astra at about one-seventh of the cost) contains no Anthropic model at all. Anthropic reports its computer-use result on OSWorld 2.1, a different version, so there is no like-for-like computer-use comparison between the two models in either vendor's material.
Anthropic's claims for Claude Opus 5.5
Anthropic's post makes a broader efficiency argument. It says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5," and that "at its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost" - Anthropic. On GDPval-AA, Artificial Analysis's test of real work across 44 occupations, "at default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task."
The chart labels on Anthropic's page carry every data point, and they reproduce the claims precisely. On Terminal-Bench 4.0, Opus 5.5 at medium scores 57.6% for $2.94 per task, while Astra at high scores 57.9% for $7.21, a ratio of 41%. On GDPval-AA, Opus at medium reaches 1,576 Elo for $0.86, against Astra at max at 1,542 Elo for $4.53, a ratio of 19%. On FrontierCode, Opus at medium scores 54.6% for $0.80 against Astra's best of 53.3% for $4.36. None of these comparisons involve GPT-6.1 Sol, which shipped a week later at one-fifth of Astra's token price.
Anthropic's own footnotes deserve equal attention. Its AutomationBench result of 40.0% came from runs "performed without fallback models, so safeguard interventions were considered failures," while Zapier's public run with default fallbacks scored 42.5%. And on Terminal-Bench 4.0, Opus is reported at xhigh effort (66.4%) against Astra at high effort "as reported by OpenAI," with the note that "Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8." Safety fallbacks, in other words, are part of what you buy with Opus 5.5, and they move both scores and bills. Section 6 prices them.
Why this matters. Neither vendor compared the two models at matched effort levels across a common benchmark, and neither could, because the benchmarks each one quotes favor its own design. OpenAI's numbers are strongest on structured tasks at moderate effort; Anthropic's are strongest on knowledge work and on long agentic runs at high effort. Both sets of numbers survive independent checking, which is the useful finding: the disagreement is about where you sit on the curve, not about whose data is wrong.
How to apply this. When you read a cost-per-task claim, write down four things before you believe the ratio: the benchmark, the effort level of each model, whether safety fallbacks were counted, and whose harness ran the test. Then look up the same benchmark on an independent board at a matched effort. The next two sections give you the price sheets and the decomposition you need to do that yourself.
2. The price sheets: every line item differs by exactly 2x
The striking fact about the two price sheets is how regular they are. On standard API pricing, every token category for Opus 5.5 costs exactly twice the GPT-6.1 Sol rate: input, cached input, cache writes and output alike. That regularity is useful because it isolates the price effect. If Opus ever costs more than twice as much as Sol for the same task, the extra is coming from tokens per task or attempts per success, not from the rate card.
The regularity breaks in four places, and each one can move a real bill. Sol charges a long-context premium above 272K input tokens that Claude does not. Claude offers a one-hour cache at a premium, while OpenAI's cache now lives at least 30 minutes by default. The two vendors' fast modes sit at different multiples. And Anthropic's safety fallbacks can route a request to an older, more expensive model. The table below lists the standard rates first, then the modifiers.
| Line item (per 1M tokens) | GPT-6.1 Sol | Claude Opus 5.5 | Ratio |
|---|---|---|---|
| Input (uncached) | $2.00 | $4.00 | 2.0x |
| Cached input (cache read) | $0.10 | $0.20 | 2.0x |
| Cache write | $2.50 (1.25x input) | $5.00 for 5 min, $8.00 for 1 hour | 2.0x (5 min) |
| Output, including reasoning | $10.00 | $20.00 | 2.0x |
| Batch | 50% off ($1/$5) | 50% off ($2/$10) | 2.0x |
| Fast mode | 2x standard ($4/$20) | $8/$40 | 2.0x |
| Prompts over 272K input tokens | 2x input and cache, 1.5x output, whole request | No surcharge up to 1M | 1.0x to 1.33x |
| Data residency | +10% | +10% (US-only inference) | 2.0x |
Sol's rates, cache rules and long-context premium are listed on its model page - OpenAI, and Opus 5.5's on Anthropic's pricing page, which states that Claude 4.6 and later models "include the full 1M token context window at standard pricing" - Claude Docs. OpenAI's flagship pricing table shows the same structure for all three GPT-6 generation tiers, with separate short-context and long-context columns - OpenAI pricing.
OpenAI's own model card graphic shows how it wants buyers to read the lineup. Astra sits at $10 and $50, Sol at $2 and $10, and the budget tier Luna at $0.10 and $0.50, with cached input at $1.00, $0.10 and $0.01 respectively.
The detail in that graphic that matters most for agents is the cached input column. Sol's cached rate of $0.10 is "95% less than standard input pricing and 50% less than GPT-6 Sol's cached input pricing," as OpenAI's post puts it. Anthropic made the same move with Opus 5.5, cutting cache reads to $0.20, 60% below Opus 5, and noting that cache reads "make up the majority of agentic and coding work costs." Both labs priced their new models for long loops that re-read the same context, and both now discount cached tokens by 95% against uncached input. CNBC's launch segment covers the announcement for a general audience, including Anthropic's claim that the new model is 40% cheaper to run than Opus 5.
Two smaller differences in the fine print shape cache economics. Opus 5.5 can cache a prefix as short as 512 tokens, while OpenAI's minimum is 1,024 tokens for GPT-5.6 and later models - OpenAI prompt caching. And the vendors count tokens differently: Anthropic notes that its tokenizer for Claude 4.7 and later "produces approximately 30% more tokens for the same text" than its previous one, which is a reminder that a million Claude tokens and a million OpenAI tokens are not the same amount of text. Dollar costs per task, which the rest of this guide uses, already absorb that difference; token counts do not.
Tool fees are close to identical. Both vendors charge $10 per 1,000 web searches plus the tokens the results add. OpenAI bills hosted containers for its shell and code interpreter at $0.03 per 20-minute session for 1 GB, while Anthropic makes code execution free when web search or web fetch is in the request and otherwise gives each organization 1,550 free container hours per month before charging $0.05 an hour. For most agents, these fees are a rounding error next to the token bill.
Why this matters. A 2x rate difference is a floor, not an estimate. If your measured cost ratio is close to 2x, the two models are using similar numbers of tokens; if it is 8x, Opus is spending four times as many tokens per task. That single diagnostic tells you where to look for savings: in effort settings and trajectory length, not in negotiating rates. Our cheapest LLM APIs price table lists the rest of the market on the same basis.
How to apply this. Budget Sol and Opus with the same spreadsheet, using the 2x factor for every rate, then add three adjustments: Sol's 272K premium for any request that can grow past it, Claude's 1-hour cache premium for loops that pause longer than five minutes, and a fallback allowance for Opus on security-adjacent or biology-adjacent work. Section 6 puts numbers on all three.
3. Cost per task from first principles: price, tokens, attempts
Strip an agent down to its mechanics and its bill has only three variables. The price per token, which the previous section showed is exactly 2x apart. The tokens per task, which depend on how many turns the agent takes, how much it thinks per turn, how much context it carries, and how much of that context is served from cache. And the attempts per success, which depend on how often the model fails and has to be rerun or rescued. Cost per finished task is the product of the three.
The first multiplier is fixed and the third is mostly fixed by the model, which makes the middle one the lever. Artificial Analysis's data shows how much it varies. At every effort level, Opus 5.5 costs between 4.2x and 8.9x as much as Sol per index task - Artificial Analysis. Divide out the 2x price difference and Opus is spending roughly 2.1x as many tokens at low effort, 3.2x at medium, and 4.2x at max. Artificial Analysis measured Opus 5.5 at max at about 119,000 output tokens per index task, against about 27,000 for GPT-6 Astra - Artificial Analysis.
What do those tokens buy? The answer is on the capability side of the same dataset. GPT-6.1 Sol's index score is remarkably flat: 48 at medium, 50 at high, 52 at max, so its default setting already delivers most of its intelligence. Opus 5.5 keeps climbing: 51 at medium, 54 at high, 58 at max, and it is the strongest model on the index at max - Artificial Analysis. At equal index scores, Sol is consistently cheaper: both reach 51 (Sol at xhigh, Opus at medium), for $0.39 versus $1.34. Above 52, there is no Sol configuration at all, and the only way up is to pay Opus prices.
Artificial Analysis's own scatter of intelligence against cost per task, published with its Gemini 4 Argon analysis on September 30, shows where both models sit in the wider market. GPT-6.1 Sol at max sits on the Pareto line, the frontier of best capability for the money, at about $0.72. Opus 5.5 at max sits at the top of the chart near $6, with the highest score of any model. Everything between those two points on the cost axis is a trade of dollars for index points.
The independent agent benchmarks confirm the token story at the level of whole benchmark runs. On the Vals Index, which averages finance, coding, legal and tax agent tasks, Opus 5.5 at max processed 28.76 billion input tokens and 491.89 million output tokens across the suite, against 4.50 billion and 104.77 million for GPT-6.1 Sol at max - Vals AI. That is 6.4x the input and 4.7x the output, which, multiplied by the 2x price, explains almost all of the 9.9x cost gap ($32.14 against $3.24 per test). Opus reads far more context per task, mostly because it takes longer trajectories: an average of 1 hour 19 minutes per test against Sol's 43 minutes.
Why this matters. The vendors' rate cards suggest a 2x difference; the measured difference at the settings people actually run is 4x to 10x. Almost all of the gap is behavioral. Opus 5.5 explores more, reads more and thinks more, which is also why it scores higher. You cannot negotiate that away, but you can control it with effort levels, task budgets and context discipline.
How to apply this. Treat effort as a price list. For each candidate model, find the lowest effort that clears your quality bar on your own tasks, then compare cost at those settings. For most teams, that means comparing Sol at medium or high with Opus at medium or high, not the maxima that vendors quote. The next section checks which of those settings actually wins on independent agent benchmarks.
4. The independent evidence, benchmark by benchmark
Vendor charts tell you where each lab thinks its model is strongest. Independent evaluators tell you what happens when someone else runs both models through the same harness. Four independent evaluators have published results that include both GPT-6.1 Sol and Claude Opus 5.5 at their maximum settings within days of Sol's launch, and a fifth, Zapier, publishes the Opus side of a benchmark for which OpenAI's chart supplies Sol's numbers. They do not all agree. That disagreement is informative, because each benchmark measures a different kind of agent work, and the pattern of wins tracks the kind of work.
Before reading the numbers, it is worth knowing what each evaluator holds constant. Vals AI runs both models in the same minimal harness and publishes cost and time per task. Artificial Analysis runs ten evaluations at every effort level with its own reference agent. Mercor publishes long-horizon software and professional-work boards. Surge AI builds document-reasoning benchmarks from real professional files. Zapier grades business workflows by the final state of a simulated company. A model that wins on all five is better; a model that wins on some is better at something specific, and that is the more common and more useful result.
Agentic coding and terminal work
On Terminal-Bench 4.0, the most widely quoted long-horizon agent benchmark, the independent runs favor Opus 5.5 but by less than Anthropic's chart suggests. Vals AI, using the minimal mini-swe-agent harness with a single bash tool and no step or cost limit, measured Opus 5.5 at max at 65.15% for $13.20 per task, first of 16 models, and GPT-6.1 Sol at max at 55.05% for $1.72, sixth - Vals AI. Artificial Analysis's run puts the gap smaller, at 59.6% against 56.1% at max. The median Terminal-Bench 4.0 task is estimated at about four hours of expert work, and every task allows the agent eight hours, so these are long, failure-prone trajectories.
On longer software-engineering work the gap closes completely. Mercor's DeepSWE v1.1 board, which uses 113 tasks across 91 repositories that each require "large, novel fixes not sourced from existing public commits," shows Opus 5.5 and GPT-6.1 Sol tied at 72.3% at max effort - Mercor. Vals' Code Migration benchmark, which ports projects between languages including COBOL modernization, has Opus at 66.65% and Sol at 65.12%, inside each other's error bars, but at $112.97 against $6.51 per test and 3 hours 19 minutes against 1 hour 36 minutes - Vals AI. On Vibe Code Bench, which asks agents to build working applications end to end, Opus scores 90.29% for $57.92 and Sol 88.93% for $6.23 - Vals AI.
The one coding-adjacent board where Sol clearly wins is SRE Bench, and the reason is instructive. Sol scores 50.76% for $2.69, while Opus 5.5 scores 33.59% for $34.65 - Vals AI. Vals explains why in its Opus 5.5 write-up: on SRE Bench, "217 of 262 tasks (82.82%) were fallback-assisted," meaning Anthropic's cybersecurity classifiers declined the work and an older model finished it - Vals AI. Infrastructure and security operations sit close enough to cybersecurity that Opus 5.5's safeguards treat much of it as out of bounds. Vals ran Opus 5.5 with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks, so if your agent does that kind of work, the model you are actually paying for is often an older Claude model, not Opus 5.5.
Knowledge work, documents and spreadsheets
The clearest Opus advantage is in open-ended professional work. On GDPval-AA v2.1, Artificial Analysis's test of real deliverables across 44 occupations graded by pairwise comparison, Opus 5.5 at max reaches 1,846 Elo against 1,575 for Sol at max - Artificial Analysis. A 271-point Elo gap implies that a grader would prefer the Opus deliverable about 83% of the time in a head-to-head comparison (by the standard Elo expectation formula). The same pattern holds on Artificial Analysis's private AA-Briefcase knowledge-work evaluation: 1,822 against 1,564. Notably, Opus at its default medium effort (1,576) already matches Sol at max on GDPval-AA.
Documents split the verdict by format. On GDP.pdf, answering professional questions from complex PDFs, Sol leads at every effort level in Artificial Analysis's runs, peaking at 32.0% at high against Opus's 28.8%. On GDP.xlsx, Surge AI's benchmark of working inside real professional spreadsheets, Opus 5.5 scores 30.3% and GPT-6.1 Sol 22.0%, an 8-point Opus lead - Surge AI. If your agent reads contracts and reports, Sol is at least as good; if it builds and edits financial models, Opus is clearly better. Vals' Excel Modeling Benchmark points the same way, with Opus at 75.9% against Sol at 70.77% - Vals AI.
On the composite Vals Index, which weights finance, coding, legal and tax agent tasks by their share of US GDP, Opus 5.5 scores 66.97% for $32.14 per test and Sol 61.15% for $3.24 - Vals AI. That is a 5.8-point gap at one-tenth of the cost, which is close to the overall shape of the comparison: Opus is better on most professional work by a margin you can measure, and Sol is an order of magnitude cheaper.
Business workflows and long professional tasks
Business automation is closest to a draw. On Zapier's strict AutomationBench, the two models track each other within about two points from low to xhigh effort, and Opus pulls ahead only at max (42.5% against 36.1%), as section 1 showed. Artificial Analysis's version, AutomationBench-AA, which reports a more forgiving headline score, has them within five points at max (69.5% against 64.9%) and Sol slightly ahead at xhigh (66.6% against 65.0%). Zapier's notes on failure modes matter more than the scores: "models declare success while actually failing" in a large share of failed runs, which means many AutomationBench failures are the expensive, undetected kind - Zapier.
On Mercor's APEX-Agents board, which tests long cross-application tasks written by investment bankers, consultants and corporate lawyers, Opus 5.5 at max scores 73.5%, but at medium effort it falls to 52.5% - Mercor. That 21-point drop is the largest effort sensitivity in this comparison, and it is a warning against assuming that Opus's default setting carries its headline capability into long professional tasks. GPT-6.1 Sol does not yet appear among the board's top results.
The table below collects the independent results that include both models, so you can see the pattern at a glance.
| Benchmark (evaluator) | What it tests | Claude Opus 5.5 | GPT-6.1 Sol | Cost per attempt | Better per dollar |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (Vals) | Long terminal tasks | 65.15% | 55.05% | $13.20 vs $1.72 | Sol, unless failures are costly |
| DeepSWE v1.1 (Mercor) | Novel fixes in real repos | 72.3% | 72.3% | not published | Sol (tie at lower price) |
| Code Migration (Vals) | Porting codebases | 66.65% | 65.12% | $112.97 vs $6.51 | Sol |
| GDPval-AA v2.1 (Artificial Analysis) | Real deliverables, 44 occupations | 1,846 Elo | 1,575 Elo | index $5.98 vs $0.72 | Opus when quality is judged |
| GDP.xlsx (Surge AI) | Spreadsheet work | 30.3% | 22.0% | not published | Opus |
| GDP.pdf (Artificial Analysis) | PDF question answering | 28.8% (high) | 32.0% (high) | $0.83 vs $0.35 (OpenAI chart) | Sol |
| SRE Bench (Vals) | Infrastructure operations | 33.59% | 50.76% | $34.65 vs $2.69 | Sol |
Why this matters. The independent evidence supports neither vendor's sweeping version. Opus 5.5 is measurably better on judged knowledge work, spreadsheets and long terminal tasks; GPT-6.1 Sol is equal on software engineering, better on PDFs and on security-adjacent operations, and far cheaper everywhere. The right model depends on which of these your agent spends its time doing. Our broader October model ranking places both models in the context of the other fourteen agent-relevant releases, and our explainer on why coding benchmarks mislead covers why harness choices move these scores.
How to apply this. Classify your agent's work into the buckets above before you choose. If more than half of its steps are judged deliverables or spreadsheet edits, start your evaluation with Opus. If they are code changes, document lookups or operational runbooks, start with Sol. Either way, the decision is not finished until you price the failures, which is the subject of the next section.
5. Cost per solved task and the break-even price of a failure
A cost per attempt is only half a number. If one model costs $1.72 per Terminal-Bench task and solves 55% of them, the honest cost of a solved task is $1.72 divided by 0.55, or $3.12. Opus 5.5, at $13.20 per attempt and 65% success, costs $20.26 per solved task. The ratio narrows slightly, from 7.7x to 6.5x, because the more capable model wastes fewer attempts. But this calculation still treats a failed task as free, and in a real business it never is.
Every failure has a price outside the API bill. Someone has to notice it, diagnose it and finish the work, or a customer has to live with the consequence. That cleanup cost is the variable that decides which model is cheaper in practice, and it is almost never on a vendor chart. So rather than guess it, it is more useful to ask the reverse question: how expensive does a failure have to be before the more accurate model becomes the cheaper one?
The break-even formula
The expected total cost of handing one task to a model is its cost per attempt plus the probability of failure times the cost of handling a failure. Call the cost per attempt c, the success rate p, and the cost of a failure H. Model A (Opus) is cheaper than model B (Sol) whenever cA + (1 - pA) H is smaller than cB + (1 - pB) H. Rearranged, Opus wins when:
H > (cost_opus - cost_sol) / (success_opus - success_sol)
The right-hand side is the break-even failure cost: the extra dollars Opus charges per attempt, divided by the extra fraction of tasks it completes. If a failed task costs your team more than that to fix, Opus 5.5 is the cheaper model overall, despite its larger bill. If it costs less, GPT-6.1 Sol is. The formula assumes failures are detected and fixed at a flat cost; undetected failures, discussed below, push the threshold down.
Applying it to the independent data produces thresholds that vary by two orders of magnitude, which is the single most useful finding in this guide.
| Workload (source) | Opus 5.5 | GPT-6.1 Sol | Cost per solved task | Break-even failure cost |
|---|---|---|---|---|
| Business workflows, max effort (Zapier via OpenAI) | 42.5% at $1.44 | 36.1% at $0.30 | $3.39 vs $0.83 | $18 |
| Terminal-Bench 4.0 (Vals) | 65.15% at $13.20 | 55.05% at $1.72 | $20.26 vs $3.12 | $114 |
| Terminal-Bench Science, max (OpenAI chart) | 63.3% at $23.21 | 57.0% at $5.47 | $36.65 vs $9.58 | $281 |
| Vals Index (Vals) | 66.97% at $32.14 | 61.15% at $3.24 | $47.99 vs $5.30 | $497 |
| Vibe Code Bench (Vals) | 90.29% at $57.92 | 88.93% at $6.23 | $64.15 vs $7.01 | $3,801 |
| Code Migration (Vals) | 66.65% at $112.97 | 65.12% at $6.51 | $169.50 vs $10.00 | $6,958 (scores within noise) |
Read the thresholds against what a failure costs you. On business workflow automation, the tasks are cheap to attempt, so even Opus at max costs only $1.44 per run; Opus wins as soon as a failed workflow costs more than $18 to catch and repair, which for anything touching customers, money or a CRM is almost certainly true. On long terminal tasks, the threshold is $114, roughly one to two hours of an engineer's time, and a failed four-hour task usually costs more than that. On app building and code migration, where Sol nearly matches Opus's success rate at a tenth of the price, a failure would have to cost several thousand dollars before Opus pays for itself.
The counterintuitive result is the business-workflow row. OpenAI's own marketing chose AutomationBench to argue that Sol is cheaper, and at medium effort it is. But the full curve from the same chart shows that for any workflow where a mistake costs more than about $18, Opus 5.5 at max is the cheaper way to get correct outcomes. Low-cost tasks with expensive failures are exactly where the more accurate model wins, because the attempt is cheap and the error is not.
Undetected failures change the math
The formula assumes you know when a task failed. Zapier's AutomationBench analysis found that in an earlier study "72% of Opus's failures, 91% of Gemini's, and 84% of GPT-5.4's involved this false confidence," where the agent reports success but the business state is wrong - Zapier. Those numbers predate both models in this guide, but the failure mode is structural, not model-specific. An undetected failure costs more than a detected one, because it surfaces later, often in front of a customer, and it erodes trust in every other result the agent produced.
When you cannot verify outcomes cheaply, the effective failure cost H rises, and the break-even point moves toward the more accurate model. When you can verify outcomes cheaply, with tests, schema checks, state assertions or a second model acting as reviewer, the effective failure cost falls toward the cost of a retry, and the cheaper model gains ground. Verification is therefore a cost lever, not just a quality practice, and it is often worth more than the model choice.
Retries: when two cheap attempts beat one expensive one
If a verifier can tell you that an attempt failed, you can retry. On Terminal-Bench 4.0, two Sol attempts cost $3.44, about a quarter of one Opus attempt at $13.20. If the attempts were fully independent, two Sol tries would succeed 80% of the time (one minus 0.45 squared), well above Opus's single-attempt 65%. In practice they are not independent: a task that defeats a model once tends to defeat it again, so the real figure sits somewhere between 55% and 80%. Even at the low end, the pattern of "cheap model, verify, retry once, then escalate" usually beats "expensive model once" on cost per solved task.
The escalation step matters as much as the retry. Sending only the tasks that Sol failed twice to Opus means Opus sees the hardest few percent of your workload, which is where its extra capability is concentrated. We covered the full routing architecture, and the savings teams report from it, in AI model routing in 2026.
Why this matters. The model that is cheaper per token is not necessarily cheaper per outcome, and the threshold where that flips is specific to the workload. A team that runs cheap business automations with expensive errors should probably be paying for Opus at max, while a team that builds apps behind a solid test suite should probably be running Sol and pocketing the difference. Most teams have both kinds of work, which is why a single model choice is usually wrong.
How to apply this. Estimate H for each of your agent's task types: the minutes a person needs to detect, diagnose and finish a failed task, times their loaded hourly cost, plus any customer or compliance cost. Measure c and p for both models on 50 to 100 of your own tasks. Compute the break-even for each task type, and route accordingly. Our guide to the true cost of LLM inference covers how to estimate the human side of the equation.
6. The hidden multipliers: caching, long context, fallbacks and speed
The independent benchmarks above measure models under clean conditions: fresh sandboxes, steady tool latency, contexts that fit comfortably in a window. Production agents are messier. They pause for slow builds and human approvals, they accumulate history until the context is enormous, they touch topics that trip safety classifiers, and they run against rate limits. Each of these interacts with the two vendors' pricing rules differently, and together they can move the cost ratio from 2x to anywhere between roughly 1.4x and 6x on an identical trajectory.
The cleanest way to see these effects is to hold the agent's behavior constant and change only the conditions. The scenarios below use an illustrative trajectory: a 25,000-token stable prefix (system prompt, tool definitions, instructions), 4,000 tokens of tool output per turn, and 1,500 output tokens per turn including reasoning. Each turn re-reads the previous context from cache and writes the new material. The prices are each vendor's published rates; the trajectory is our model, not a measurement, and it deliberately ignores the fact, shown in section 3, that Opus tends to take more tokens for the same work.
| Scenario | GPT-6.1 Sol | Claude Opus 5.5 | Ratio | What drives it |
|---|---|---|---|---|
| 40 turns, context ends at 243.5K, all cache hits | $1.73 | $3.46 | 2.0x | Pure rate card |
| 80 turns, context ends at 463.5K, all cache hits | $6.30 | $8.56 | 1.36x | Sol's 272K long-context premium |
| 40 turns, every fourth turn misses the cache | $1.73 | $10.13 | 5.9x | Claude's 5-minute default TTL lapses |
| 40 turns, Opus writes everything with 1-hour TTL | $1.73 | $4.19 | 2.4x | Insurance against TTL lapses |
The first row is the baseline: on an identical trajectory with perfect caching, Opus costs exactly twice as much, split roughly evenly between cache reads ($1.04), cache writes ($1.22) and output ($1.20). That even split is itself worth noticing. In a long agent loop, cache writes cost as much as output, because every turn writes its new tool results and the previous reply at 1.25x the input rate.
Long context erodes Sol's price advantage
GPT-6.1 Sol bills any request with more than 272K input tokens at "2x input and cache rates and 1.5x output for the full request" - OpenAI. Above that line, Sol's input and cache rates equal Opus's standard rates, and its output costs three-quarters of Opus's. Claude bills the full 1M window at standard prices. In the 80-turn scenario, 35 of the 80 turns cross the line, and Sol's advantage falls from 2.0x to 1.36x. An agent that lets its context drift past 272K and stay there pays nearly Opus prices for Sol intelligence.
The fix is architectural, not contractual. Compaction, summarizing older history into a shorter block, keeps requests under the line; both vendors now offer it natively, OpenAI through the Agents API and Responses compaction and Anthropic through a compact-on-demand beta that returns a signed summary block - Claude Docs. Our guide to context engineering for agents covers the techniques. For Sol, a compaction policy that triggers around 250K input tokens is one of the highest-return optimizations available.
Cache lifetimes punish agents that wait
Anthropic's cache has a 5-minute default lifetime, refreshed for free each time it is read; a 1-hour lifetime costs 2x the base input price to write - Claude Docs. OpenAI's cache for GPT-5.6 and later "remains eligible for reuse for 30 minutes after its most recent write or reuse," and organizations without Zero Data Retention default to extended retention of up to 24 hours - OpenAI. An agent that waits six minutes for a CI run, a slow API or a human approval loses its Claude cache and pays to write the entire context again; the same pause costs nothing extra on Sol.
In the scenario where every fourth turn waits long enough to miss the cache, the Opus bill nearly triples, from $3.46 to $10.13, while Sol's is unchanged. Writing everything with the 1-hour TTL instead costs $4.19, a 21% premium that buys immunity from pauses shorter than an hour. If your agents wait on anything slow, the 1-hour TTL on Claude is cheap insurance; our prompt caching guide walks through where to place breakpoints so the prefix stays stable.
Safety fallbacks cost money and change behavior
Opus 5.5 ships with classifiers for cybersecurity and biology, and when they decline a request, Anthropic can retry it on a fallback model. With fallbacks set to "default" (a beta on the Claude API), "the API retries a declined request on the fallback model Anthropic recommends for its refusal category" - Claude Docs. For cybersecurity, that model is Claude Opus 4.8, which costs $5/$25 per million tokens, more than Opus 5.5 itself. The refusal that triggered the fallback is also billed when it arrived mid-stream or falls into a billed category, although a "fallback credit" compensates for the fallback request's cache miss.
How often does this happen? On Vals' full suite, Opus 5.5 had a 3.99% fallback rate and a 0.79% refusal rate - Vals AI. In Anthropic's own system card, during the Gray Swan prompt-injection evaluation, "18% of rollouts were served by Claude Opus 4.8 after a classifier-triggered fallback" - Anthropic system card. And on SRE Bench, as section 4 showed, the figure was 83%. For general coding and knowledge work, fallbacks are a small tax; for security tooling, infrastructure operations and anything near biology, they can dominate both the bill and the behavior. GPT-6.1 Sol runs OpenAI's safeguards stack for a model it treats as "Critical in cybersecurity," but it does not reroute to a different model.
Speed, fast modes and rate limits
Opus 5.5 generates tokens faster but takes longer per task. Artificial Analysis measures 93 tokens per second for Opus at max against 64 for Sol at max, but 13.7 minutes per index task against 9.9 - Artificial Analysis. At high effort the gap is smaller: 5.0 minutes for Opus and 3.8 for Sol. On Vals' Terminal-Bench run the averages are 1 hour 4 minutes for Opus and 45 minutes for Sol. For interactive agents, where a person waits on each step, Sol's shorter trajectories are a real advantage.
Both vendors sell speed for money. Sol's fast mode costs 2x its standard rate, which puts it at $4/$20, exactly Opus 5.5's standard price; Opus's fast mode runs "up to 2.5x speed" for $8/$40 and is only available on the Claude API, not on Bedrock, Google Cloud or Foundry. OpenAI's new Ultrafast tier, up to 8x faster in Codex, is live for Astra at $60/$300 and promised for Sol "in the coming days." For batch work that nobody watches, both vendors' Batch APIs halve the bill, and Opus 5.5's batch mode allows up to 300K output tokens per request in beta.
Rate limits favor caching-heavy workloads on Claude. Anthropic's limits for Opus 5.5 are 1,000 requests, 2 million input tokens and 400,000 output tokens per minute, and "for most Claude models, only uncached input tokens count toward your ITPM rate limits" - Claude Docs. With an 80% cache hit rate, that is effectively 10 million input tokens per minute. Sol's limits scale with usage tier, from 500,000 tokens per minute at Tier 1 to 40 million at Tier 5. A new account running a busy agent fleet on Sol can hit Tier 1 limits quickly.
Why this matters. The same agent can cost 1.4x or 6x more on Opus than on Sol depending on how long its contexts grow and how long it pauses, before any difference in model behavior. These conditions are invisible on benchmark leaderboards, which run clean trajectories, and they are where real bills diverge from forecasts.
How to apply this. Log, per task, the peak input context, the number of cache misses, the longest pause between turns, and any fallback events. Then apply the fixes that match: compaction under 272K for Sol, 1-hour cache writes for Claude agents that wait, a non-Anthropic or explicitly named fallback for security-adjacent work, and batch mode for anything that does not need an answer within minutes.
7. Safety, hallucination and liability are costs too
An agent that reads the open web, opens email attachments or operates a browser is exposed to text written by people who want it to misbehave. When such an attack succeeds, the cost is not a failed task. It is a leaked credential, a deleted database, an unauthorized payment or a message sent in your name. Those events are rare per task and expensive per event, which makes them hard to put in a spreadsheet and easy to ignore until one happens. Our guide to AI agent sandbox security documents the incidents of the past year; this section asks what the two models' published evidence says about the probability side of that equation.
The short answer is that the evidence is asymmetric. Anthropic publishes third-party prompt-injection results with absolute attack success rates; OpenAI publishes its own robustness metrics on its own evaluations. There is no benchmark in either vendor's material, or in the independent boards above, that measures both models under the same attack. That gap is itself a cost: if you cannot compare the risk, you have to buy the controls as if both models were the weaker one.
What Anthropic publishes for Opus 5.5
On Gray Swan's indirect prompt-injection benchmark, which uses 1,804 attacks selected for transferability across models, Opus 5.5 has "an attack success rate of 0.1% at k=1, 0.7% at k=10, and 1.0% at k=15," matching Fable 5.1 as Anthropic's most robust result to date - Anthropic system card. The rate varies by setting: at 15 attempts it is 2.8% in GUI computer use, 0.5% in coding and 0.4% in tool use. In a separate test of Gray Swan's Shade attacker against computer-use environments, Opus 5.5 was broken in 2 of 2,800 attempts (0.07%), both in a single scenario.
The same system card contains the caveat that matters most for document-processing agents. Opus 5.5 "is more likely than previous models to follow malicious instructions in text that a user pastes into their own prompt." The example Anthropic gives is a user pasting the output of a package install whose last line, written by an attacker, tells AI assistants to run a remote script. If your agent accepts documents, logs or tickets from users and places them in the user turn, that content deserves the same isolation as web content. The fix costs a few lines of harness code: put untrusted material in tool results or clearly delimited data blocks, not in the user's own message.
What OpenAI publishes for GPT-6.1 Sol
OpenAI's system card addendum states that "GPT-6 Astra, GPT-6.1 Sol, and GPT-6 Sol are the most robust frontier models across both our static and multiturn robustness evaluations," and reports instruction-hierarchy robustness of 99.99% for Sol - OpenAI system card. Its prompt-injection chart, reproduced below, shows GPT-6.1 Sol at 98.983% robustness, slightly below GPT-6 Sol and below GPT-6 Astra's 99.789%.
The two metrics cannot be subtracted from each other. OpenAI's figure is a defender success rate averaged per query on its internal attacks; Anthropic's is an attacker success probability after repeated attempts on a third party's attacks. Read naively, "98.98% robust" sounds like a 1% failure rate, close to Opus's 1.0% at 15 attempts, but the attack sets, the number of attempts and the scoring are all different. What you can say is that OpenAI's own chart places 6.1 Sol below its flagship on this axis, and that OpenAI treats Sol as Critical capability in cybersecurity and High in biology and chemistry under its Preparedness Framework, applying "the same safeguards stack to GPT-6.1 Sol as GPT-6 Astra."
Hallucination: closer than the reputations suggest
For agents that look up facts and act on them, confident wrong answers are a direct cost. On Artificial Analysis's AA-Omniscience test, GPT-6.1 Sol at max has a 54% hallucination rate - Artificial Analysis, and Opus 5.5 at max has 59%, according to Artificial Analysis's comparison in its Sonnet 5.5 analysis - Artificial Analysis. Opus answers more questions correctly, which is why its overall AA-Omniscience index is higher (46 against 42), but it is not the more cautious of the two. Neither comes close to Gemini 4 Argon's 15%. For both models, a retrieval-and-verify step is the reliable control, not model choice.
Liability is becoming part of the price
On September 30, a Senate Homeland Security subcommittee held a hearing on "rogue" AI agents, at which Senator Josh Hawley called for clarifying who is liable for AI agent hacking incidents; OpenAI's chief executive was invited and declined to attend - Roll Call. Hawley and Senator Chris Murphy announced the AI Agent Accountability Act, which would make AI firms liable for reckless design and users liable for reckless deployment - Senator Hawley. It is a proposal, not law, but it signals that the "cost of a failure" in the break-even formula may soon include legal exposure for the deployer, not just cleanup time.
Why this matters. Security incidents are the tail of the failure distribution, and the tail is where model differences become expensive. Opus 5.5 has the stronger published evidence on indirect injection; neither model has evidence you can compare directly against the other's. For any agent with outward-facing or irreversible actions, the control layer around the model matters more than the model's own robustness.
How to apply this. Price the controls into your cost per task rather than treating them as overhead: an approval gate for irreversible actions, an egress allowlist, separate credentials per agent, and a sandbox the agent cannot modify. If your agent does security operations or infrastructure work, test both models with Anthropic's fallbacks enabled and disabled, because the cheapest model on paper may be the one that refuses half of your tasks.
8. Verdicts by workload
The evidence so far adds up to a pattern that is more useful than a single winner. GPT-6.1 Sol is the better value wherever success is easy to verify, tasks are bounded, and failures are cheap; Claude Opus 5.5 is the better value wherever output quality is judged rather than tested, tasks are long and open-ended, and failures are expensive. Most real agents mix both kinds of work, so the verdicts below are for task types rather than for whole products.
Each verdict names a default configuration, an escalation path and the evidence behind it. The defaults lean toward the cheaper setting that clears the quality bar on independent data, and the escalations reserve the expensive settings for the cases where the break-even analysis says they pay. Think of them as starting points for your own bake-off, not as conclusions to adopt untested.
Coding agents
For software engineering with a test suite, GPT-6.1 Sol at high effort is the default. It ties Opus 5.5 on Mercor's DeepSWE, sits within noise on Vals' Code Migration, and trails by 1.4 points on Vibe Code Bench, all at roughly a tenth of the cost and in less time. Tests act as a cheap verifier, which makes the retry-then-escalate pattern from section 5 work well. Escalate to Opus 5.5 at high for tasks that fail twice, for large refactors where a reviewer will judge design quality, and for long terminal work where Opus's 10-point Terminal-Bench lead and a $114 break-even favor it.
There is one important exception. Security-adjacent work, such as vulnerability triage, infrastructure operations or exploit reproduction, runs into Opus 5.5's cyber safeguards and can end up on Opus 4.8. For that work, Sol's SRE Bench result (50.76% against 33.59%) is the stronger evidence, and Anthropic's own route for verified defenders is its Cyber Verification Program. For a broader comparison of the coding tools that wrap these models, see our AI coding CLI comparison.
Business workflow automation
For CRM updates, routing, invoicing and support operations, the decision turns on the cost of a wrong outcome. If an error is cheap and reversible, Sol at medium is the efficient default, which is the configuration OpenAI quoted. If an error reaches a customer, moves money or corrupts records, Opus 5.5 at max is cheaper per correct outcome once a failure costs more than about $18, as the break-even table showed. The middle path, Sol with hard state assertions after each workflow and Opus for the ones that fail them, captures most of both.
Document-heavy knowledge work
For reading and answering questions from PDFs, contracts and reports, Sol at high is at least as accurate as Opus on Artificial Analysis's GDP.pdf runs and costs less than half as much. For producing deliverables that a person will judge, such as memos, analyses, board decks and research reports, Opus 5.5 has the largest lead in this comparison: 271 Elo points on GDPval-AA at max, and a match for Sol's best at its own medium default. For spreadsheets and financial models, Opus leads by 8 points on Surge's GDP.xlsx. We covered what GDPval measures, and how much of real work it captures, in Can AI do my job?.
Computer use agents
There is no clean comparison here, because OpenAI reports Sol on OSWorld 2.0 (71.4% at max, within 2.1 points of Astra at about one-seventh of the cost) and Anthropic reports Opus on OSWorld 2.1 (81.8% partial score). What you can compare is cost structure. Screen-driven agents generate long trajectories with many screenshots, which favors Sol's lower token price and penalizes Claude's cache if steps are slow; but they are also the setting where injection risk is highest, and Opus 5.5's 2.8% attack success at 15 attempts in GUI use is the only published third-party figure. Our analysis of OSWorld 2.0 failure modes explains why most computer-use tasks still fail on both.
Research and scientific workflows
On Terminal-Bench Science, OpenAI's chart puts Opus 5.5 at 63.3% for $23.21 per task and Sol at 57.0% for $5.47, with Astra still ahead of both at 68.1%. The break-even is $281 per failed task, which most research workloads exceed, since a failed analysis usually costs a researcher hours. Run Sol at xhigh for exploratory and high-volume work and Opus 5.5 or Astra for the analyses you will publish or act on; our GPT-6 Astra pricing analysis covers when Astra's higher token price is worth paying.
Always-on agents
The new class of persistent agents makes per-task cost a monthly line item. OpenAI's dots run on GPT-6 Astra, not Sol, and connect to "over 4,000 apps" through plugins - OpenAI; xAI's Grok Bot gives bots "their own computer" - xAI. For teams building their own always-on agent, the arithmetic is simple: an agent that runs 200 tasks a day at Sol's $0.32 per index-task-equivalent spends about $64 a day, and at Opus high's $1.82 about $364. Most of an always-on agent's work is monitoring, triage and routine follow-up, which is Sol's strength; the occasional judgment call is where Opus earns its place.
The table below condenses these verdicts into a starting allocation. Each row pairs a default configuration with an escalation path and the single strongest piece of independent evidence behind it, so you can see at a glance which assumption to test first in your own bake-off. Where the evidence is a tie or within noise, the default goes to the cheaper model, because a tie at one-tenth of the price is not a tie in cost per task.
| Workload | Default | Escalate to | Key evidence |
|---|---|---|---|
| Coding with tests | GPT-6.1 Sol (high) | Opus 5.5 (high) after two failures | DeepSWE tie at 72.3%; Code Migration within noise at 1/17 the cost |
| Security and infrastructure ops | GPT-6.1 Sol (high) | Specialist or verified-access models | SRE Bench 50.76% vs 33.59%; Opus fallbacks on 83% of tasks |
| Business workflows, costly errors | Claude Opus 5.5 (max) | Human review on failed assertions | AutomationBench 42.5% vs 36.1%; break-even about $18 |
| Judged deliverables and spreadsheets | Claude Opus 5.5 (medium to high) | Opus 5.5 (max) for final drafts | GDPval-AA 1,846 vs 1,575 Elo; GDP.xlsx 30.3% vs 22.0% |
| PDF and document Q&A | GPT-6.1 Sol (high) | Opus 5.5 (high) when answers feed decisions | GDP.pdf 32.0% vs 28.8% at under half the cost |
Why this matters. The difference between a good and a bad model allocation is larger than the difference between the two models. A team that runs Opus at max on PDF lookups overpays several times over; a team that runs Sol at medium on customer-facing workflows with expensive errors underpays on tokens and overpays on cleanup.
How to apply this. Tag each step in your agent with one of the workload types above, assign the default configuration, and wire the escalation path into the harness rather than leaving it to a person. Review the tags monthly against your logs: steps that escalate often belong one tier up by default.
9. Subscriptions, managed runtimes and where the bill lands
Not every agent bill comes from the API. A growing share of agent work runs inside Codex and Claude Code, billed through flat subscriptions, and inside managed runtimes where the vendor runs the loop. Each of these changes what "cost per task" means, because the marginal cost of a task on a subscription is zero until you hit the plan's limits, and the managed runtimes bundle compute, sandboxes and orchestration with the tokens.
On the OpenAI side, GPT-6.1 Sol is available "to all Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex," though "not yet available in Chat" - OpenAI. ChatGPT Plus costs $20 a month, and Pro now comes in three tiers: Pro 100, Pro 200 and the new Pro 500, which includes Astra Ultrafast - OpenAI Help Center. For ChatGPT Work and Codex, the updated Pro 200 allowance is "10 times the Plus allowance, while Pro 500 includes 25 times the Plus allowance." Fast mode consumes the allowance at 2.5 times the standard rate.
On the Anthropic side, Claude Pro costs $20 a month ($17 on annual billing) and Max starts at $100 for 5x or 20x Pro's usage, both with Claude Code included - Anthropic. With the Opus 5.5 launch, Anthropic increased "five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans" and gave subscribers a saved rate-limit reset. Because Opus 5.5 uses several times as many tokens per task as Sol, the same subscription allowance buys fewer Opus tasks than Sol tasks, which is the subscription version of the per-task gap.
Managed runtimes move the cost boundary
DevDay pushed OpenAI further into running agents rather than selling tokens. The Agents API now supports computer use and "brings Codex's multi-agent capabilities, tool search, tool calling, and context compaction into your application," and Bedrock Managed Agents lets teams "build OpenAI agents that run entirely in AWS" - OpenAI DevDay recap. Anthropic's equivalent is Claude Managed Agents; we covered it in Claude Managed Agents and OpenAI's runtime in OpenAI Agents API. In both, the vendor's compaction and caching policies decide a large part of the bill, which makes the hidden multipliers from section 6 the vendor's problem as much as yours.
Microsoft's new Copilot illustrates where enterprise pricing is going. Its Auto mode "weighs accuracy, speed and cost on each request to route to the model best suited for the job," while Cowork, Code, Autopilot "and frontier models like Astra and Fable all run on UBB," Microsoft's usage-based billing - Microsoft. Governance is following the money: Dataiku launched a standalone Agent Management product on September 24 that inventories agents across vendors' stacks - SiliconANGLE. The direction is clear: model choice is moving inside routers, and cost per task is becoming a metric that platforms report rather than one that teams compute.
At the far end of that spectrum sit platforms that take the model decision away from the buyer entirely. Investors are betting heavily on this layer: Instinct, a personal agent that uses its own computer and phone to finish tasks, raised $1 billion at a $10 billion valuation - Tech Startups. Platforms such as O-mega, which builds and operates an entire autonomous company through AI agents, sit in the same place for businesses: the user describes outcomes, and the platform chooses which model handles which step. Whether you build the router yourself or buy it inside a platform, the economics in this guide still apply; the difference is who measures them. Our ranking of enterprise AI agent platforms compares the options.
Why this matters. A subscription hides marginal cost until you hit a limit, and a managed runtime bundles costs you used to control. Both are reasonable trades, but both make it harder to see the per-task differences this guide measures, and both make switching models harder later.
How to apply this. Use subscriptions for interactive, human-in-the-loop agent work where the allowance covers your volume, and the API for automated pipelines where you need per-task metering and routing. If you adopt a managed runtime, ask the vendor for per-session token and cache reports before you commit, so you can still compute cost per solved task.
10. How to run your own cost-per-task bake-off
Everything in this guide is a starting point because your tasks are not the benchmarks' tasks. The good news is that a credible bake-off between GPT-6.1 Sol and Claude Opus 5.5 is small: 50 to 100 representative tasks, a verifier for each, and a cost calculator that reads the usage fields both APIs return. Most teams can run one in a day, and it will tell them more than every chart above.
The procedure has five steps, and the order matters because each step narrows the next. Skipping the effort sweep, in particular, is the most common reason bake-offs reach the wrong conclusion: comparing both models at their defaults compares Sol at medium with Opus at medium, which section 3 showed is not where either model is most efficient for many tasks.
- Sample real tasks from your logs, weighted by volume, and write a pass/fail check for each
- Sweep effort levels for both models: low, medium, high, xhigh and max
- Record usage per attempt: uncached, cached and cache-write input, output, and wall time
- Compute cost per solved task and the break-even failure cost per task type
- Choose a default and an escalation for each task type, then re-run monthly
The checks in step 1 are what make the rest possible. If a task cannot be checked automatically, write a rubric and have a person score a sample, because without a verifier you cannot measure success rates, and without success rates cost per task is meaningless. Step 3 is where most teams under-instrument: logging only total tokens hides the cache behavior that section 6 showed can triple a Claude bill. The calculator below reads the usage objects from both APIs. It is illustrative, using the published standard rates as of this writing, and it ignores the long-context premium and batch discounts, which you should add for your own workload.
# Illustrative cost-per-attempt calculator using standard published rates (October 2026).
# OpenAI Responses usage: input_tokens includes cached and cache-write tokens.
# Anthropic Messages usage: input_tokens excludes cache reads and cache writes.
SOL = {"input": 2.00, "cache_read": 0.10, "cache_write": 2.50, "output": 10.00}
OPUS = {"input": 4.00, "cache_read": 0.20, "cache_write_5m": 5.00,
"cache_write_1h": 8.00, "output": 20.00}
def sol_cost(usage):
details = usage.get("input_tokens_details", {})
cached = details.get("cached_tokens", 0)
written = details.get("cache_write_tokens", 0)
uncached = usage ["input_tokens"] - cached - written
return (uncached * SOL ["input"] + cached * SOL ["cache_read"]
+ written * SOL ["cache_write"]
+ usage ["output_tokens"] * SOL ["output"]) / 1e6
def opus_cost(usage, one_hour_cache=False):
write_rate = OPUS ["cache_write_1h"] if one_hour_cache else OPUS ["cache_write_5m"]
return (usage ["input_tokens"] * OPUS ["input"]
+ usage.get("cache_read_input_tokens", 0) * OPUS ["cache_read"]
+ usage.get("cache_creation_input_tokens", 0) * write_rate
+ usage ["output_tokens"] * OPUS ["output"]) / 1e6
def cost_per_solved(attempt_costs, passed):
return sum(attempt_costs) / max(sum(passed), 1)
def break_even_failure_cost(c_opus, p_opus, c_sol, p_sol):
return (c_opus - c_sol) / (p_opus - p_sol) if p_opus > p_sol else None
OpenAI's usage object reports cached and written tokens inside input_tokens_details, as its caching guide shows, while Anthropic reports cache_read_input_tokens and cache_creation_input_tokens separately from input_tokens. Reasoning and thinking tokens are billed as output on both platforms. Sum the per-request costs across every request in a task to get cost per attempt, and keep failed attempts in the total: they are part of what a solved task costs.
Controls worth testing during the bake-off
Both vendors now let you change effort in the middle of a conversation without losing the cache, which opens a cheaper pattern than picking one setting for a whole task. OpenAI's configuration_update input items "increase reasoning effort for difficult work or reduce it for routine follow-ups without rewriting the original prompt prefix" - OpenAI. Anthropic supports per-message effort on Opus 5.5 "with a per-message output_config, which preserves the prompt cache" - Claude Docs. Test a variant of each model that plans at high effort and executes at low or medium.
Anthropic's task budgets are the other control worth testing on Opus. A task_budget in output_config caps "the number of tokens Claude can spend across the agentic loop, including thinking, tool calls, tool results, and output," and the model sees a countdown so it can "finish gracefully (summarize findings, report progress) as it approaches the budget rather than cutting off mid-action" - Claude Docs. Because most of Opus's cost gap is tokens per task, a budget set at the 75th percentile of successful runs is a direct way to cap the long tail of runaway trajectories. The snippet below shows the shape of a budgeted request; check the current beta header in the docs before using it.
# Sketch: an Opus 5.5 request with explicit effort and a task budget (beta).
import anthropic
client = anthropic.Anthropic()
response = client.beta.messages.create(
model="claude-opus-5-5",
max_tokens=64000,
betas= ["task-budgets-2026-03-13"],
output_config={
"effort": "high",
"task_budget": {"type": "tokens", "total": 150000},
},
tools=my_tools, # your tool definitions
messages=conversation, # the running agent conversation
)
On the Sol side, the equivalent discipline is context management. Track peak input tokens per request and trigger compaction before 272K; route classification and extraction steps to a smaller tier rather than to Sol at low; and use Flex or Batch processing, at half price, for anything that can wait. Our LLM cost reduction guide covers these techniques in more depth, and the best AI agent evals guide covers how to build the task set itself.
Why this matters. Every number in this guide comes from someone else's tasks, harness and verifier. The break-even thresholds in section 5 span from $18 to nearly $7,000, which means small differences between your work and a benchmark's can change the answer. A one-day bake-off removes that uncertainty for the price of a few hundred dollars in tokens.
How to apply this. Run the five steps on your highest-volume task type first, since that is where the savings or the quality gains compound fastest. Repeat whenever either vendor ships a new model or changes a price, which, judging by September, will be often.
11. The bottom line
Both cost-per-task claims survive scrutiny, and neither answers the buyer's question on its own. GPT-6.1 Sol is the best value in this comparison for most agent steps: on independent data it ties Opus 5.5 on long software engineering, leads on PDF work and security-adjacent operations, and costs between 2x and 17x less per task depending on the workload. Claude Opus 5.5 is the most capable model you can buy, with a lead that is largest on judged knowledge work (271 Elo on GDPval-AA), spreadsheets and long terminal tasks, and it is the cheaper choice per correct outcome whenever failures are expensive enough.
The decision framework that falls out of the evidence is short. Start from cost per solved task, not price per token. Compute the break-even failure cost for each task type and compare it with what a failure actually costs you. Default to GPT-6.1 Sol at high or medium effort where outcomes can be verified cheaply, and to Opus 5.5 at high effort where outputs are judged or errors are costly. Use Opus 5.5 at max for the small set of tasks where a mistake is expensive and the task itself is cheap to attempt. Avoid Opus for security-adjacent work unless you have verified access, because its safeguards will route that work elsewhere.
Then manage the multipliers that the leaderboards do not show. Keep Sol requests under 272K tokens with compaction, write Claude caches with a 1-hour TTL when your agent pauses, cap Opus trajectories with task budgets, and log fallbacks and cache misses per task so you can see them. Those four habits can change your bill by more than the model choice itself.
Finally, expect the comparison to move. September produced nine agent-relevant model releases in ten days, Anthropic has promised Haiku 5.5 "in the coming weeks," and OpenAI has promised Sol Ultrafast "in the coming days." A routed system that can re-run a bake-off and change a default in an afternoon will capture each of those shifts; a system wired to one model will need a migration project every time. The labs are now competing directly on cost per task, which is good news for buyers, as long as buyers measure it themselves.
This guide reflects pricing, benchmarks and product availability as of October 2, 2026. Model prices, effort defaults, safeguard behavior and leaderboard scores change frequently, so verify current figures with each vendor and evaluator before making purchasing decisions.