What Claude Code actually costs in August 2026, verified against the live pricing pages this week, plus the honest math on limits, models, and every alternative worth pricing.
Claude Code passed $2.5 billion in run-rate revenue by February 2026, and roughly 4% of all public GitHub commits are now authored by it - Anthropic.
Here is the fast answer, because you came here for a price. Claude Code starts at $20 per month on the Pro plan ($17 billed annually), scales through Max at $100 and up, and every plan shares one usage pool across Claude Code, claude.ai chat, and Cowork, with overflow billed at API rates. The free plan does not include Claude Code at all - Claude pricing. Everything past that sentence is where the real budgeting decisions live: which model your plan actually runs, what the usage limits are this month (they have changed six times since May), and when a competitor or the raw API is the cheaper meter.
A word on why this page was rewritten rather than patched. The previous version of this guide, dated July 8, was accurate for exactly five days. Then Anthropic extended its limit boost twice, restructured Fable 5 from a free trial into a plan feature, shipped Claude Opus 5 as the new Max default on July 24, and suffered a capacity outage that ended in usage-limit resets. Four load-bearing claims died in three weeks. We pay Anthropic invoices ourselves at O-mega, we watch these meters daily, and the only way to keep a pricing page honest at this market's tempo is to re-verify everything and say plainly what changed. Every number below was checked against a live source during the first week of August 2026, and where a price has become unverifiable (two vendors now hide tier prices behind checkout), we say that instead of quoting a stale figure.
This guide covers every Claude plan and what it includes, the August 2026 model lineup and the version requirements that silently decide what you run, the Fable 5 restructuring nobody's July articles caught, the full API rate card with caching and batch mechanics, the usage-limit timeline through the current August 19 expiry, subscription-versus-API break-even math, and a head-to-head against Codex, Cursor, GitHub Copilot, Gemini CLI, Kiro, Devin, and (new this edition) the open-source agents that the alternatives conversation now starts with.
Contents
- Claude Code Pricing at a Glance: August 2026
- Every Claude Plan in Detail
- The August 2026 Model Lineup: Opus 5 Changes the Defaults
- The Fable 5 Restructuring: From Free Window to Plan Feature
- The Full API Rate Card: Caching, Batch, Long Context, and Tools
- The 30% Tokenizer Effect
- The Usage-Limit Timeline: Boosts, Outages, and Resets
- Subscription vs API: The Break-Even Math
- Token Efficiency, Benchmarks, and the Three Answers Problem
- What Your Subscription Buys Beyond the CLI
- Alternatives Head-to-Head: Codex, Cursor, Copilot, Gemini CLI, Kiro, Devin
- The Open-Source Path: OpenCode, Cline, Aider, and Antigravity
- Enterprise Buying: Seats, CCUs, and Data Residency
- The Cost-Optimization Playbook
- Decision Framework: Which Option Fits You
The Verdict Table: Every Option Scored
Before the detail, the summary. We scored the eight realistic ways to run a coding agent in August 2026 on the four criteria that determine what a dollar of spend actually buys, and every score's justification sits in its cell because a bare number tells you nothing. Value per dollar (30%) weighs entry price against what that price buys in real work, including token efficiency. Frontier capability (30%) reflects the current independent benchmark picture (section 9 explains why self-reported scores are excluded). Limits and predictability (20%) measures how transparent and stable the caps, meters, and prices are, a criterion this quarter's price-hiding trend makes newly important. Product surface (20%) measures how far the subscription reaches beyond one terminal or editor.
Every data point in the table is verified in the sections that follow; the two ties and the ordering are explained there too. Google's free Antigravity platform is deliberately absent: it is priced at $0 and unbenchmarked, so it appears in section 12 rather than in a scored ranking it would distort.
| # | Tool | What It Does | Value per Dollar (30%) | Frontier Capability (30%) | Limits & Predictability (20%) | Product Surface (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | Claude Code | Terminal-first agent, $20 up, leads agentic benchmarks | 8 - $20 runs Opus 5-class work; Max beats $800-1,500 API-equivalent full-time spend | 10 - Fable 5 leads Terminal-Bench 2.1 at 83.8% | 6 - shared pool, boost expires Aug 19, six limit changes in four months | 9 - CLI, web, desktop, Cowork, Design, Science, agent teams | 8.4 |
| 2 | OpenAI Codex | ChatGPT-bundled agent from $8/mo, best token efficiency | 9 - $8 Go entry; 3-4x fewer tokens per task (directional) | 9 - GPT-5.5 at 83.1% Terminal-Bench 2.1 | 8 - published per-model windows and credit schedule | 7 - ChatGPT app, CLI, IDE extension, cloud tasks | 8.4 |
| 3 | OpenCode (open source) | Free harness on your own API key or subscription | 7 - $0 harness; Sonnet 5 intro + Haiku undercut plans for bursty use | 8 - rides any frontier model via 75+ providers | 9 - no weekly caps or quota politics, just the token meter | 5 - terminal/IDE/desktop, no managed sandboxes or support | 7.3 |
| 4 | Gemini CLI | Free terminal agent, 1,000 requests/day at $0 | 10 - 1,000 free requests/day with a Google account | 6 - model version now at Google's discretion, trails leaders | 8 - simple daily quota, generous free tier | 4 - CLI-centric, thinner agent ecosystem | 7.2 |
| 5 | Cursor | AI-native IDE, $20/mo Pro, frontier-model credits | 7 - $20 Pro with frontier credits | 8 - routes frontier models; Cursor CLI + Grok 4.5 hits 79.3% | 6 - Pro+/Ultra prices now hidden behind checkout | 6 - IDE-centric with background agents | 6.9 |
| 6 | GitHub Copilot | IDE assistant on $0.01 AI credits, unlimited completions | 7 - $10 entry with $15 credits, completions unlimited | 7 - multi-model access incl. Sonnet 5, no in-house frontier model | 7 - transparent dollar-denominated credits | 6 - IDE plus GitHub platform integration | 6.8 |
| 7 | Kiro (AWS) | Spec-driven agent IDE on credit packs | 6 - $20 Pro for 1,000 credits, $0.04 overage | 6 - Sonnet 4.5-generation included models | 6 - clear pricing but no monthly rollover | 4 - IDE-only surface | 6.0 |
| 8 | Devin | Autonomous cloud engineer, $20 Pro to $200 Max | 5 - $20 entry but opaque work-per-dollar | 6 - capable autonomous runs, no current leaderboard entries | 4 - allowance refreshes daily/weekly with no visible ledger | 6 - autonomous cloud sessions, Desktop, Teams | 5.4 |
How to read it: the weights (30/30/20/20) sum to 100%, and each final score is the weighted average of its row, sorted descending. The Claude Code versus Codex tie at 8.4 is the honest headline of this market: the capability leader and the efficiency leader currently price out even, and which one wins for you depends on the workload shape that sections 8 and 9 teach you to measure. The scores lean on independent numbers only; where a vendor hides a price (Cursor's top tiers, Anthropic's Max 20x) or a meter (Devin's allowance), the predictability score says so.
1. Claude Code Pricing at a Glance: August 2026
Before any analysis, the raw numbers. This table is the answer to "what does Claude Code cost," pulled from the official pricing page this week, and the two paragraphs after it flag the three things the table cannot show: the shared pool, the moving limits, and the one price Anthropic no longer prints.
Claude subscription pricing (verified August 2026) - Claude pricing:
| Plan | Monthly | Annual (per month) | Claude Code | Usage |
|---|---|---|---|---|
| Free | $0 | - | No | Chat only |
| Pro | $20 | $17 ($200 up front) | Yes | 1x baseline |
| Max 5x | $100 | - | Yes | ~5x Pro |
| Max 20x | See note | - | Yes | ~20x Pro |
| Team Standard | $25/seat | $20/seat | Yes | Per-seat baseline |
| Team Premium | $125/seat | $100/seat | Yes | 5x Standard |
| Enterprise | $20/seat + usage | - | Yes | Billed at API rates |
The note on Max 20x is itself a finding: the official page now lists Max only as "From $100" per month, and no $200 figure appears anywhere on it; the 20x tier's price surfaces at checkout - Claude pricing. When we documented the tier directly in our June 2026 cost snapshot, it was $200 per month, and nothing suggests that changed, but a pricing page that stops printing its top price is worth noticing. Vendors bury numbers that are in motion.
Two structural facts frame every row. First, quota is shared: Claude Code, claude.ai chat, and Cowork all draw from one pool per account, so the table's "usage" column describes a household budget, not a per-tool one - Anthropic support. Second, limits are currently boosted 50% through August 19, 2026, a promotion that has already been extended at least four times - Digital Applied. Any capacity plan built on this month's subjective experience is built on a promotion, and section 7 gives you the full timeline so you can budget on the baseline instead.
2. Every Claude Plan in Detail
Claude Code is not sold standalone. It ships inside Claude subscriptions from Pro upward, and the Pro bundle now includes far more than a CLI: Claude Code, Claude Cowork, Claude Design, and Claude Science all ride the same $20 - Claude pricing. That bundling is deliberate competitive architecture. No single-surface rival can match a four-product bundle at $20, so understanding what the bundle displaces in your own stack is half of evaluating the price.
The individual ladder has three rungs. Pro at $20 per month (or $17 on annual billing, $200 up front) is the entry point and covers most working developers most of the time. Max 5x at $100 multiplies the usage pool roughly five times, and the Max 20x tier roughly twenty times, for people whose agents run more hours than they do. On the team side, Team Standard runs $25 per seat monthly ($20 annual) and Team Premium $125 per seat monthly ($100 annual) with 5x the usage of Standard; teams span 2 to 150 people, and Enterprise converts the whole structure into $20 per seat plus usage billed at API rates - Claude pricing. Several competing pricing pages claim a five-seat minimum on Team Premium; we could not confirm any seat minimum on the official page this run, so we are not printing one.
The annual-versus-monthly arbitrage on Pro is straightforward: $204 per year versus $240 is a 15% discount on a product whose value is its quota, so the only rational reason to stay monthly is genuine doubt about whether the tool sticks. Notice that Anthropic publishes annual pricing for Pro and the Team seats but keeps Max month-to-month on the public page. A vendor that will not sell you a twelve-month lock on its heaviest tier is telling you it expects its own price structure to keep moving, and the past four months of limit changes (section 7) say the same thing.
The mechanism that changed the shape of all these plans is extra usage. Hitting a cap used to mean a hard stop; now every plan behaves like a flat rate with a metered overflow, billing additional work at standard API rates from the same account, and Anthropic is explicit that all transitions to API-rate billing require your consent: you can always decline and wait for the reset instead - Anthropic support. The practical consequence is that plan choice is no longer a capacity ceiling, only a price break point. A Pro subscriber who overflows roughly $30 of API-rate work each month is paying $50 all-in, and should move to Max 5x only when the overflow regularly exceeds the $80 gap between tiers. One month of honest /status checks plus your overflow line item beats any pricing page's persuasion architecture.
There is a quieter consequence too: because the pool is shared, a heavy Cowork afternoon eats your evening's Claude Code budget, and a remote Cowork session you brief at 6pm can spend quota all night (section 10). Plan-tier math is household math. If Cowork is a meaningful part of your usage, our Claude Cowork pricing and ecosystem guide breaks down how session-based agent work drains the pool differently from interactive coding, and if you are starting from zero, the Claude Code beginner's guide covers setup before any of this arithmetic matters.
3. The August 2026 Model Lineup: Opus 5 Changes the Defaults
This is the section that rots fastest in every pricing article, including the last edition of this one, so here is the current state plainly. Claude Opus 5 launched July 24, 2026 at $5/$25 per million tokens, the same price as Opus 4.8, and it is now the default model on Max and the strongest model on Pro - Anthropic. Anthropic's own positioning is unusually direct: Opus 5 comes "close to the frontier intelligence of Claude Fable 5 at half the price," landing within 0.5% of Fable 5 on CursorBench 3.2 at half the cost per task, beating Fable 5's best OSWorld 2.0 result at roughly a third of the cost, and posting an ARC-AGI 3 score three times the next-best model - Anthropic. Those are vendor benchmarks and deserve vendor-benchmark skepticism, but the pricing implication does not depend on them: the flagship tier got a generation better without getting a dollar more expensive.
The strategic detail most coverage missed is the safety architecture. Fable 5's safety classifiers, which trigger most often in cybersecurity and biology domains, are expected to engage 85% less often on Opus 5, and a new automatic fallback feature routes classifier-flagged API requests to another model instead of returning an error - TechCrunch. For agentic coding, where a refusal mid-pipeline is a broken pipeline, "slightly less capable but far less likely to stop" is a real economic property, not a footnote. Opus 5 also escapes the 30-day data-retention policy that applies to Fable and Mythos models, which matters to compliance reviews - TechCrunch.
Here is the model-to-price mapping as it stands, from the live rate card - Anthropic pricing docs:
| Model | API price (in/out per MTok) | Status in August 2026 |
|---|---|---|
| Fable 5 | $10 / $50 | Frontier tier; included in Max and Team Premium up to 50% of weekly limits |
| Mythos 5 | $10 / $50 | Limited availability (approved partners) |
| Opus 5 | $5 / $25 | Default on Max, strongest on Pro |
| Opus 4.8 / 4.7 / 4.6 / 4.5 | $5 / $25 | Previous Opus generations, selectable |
| Sonnet 5 | $2 / $10 intro, then $3 / $15 | Intro pricing ends August 31, 2026 |
| Sonnet 4.6 / 4.5 | $3 / $15 | Prior Sonnet generation |
| Haiku 4.5 | $1 / $5 | Cheapest current model, still no 5-series |
| Opus 4.1 | $15 / $75 | Deprecated; triple the price of its successors |
The chart's shape is the story: five Opus generations now sit at one price point while the deprecated 4.1 lingers at triple the cost, and a $10/$50 frontier tier floats above a $2/$10 Sonnet floor. One curiosity worth naming: Haiku is the only family without a 5-series release. Every other tier turned over in 2026; the budget tier did not, which suggests either a Haiku 5 is coming or Anthropic considers Sonnet 5's intro pricing its real budget play until September.
For subscribers, the operational takeaway is version hygiene. Opus 5 requires Claude Code v2.1.219 or later, Sonnet 5 requires v2.1.197+, Opus 4.8 requires v2.1.154+, and Fable 5 requires v2.1.170+; a months-stale CLI silently runs an older default - model configuration docs. The alias system decides the rest: opus now resolves to Opus 5 on the Anthropic API, a best alias selects Fable 5 where your organization has access (falling back to the latest Opus), and the opusplan mode runs Opus during planning then drops to Sonnet for execution, which is a cost-control feature disguised as a convenience - model configuration docs. Run claude update before you evaluate anything about cost or quality, and see our Opus 5 versus 4.8 benchmark and cost breakdown for the upgrade-or-not decision in detail.
Two smaller rate-card notes complete the picture. Fast mode, still a research preview, now covers Opus 5 and Opus 4.8 at $10/$50, roughly 2.5x the default speed at double the base price, and it is API-only: Opus 4.7 fast-mode requests simply return an error now - Anthropic pricing docs. And the release cadence itself is a planning input: four flagship releases between late May and late July (Opus 4.8, Fable 5, Sonnet 5, Opus 5) is the current tempo. Treat model IDs as configuration, review pins monthly, and treat any release week as a week to watch /status more closely, because a default-model change can shift your quota burn overnight.
4. The Fable 5 Restructuring: From Free Window to Plan Feature
If you read anything about Claude pricing in June or early July, you read that Fable 5's free inclusion was temporary and would lapse into pay-as-you-go for everyone. That is not what happened, and the actual outcome redraws the value of every plan tier, which is why it gets its own section.
The sequence, reconstructed from the announcements: free Fable 5 access for paid subscribers, originally scheduled to end June 22, was extended to July 12, then July 19, with the final window closing at 11:59:59 PM PT on July 19 - Digital Applied. From July 20, the permanent structure took effect. Max and Team Premium subscribers keep Fable 5 included, capped at up to 50% of weekly limits and drawn from the shared usage pool, not as an additive allowance. Pro and Team Standard moved to pay-as-you-go usage credits at Fable's $10/$50 API rates, softened by a one-time $100 transition credit that expires September 19, 2026 - Just Being Resourceful.
Read the economics of that split honestly. For a Max subscriber, "half your weekly pool can be Fable" converts the $100-and-up tiers from usage multipliers into frontier-model access tiers: that is now the cheapest way to run sustained Fable 5 workloads, full stop. For a Pro subscriber, the $100 credit is a metered trial with a 60-day fuse: at $10/$50 rates, a single heavy agentic task can consume $30-50 of it (a realistic long session with roughly 3M input and 300K output tokens prices at about $45 at list rates), so the credit funds a handful of serious experiments, not a habit. The rational Pro pattern is unchanged from what we advised in July: escalation, not adoption. Run everything on your plan's default, and reach for Fable 5 on the specific problems where Opus 5 has already failed twice, because a third Opus attempt costs quota and time while a first Fable attempt costs real dollars and often ends the loop.
The July 24 Opus 5 launch then quietly undercut the product it had just been restructured around. When the mid-tier model is marketed as Fable-class intelligence at half the price with far fewer safety refusals (section 3), the population of tasks that genuinely need Fable shrinks to long-horizon, maximum-difficulty work. In Claude Code, Fable 5 is never the default even where included: you select it with /model fable, its classifier-flagged requests trigger automatic fallback, and it is unavailable under zero-data-retention configurations - model configuration docs. For a current read on whether the escalation is worth it for your workload class, see our guides to Fable 5 availability and access in practice and the Fable 5 and Mythos 5 benchmark picture.
5. The Full API Rate Card: Caching, Batch, Long Context, and Tools
If you build on the API directly, or you enable extra usage on a subscription (same rates), the per-token price is only the starting point. Four multiplier systems stack on base pricing, and using them well is routinely the difference between a painful bill and a trivial one. Everything in this section comes from the official pricing documentation, re-verified this week - Anthropic pricing docs.
Prompt caching is the biggest lever. A 5-minute cache write costs 1.25x base input, a 1-hour write costs 2x, and a cache hit costs 0.1x: on Opus 5 that means cache reads at $0.50 per million tokens. The docs now state the payoff rule plainly: caching pays for itself after one read on the 5-minute tier and two reads on the 1-hour tier. Coding agents are cache-friendly by construction (system prompts, CLAUDE.md, repository context repeat every turn), so real sessions serve most input from cache. Batch API stacks on top for async work at 50% off both input and output, which prices Opus 5 batch at $2.50/$12.50 and turns Haiku 4.5 into a $0.50/$2.50 workhorse.
To see the multipliers combine, price a realistic two-hour Opus 5 session: 2M input tokens, of which 1.6M is repeated cached context and 0.4M fresh, plus 150K output. Naive list-price math says $13.75. With caching, the repeated context bills at $0.50/MTok ($0.80), fresh input carries the 1.25x write premium ($2.50), output is unchanged ($3.75): the session lands near $7, roughly half the naive figure. This is why cost comparisons between tools that cache well and tools that do not are meaningless at list price, and why section 14 treats caching discipline as the first optimization rather than an advanced one.
Tool and platform pricing (August 2026) - Anthropic pricing docs:
| Item | Price | The catch or the gift |
|---|---|---|
| Code execution | 1,550 free hours/mo, then $0.05/hr | Free when paired with current web search or web fetch tools |
| Web search | $10 per 1,000 searches | Plus tokens for retrieved content |
| Web fetch | $0 | Standard token costs only |
| Managed Agents runtime | $0.08 per session-hour | Metered only while status is running; replaces container-hour billing |
| US data residency | 1.1x multiplier on all token categories | Claude 4.6+ via inference_geo: "us" |
There is also a class of costs the rate card documents and almost nobody reads: tool definitions bill as input tokens, on every request, forever. Enabling any tool adds a per-model tool-use system prompt (286 tokens on Opus 5 with auto tool choice, and notably down from 675 on Opus 4.7, so the current generation actually got cheaper here), the bash tool adds 325 tokens on Opus 5/4.8/4.7 versus 244 on earlier models, the text editor adds 700, and computer use adds 735 plus screenshot vision costs - Anthropic pricing docs. Each number is trivial alone and none of them is trivial multiplied by every turn of every session of every agent in a fleet. The same page's web-fetch guidance is the right mental model for content costs: an average web page runs ~2,500 tokens, a large documentation page ~25,000, and a research PDF ~125,000, which is why unbounded fetching without the max_content_tokens limit is the classic surprise line item on agent bills.
The long-context story stays flipped from surcharge to standard: Claude 4.6 and later models include the full 1M-token context window at standard pricing, with caching and batch discounts applying across it, so a 900K-token request bills at the same rate per token as a 9K one. You still pay for every token you send, and a full context window is real money on any model, but the punitive rate tier that older guides warn about is gone. The Managed Agents meter also got worked examples in the docs this cycle: a one-hour Opus 5 session consuming 50K input and 15K output tokens totals $0.705 including the runtime fee, dropping to $0.525 with caching, which is a useful sanity anchor for anyone modeling session-based economics. If you are building on the SDK rather than consuming Anthropic's packaged agents, our Claude Agent SDK deep dive traces the token flow end to end.
One genuinely free corner deserves emphasis because it reads like a typo and is not: code execution is entirely free when combined with web search or web fetch (the current tool versions), versus burning the 1,550 free monthly hours or paying $0.05 per container-hour standalone. An agent that runs code as part of a search-driven task pays nothing for the sandbox. Loopholes this clean rarely survive repricing cycles; architect for it while it lasts, but do not build a business case on it.
6. The 30% Tokenizer Effect
Here is the fact that quietly invalidates most per-token comparisons published this year: Claude 4.7 and later models, plus Claude Mythos Preview, use a newer tokenizer that produces approximately 30% more tokens for the same text, while Sonnet 4.6 and earlier use the previous tokenizer - Anthropic pricing docs. The list price per million tokens is real; the number of tokens in your million changed. A codebase that tokenized to 100K tokens under the old scheme tokenizes to roughly 130K under the new one, and you pay for 130K.
Run the worked example. A refactoring task that flowed 520K input and 52K output tokens through the old tokenizer becomes roughly 676K input and 68K output under the new one. On Opus 5 at $5/$25, that is $5.08 instead of $3.90: same task, same list price, about 30% higher effective cost. Now compare against deprecated Opus 4.1 at $15/$75 on the old tokenizer, where the same task cost $15.60. The honest generational statement is a 3.1x real-cost improvement, not the naive "3x cheaper and same tokens" a list-price table implies. The current generation is better and cheaper even after the adjustment; the point is calibration, not complaint.
The shift matters in three places. On the API, pad any budget model built on list prices by 30% for current-generation models. On subscriptions, the same hours of work burn more tokens from the shared pool, which is part of why Anthropic's 2026 limit increases were necessary to keep the subjective experience stable. And in competitor comparisons, "Claude charges $2 and vendor X charges $2.50 per MTok" is not a like-for-like statement unless you normalize for tokens per completed task, which is exactly the fight section 9 covers.
For teams that track internal metrics, the practical move is to re-baseline. Any dashboard built in 2025 that reports tokens per pull request or cost per completed task now mixes two tokenizer regimes, and quarter-over-quarter comparisons will show a mysterious 30% regression no engineering change explains. Recompute historical baselines where you can, or at minimum annotate the cutover on every chart, and when a vendor quotes tokens-per-task at you, ask which tokenizer generation produced the number. Independent analysis of the 4.7-era rollout measured the inflation at up to 35% on some inputs, so treat 30% as a floor estimate, not a ceiling - Duet.
7. The Usage-Limit Timeline: Boosts, Outages, and Resets
Usage limits are the most-searched Claude Code pricing topic, and the story now has more plot than most products' entire pricing history. The previous edition of this guide predicted that the +50% weekly boost would expire July 13 and revert. That prediction was wrong within days, and the way it was wrong is instructive: Anthropic has moved the deadline at least four times, and has now twice used limits as outage compensation, not just as marketing.
The verified sequence. Weekly caps arrived in August 2025 on top of the 5-hour rolling windows. On May 13, 2026, Anthropic boosted weekly limits 50%, explicitly time-boxed to July 13 and widely read as a defensive answer to Codex's token-efficiency advantage - Pasquale Pillitteri. After a week of degraded service in early July, Anthropic reset 5-hour and weekly limits for all users on July 10 - Startup Fortune. On July 13 the boost was extended through July 19 for Pro, Max, Team, and legacy seat-based Enterprise plans - Help Net Security. Then on July 18, alongside the Fable restructuring, the +50% weekly boost was extended through August 19, 2026 - Digital Applied. Cowork got its own parallel gift, a doubled 5-hour limit that expires August 5, 2026, literally the day this guide was re-verified - Just Being Resourceful. And on July 29, a 529 "Overloaded" outage took down the web app, API, and Claude Code for roughly 44 minutes at peak, the fourth capacity incident of 2026 after March 2, June 2, and June 18 - DeployFlow.
The July 10 reset deserves a second look, because it is the newest behavior in the system. Anthropic wiped the 5-hour and weekly meters for every user, with no stated cause, three days before the promotional boost was then scheduled to lapse - Startup Fortune. Whatever the internal reason, the external lesson is that limits now function as an operational pressure valve as much as a pricing mechanism: they flex upward for competition, and they flex again when infrastructure wobbles. For buyers this cuts both ways. It means degraded weeks tend to get compensated, and it means your effective monthly quota is not a fixed contractual quantity but a managed variable that Anthropic tunes in near-real time.
Three mechanics govern how the timeline feels in practice. The quota is one shared pool across Claude Code, chat, and Cowork, so limit math is household math. At a cap you either wait for the reset or consent to extra usage at API rates - Anthropic support. And the /status command shows your remaining allocation; making it a habit is the cheapest usage-management tool that exists.
What should you assume happens after August 19? The July prediction failed by assuming Anthropic would let a deadline bite, so this edition predicts the pattern instead of the date: limits have moved six times in four months, every move has been user-favorable, and two of the moves were compensation for capacity incidents. The base case is that boosts keep rolling while the competitive war with OpenAI runs hot, and the tail risk is the opposite: a capacity-driven tightening after the next outage cluster, because the July 29 incident shows demand is pressing against infrastructure - DeployFlow. The planning rule survives either way: budget on the permanent baseline, treat boosts as free upside, and rehearse the reversion now. If your team is tuned to boosted limits, cut agent parallelism or shift overflow to metered billing for one week and measure the delta; learning it costs $40 in a controlled experiment beats discovering it as a stalled sprint on August 20.
It also helps to translate the two leading vendors' cap systems into the same units, because they meter different things. Anthropic meters hours of model work through a shared token pool; OpenAI meters messages per rolling 5-hour window, per model tier (section 11 has the current numbers). Message metering is easier to reason about interactively but punishes long agentic runs, where one task fans out into many counted turns; token-pool metering does the reverse. Match the meter to your workload's shape and half the "whose limits are more generous" arguments dissolve: they are generous along different axes.
8. Subscription vs API: The Break-Even Math
The question underneath every Claude Code pricing search is really this one: subscription or pay-as-you-go? The most-cited usage analysis, published by Duet in April 2026 against the Opus 4.7 generation, put a full-time user (six to eight hours daily with heavy Opus usage) at roughly $800 to $1,500 per month in API-equivalent consumption, and documented one developer case of $15,000 in API charges over 8 months for work that would have cost about $800 on Max across the same period, a 93% swing - Duet. Those figures predate Opus 5 and the current limits, so treat them as shape rather than gospel, but the shape is unambiguous: for a human working full days in an agent, subscriptions crush the API on price, by an order of magnitude at the heavy tail.
So when does the API win? Three real scenarios. Intermittent use: a few agent-days a month on Sonnet 5 at intro pricing ($2/$10 through August 31) can land under $20, beating Pro, and drops further with batch and caching. Cheap-model pipelines: bulk transformations on Haiku 4.5 with Batch at 50% off and cache hits at 0.1x produce per-task costs no seat-based plan can express; a nightly job rewriting thousands of files has no business on a subscription. Production automation: anything headless, unattended, or at scale belongs on the API or on Managed Agents at $0.08 per session-hour, because subscription limits exist precisely to prevent that pattern - Anthropic pricing docs.
Make it concrete with three profiles. The weekend builder ships a side project two weekends a month: pay-as-you-go Sonnet 5 plausibly costs $8-15 in build months and $0 otherwise, while Pro costs $20 every month; the API wins. The full-time engineer lives in the agent daily: even the conservative end of Duet's full-time range is 8x the Max 5x price, so the subscription wins without a close call. The platform team runs agents in CI on every pull request: that is machine work, excluded from subscription terms by design, and it belongs on Haiku with batch and caching, where a whole fleet can cost less than one Max seat. Same product, three rational meters.
If you genuinely cannot tell which side of the line you are on, run a measurement month. Keep your current plan, enable extra usage with a hard budget cap, and let one normal month of work write the answer: the overflow line item is your personal API-equivalent price, measured on your actual repositories with your actual habits, at zero methodology cost. A developer whose overflow reads $0 is over-provisioned and can test a downgrade; one whose overflow reads $60 has just learned the Max upgrade pays for itself; a team whose overflow is dominated by scheduled jobs has found the workloads that belong on a real API key. Every consent prompt before API-rate billing is Anthropic handing you a free experiment - Anthropic support. Most buyers argue about meters in the abstract; the overflow ledger settles it with your own data.
The clean rule: humans on subscriptions, machines on the API. Human usage is bursty and capped by attention, which is what flat-rate pricing is designed around; automated pipelines are limited only by budget, which is what metered pricing is designed around. The hybrid that sophisticated teams converge on is a Max plan per heavy developer, an API key for CI and production agents, and subscription extra usage as the overflow valve rather than the primary meter. For the deeper machine-side economics across providers, our guide to the true cost of LLM inference in 2026 prices the same workloads across the whole market.
9. Token Efficiency, Benchmarks, and the Three Answers Problem
Every pricing table in this space, including ours, prices tools by monthly fee. The metric that actually determines value is tokens burned per completed task, and the best documented head-to-head remains Pasquale Pillitteri's testing: Codex consumed 3-4x fewer tokens than Claude Code on equivalent builds, including a Figma plugin at 1.5M tokens for Codex versus 6.2M for Claude Code, a React scheduler at 72.5K versus 234.7K, and an OAuth REST API at roughly 180K versus 650K - Pasquale Pillitteri. Two honesty flags on that data. It was measured against the GPT-5.3-Codex generation versus Opus 4.6/4.7, and neither the GPT-5.6 family nor Opus 5 has been re-tested by that source, so the ratio is directional, not current. And part of the gap is the tokenizer effect from section 6: current Claude models emit ~30% more tokens for identical text, so even identical behavior reads as less efficient in raw counts.
Why the gap exists is partly philosophy. Claude Code's agentic style is exploratory: it reads widely, re-verifies, and self-reviews, which is token-hungry by design; OpenAI has visibly tuned Codex for economy because economy is the axis where it can win against a capability leader. The consequence ripples through every plan comparison: a Codex Plus user at $20 hits walls less often than a Claude Pro user at $20 doing the same work, even when the plans look symmetric on paper, and that asymmetry is precisely what Anthropic's limit boosts have been compensating for since May - Pasquale Pillitteri.
Efficiency is not quality, though, and here 2026 delivered a mess nobody's comparison table admits: the industry's flagship coding benchmark now returns three different winners depending on who you ask. On the SWE-bench Pro aggregate leaderboard, which is entirely self-reported by vendors, Claude Fable 5 leads at 80.0%, with Opus 4.8 at 69.2% and GPT-5.6 Sol at 64.6% - LLM Stats. On Scale's independently evaluated public set, the same benchmark family tops out at 61.5% (Muse Spark 1.1), with GPT-5.4 at 59.1% and the best listed Claude, Opus 4.6, at 51.9% - Scale. Fifteen to twenty points of divergence between self-reported and independently measured scores, on the same benchmark name, is the single most useful fact in this section: a vendor quoting "SWE-bench Pro" without naming the evaluation set is quoting marketing.
The terminal-native benchmark tells a tighter story. On Terminal-Bench 2.1, the current live version, Claude Code with Fable 5 leads at 83.8%, with Codex plus GPT-5.5 at 83.1%, Claude Code with Opus 4.8 at 78.9%, Codex with GPT-5.6 Terra at 78.4%, and Claude Code with Sonnet 5 at 74.6% - Terminal-Bench. Neither Opus 5 nor GPT-5.6 Sol had a listed run at re-verification time, which is its own lesson: leaderboards lag launches by weeks, and the gap between "model exists" and "model has independent numbers" is exactly when marketing does its best work. On the older SWE-bench Verified, Fable 5 posts 95.0% and the top five slots are all Anthropic models, which mostly tells you that benchmark is saturated as a differentiator - LLM Stats.
Given that mess, the cheapest reliable benchmark is the one you run yourself. Assemble a fixed task portfolio: five real items from your own backlog (a bug fix, a refactor, a greenfield feature, a test-writing job, a documentation pass), run the identical five through each candidate's free tier or trial, and score completion, cleanup time, and where each tool gave up. Five tasks fit inside Kiro's 50 free credits, one Gemini CLI afternoon, and a Codex Plus window, and the resulting comparison is weighted by your codebase, your conventions, and your definition of done, which no leaderboard can be. The free tiers exist to sell subscriptions; used deliberately, they are the best procurement research available at $0.
The synthesis that survives all three answer sets: for hard, long-horizon tasks in messy codebases, Claude's capability premium tends to pay for its token appetite; for high-volume routine work, Codex's efficiency compounds into real savings. But the only number that generalizes to your budget is cost per completed task measured on your own repository, and the four-variable math (rate x tokens-per-task x attempts x cleanup time) flips winners between workloads. We maintain a continuously updated cross-model view in our best LLM for AI agents ranking, and the model-by-model head-to-head in GPT-5.6 versus Claude Opus 5 for agents.
10. What Your Subscription Buys Beyond the CLI
A 2025-vintage guide treats Claude Code as a terminal program and prices it accordingly. In August 2026 the same subscription buys a materially larger surface, which changes what "is $20 worth it" even means, and which makes single-surface price comparisons subtly unfair in Anthropic's favor.
The headline expansion is Claude Cowork, the agentic workspace for non-coding work, which reached web and mobile on July 7, 2026, Max subscribers first, with remote sessions that keep working after you close the laptop - TechCrunch. The same TechCrunch analysis of 1.2 million anonymized sessions across 600,000 organizations found business-process operations at 33.4% of usage versus 8.7% for software development: the "coding agent" subscription is now majority-used for work around the work. The quota implication from section 2 bears repeating: remote sessions decouple consumption from hours at a desk, so budget them like cloud instances, not like an app you open and close. Our Cowork desktop deep dive covers what those sessions can actually do.
For engineering teams, the most consequential addition is agent teams, still experimental and enabled via the CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 setting: multiple coordinated Claude Code instances share a task list and message each other directly, with one lead session coordinating - agent teams docs. The pricing fine print is the part to internalize: each teammate is a separate instance with its own context window, and token usage scales with active teammates. The docs themselves recommend 3-5 teammates and note that subagents (which summarize results back instead of running independently) are the cheaper pattern when workers do not need to talk to each other. On Max, a team is a powerful trade; on Pro, it is a fast route to the weekly cap. Treat team size as a literal budget dial.
The architecture difference is easier to see than to describe, and Anthropic's own documentation diagram draws the cost model precisely: subagents funnel results back to one context, while teammates each carry a full context of their own.
The diagram is the budgeting model to keep in your head: every surface drains one pool, Fable draws from the same pool where included, and the overflow valve is opt-in API-rate billing. This is also where cross-vendor comparisons get slippery. Cursor's $20 buys an IDE. Copilot's $10 buys IDE assistance plus credits. Claude's $20 buys a coding agent plus a general agentic workspace plus design and science tools on one quota. Whether the bundle is worth more depends entirely on whether you use the non-coding surfaces; a developer who lives only in the terminal is paying for breadth they never touch, which is a genuine argument for Codex or Cursor at the same price. And for teams whose real goal is agents running whole business workflows rather than code specifically, that logic extends past Anthropic's bundle entirely: platforms like O-mega provision an AI agent workforce that operates browsers, tools, and processes end to end, a different altitude of automation than a per-developer coding subscription.
11. Alternatives Head-to-Head: Codex, Cursor, Copilot, Gemini CLI, Kiro, Devin
The paid-alternative landscape repriced again since July, and one vendor moved its documentation mid-quarter: OpenAI's Codex pricing docs now live on the ChatGPT developer-docs domain rather than the old developers.openai.com path - Codex pricing. Here is the verified August 2026 head-to-head, followed by what actually changed per tool.
| Tool | Entry price | Mid tier | Top tier | Billing model |
|---|---|---|---|---|
| Claude Code | $20/mo Pro ($17 annual) | $100/mo Max 5x | Max 20x (checkout-priced) | Shared quota + opt-in API-rate extra usage - source |
| OpenAI Codex | $8/mo ChatGPT Go | $20/mo Plus | $100-200/mo Pro 5x/20x | Per-model 5-hour message windows + credit overage - source |
| Cursor | $0 Hobby | $20/mo Pro | Pro+ (3x) and Ultra (20x), checkout-priced | Included usage + frontier-model credits - source |
| GitHub Copilot | $10/mo Pro ($15 credits) | $39/mo Pro+ ($70 credits) | $100/mo Max ($200 credits) | AI credits at $0.01 each; completions unlimited - source |
| Gemini CLI | $0 (1,000 req/day) | Code Assist Standard (1,500/day) | Enterprise (2,000/day) | Free daily quota, then per-seat plans - source |
| Kiro (AWS) | $0 (50 credits) | $20/mo Pro (1,000 credits) | $200/mo Power (10,000) | Credit packs, $0.04 overage, no rollover - source |
| Devin | $0 Free | $20/mo Pro | $200/mo Max; Teams $80 + $40/seat | Refreshing daily/weekly allowance - source |
OpenAI Codex remains the sharpest price attack, and its metering got more granular with the GPT-5.6 generation. Codex now runs the GPT-5.6 family: Sol (flagship), Terra (balanced), and Luna (fast and cheap), with GPT-5.3-Codex-Spark as a Pro-only research preview - Codex pricing. Plus subscribers get per-model 5-hour windows of roughly 10-100 Sol messages, 25-200 Terra, and 250-2,000 Luna, with Pro multiplying those 5x or 20x at $100 and $200, and overage billed in credits at 125/750 credits per million input/output tokens for Sol, 50/300 for Terra, and 5/30 for Luna - Codex pricing. The $8 Go tier is still the cheapest paid on-ramp to a frontier coding agent on the market, and the credit schedule quietly rewards the same discipline Anthropic's cache multipliers do: cached input bills at a tenth of fresh input (12.5 credits per million on Sol versus 125), so a Codex workflow with stable context stretches its window several times further than a naive one - Codex pricing. On the API side, the July 30 cuts moved the family to Sol $5/$30, Terra $2/$12, and Luna at $0.20/$1.20 per MTok (Terra down 20%, Luna down 80%, a 25x spread across one family), with cached input discounted 90%, Batch and Flex tiers at 0.5x, and one sharp edge Anthropic dropped and OpenAI kept: beyond 272K input tokens, the entire request reprices at 2x input and 1.5x output, not just the overflow - eesel. Our GPT-5.6 benchmark and pricing guide covers the model family in depth.
Cursor now mirrors Anthropic's price opacity. The public page verifies Hobby free, Pro $20, and Teams from $40 per user with a Standard/Premium split where Premium carries 5x Standard's agent limits, but the Pro+ (3x) and Ultra (20x) tiers no longer display dollar prices on the public page; they surface in-product - Cursor pricing. Two of the seven vendors in this table now hide their top-tier prices behind checkout, which is a genuinely new pattern this quarter and consistent with a market where usage economics are being repriced faster than marketing pages update. Cursor's year has been eventful well beyond pricing, and we covered the ownership question separately in our analysis of the SpaceX-Cursor deal.
GitHub Copilot has fully settled into its credits era: 1 AI credit equals $0.01, plans span Pro at $10 with $15 in monthly credits, Pro+ at $39 with $70, and Max at $100 with $200, organizations pay $19 (Business) and $39 (Enterprise) per seat, and code completions plus next-edit suggestions remain unlimited on every paid plan - GitHub Copilot plans. The cross-vendor detail worth knowing: Copilot passes through Claude Sonnet 5 at its $2/MTok promotional input pricing through August 31, 2026, the same intro window as Anthropic's own API, so the Sonnet 5 discount expires everywhere at once - GitHub docs.
Gemini CLI is still the free-tier champion: 1,000 model requests per day and 60 per minute at $0 with a personal Google account, 250 Flash-only requests daily on an unpaid API key, and 1,500/2,000 daily on paid Code Assist Standard/Enterprise - Google. Note that Google's docs no longer promise a specific model version, only requests "across the Gemini model family as determined by Gemini CLI," which means the free tier's capability floor is explicitly at Google's discretion. Kiro, AWS's spec-driven agent IDE, runs credit packs from Free (50 credits, open-weight models plus Sonnet 4.5) through Pro $20, Pro+ $40, Pro Max $100, and Power $200, with $0.04 per credit overage and no rollover, though separately purchased add-on packs keep a 12-month life and GovCloud deployments run about 20% higher - Kiro. Devin spans Free, Pro $20, Max $200, and Teams at $80 plus $40 per developer seat, now metering through a usage allowance that refreshes daily and weekly rather than a visible work-unit ledger, with Devin Desktop included across tiers - Devin.
The pattern across the whole table is convergence on three price points ($0, $20, $100-200) and divergence on meters: token pools, message windows, dollar credits, opaque allowances. When every vendor charges $20 for entry, the meter is the product, and sections 7 through 9 are the tools for reading yours.
12. The Open-Source Path: OpenCode, Cline, Aider, and Antigravity
The previous edition of this guide compared only paid products, and that is no longer how the alternatives conversation starts. The "Claude Code alternatives" search results of August 2026 lead with open-source agents, for a structural reason worth stating plainly: when the intelligence lives in a rented model anyway, the agent harness around it is a commodity, and commodities go to zero. An open-source harness plus your own API key (or even your existing subscription) decouples the tool from the vendor's quota politics entirely: no weekly limits, no boost expiry dates, no plan restructuring, just the raw token meter from section 5.
OpenCode is the flagship of this wave: an open-source terminal/IDE/desktop coding agent with over 160,000 GitHub stars, 900 contributors, and a claimed 7.5 million monthly developers, connecting to 75+ providers including Anthropic, and able to ride existing subscriptions rather than API keys alone - OpenCode. It also ships a curated model endpoint called Zen, a hand-benchmarked model set for coding agents, which is the tell for where this category monetizes: the harness stays free, and the money moves to model curation and routing. Cline takes the IDE-native lane: Apache 2.0 licensed, running as a VS Code extension, CLI, and JetBrains early access, with 8 million+ installs and 65.7K GitHub stars, and a bring-your-own-key model spanning every major provider plus local models - Cline. Aider, the original terminal pair-programmer with 6.8 million installs, still works with almost any LLM including local ones, though its homepage still recommends model pairings from two generations back, a maintenance signal worth weighing before you standardize on it - Aider.
The cost profile of this path is honest but not free. You trade the subscription's flat rate for the API meter, which means the section 8 logic applies in reverse: an open-source agent driving Opus 5 at $5/$25 full-time costs API full-time money (hundreds to low thousands monthly), while the same agent on Sonnet 5 intro pricing or Haiku 4.5 with caching can undercut every subscription in this guide for bursty use. You also inherit the operational work the vendors do for you: no shared-pool accounting, no managed sandboxes, no support line when a release breaks your workflow. The honest framing is that open-source agents are the API path with a better steering wheel, and they reward exactly the teams that section 8 already routed to the API. Our full ranked review of the field lives in the top 10 open-source AI coders guide.
Google complicated the taxonomy by shipping something that is neither open source nor conventionally priced. Antigravity 2.0, announced around Google I/O 2026, is an agentic development platform (a command center for parallel local agents, plus a CLI, IDE, and Python SDK) offered at no charge and wired to current Gemini models, with Gemini 3.6 Flash integration landing in July - Google Antigravity. A free, first-party agent platform from the vendor with the most spare inference capacity is a price-floor statement aimed at everyone in section 11's table: it does to the $20 entry tier what Gemini CLI's 1,000 free requests did to hobbyist budgets. Whether its capability matches the paid leaders is a separate question we examine in our complete Antigravity 2.0 guide, but on pure pricing, the bottom of this market now belongs to Google.
13. Enterprise Buying: Seats, CCUs, and Data Residency
Enterprise Claude buying in 2026 is a three-lever negotiation: seat mix, consumption mechanics, and compliance multipliers. The levers compound across hundreds of seats, which is why modeling them before the negotiation beats discovering them on the first invoice.
The seat mix decision is Standard versus Premium at $20 versus $100 per seat annually, with Premium carrying 5x the usage plus permanent Fable 5 inclusion since July 20 (section 4) - Claude pricing. The naive read is "same price per unit of usage"; the operational read is that Premium seats belong to engineers whose agent usage is a daily primary tool and who benefit from frontier-model escalation, while Standard fits reviewers, occasional users, and Cowork-first staff. Price a concrete organization: 100 seats at a 25% Premium mix is $4,000 per month on annual billing, versus $10,000 all-Premium or $2,000 all-Standard with predictable cap pain in the heaviest quartile. The mix is a $96K annual swing on identical headcount, which is why usage forecasting beats feature-grid reading. Above Team, Enterprise at $20 per seat plus API-rate usage converts the plan from a quota bundle into a metered utility with a seat-license floor, better for organizations that want per-team chargeback data.
Procurement routing has its own economics. Buying through AWS Marketplace or Microsoft Foundry bills in Claude Consumption Units at $0.01 per CCU (100 CCU = $1.00 of usage at standard rates, after discounts), metered hourly, invoiced monthly, and arrears-only with no prepaid credits on both marketplaces - Anthropic pricing docs. The practical value is budget plumbing: CCU billing draws Claude spend down existing cloud committed-spend agreements, often the difference between a new budget line (slow) and an existing one (fast). Compliance adds the last multiplier: US-only inference (inference_geo: "us") costs 1.1x on every token category for Claude 4.6+ models, including cache reads and writes, on the Claude API, Claude Platform on AWS, and Foundry's US Data Zone deployments - Anthropic pricing docs. Ten percent for jurisdictional certainty is cheap against a failed vendor review, but on a $1M annual commitment it is a real $100K line that belongs in the model.
One governance capability worth knowing exists before the negotiation, because it interacts with cost: enterprise admins can pin an availableModels allowlist and an organization default model in managed settings, and Claude Code resolves aliases like opus to the newest version the allowlist permits - model configuration docs. That is the mechanism for enforcing "no Fable without approval" or "Sonnet only in CI," but it has a sharp edge: a stale allowlist silently pins the whole organization to an older generation at the same price as the newer one, and an automatic-fallback target that the allowlist excludes turns a would-be fallback into a refusal. Model governance and cost governance are the same file; someone should own it.
For the committee that asks about counterparty risk: Anthropic raised a $30 billion Series G on February 12, 2026 at a $380 billion post-money valuation led by GIC and Coatue, reported a ~$14 billion revenue run rate, and disclosed that enterprise is over half of Claude Code revenue, business subscriptions quadrupled since the start of the year, and customers spending $1M+ annually grew from 12 to over 500 in two years - Anthropic. Vendor solvency is not the risk to model. The genuinely open enterprise questions are the ones this guide has already priced: tokenizer-inflated consumption forecasts (section 6), the post-August-19 limit baseline (section 7), capacity-outage exposure for production workflows (July 29 was the fourth incident of the year), and agent-team token multipliers (section 10).
14. The Cost-Optimization Playbook
Every mechanism here follows from the verified rate card in section 5; none of it is folklore. Applied together, these routinely cut effective Claude spend by half or more, whether your meter is dollars (API) or quota (subscription).
Start with the two structural discounts. Prompt caching is the highest-leverage single change for agentic work: cache reads at 0.1x mean stable system prompts and repo context cost a tenth on every turn after the first, and the official break-even is brutal in caching's favor, one read repays a 5-minute write and two reads repay a 1-hour write - Anthropic pricing docs. The direction of error matters: over-writing throwaway context to the 1-hour tier wastes a little; under-caching stable context wastes ten times more. When in doubt, cache. Batch API at 50% off everything is the second: any workload that tolerates async (test generation, migration sweeps, evaluation runs) should never run at interactive rates.
Then the model-selection plays, in rough order of impact:
- Ride the Sonnet 5 window: $2/$10 becomes $3/$15 on September 1, 2026, so heavy Sonnet workloads run before month-end carry a built-in 33% discount - Anthropic pricing docs
- Route by difficulty: default to Sonnet 5, escalate to Opus 5 for architecture and gnarly debugging, drop to Haiku 4.5 for mechanical edits, and use
opusplanto get Opus-quality plans with Sonnet-cost execution - Pair code execution with web search or fetch: the sandbox becomes entirely free instead of consuming the 1,550 monthly hours - Anthropic pricing docs
- Watch
/statusand effort levels: check remaining quota before big runs, and match reasoning effort to task difficulty rather than defaulting to maximum - Size agent teams deliberately: each experimental teammate is a separate instance whose tokens scale with the team (section 10), so parallelism is a budget dial, not a free speedup
The common thread is that defaults are expensive: default model, default effort, default single-shot execution, and default caching-off all leave money on the table. Model routing as an engineering discipline is worth a guide of its own, and ours on AI model routing to cut agent costs covers the decision logic; the broader toolbox lives in how to cut LLM costs. Reusable skills also compound here, since a well-built skill replaces exploratory token burn with a known-good procedure; the top 100 Claude Code skills ranking is the shortcut catalog.
One warning against over-optimizing: the failure mode of aggressive cost tuning is spending Opus-grade engineer time babysitting Haiku-grade output. Optimize machine-scale spend (API pipelines, batch jobs, agent fleets) ruthlessly, and let humans on subscriptions mostly just work. Duet's full-time figures from section 8 set the frame: when the subscription already captures $800-1,500 of monthly API-equivalent value for $100-200, shaving pennies off the human's workflow is value destruction dressed as diligence - Duet.
15. Decision Framework: Which Option Fits You
Strip away the tables and the decision compresses to a handful of forks. If you are an individual developer using an agent daily, Claude Pro at $20 ($17 annual) is the default answer, upgrading to Max the moment limits interrupt you more than once a week, with the added pull that Max now bundles Fable 5 at up to half your pool and defaults to Opus 5. If price floor beats capability ceiling, Codex via ChatGPT Go at $8 or Plus at $20 delivers the market's most granular metering and a documented efficiency edge (with the caveat that the current model generations have not been independently re-measured against each other). If your usage is bursty or pipeline-shaped, skip subscriptions: Sonnet 5 intro pricing and Haiku 4.5 with caching and batch make the API the rational meter, with an open-source harness from section 12 as the steering wheel. If you are spending nothing, Gemini CLI's 1,000 daily requests and Google's free Antigravity platform now bracket the $0 tier from both ends.
Whatever you choose, put a review date on the decision, and make it soon: this market's numbers now expire on named dates. August 19 is the current limit-boost expiry, and the base case is another extension rather than a reversion, but only the baseline is safe to budget on. August 31 ends Sonnet 5's intro pricing simultaneously on the Anthropic API and inside GitHub Copilot, so front-load heavy Sonnet workloads this month. September 19 expires the $100 Fable transition credit for Pro and Team Standard subscribers who have not spent it. And the pattern behind all three dates is the durable lesson of this refresh: the July edition of this guide bet on a deadline holding, and the market taught it that deadlines here are negotiating positions. Budget on baselines, treat promotions as upside, and re-quote quarterly the way you would any cloud line item.
The bigger arc is worth naming as you budget. A third of Claude sessions are business-process operations, enterprise is over half of Claude Code's revenue, and the agent surface now spans terminal, web, mobile, and autonomous background sessions. Coding agents stopped being a developer-tools line item and became general work infrastructure, and the buyers who win are the ones who model work completed per dollar, not tokens per dollar or seats per dollar. That is the number every section of this guide has been steering you toward.
This guide was researched and written by Yuma Heymans ( @yumahey), founder and CEO of O-mega and co-founder of HeroHunt.ai, who re-verified every price in this article against the live pages during the first week of August because last month's edition taught him what happens when you trust a vendor's deadline.
This guide reflects pricing and plans verified during the first week of August 2026. AI pricing changes weekly: three of the numbers above carry expiry dates this quarter, so verify current rates on the official pricing pages linked throughout before committing budget.