title: "Top 10 Browser Use Agents in 2026: Ranked on Real Data" slug: "top-10-browser-use-agents-full-review-2026" excerpt: "The August 2026 review of browser use agents: verified pricing, live benchmark data, the Atlas shutdown, and the first peer-reviewed security evidence." author: "O-mega Team" type: "blog" category: "Artificial Intelligence" tags: ["Browser Agents", "AI Automation", "Web Tools", "Benchmarks", "Cybersecurity"] date: "2026-08-05"
The August 2026 re-review of browser use agents: verified pricing, measured benchmarks, one scheduled product death this week, and the honest calls no vendor listicle will make.
The number five product in our previous ranking dies this week. OpenAI is shutting down the ChatGPT Atlas browser on August 9, 2026, less than a year after its October 2025 launch, folding agentic browsing into ChatGPT itself and the new ChatGPT Work desktop experience that launched July 9, 2026 - BigGo. That makes OpenAI two-for-two on killing its own browser-agent flagships: Operator lasted seven months, Atlas lasted under ten. If your evaluation shortlist still contains either, you are shopping in a graveyard, and most of the roundups ranking this category have not noticed.
This is the third full edition of this review. The first ran in December 2025, the second in July 2026, and we rewrite rather than patch because the category keeps invalidating its own reviews: since the July edition alone, Atlas got a death date, the desktop-automation benchmark that anchored our capability analysis was replaced by a dramatically harder successor, a hyperscaler entered the browser-infrastructure tier at effectively bundled pricing, and the first peer-reviewed security evidence on agentic browsers arrived from the University of Washington. Everything below reflects the state of August 2026: every price pulled from an official pricing page this run, every benchmark number read from the live leaderboard this run, and every correction to our own previous editions stated in the open. A ranking that never admits its own staleness is not a review, it is an advertisement.
The subject is unchanged: browser use agents, AI systems that operate a real web browser the way a person does. They click, type, scroll, log in, fill forms, extract data, and complete multi-step workflows on websites that were never designed for machines. The market still splits into three tiers (consumer agentic browsers, developer browser infrastructure, and enterprise outcome agents), and the fastest way to waste money remains buying from the wrong tier. This edition ranks across tiers with a transparent scoring model, grounds every score in a named data point, and then goes deep on the products, the benchmarks, the infrastructure data, and the security problem that now has academic confirmation.
This guide was researched and written by Yuma Heymans ( @yumahey), founder of O-mega, who runs production browser-agent fleets and has personally hit most of the failure modes this review describes, including the one where a vendor kills the product mid-quarter.
Contents
- The Week the Number Five Pick Died
- What Our July Edition Got Wrong: A Revision Log
- The Three Tiers of Browser Agents in 2026
- Benchmarks: Online-Mind2Web and the OSWorld 2.0 Reset
- The Infrastructure Shootout: Measured Latency and Reliability
- Browser Use: Open Source Holds the Benchmark Crown
- Browserbase: From Browser-Hours to Full Agent Platform
- TinyFish: Enterprise Outcomes, Now With Self-Serve Pricing
- Steel.dev: Still the Fastest Session Starts on Record
- Cloudflare Browser Run: The Hyperscaler Enters Tier 2
- Manus: Independent, Growing, and Still Under a Cloud
- O-mega: A Multi-Agent Workforce That Browses
- Airtop: No-Code Automation With Deployed Agents
- Kernel: The Marketing-vs-Benchmark Gap, Continued
- Anchor Browser: Deterministic Replay and Verified-Bot Credentials
- The Consumer Wave After Atlas
- Security: The First Peer-Reviewed Evidence Arrives
- Plan Once, Replay as Code
- The Graveyard: Now Two OpenAI Browser Agents Deep
- New Entrants and Near Misses
- Decision Framework: Picking by Job-to-be-Done
- FAQ
- Final Take
The August 2026 Ranking at a Glance
Four criteria, weights summing to 100 percent, every cell carrying the verified data point that produced its score. Capability (30 percent) measures whether the agent completes real tasks on the live web, anchored to independent benchmark results where they exist. Reliability and scale (25 percent) covers measured session behavior and concurrency ceilings. Pricing value (25 percent) rewards transparent, verifiable pricing and low entry cost. Security and governance (20 percent) covers compliance posture, credential handling, and vendor risk, a criterion the Atlas shutdown just re-taught the whole market to weight.
| # | Tool | Category | Capability (30%) | Reliability & Scale (25%) | Pricing Value (25%) | Security & Governance (20%) | Final |
|---|---|---|---|---|---|---|---|
| 1 | Browser Use | Open-source framework + cloud | 10 - 97.0% Online-Mind2Web, still the record | 8 - cloud tops leaderboard, 107.9k GitHub stars | 9 - free OSS, cloud from $29/mo, $0.02/browser-hr | 7 - self-host escape hatch, younger cloud | 8.7 |
| 2 | Browserbase | Browser infrastructure | 8 - Stagehand + Director, now bundles agent runs | 9 - 99.96% measured reliability, 100 concurrent at $99 | 8 - $20 entry now includes search/fetch/gateway | 9 - HIPAA BAA, DPA, SSO on Scale | 8.5 |
| 3 | TinyFish | Enterprise outcome agent | 9 - 90.0% Online-Mind2Web, 4th worldwide | 8 - enterprise fleets, free Search/Fetch APIs | 8 - NEW self-serve: $0.015/credit, 500 free, failed runs free | 9 - enterprise governance, no-charge failures | 8.5 |
| 4 | Steel.dev | Browser infrastructure | 7 - infra only, bring your own agent brain | 10 - 894ms starts, 100% reliability, best measured | 9 - $0 + usage entry, $30 starter credits | 7 - open source, HIPAA BAA only at $250 tier | 8.3 |
| 5 | Cloudflare Browser Run | Browser infrastructure | 6 - infra only, no public benchmark score yet | 8 - 120 concurrent, global network, unmeasured independently | 10 - bundled into Workers Free and Paid, no separate SKU | 8 - Web Bot Auth signed agents, robots.txt compliance | 7.9 |
| 6 | Manus | Autonomous general agent | 9 - full task autonomy, Manus 1.6, $125M run rate | 7 - 20 concurrent on every paid plan | 8 - free 300 daily credits, $20 entry | 6 - blocked acquisition, jurisdiction overhang | 7.7 |
| 7 | O-mega | Multi-agent workforce | 8 - orchestrated agents run browser sessions + tools | 7 - parallel agents, per-session isolation | 7 - free tier, credit-based paid plans | 8 - managed credentials, approval gates | 7.5 |
| 8 | Airtop | No-code agent platform | 7 - natural-language automation + deployed Mark agents | 7 - 3 to 100 sessions by tier | 8 - free 1,000 + 10,000 bonus credits, $26 entry | 7 - SOC 2 Type 2 on Enterprise, no-training pledge | 7.3 |
| 9 | Kernel | Browser infrastructure | 7 - infra only, unikernel Chromium | 8 - 100% measured reliability, 1,519ms starts | 7 - usage-based, thin public rate card this cycle | 7 - young GA product | 7.3 |
| 10 | Anchor Browser | Browser infrastructure | 7 - Web Action Cache, 80x token claim, 1Password auth | 6 - 8,001ms starts, 97.34% measured reliability | 7 - $50 entry, 0.1 credits/task, steep jump to $500 | 8 - SOC2/ISO27001 on Growth, Cloudflare verified-bot | 7.0 |
Two notes on reading this honestly. First, the top four are separated by four tenths of a point, and the right pick depends on whether you want a full agent (Browser Use), the deepest managed platform (Browserbase), delivered outcomes (TinyFish), or the fastest raw session layer (Steel). Second, a conspicuous absence: no OpenAI product is ranked. Atlas held fifth place in July; a product with a published shutdown date four days out cannot hold a slot in a buying guide, and its replacement inside ChatGPT Work is too new to have independent data. Section 16 and Section 19 cover that story in full, because the story itself is the single most important vendor-risk lesson in this market.
1. The Week the Number Five Pick Died
Reviews of fast-moving categories usually age gradually. This one aged overnight. OpenAI announced that ChatGPT Atlas shuts down August 9, 2026, with the company saying it will use what it learned to "deliver a more powerful browser experience within ChatGPT" - BigGo. Atlas launched in October 2025 as OpenAI's Chromium browser with agent mode built in. It was the most downloaded, most discussed, most reviewed product in this entire category. It survived less than ten months.
The strategic logic matters more than the obituary. OpenAI concluded that as chat-native search and task execution improved, users did not need a standalone AI browser at all: the work a dedicated browser was built for now happens inside the assistant. The successor surface is ChatGPT Work, launched July 9, 2026, which consolidates agentic capabilities and the Codex programming agent into a unified entry point in the ChatGPT desktop app. For what that consolidation means for office workflows specifically, our guide to OpenAI workspace agents covers the Work surface in depth.
For the record, and because dead products vanish from the web with surprising speed, here is the moment being unwound this week: OpenAI's own launch presentation of Atlas from October 2025, preserved as the historical document it has now become. Ten months separate this video from the shutdown notice.
For buyers, the pattern is now unmistakable, because it has happened twice. Operator launched January 23, 2025 and was shut down August 31, 2025, succeeded by ChatGPT Agent - Wikipedia. Atlas launched October 2025 and dies August 9, 2026. Two flagship browser agents, two shutdowns, an average lifespan under nine months. This is not incompetence: it is what it looks like when a platform company treats browser agents as experiments in search of a form factor rather than as products. But your automation stack does not care about the distinction. Anything you built on Atlas needs a migration path this week, and anything you build on the next OpenAI browsing surface should be built with an exit plan from day one. Our archived pieces on what Operator was and what Operator-era access cost now document the first funeral; this edition documents the second.
There is one genuinely useful data point Atlas leaves behind, and we put it to work in Section 4: before its death, Atlas Agent Mode scored 71.0 percent on the independent Online-Mind2Web leaderboard, eighth place, a full 26 points behind the open-source leader - Steel Leaderboard. The consumer agent everybody actually used was never close to the capability frontier. That gap between mindshare and measured performance is the quiet theme of this entire review.
2. What Our July Edition Got Wrong: A Revision Log
No competing roundup publishes a revision log, so here is ours, and this cycle it cuts against us. The July 2026 edition was a genuine full rewrite with verified pricing and independent benchmarks, and it still shipped with defects a reader deserved better than. Naming them is the fastest way to explain what this edition fixes.
First, it missed facts that were already public on its publish date. Cloudflare had rebranded and relaunched its browser platform as Browser Run on April 15, 2026, a structural event for the infrastructure tier, and our July edition did not mention it - Cloudflare. The University of Washington's agentic-browser security study went out through UW News on June 30, 2026, and our security section cited only vendor blogs - UW News. OSWorld 2.0, which resets the entire desktop-automation capability picture, was live before we published, and we cited the old version's numbers - OSWorld 2.0. All three are corrected in full below.
Second, it understated freshness decay in its own facts. We printed a GitHub star count of roughly 97,000 for Browser Use; the repository stands at 107,928 stars as checked this run - GitHub. We scored TinyFish down for having "custom enterprise pricing only," which was true in July and is false now: TinyFish publishes a full self-serve rate card - TinyFish. We printed Manus plan names that no longer match the current tier structure - NoCode MBA. Each of these is rescored in this edition's table.
Third, and least visible to readers but most damaging: several of the July edition's internal links were broken, pointing at articles that had moved to a different product domain or never existed under the slugs we used. Every internal link in this edition was verified to resolve on this site before publish. It is a small discipline, and skipping it is exactly the kind of quiet sloppiness that separates a maintained reference from a generated one. Finally, on scoring: we keep the transparent weighted model, but this edition grounds every score cell in a named, verified data point, because a decimal without a data point behind it is false precision wearing a spreadsheet costume.
3. The Three Tiers of Browser Agents in 2026
The market stratification we introduced in earlier editions has only hardened, and it remains the single most useful lens for not wasting money. Three tiers, three buyers, three failure modes: rank a consumer copilot against raw infrastructure on the same criteria and you get rankings that mislead everyone. The tiers are converging from opposite directions (consumer products adding governance, infrastructure vendors adding no-code layers, outcome vendors wrapping both), but the buyer distinction still holds.
What changed in this cycle is the pressure inside tier 2. When browser infrastructure was sold only by venture-backed specialists, metered browser-hours were the natural business model. A hyperscaler bundling the same capability into an existing compute plan attacks that model at its root, and that is exactly what Cloudflare did in April. The tier taxonomy below is unchanged; the economics inside the middle tier are not.
Tier 1, consumer agentic browsers, sell to individuals. The product is a browser or extension with an agent acting inside your logged-in session, using your cookies and your identity. Economics are free or subscription, concurrency is one task at a time, and the defining risk is prompt injection against your personal accounts, which now has peer-reviewed documentation (Section 17). This tier also just demonstrated the second defining risk: vendor mortality, since its most prominent product dies this week.
Tier 2, developer browser infrastructure, sells browser sessions to engineers building agents. Browserbase, Steel, Cloudflare Browser Run, Anchor, Hyperbrowser, and Browser Use Cloud live here, measured in browser-hours, concurrency, cold-start milliseconds, and anti-bot pass rates. Nobody in this tier supplies the intelligence by default (Browser Use is the exception: the framework is the brain, the cloud is the body). The defining risks are reliability at scale and cost predictability, which is why Section 5's measured data matters more than feature lists.
Tier 3, enterprise outcome agents, sell completed work. TinyFish sells monitored web operations, Amazon Nova Act sells compiled workflow execution as an AWS service, Manus sells general task completion, and O-mega sells an orchestrated workforce of agents where browsing is one capability among several. Buyers here evaluate outcomes per dollar, auditability, and governance. Ranking tier 3 on infrastructure criteria (or vice versa) is meaningless, which is why our master table carries a Category column and every entry names its tier first.
4. Benchmarks: Online-Mind2Web and the OSWorld 2.0 Reset
Benchmark data is what separates this review from the vendor listicles that currently crowd this query, so the numbers get a full section and their caveats get printed alongside them. Two benchmarks matter in August 2026, and they now tell opposite stories: live-web browsing is close to solved for well-specified tasks, while long-horizon computer work just got formally reset to roughly one-in-five reliability.
The live-web standard remains Online-Mind2Web: 300 realistic tasks across 136 real websites, maintained by the OSU NLP Group with UC Berkeley collaborators, scored on actual task completion with three published judging modes (third-party human raters, the WebJudge LLM judge, and a fine-tuned WebJudge-7B) - GitHub. The independent leaderboard, hosted by Steel and last updated June 29, 2026, is the closest thing this industry has to ground truth - Steel Leaderboard.
The top of the table is stable since March: Browser Use Cloud at 97.0 percent, the highest score ever recorded, then GPT-5.4 Native Computer Use at 93.0, a hosted-browser pipeline with Claude Opus 4.6 at 90.53, TinyFish at 90.0, and ByteDance's UI-TARS-2 at 88.2. The new information is at the bottom of our chart: ChatGPT Atlas Agent Mode at 71.0 percent, eighth place, a score the leaderboard notes was OpenAI-reported in the GPT-5.4 announcement with underlying run data not public. Sit with that spread for a moment. The open-source framework anyone can run beat the most famous consumer browser agent on earth by 26 points on the same benchmark, and that consumer agent is now being discontinued. Mindshare and capability are not the same axis, and in this category they have rarely even been correlated. For how these evaluation methodologies work and where they break, our guide to AI agent evals and benchmarks goes deeper than this section can.
The judging caveat still applies to every number above: Online-Mind2Web scores shift depending on whether completion is graded by human raters or by WebJudge, and cross-vendor comparisons are only clean when the judging methodology matches - GitHub. A vendor quoting a score without naming the judge is quoting marketing.
The desktop story is where this edition corrects itself hardest. Our July edition cited OSWorld version 1 numbers, where the frontier had climbed above 85 percent, and framed desktop automation as trailing but closing. OSWorld 2.0 resets that picture entirely: 108 long-horizon workflows with a median human completion time around 1.6 hours, 69.6 percent of tasks taking skilled humans over an hour, and agent runs averaging 318 tool calls against roughly 30 in the prior version - OSWorld 2.0. On this harder target, the current frontier is Claude Opus 4.8 at 20.6 percent binary completion (54.8 percent partial credit, 244K output tokens per run), Claude Opus 4.7 at 18.2 percent, and GPT-5.5 at roughly 14 percent while spending only about 39K output tokens, by far the most token-efficient frontier entry.
Read the two charts together and the strategic picture is sharper than any vendor will state it: browsing is near-solved, computering is not. If your workflow lives in web apps with well-specified tasks, you can buy near-human reliability today, and the rest of this review is about choosing whom to buy it from. If your workflow spans desktop software, multi-hour horizons, and hidden state, the honest frontier number is one-in-five, and you should design for supervision, checkpoints, and partial credit. That boundary, not any vendor's roadmap slide, is where automation budgets should draw the line this year. We map the desktop side of that line in our deep guide to agentic computer use and in our computer-use benchmark guide, both of which this section supersedes on the specific numbers.
5. The Infrastructure Shootout: Measured Latency and Reliability
The browser infrastructure tier still has exactly one large independent-ish performance study, and this edition adds the caveat that time has been quietly attaching to it: it is now nine months old, and it predates the tier's newest structural entrant. Steel's browserbench study, published November 7, 2025, ran 5,000 sessions per provider from AWS us-east-1 against five platforms and published every number, including the ones flattering its competitors - Steel.
Yes, Steel benchmarked its own market, and you should discount accordingly. The mitigations are unchanged: the methodology is public, the sample is large, and no vendor has published a credible rebuttal in the nine months since. But two honesty notes now belong in any citation of it. First, nine months is a long time in this tier, and any provider could have materially improved (or regressed) since. Second, the study does not include Cloudflare Browser Run, which did not exist in its current form when the tests ran, so the tier's newest entrant is invisible in the tier's only measured comparison. Until someone reruns the methodology, this is still the best available evidence, cited here with its expiry date showing.
The measured spread remains enormous: 894 milliseconds for Steel (p95 1,090ms) against 8,001 milliseconds for Anchor (p95 11,561ms), a 9x difference in how long your agent waits before doing anything. Reliability separated less dramatically but meaningfully: Steel, Kernel, and Hyperbrowser completed 100 percent of sessions, Browserbase 99.96 percent, Anchor 97.34 percent - Steel. A 97.34 percent session success rate sounds fine until 10,000 daily sessions hand you 266 failures before your agent logic even gets a chance to fail on its own.
How to apply the numbers has not changed. Long sessions amortize start latency into irrelevance, so reliability dominates. Thousands of short tasks make start latency your biggest cost line, and the 9x spread is the whole ballgame. Authenticated workflows where a mid-session failure leaves a half-submitted form make the reliability column worth more than every other spec combined. What has changed is the question buyers should now append to every infrastructure evaluation: how does the vendor's number compare with running the same workload on a bundled hyperscaler platform at effectively zero marginal cost? Section 10 takes that question seriously.
6. Browser Use: Open Source Holds the Benchmark Crown
Tier: open-source framework plus cloud. Final score: 8.7, first place. Browser Use keeps the top slot for the second consecutive edition, and the case has only strengthened. The repository now stands at 107,928 GitHub stars, checked against the GitHub API this run - GitHub. Its cloud still holds the highest Online-Mind2Web score ever recorded at 97.0 percent, achieved March 25, 2026 and unbeaten through the leaderboard's June 29 update - Steel Leaderboard.
The framework's approach is the one the category converged on: expose the browser to a language model as structured DOM state plus screenshots, let the model plan and act in a loop, and handle the translation between "click the submit button" and actual browser commands. The record-setting run is worth understanding accurately, because it says something about where capability now comes from. The team gave a coding agent CLI access to their evaluation platform with instructions to run in a loop, iterating over agent configurations across 20 cycles per session, effectively building a search tree over possible agents, and letting the agent execute Python for HTML parsing on edge cases - Browser Use. The best browser agent was not hand-designed; it was searched for by another agent. If that methodology sounds familiar, it is the same pattern reshaping agent frameworks generally, a landscape we map in our comparison of the top agent frameworks.
The economics are the second reason it stays first. The framework is free and open source, so the floor cost of a serious browser agent in 2026 is an API key and your own compute (and choosing that API key well matters: see our August 2026 ranking of the best LLMs for agents). The hosted cloud is published-and-cheap, with the pricing page itself stating verification as of July 27, 2026 - Browser Use Pricing:
| Plan | Monthly Cost | Concurrent Sessions | Notes |
|---|---|---|---|
| Free | $0 | 3 | 10 agent tasks/month |
| Dev | $29 | 25 | credits match plan price |
| Business | $299 | 200 | credits match plan price |
| Scaleup | $999 | 500 | credits match plan price |
Pay-as-you-go top-ups run $5 to $500 without a subscription, and browser sessions meter at $0.02 per browser-hour, an order of magnitude below the browser-hour rates elsewhere in the tier - Browser Use Pricing.
Limitations. You still assemble the body: credential management, proxy strategy, observability, and guardrails are your job on the open-source path, and the cloud remains younger and less battle-hardened than Browserbase's platform. The rapid release cadence lands breaking changes more often than enterprise change management enjoys. And the 97.0 benchmark number was achieved with the vendor's own best-searched configuration; your prompts, your sites, and your edge cases will land lower. Best for: any developer starting a browser-agent project in 2026, and any team that wants benchmark-leading capability without per-seat platform pricing.
7. Browserbase: From Browser-Hours to Full Agent Platform
Tier: developer browser infrastructure. Final score: 8.5, second place. Browserbase keeps its podium slot, but the product it defends it with has quietly changed shape, and the change is strategically revealing. The pricing page no longer sells browser-hours alone: every plan now bundles agent runs, search calls, fetch calls, proxy bandwidth, a free Model Gateway, and a runtime - Browserbase Pricing. That is not a pricing tweak. That is an infrastructure vendor concluding that raw browser-hours are becoming a commodity and climbing the stack toward the agent itself before someone bundles it from below. Given what Cloudflare did in April (Section 10), the timing looks less like ambition and more like self-defense, executed early and well.
The core platform remains the most complete managed browser layer in the tier: headless Chromium fleets with persistent sessions, stealth and captcha handling, live session replay for debugging, and the Stagehand SDK that turns natural-language intents into browser actions and can emit deterministic code from successful runs, with the Director product packaging the same power for non-developers. In the tier's only independent measurement, Browserbase posted 99.96 percent reliability and 1,677ms average session starts, third fastest of five - Steel. The verified August 2026 ladder - Browserbase Pricing:
| Plan | Monthly Cost | Browser Hours | Concurrency | Now Bundled |
|---|---|---|---|---|
| Free | $0 | 1 | 3 | 3 agent runs, 1k search + 1k fetch calls, $5 gateway tokens |
| Developer | $20 | 100, then $0.12/hr | 25 | 15 agent runs, 1GB proxies then $12/GB |
| Startup | $99 | 500, then $0.10/hr | 100 | 50 agent runs, 10k fetch calls, 5GB proxies then $10/GB |
| Scale | Custom | usage-based | 250+ | HIPAA BAA, DPA, SSO, verified agents |
Overage rates on the bundled services are published too (search at $7 per 1,000 calls, fetch from $0.50 to $7 per 1,000 depending on proxy and extraction options), which matters because bundle pricing is where surprise bills usually hide - Browserbase Pricing.
Limitations. Browserbase still does not supply the intelligence: Stagehand helps, but you bring the agent logic, and pure capability now lives elsewhere in this table. It was edged on raw session-start speed by Steel in the independent test, so the premium buys maturity, ecosystem, and debugging depth rather than any single measured number. For scraping-shaped workloads specifically, extraction-first stacks like the one we profiled in Firecrawl, the scraper made for the AI web can be simpler than a general browser fleet. Best for: production teams that want the most proven managed browser platform, now with enough bundled agent tooling to defer several build-vs-buy decisions at once.
8. TinyFish: Enterprise Outcomes, Now With Self-Serve Pricing
Tier: enterprise outcome agent. Final score: 8.5, third place, up from fourth. TinyFish gets the single largest rescore of this edition, and it earned it by fixing the exact thing we docked it for. In July we scored its pricing a 6 with the note "custom enterprise pricing only, no self-serve tier, no published rate card." That is no longer true. TinyFish now publishes a full self-serve ladder - TinyFish Pricing:
| Plan | Cost | Credits | Overage |
|---|---|---|---|
| Pay as You Go | $0.015/credit | 500 free to start, no card | n/a |
| Starter | $15/mo | 1,650/mo | $0.014/credit |
| Pro | $150/mo | 16,500/mo | $0.012/credit |
| Enterprise | Custom | Custom | Custom |
Two details in that rate card deserve the attention they were designed to attract. The Search and Fetch APIs are free on every plan (0 credits per request), and failed runs are not charged - TinyFish Pricing. The second one is quietly radical for this category. Every metered browser platform in this review charges you for the session whether or not the task succeeded; TinyFish pricing completed outcomes rather than attempts is the tier-3 philosophy expressed in billing, and it puts genuine pressure on every competitor whose failed runs still show up on the invoice.
Capability was never the question. TinyFish holds 90.0 percent on the independent Online-Mind2Web leaderboard, fourth worldwide and the highest of any pure enterprise vendor, a score from February 2026 that still stands in the June 29 update - Steel Leaderboard. The product is composed of four APIs (Search, Fetch, Agent, Browser) that let an enterprise assemble anything from a single monitored extraction to a fleet of long-running web operations, measured on completed workflows rather than browser-hours consumed.
Limitations. The self-serve tier is new, and the free-failed-runs promise has not yet been stress-tested publicly at scale, so treat the economics as verified-but-young. The benchmark score, excellent as it is, trails the open-source leader by seven points, a gap TinyFish would argue it buys back through governance and anti-bot robustness on hostile production sites. And enterprise deployments still involve compliance and IT collaboration, so fleet-scale time-to-value is measured in weeks even if your first credit is spent in minutes. Best for: teams that want delivered web operations with enterprise governance, and, as of this cycle, anyone who wants to try that proposition for the cost of a lunch instead of a procurement cycle.
9. Steel.dev: Still the Fastest Session Starts on Record
Tier: developer browser infrastructure. Final score: 8.3, fourth place. Steel drops one rank without getting one measurable thing worse, which says more about the tier than about Steel: TinyFish's pricing fix leapfrogged it, and the measured numbers that justify Steel's position are aging. Its browserbench study still shows the best figures ever recorded in the tier: 894ms average session starts (p95 1,090ms) and 100 percent reliability across 5,000 runs, with control-plane operations around 229ms - Steel. It also still operates the independent Online-Mind2Web leaderboard the whole industry (this review included) treats as ground truth - Steel Leaderboard.
The product philosophy is unchanged and remains its charm: an open-source browser API that gives an agent a full virtual browser (mouse, keyboard, pixels, DOM) through clean API calls, with stealth, proxies, and session management underneath. Self-host the core or use the hosted cloud. It stays deliberately unopinionated about the agent brain, pairing equally well with Browser Use, Stagehand-style tooling, or hand-rolled loops against any frontier model, and that neutrality is precisely what makes its leaderboard credible. Verified pricing - Steel Pricing:
| Plan | Monthly Cost | Notes |
|---|---|---|
| Launch | $0 + usage | $30 one-time starter credits |
| Scale | $250 + usage | $100 monthly credits, HIPAA-ready BAA, enterprise SSO |
| Enterprise | Custom | 1,000+ concurrent sessions |
Limitations. Steel is the rawest of the top four: no no-code layer, no bundled agent, a smaller team than Browserbase, and community-plus-core-team support below the $250 tier. The conflict-of-interest asterisk on its self-published benchmark stands, mitigated by open methodology, and this edition adds the sharper caveat: those winning numbers are nine months old, and the study that produced them predates Cloudflare's entry into the tier. Steel's strongest move for the next edition would be rerunning browserbench with Cloudflare in the field; a vendor confident enough to do that twice would be making the rest of the tier's marketing departments visibly nervous. Best for: engineering teams optimizing for raw session speed, measured reliability, and open-source flexibility, especially high-volume short-task workloads where 894ms versus 8 seconds compounds into real money.
10. Cloudflare Browser Run: The Hyperscaler Enters Tier 2
Tier: developer browser infrastructure. Final score: 7.9, fifth place, new entry. The biggest structural event in the infrastructure tier since our first edition happened on April 15, 2026, and our July edition inexcusably missed it: Cloudflare relaunched its Browser Rendering product as Browser Run, a global headless-Chrome platform for AI agents, included on both Workers Free and Workers Paid plans with no separate SKU - Cloudflare.
Read that pricing sentence again, because it is the strategic earthquake. Every specialist in this tier sells browser sessions as the product; Cloudflare ships them as a feature of a compute plan developers already pay for (or do not pay for at all). The concurrency ceiling quadrupled to 120 concurrent browsers, quick actions handle 10 requests per second, and control comes through every interface the ecosystem actually uses: Puppeteer, Playwright, a directly exposed Chrome DevTools Protocol endpoint, MCP clients (Claude Desktop, Cursor, OpenCode), and WebMCP for agent-discoverable site tools - Cloudflare. The observability story arrived complete rather than promised: Live View streams real-time browser activity with DOM and network detail, session recordings capture interactions as JSON for replay, and a human-in-the-loop feature lets a person take over an active session when automation hits an authentication wall.
One more feature deserves separate attention because it points at where the whole crawling economy is going. Browser Run's /crawl endpoint discovers and scrapes entire sites into HTML, Markdown, or structured JSON, and it does so as a cryptographically signed agent under the Web Bot Auth standard, respecting robots.txt and AI Crawl Control directives - Cloudflare. The company that fronts a huge share of the web's traffic is simultaneously selling the browser agents and defining the compliance regime they operate under. That dual position is either reassuring or alarming depending on your politics, but it is unambiguously leverage no specialist vendor can match.
Limitations. The scoring is honest about what is unproven: Browser Run has no public benchmark score, no independent latency or reliability measurement (browserbench predates it), and its capability score reflects that it supplies infrastructure only, with no agent brain and a younger stealth/anti-bot track record than the specialists whose entire business depends on it. Its concurrency ceiling of 120 also sits below what Browserbase, Steel, and Anchor sell at their top tiers. The bet you make with Cloudflare is that bundled-and-good-enough beats specialist-and-measured; for workloads already on Workers, that bet is close to free to test. Best for: teams already building on Cloudflare, cost-sensitive workloads where a bundled 120-concurrent browser fleet at no separate price beats any metered rate card, and anyone whose crawling needs to be verifiably compliant.
11. Manus: Independent, Growing, and Still Under a Cloud
Tier: autonomous general agent. Final score: 7.7, sixth place. Manus remains the purest expression of the "digital employee" pattern: hand it a high-level goal and it plans, browses, writes and executes code in a sandbox, and delivers finished artifacts with minimal supervision. The corporate saga that defined its 2026 is unchanged since our last edition and worth restating as the risk disclosure it is: Chinese regulators blocked Meta's agreed acquisition of parent company Butterfly Effect on April 27, 2026, and Meta formally cut ties on June 15, 2026, leaving Manus independent under Butterfly Effect Pte. Ltd., incorporated in the Cayman Islands with Singapore operations - Wikipedia. The business itself kept compounding through the turbulence: from a $90 million annualized run rate in August 2025 to $125 million by December 2025, with Manus 1.6 shipping December 15, 2025 - Wikipedia.
What did change since July is the plan structure, which our previous edition printed under now-outdated names. The current ladder, per third-party pricing documentation updated June 26, 2026 - NoCode MBA:
| Plan | Monthly Cost | Credits | Concurrent Tasks |
|---|---|---|---|
| Free | $0 | 300 daily refresh | 5 |
| Standard | $20 | 4,000/mo | 20 |
| Customizable | $40 | 8,000/mo | 20 |
| Extended | $200 | 40,000/mo | 20 |
Two corrections to our July table are embedded there: every paid plan now gets 20 concurrent tasks (concurrency is no longer the top-tier carrot), and the 300 daily refresh credits apply across plans, not just free. Annual billing saves roughly 17 percent - NoCode MBA. One transparency note: Manus's own pricing page resists server-side verification, so tier names here rest on the cited third-party documentation; check manus.im before purchasing.
Limitations. Autonomy still cuts both ways: Manus takes initiative, which means it sometimes takes routes you did not anticipate, and supervising a highly autonomous agent on sensitive accounts requires the guardrail discipline of Section 17. Credit burn on complex tasks is real, and power users routinely outrun mid-tier allowances. The governance score carries the overhang: a blocked acquisition, a jurisdiction migration, and a revenue engine that will keep attracting regulatory attention constitute the exact vendor-risk profile this category just punished twice elsewhere. Best for: individuals and teams that want maximum end-to-end task autonomy available today and can tolerate vendor turbulence in exchange.
12. O-mega: A Multi-Agent Workforce That Browses
Tier: multi-agent workforce platform. Final score: 7.5, seventh place. Full disclosure first, as in every edition: O-mega is our platform, scored with the same rubric as everyone else, which lands it mid-table again. O-mega is a platform where you run an AI workforce: multiple persistent agents with distinct roles that plan work, browse the live web in isolated sessions, execute code, use connected tools, and coordinate under one orchestration layer.
The browser capability sits inside that larger structure rather than being the whole product. An O-mega agent spins up isolated browser sessions for research, monitoring, outreach, and operational tasks, with managed credentials, per-agent identities, and human approval gates on sensitive actions. Where a single-agent product gives you one very capable operator, O-mega's bet is that real workloads decompose: one agent gathers, another verifies, another acts, with the orchestration layer handling handoffs, scheduling, and memory. The security architecture is the part of that bet this year's evidence most directly supports: the UW findings in Section 17 all concern agents acting inside a user's personal logged-in identity, and scoped per-agent credentials in isolated sessions are the structural alternative to that exposure.
Pricing is credit-based and tiered: a free tier to start, with paid plans that scale monthly credit allowances, model quality, and concurrency, listed at o-mega.ai/plans. Credits meter actual agent work (browsing, generation, tool calls) rather than seats, which fits workloads where a small team supervises a larger fleet of agents.
Limitations. O-mega is not the right buy if you need raw metered browser infrastructure (tier 2 does that cheaper and lower-level) or a single free consumer copilot (tier 1 does that at zero). Multi-agent setups carry inherent coordination overhead: decomposing work across agents takes more upfront thought than pointing one agent at one task, and the platform is younger than the infrastructure incumbents ranked above it. It also publishes no score on the public benchmark this review anchors capability to, the same deduction we apply to every unbenchmarked vendor. Best for: operators and small teams that want several specialized agents working in parallel, with browsing as one capability among many, managed under one roof.
13. Airtop: No-Code Automation With Deployed Agents
Tier: no-code agent platform. Final score: 7.3, eighth place. Airtop keeps its identity: cloud browsers controlled in natural language, aimed at operators, marketers, and researchers who will never write Playwright, with the packaged Mark agent running recurring workflows (continuous lead research, standing monitoring) rather than one-off sessions. The free tier remains one of the most generous on-ramps in the category, and the verified August 2026 ladder shows the credit allowances have grown substantially at the paid tiers - Airtop Pricing:
| Plan | Monthly Cost | Credits | Sessions / Agents |
|---|---|---|---|
| Free | $0 | 1,000 + 10,000 one-time bonus | 3 sessions, 1 deployed agent |
| Starter | $26 | from 30,000 | 3 sessions, 10 agents |
| Professional | $170 | from 225,000 | 30 sessions, 30 agents |
| Enterprise | $502 | from 775,000 | 100 sessions, unlimited agents, SOC 2 Type 2 |
Two governance details differentiate Airtop at its price point: the platform states plainly that it never uses your data for AI training, and it maintains HIPAA alongside SOC 2 Type 2 compliance - Airtop Pricing. For a tier where credential handover to a cloud platform is the whole operating model, a written no-training pledge is worth more than most feature checkboxes.
Limitations. Natural-language automation still inherits natural-language ambiguity: vague instructions produce confident wrong behavior, and non-technical users hit a ceiling when a workflow needs exception handling plain English cannot express. Airtop publishes no score on the public benchmarks this review leans on, so capability assessment rests on product surface rather than measured task completion, which caps its capability score at 7. Best for: non-technical teams that want recurring web workflows running this week, with a real free tier and the lowest paid entry point in this ranking.
14. Kernel: The Marketing-vs-Benchmark Gap, Continued
Tier: developer browser infrastructure. Final score: 7.3, ninth place, down from seventh. Kernel's ranking slips this cycle for a reason we want to state precisely, because it is a methodology point as much as a product point: we could verify less about Kernel this run than last run. Its public site rendered almost no product or pricing detail to our verification pass this cycle, which leaves its independent benchmark data as the primary verifiable surface, and costs it pricing-transparency points in a review that scores what can be checked, not what is remembered.
What the independent data shows remains genuinely strong: in browserbench's 5,000-run measurement, Kernel posted 100 percent reliability, tied for best in the tier, with 1,519ms average session starts (p95 1,752ms), second fastest of the five providers tested - Steel. The instructive tension we documented in July also stands as this tier's cautionary tale: Kernel's marketing has advertised cold starts under 150 milliseconds on its unikernel architecture, while the end-to-end measured session start was ten times that. Both numbers can be simultaneously true (a unikernel can boot in 150ms while networking, profile loading, and API overhead take the rest), but only one of them describes what your agent experiences. The gap between the microbenchmark a vendor quotes and the end-to-end number a customer feels is endemic in this tier, and it is exactly why this review weights measured data over claimed data everywhere.
Limitations. Kernel is infrastructure only (the racecar, not the driver), the youngest GA product among the measured five, and, this cycle, the hardest to verify: thin public pricing, minimal self-serve documentation surfaced, and a marketing site that resists inspection. None of that means the product regressed; it means a buyer doing diligence in August 2026 has less public evidence to stand on than for any other ranked infrastructure vendor, and rankings built on evidence must say so. Best for: teams that want isolation-first browser infrastructure with measured top-tier reliability and are comfortable doing their evaluation through direct vendor contact rather than public documentation.
15. Anchor Browser: Deterministic Replay and Verified-Bot Credentials
Tier: developer browser infrastructure. Final score: 7.0, tenth place. Anchor remains the clearest strategic differentiation in the infrastructure tier attached to the weakest independent performance numbers, and this cycle adds real substance on the strategy side. The core bet is unchanged: the Web Action Cache plans a workflow once with full LLM reasoning, then maintains it as deterministic code that replays on every subsequent run, with Anchor claiming 80x fewer tokens than runtime AI, 12x faster execution than browser agents, and 23x lower error rates than the best-performing browser agents - Anchor Browser. Those are vendor claims without independent verification, printed here as such, but the architecture they describe is the subject of Section 18 because it is bigger than one vendor.
What is new and verified this cycle: Anchor now lists an official 1Password partnership for handling the authentication lifecycle across cloud computer-use agents, and its Team tier includes Cloudflare Verified Browser Agents, plugging Anchor sessions into the same Web Bot Auth verified-crawling regime Cloudflare is building around Browser Run - Anchor Browser. For Anchor's core audience (authenticated workflows on internal tools, SSO, MFA), first-class credential infrastructure and verified-bot identity are exactly the right things to build. One logistics note: Anchor's separate pricing page is gone; pricing now lives on the homepage, and the numbers are confirmed unchanged with new specificity - Anchor Browser:
| Plan | Monthly Cost | Credits | Concurrency |
|---|---|---|---|
| Free | $0 | 5 | 5 |
| Starter | $50 | 50 | 25 |
| Team | $500 | 500 | 50, Cloudflare verified agents |
| Growth | $2,000 | 2,000 | 200, SOC2 + ISO27001 |
Completed tasks meter at 0.1 credits each, with overage at $1.00 per credit across tiers - Anchor Browser.
Limitations. The independent data is still the headline deduction: slowest measured session starts (8,001ms average, p95 11,561ms) and lowest reliability (97.34 percent) of the five providers in browserbench, with the caveat that the study is nine months old - Steel. The ladder still jumps steeply from $50 to $500 with nothing between, and replay economics erode on sites that change frequently, forcing re-planning. For teams evaluating this segment, we maintain dedicated comparisons of Anchor Browser alternatives and stealth-focused alternatives to Anchor. Best for: teams running stable, repeated, authenticated workflows where replay economics, 1Password-managed credentials, and verified-bot identity outweigh raw session performance.
16. The Consumer Wave After Atlas
Seven months ago this tier defined public perception of browser agents; this week its most famous product dies. The consumer wave is not receding, but it is changing shape in a way the Atlas shutdown makes legible: the standalone AI browser is losing to the assistant that contains a browser, and the assistant is increasingly a cloud agent that does not need your machine on at all.
The survivors first. Perplexity Comet completed its journey from $200-per-month exclusive to free product: the browser is free to download, reaching iOS on March 18, 2026 after starting life behind the Max subscription, with Pro and Max upsells from $20 per month - MacRumors. Claude for Chrome is Anthropic's wager that the extension beats the browser: in beta on all paid Claude plans, with pre-approved action lists per site and mandatory confirmation before irreversible actions like purchases - Claude. (Our July edition listed per-plan model assignments for the extension; the current official page no longer specifies models per tier, so that claim is retired rather than repeated.) Gemini in Chrome continues Google's strategy of folding agentic browsing into the browser 3 billion people already have, and it is one of the seven products the UW security study examined in Section 17.
The genuinely new entrant points at the next form factor. Gemini Spark, which surfaced in the August 3, 2026 news cycle, is a 24/7 cloud-based agent that can operate desktop Chrome using the user's logged-in accounts and saved passwords, maintain state for several days, chain browser automation with phone calls and backend APIs, and hand control back to the user for payments - AI Agent Store. Read against the Atlas shutdown, the direction is clear: the browser agent is dissolving into the assistant on one side and into the persistent cloud worker on the other, and the desktop browser is becoming a surface either kind of agent reaches into, rather than a product either needs to be. Opera Neon and Atlassian's Dia (the knowledge-worker browser Atlassian is building from its Browser Company acquisition - Atlassian) continue as subscription browsers, but the energy has visibly moved.
What the wave means for buyers of the other tiers is sharpened by this cycle's news. Expectation setting still cuts both ways: executives experience agents personally and then import consumer mental models into procurement. Identity risk now has academic documentation rather than only vendor blog posts, and it attaches precisely to this tier's defining feature of acting as you in your session. And platform commitment turns out to be revocable: the deepest-pocketed vendor in the wave just retired its entry, which is the strongest argument this review can make for keeping your automation stack portable across surfaces.
17. Security: The First Peer-Reviewed Evidence Arrives
Every previous edition of this review had to build its security section from vendor disclosures and security-team blog posts. This edition gets to cite something better: the first peer-reviewed-venue study of agentic browser security, from David Kohlbrenner and Franziska Roesner's teams at the University of Washington's Paul G. Allen School, presented April 26, 2026 at the Agents in the Wild Workshop in Rio de Janeiro and published through UW News on June 30 - UW News.
The findings deserve exact statement. The team examined seven popular agentic browsers. Four of the seven create ways for attackers to bypass the same-origin policy, the foundational browser rule that stops one website from reading another's data. The researchers demonstrated a successful proof-of-concept against ChatGPT Atlas in which one website could steal information from another embedded site, and found conditions enabling similar attacks in Chrome with Gemini, Claude for Chrome, and Perplexity Comet. Firefox's AI Mode presented the least risk while offering the fewest capabilities, the trade-off in its purest form. Kohlbrenner's summary quote is the bluntest sentence any credible source has produced about this category: "Browser agents aren't ready for the public... you should not trust that these systems are ready to truly protect your information" - UW News.
The academic result lands on top of the practitioner evidence that was already decisive. Brave's security team found the canonical attack on July 25, 2025 and disclosed it that August: instructions hidden behind a Reddit spoiler tag hijacked Comet's summarize function, walked the agent into the victim's account details and Gmail, and exfiltrated a one-time password by posting it as a Reddit reply, all under the user's own authenticated privileges - Brave. OpenAI, before retiring Atlas, called prompt injection "one of the most significant risks we actively defend against" and shipped adversarially trained defenses discovered through automated red-teaming, while the UK's National Cyber Security Centre has warned that prompt injection against generative AI may never be fully mitigated and that organizations should plan for risk reduction rather than elimination - CyberScoop. The structural cause has not moved: a language model cannot reliably distinguish instructions from data when both arrive as text in one context window, so every page an agent reads is a potential command channel.
The practical guardrails follow from the structure, and the tier distinction does most of the work. For personal use of consumer agents: run agent mode logged out where possible, keep approval gates on for purchases, credentials, and email, and treat "browse my inbox" as the highest-risk instruction you can give. For production fleets, the calculus inverts: infrastructure-tier and workforce-tier agents run in isolated sessions with scoped, purpose-built credentials rather than your personal identity, so a successful injection compromises one sandboxed workflow rather than your life; this is the pattern tier 2 and tier 3 platforms increasingly ship by default, and the one we build around at O-mega. For the full defense-in-depth playbook (input isolation, least-privilege credentials, output filtering, approval tiers), our dedicated guide to prompt injection defense for AI agents goes considerably deeper than this section can. The honest bottom line for August 2026: the academics have now said in peer review what the attackers demonstrated in proof-of-concept, and the vendors saying it out loud remain more trustworthy than the ones whose security pages are silent.
18. Plan Once, Replay as Code
Underneath the product churn, the architectural trend we flagged in July has kept compounding, and it will shape this category's pricing more than any funding round: deterministic replay. An LLM plans and executes a workflow once, expensively, with full reasoning; the successful run is captured as executable code that replays deterministically on every subsequent execution at code prices instead of token prices.
Three vendors in this review converge on the pattern from different angles. Anchor's Web Action Cache is the most explicit, with its claim of 80x fewer tokens on replayed workflows - Anchor Browser. Browserbase's Stagehand generates deterministic automation code from natural-language runs, with Director packaging the idea for non-developers - Browserbase Pricing. Amazon Nova Act, generally available in us-east-1 on its custom Nova 2 Lite model, has enterprises defining workflows in natural language plus Python and running them at serious scale: Hertz reports 5x faster shipping velocity on automation projects, and customer Sola runs mission-critical processes hundreds of thousands of times per month - AWS. Different marketing, same insight: most business web automation is not novel exploration, it is the same fifty workflows run ten thousand times, and paying frontier-model prices for run 9,999 is waste.
The first-principles economics have not changed since we first wrote them, but the frontier pricing data keeps making them sharper. An LLM-driven browser action costs tokens on every step of every run forever; a replayed script costs compute rounding error. As run count grows, amortized cost collapses toward zero and the LLM's role shifts from executor to compiler: the model writes and repairs the automation rather than performing it. The same logic drives the model-selection layer, where routing cheap models to routine steps and frontier models to planning produces the cost curves we quantified in our guide to AI model routing. And the OSWorld 2.0 token data from Section 4 gives the compiler framing new urgency: when a single frontier-model desktop run consumes 244K output tokens, nobody can afford exploration pricing for repetition workloads.
For buyers, the operative question remains: what happens on run two? If the answer is "the same tokens get spent again," you are buying exploration pricing for repetition work. The counterweight is brittleness: replayed code inherits traditional automation's fragility when sites change, so the winning architectures pair cheap replay with automatic LLM-powered repair on failure. Vendors that close that loop (detect breakage, re-plan once, resume replaying) will structurally undercut everyone still charging per-token per-run, and TinyFish's failed-runs-free billing shows a second way to attack the same waste from the pricing side.
19. The Graveyard: Now Two OpenAI Browser Agents Deep
Most roundups only add products; honest ones also subtract, and this category keeps supplying material. The actuarial table gains its most prominent row this week, and the pattern behind it has held with uncomfortable consistency across three editions of this review.
The confirmed casualties and absorptions. OpenAI Operator: launched January 23, 2025, shut down August 31, 2025, seven months from flagship launch to sunset - Wikipedia. ChatGPT Atlas: launched October 2025, shuts down August 9, 2026, with agentic browsing folded into ChatGPT and the ChatGPT Work desktop experience - BigGo. The Browser Company: absorbed into Atlassian, with Dia repositioned as the knowledge worker's browser - Atlassian. Project Mariner: Google's standalone browser-agent experiment, discontinued and folded into Gemini's browsing surface, as we documented in the July edition. MultiOn, one of the earliest venture-backed autonomous browsing agents, is gone as an independent product. And the Meta-Manus acquisition: agreed, blocked by Chinese regulators on April 27, 2026, formally abandoned June 15, 2026 - Wikipedia.
The pattern is unchanged and now has a larger sample size: nothing on this list died from lack of capability. Everything died from strategic redundancy. Operator, Atlas, and Mariner were standalone experiments made obsolete the moment their parent companies decided the agent belongs inside the assistant, not beside it; OpenAI even said so in nearly those words when it retired Atlas to build "a more powerful browser experience within ChatGPT" - BigGo. The buying lesson compounds with each funeral: a big-tech browser agent is a feature of a corporate strategy that can pivot inside a budget cycle, while an independent vendor is a company whose survival depends on the product you are buying. Neither is safe; they fail differently; your migration plan should match. The three diligence questions from our July edition survive unedited because they keep being the right ones: does my workflow export (scripts, recorded workflows, data)? Is there an open-source escape hatch (the reason Browser Use and Steel score well on governance despite their youth)? And if this product disappeared in ninety days, what would the rebuild cost? For three products this review has covered across its editions, that last question stopped being hypothetical within a year of us asking it.
20. New Entrants and Near Misses
Ten slots force choices, and the bar for entry keeps rising: in 2025 a browser agent made lists by existing, in 2026 it needs a public benchmark score, an independent measurement, or verified pricing, ideally all three. This section names who missed the cut and exactly why, because a review that silently omits products is hiding its own judgment calls.
Amazon Nova Act remains the closest call and the most likely next entrant. It is generally available as an AWS service in us-east-1, built on the custom Nova 2 Lite model, with workflows defined in natural language plus Python, deployed through Bedrock AgentCore with CloudWatch monitoring and IAM auth, and named customers at meaningful scale: Hertz (5x shipping velocity on automation), PGA TOUR, 1Password, and Sola running processes hundreds of thousands of times per month - AWS. What keeps it out this cycle is the same evidence gap as last cycle: no score on the public leaderboard this review anchors capability to, and a workflow-compilation shape that makes head-to-head comparison with general browsing agents partially apples-to-oranges.
Hyperbrowser stays a near miss on measured grounds: in the tier's independent benchmark it posted 100 percent reliability, a genuine bright spot, but 3,657ms average session starts (p95 5,338ms), fourth of five - Steel. Its pricing page did not render details to our verification pass this run, so we drop rather than repeat the rate card we printed in July; the reliability number keeps it on the watchlist. Skyvern remains the most interesting open-source no-code entrant, but it still publishes no Online-Mind2Web score, and this review no longer ranks on the retired WebVoyager benchmark that top agents saturated. And a procedural note on scope: a July press announcement for a prompt-to-scraper agent product crossed our research net this cycle, but its source link went dead before we could verify it, which under this review's rules means it does not get named. A claim we cannot re-open is a claim we do not print.
The meta-point stands from July and strengthens with each cycle: this is what a maturing market looks like. Products are one leaderboard submission away from displacing the bottom of our table, and the displacement pressure now comes from two directions at once: specialists climbing the evidence ladder, and hyperscalers like Cloudflare entering with bundled economics that reset what the bottom of the market costs.
21. Decision Framework: Picking by Job-to-be-Done
Rankings compress; decisions need the job. The same ten products reorder completely depending on what you are hiring a browser agent to do, so here is one recommendation and one runner-up for each of the four jobs that cover nearly every real inquiry we see. The tier taxonomy from Section 3 does most of the work: name your job, and the tier (then the pick) follows.
Job 1: personal browsing copilot. This is the recommendation the Atlas shutdown forces us to rewrite, and the honest version has a caveat attached. Pick: Perplexity Comet, free to download with paid tiers from $20 per month, the strongest zero-cost answer now that Atlas is gone - MacRumors. Runner-up: Claude for Chrome on any paid Claude plan, whose per-site permission model is the most conservative default in the wave - Claude. The caveat that did not exist in July: both products appear in the UW study's findings on same-origin exposure, so whichever you pick, run it with approval gates on and read Section 17 first. There is no "safe" pick in this tier yet; there are only products whose vendors are honest about that.
Job 2: scraping and data extraction at scale. You want thousands of short sessions, fast starts, predictable cost per page. Pick: Steel, whose measured 894ms starts and 100 percent reliability are exactly the specs this job stresses, from $0 plus usage - Steel Pricing. Runner-up: Cloudflare Browser Run if you are already on Workers, where 120 concurrent browsers and a compliant /crawl endpoint at no separate SKU make the build-vs-buy math trivial - Cloudflare.
Job 3: authenticated internal-app automation. The same logged-in workflow, run reliably every day, across internal tools, SSO, and MFA. Pick: Anchor Browser, whose 1Password-integrated credential handling and Web Action Cache replay economics fit exactly this shape, provided the workflow is stable - Anchor Browser. Runner-up: Amazon Nova Act for AWS shops whose workflows can live in its compile-and-monitor model - AWS.
Job 4: fleet-scale production workflows and delegated work. You want outcomes, not sessions. Pick: TinyFish, which now pairs its 90.0 percent benchmark score with a self-serve rate card and failed-runs-free billing, collapsing what used to be a procurement-cycle decision into a $15 experiment - TinyFish Pricing. Runner-up: O-mega for teams that want orchestrated multi-agent work (browsing, code, tools) under one roof with a free tier to validate against. And beneath all four jobs, Browser Use remains the constant: the right open-source default for anything you intend to build rather than buy, which is why it tops the overall table. If your workload leans more desktop than browser, start instead from our guide to AI agents for desktop automation, because Section 4's OSWorld reset changes the calculus there entirely.
22. FAQ
Is ChatGPT Atlas shutting down? Yes. OpenAI shuts Atlas down on August 9, 2026, less than a year after its October 2025 launch, and says it will build browsing into ChatGPT itself; agentic capabilities consolidate into the ChatGPT desktop experience and ChatGPT Work, which launched July 9, 2026 - BigGo. This follows Operator, which OpenAI shut down August 31, 2025 after seven months - Wikipedia.
What is the best browser agent on real benchmarks? On the independent Online-Mind2Web leaderboard (June 29, 2026 update), Browser Use Cloud leads at 97.0 percent, followed by GPT-5.4 Native Computer Use at 93.0 and TinyFish at 90.0; ChatGPT Atlas Agent Mode sat at 71.0 percent, eighth - Steel Leaderboard.
Are browser agents safe to use with my accounts? Not unconditionally. A University of Washington study of seven agentic browsers found four create same-origin-policy bypasses, with a successful proof-of-concept against Atlas and risk conditions in Chrome with Gemini, Claude for Chrome, and Comet; the lead researcher's verdict was "browser agents aren't ready for the public" - UW News. Use scoped credentials in isolated sessions for automation, and keep approval gates on for anything consumer-grade.
Is desktop automation as good as browser automation now? No, and the gap just got formally re-measured. Live-web task success tops out at 97.0 percent, while on OSWorld 2.0 (108 long-horizon desktop workflows, median human time ~1.6 hours) the frontier is Claude Opus 4.8 at 20.6 percent binary completion - OSWorld 2.0. Buy browser automation with confidence; supervise desktop automation.
What is the cheapest way to run browser agents at scale? Two candidates depending on where you build. Browser Use Cloud meters sessions at $0.02 per browser-hour with plans from $29 - Browser Use Pricing. If you already run on Cloudflare Workers, Browser Run is included on Free and Paid plans with 120 concurrent browsers and no separate SKU - Cloudflare.
What is the best open-source browser agent in 2026? Browser Use, on the evidence: 107,928 GitHub stars and the highest Online-Mind2Web score ever recorded - GitHub. Steel is the leading open-source choice at the infrastructure layer, and Skyvern the most notable open-source no-code option, though it remains unbenchmarked on the current standard.
23. Final Take
Three editions in, this review's most reliable finding is not any ranking row: it is that the category invalidates its own reviews faster than any software market we track. Between editions, the most famous product died (again), the pricing objection we raised against our fourth-place vendor was fixed, the desktop benchmark reset from 85 percent to 20.6, a hyperscaler turned the infrastructure tier's business model into a bundled feature, and academia confirmed in peer review what security researchers had shown in proof-of-concept. Any page ranking these tools that has not been substantially rewritten in the last quarter is describing a market that no longer exists.
The three conclusions that survive the churn. First, open source won the capability race and kept it: the best-measured browser agent on earth is a free framework, now at 107,928 stars, which permanently caps what pure capability can charge. Second, evidence beats mindshare by embarrassing margins: the consumer agent everyone used scored 71.0 on the benchmark where the open-source leader scored 97.0, and it is the one being discontinued. Third, the durable risks are structural, not technical: prompt injection that may never be fully mitigated, and vendor strategy pivots that have now killed three big-tech browser agents in eighteen months. Capability is no longer the constraint. Governance, portability, and honest evidence are.
If you take one action from this review, take the tier test from Section 21: name your job before you name your vendor. A personal copilot, a scraping fleet, an internal-app workflow, and a delegated workforce are four different purchases wearing one category label, and every expensive mistake we have seen in this market, including the ones that ended in migration scrambles this week, came from buying across that boundary.
This review reflects the browser agent landscape as of August 2026. Every price and benchmark figure was verified against the cited source during this revision. This category changes faster than any other in software: pricing, model versions, benchmark leaders, and even vendor existence shift monthly, so verify current details on official pages before purchasing. Previous editions: December 2025 and July 2026; corrections to both are documented inline throughout.