title: "Best AI Agents for Desktop Automation 2026: Mac & Windows" slug: "top-10-ai-agents-for-desktop-automation-2026-mac-windows" date: "2026-08-05" excerpt: "Compare the best AI desktop automation agents for Mac and Windows in August 2026: Claude Cowork, ChatGPT Work, Manus, Lapu and more, with live-verified pricing." author: "Yuma Heymans"
The August 2026 guide to AI agents that actually operate your Mac or Windows machine, with every price and benchmark re-verified this week
Three load-bearing facts in the version of this guide we published 30 days ago are already false. ChatGPT Atlas, which we ranked as half of the #2 entry, shuts down on August 9, 2026 - Wikipedia. The "GPT-5.5 flagship" we cited was replaced by the GPT-5.6 family on July 9 - TechCrunch. And Claude Opus 4.8, whose OSWorld score anchored our capability analysis, was superseded by Claude Opus 5 on July 24 - Anthropic. We are telling you this in the first paragraph because it is the single most important thing to understand about this market: a thirty-day-old ranking of desktop AI agents is not slightly stale, it is materially wrong.
So this is a full refresh, not a touch-up. Every model name, every price, and every benchmark below was checked against a primary source in the first week of August 2026. We also restructured the ranking itself, because the market restructured underneath it. The intent behind "best AI agents for desktop automation" has shifted hard toward agents that run on your actual machine, touch your actual files, and ideally keep your data local. Our previous list spent three of ten slots on tools that never touch a desktop at all (a browser SDK, a web-only workflow engine, an enterprise platform). Those have been moved to an honest "adjacent tools" section, and a new desktop-native generation (Lapu, goose, OpenOwl, Bytebot, UI-TARS Desktop, Fazm) that barely existed at our last rewrite now gets the coverage the category deserves.
One more thing about method. We build and operate production browser-session and computer-session agents at O-mega, which means the judgments here come from running agents against real websites, real logins, and real failure modes every day, not from reading vendor pages. Where we state observed behavior, it is qualitative and first-hand; where we state numbers, every one carries a source we opened this week. We do not fabricate timed lab results, and we flag vendor self-reported figures as exactly that.
This guide was researched and written by Yuma Heymans (@yumahey), founder of O-mega and co-founder of HeroHunt.ai, who has spent the last two years debugging in production the exact agent failure modes described in section 17.
Contents
- What died in the 30 days since our July refresh
- Where the agent runs: the question that now decides everything
- Benchmark reality: OSWorld is saturated, OSWorld 2.0 is brutal
- Claude Cowork (Anthropic) - the default desktop agent, now on Opus 5
- ChatGPT after Atlas - Agent Mode, ChatGPT Work, and the desktop superapp
- Manus - independent, local-capable, and still credit-shaped
- Google's Gemini stack - auto browse goes mainstream, Spark goes global
- Lapu AI - the local-first native desktop agent
- O-mega - an AI workforce instead of a single agent
- goose (Block) - the open-source local agent that grew up
- Microsoft Copilot Actions, Agent Workspace, and Fara-7B
- Simular - research frontier with a consumer beta
- Perplexity Computer - a 19-model cloud agent behind a $200 wall
- The new desktop-native generation - OpenOwl, Bytebot, UI-TARS, Fazm
- What left the top 10 - Nova Act, Skyvern, and Context
- Verified pricing across the board (August 2026)
- What actually breaks in production: an operator's view
- Which agent for which person: the decision framework
- Future outlook
- Conclusion
The August 2026 Ranking at a Glance
We score each agent on five criteria built around what desktop-automation buyers actually ask for in 2026. Desktop reach & local execution (30%): does it run natively on Mac and Windows, and can it touch your real files and applications rather than a remote sandbox. Verified capability (25%): published benchmark results plus observed reliability, with vendor-only numbers discounted. Adoption friction (20%): how fast a non-engineer gets real value. Price transparency & value (15%): published pricing scores higher than opaque credits or contact-sales walls. Accountability (10%): approval gates, audit trails, and admin controls, because an agent you cannot audit is a liability. Scores run 0-10 and every cell states the reason for the number.
| # | Agent | What It Does | Desktop Reach (30%) | Capability (25%) | Friction (20%) | Price (15%) | Accountability (10%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Cowork | Autonomous desktop agent for general knowledge work | 10 - native macOS, Windows, Linux, ChromeOS; real file access | 10 - Opus 5 leads OSWorld 2.0 per-cost per Anthropic | 9 - install, point at folders, delegate | 8 - from $17/mo Pro, published tiers | 8 - enterprise admin controls, OpenTelemetry streams | 9.3 |
| 2 | ChatGPT (Agent Mode + Work) | Agentic browsing plus hours-long background work | 7 - desktop app on every plan, but execution is a remote virtual computer | 9 - GPT-5.6 Sol; Work ships finished files | 9 - toggle inside the app you already use | 8 - Plus $20/mo, but Work usage is metered | 7 - confirmation prompts, admin tiers | 8.1 |
| 3 | Manus | Cloud VMs plus local "My Computer" execution | 8 - desktop app runs terminal commands on your machine, gated | 8 - strong long-task persistence, watchable replays | 8 - prompt-to-task simplicity | 6 - credits are hard to predict at heavy usage | 7 - per-command approval gate | 7.6 |
| 4 | Google Gemini stack | Chrome auto browse, Spark background agent, Computer Use API | 6 - powerful but Chrome-bound, no native file agent | 8 - auto browse GA on desktop since January | 9 - lives in the browser you already have | 8 - AI Pro $19.99/mo | 7 - purchase/social confirmation gates | 7.5 |
| 5 | Lapu AI | Local-first native desktop agent for Mac and Windows | 9 - native app, files/terminal/desktop, nothing leaves the machine | 5 - no public benchmark record yet | 6 - young product, small ecosystem | 8 - free tier plus $20/mo Premium, transparent | 9 - every action approved and logged with timestamps | 7.3 |
| 6 | O-mega | Autonomous AI workforce with browser + computer sessions | 6 - OS-agnostic cloud sessions, no local file access | 7 - multi-agent orchestration over real browser/VM work | 8 - describe outcomes, agents execute | 6 - workforce pricing above single-agent tools | 9 - roles, credentials, audit trail per agent | 7.0 |
| 7 | goose (Block) | Open-source local agent: desktop app, CLI, API | 8 - native app on macOS, Linux, Windows; runs on your machine | 6 - capability tracks whichever model you bring | 5 - BYO API keys, technical setup | 9 - free, Apache 2.0, pay only model tokens | 6 - self-managed guardrails | 6.9 |
| 8 | Microsoft Copilot Actions + Fara-7B | OS-level agent features and an on-device model | 7 - deepest Windows integration, nothing for Mac | 5 - experimental, bounded task set, off by default | 6 - buried in Settings, opt-in preview | 8 - included with Windows 11, Fara MIT-licensed | 9 - OS-level agent identity and containment | 6.7 |
| 9 | Simular | Agent S3 open framework plus consumer Mac beta | 7 - S3 drives real machines; consumer app in public beta | 8 - 69.9% OSWorld with bBoN, near human baseline | 4 - beta app, framework needs engineers | 7 - framework free, app pricing unpublished | 7 - approvals before critical actions | 6.7 |
| 10 | Perplexity Computer | Cloud agent orchestrating 19 models | 5 - cloud-only, nothing runs on your machine | 7 - multi-model orchestration on long workflows | 8 - just prompt, if you have Max | 4 - $200/mo Max only, credits expire monthly | 7 - default spending caps | 6.2 |
Read the table with two things in mind. First, the sort is by final score and the tie at 6.7 is broken alphabetically; Microsoft and Simular genuinely deserve the same number for opposite reasons (one is containment without capability, the other capability without packaging). Second, rank #5 for Lapu is a bet flagged as a bet: its desktop reach and accountability scores are the best-in-class shape for where this market is going, while its capability score honestly reflects that no independent benchmark record exists yet. If you want the deep head-to-head on the two frontier model families behind the top two entries, our GPT-5.6 vs Claude Opus 5 agents comparison runs that analysis in full.
1. What Died in the 30 Days Since Our July Refresh
We keep a changelog of product deaths at the top of this guide because it is the most useful information in it, and because no aggregator will publish one about its own mistakes. The July version of this article opened by noting that Operator, Mariner, and the Meta-Manus deal had all died since the original ranking. That changelog now needs three new entries, and they all happened within 30 days of our last rewrite.
ChatGPT Atlas is dead. OpenAI's agentic browser, launched October 21, 2025 on macOS, shuts down on August 9, 2026, and the long-promised Windows version will now never ship - Wikipedia. The shutdown follows the consolidation OpenAI announced in March 2026: Atlas, the ChatGPT desktop application, and Codex are being combined into one desktop app. Our July ranking treated Atlas as a pillar of the OpenAI entry and as the anchor of an entire "agentic browser takeover" section. Both are gone from this version, and section 5 covers what Atlas users should actually do before the deadline. Our Atlas hands-on guide from launch week now reads as a complete product biography: birth to death in under ten months, and our agentic browser comparison documents the category at its peak.
The GPT-5.5 era ended one day after our refresh shipped. On July 9, 2026, OpenAI launched the GPT-5.6 family: Sol (flagship, $5/$30 per million input/output tokens), Terra (mid-tier, $2.50/$15), and Luna (budget, $1/$6) - TechCrunch. The same day, OpenAI shipped ChatGPT Work, a background agent that takes a brief and works for hours across connected apps, which restructures the entire OpenAI entry below. And on July 24, Anthropic shipped Claude Opus 5 at the same $5/$25 pricing as its predecessor, making it the new default on Max plans - Anthropic.
The older deaths still matter for context, so here is the compressed history. OpenAI Operator launched January 23, 2025, scored 38.1% on OSWorld, and was shut down on August 31, 2025 after ChatGPT agent absorbed it - Wikipedia. Google Project Mariner launched December 11, 2024 and was discontinued on May 4, 2026, its capabilities dispersing into Chrome auto browse and the Gemini API - Wikipedia. And Meta's attempted $2-3 billion acquisition of Manus, announced in late December 2025, was blocked by China's National Development and Reform Commission on April 27, 2026 before it could close, leaving Manus independent - Codersera. What Operator's pricing collapse looked like from the inside is preserved in our Operator pricing breakdown.
The structural pattern across all six deaths is the same: standalone agent surfaces get absorbed into platforms. Operator became a ChatGPT toggle. Mariner became a Chrome feature. Atlas is becoming a tab in a unified desktop app. The buyer's question in 2026 is not "does this agent work" but "which platform will this agent be a feature of in twelve months," and the ranking below rewards products whose answer is either "it is the platform" or "it is open source, so it cannot be shut down out from under me."
2. Where the Agent Runs: The Question That Now Decides Everything
Our July version buried its most important analytical distinction in sections 15 and 16. This version leads with it, because it has become the primary axis on which buyers actually choose, and the axis on which the new generation of products is attacking the incumbents. The question is deceptively simple: where does the agent execute, and therefore what can it actually touch, and who can actually see your data.
Three architectures now compete. Remote execution runs the agent in a cloud sandbox or virtual machine: ChatGPT's Agent Mode and ChatGPT Work, Perplexity Computer, and workforce platforms like O-mega all work this way. The strengths are scale (many agents in parallel), persistence (work continues when your laptop sleeps), and safety-by-distance (the agent cannot delete your local files because it cannot see them). The structural weakness is the same fact inverted: a remote agent cannot open the spreadsheet on your desktop, use your licensed native apps, or work behind your corporate VPN without plumbing. Local execution runs the agent directly on your machine: Lapu, goose, OpenOwl, Fara-7B, and UI-TARS Desktop live here. Everything is reachable and nothing leaves the device, at the price of consuming your compute and demanding much stronger containment, because a misfiring local agent is misfiring inside your real environment. Hybrid execution splits the difference: Claude Cowork reasons in the cloud but reads and writes your real folders; Manus runs cloud VMs by default and gated local commands on request.
The reason this axis now dominates is intent-driven, and we can see it in what buyers search for and what the winning products ship. The 2024-2025 generation of agents proved that models can operate computers; the 2026 question is whose computer, with whose data, under whose audit log. That is why an unbenchmarked newcomer like Lapu earns a top-five slot on architecture and accountability while a genuinely clever cloud product like Perplexity Computer ranks last in a desktop list: for the desktop-automation buyer, "it runs on my machine and shows me every action" has become a feature worth more than raw orchestration sophistication. It is also why every profile below opens by stating the product's execution locus before anything else.
3. Benchmark Reality: OSWorld Is Saturated, OSWorld 2.0 Is Brutal
In July we wrote that "OSWorld solved" headlines would arrive within a year and a new reference benchmark would emerge by 2027. That prediction came true early, and it is worth saying plainly: the field's own reference benchmark rolled over between our refreshes. Classic OSWorld (369 real desktop tasks across Ubuntu, Windows, and macOS) is now effectively saturated at the top: the verified leaderboard tracked by Steel, last updated July 10, 2026, shows Claude Mythos Preview at 85.4%, Claude Mythos 5 and Claude Fable 5 at 85.0%, Claude Opus 4.8 at 83.4%, Claude Sonnet 5 at 81.2%, and GPT-5.4 at 75.0%, all comfortably above the 72.36% human baseline, with the best open entry, Qwen3 VL 235B, at 66.7% - Steel OSWorld Leaderboard. When the top five entries beat the human reference and cluster within four points, a benchmark stops discriminating.
The frontier moved to OSWorld 2.0, published in June 2026 by the XLANG Lab at the University of Hong Kong with 36 authors across a dozen collaborating institutions, and it was designed to be brutal: 108 long-horizon workflows that take humans a median of about 1.6 hours each and require an average of 318 tool calls for an agent to complete, roughly ten times the action depth of the original benchmark - OSWorld 2.0. The scores reset accordingly. On OSWorld 2.0's binary completion metric, Claude Opus 4.8 manages 20.6%, Opus 4.7 scores 18.2%, and GPT-5.5 lands around 14%, with the benchmark authors concluding that current agents "are still far from professional-level computer use," struggling specifically with hidden state recovery and mid-task changes - arXiv. Anthropic reports that the new Claude Opus 5 "outperforms every other model at any given cost" on OSWorld 2.0, surpassing Fable 5's best result at just over a third of the cost, though a public per-model leaderboard entry was not yet posted when we checked - Anthropic.
Anthropic published its own cost-versus-score curve for OSWorld 2.0 in the Opus 5 announcement, and it is worth seeing because the cost axis is the part vendors usually hide.
How should a buyer read this pair of numbers, 85% on the old benchmark and 20% on the new one? Both are true, and together they describe the actual product experience better than either alone: today's best agents are superhuman at bounded tasks (fill this form, reconcile these two files, extract this data) and unreliable at professional-length workflows (run this two-hour process end to end without supervision). Every profile below inherits that frame. The capability curve is still steep (the best published classic-OSWorld score went from 20.6% to 85.4% in roughly eighteen months per the Simular and Steel records), so treat OSWorld 2.0's low scores as a countdown, not a ceiling - Simular. For the model-by-model view, our Opus 5 vs Opus 4.8 benchmark breakdown and Fable 5 and Mythos 5 analysis go deeper, and the wider evaluation landscape lives in our computer-use benchmarks guide.
4. Claude Cowork (Anthropic): The Default Desktop Agent, Now on Opus 5
Execution locus: hybrid. Cowork reasons in Anthropic's cloud but reads and writes the real folders you grant it, on the broadest native footprint in this ranking: macOS, Windows (x64 and arm64), Linux, and ChromeOS, with web and mobile in beta - Claude Cowork. That footprint, plus the model advantage described in section 3, is why Cowork holds #1 for the second refresh running, and holds it by a wider margin than before.
The July development that matters most is under the hood: Claude Opus 5, released July 24 at $5/$25 per million tokens, is now the default model on Max plans and the strongest model on Pro, and it is the model Anthropic says leads OSWorld 2.0 at every cost point while more than doubling Opus 4.8's Frontier-Bench performance - Anthropic. In practice, the model swap changes what you can safely delegate. The bounded-task tier (clean this data, redline this contract, draft this update from these notes) was already reliable on Opus 4.8; the difference we observe with Opus 5-class models is on multi-file, multi-hour jobs, where the agent recovers from its own dead ends more often instead of confidently shipping a half-finished result.
The enterprise story matured as well, and it addresses the accountability axis directly. Enterprise admins can now manage feature access, control spend, and track Cowork usage across the organization, stream activity to a SIEM through OpenTelemetry, and, since February 24, 2026, build private plugin marketplaces that bundle skills, connectors, and sub-agents for specific roles - Claude Cowork. That last item is quietly important: it turns Cowork from a personal assistant into a distributable, governed internal product, which is exactly the shape IT departments have been asking vendors for.
Pricing - Claude Cowork:
| Plan | Monthly Cost | What You Get |
|---|---|---|
| Pro | $17/mo annual ($20 monthly) | Cowork included, standard limits |
| Max 5x | $100/mo | 5x usage for heavier delegation |
| Max 20x | $200/mo | 20x usage, Opus 5 default |
| Team | $20/seat/mo | Teams of 2 to 150, shared workspaces |
| Enterprise | Custom | Admin controls, SIEM streaming, private marketplaces |
The honest limitations have not moved: Cowork is a single-agent experience, so delegating a whole function rather than a task still means you become the coordinator between sessions, and heavy users on Pro will feel the pull toward the $100-200 Max tiers. It is also a hybrid, not a local product: your files are readable by a cloud service, which some data-custody postures cannot accept, and for those buyers sections 8 and 10 exist. For setup and workflow patterns, our Cowork starter guide and Cowork pricing and ecosystem breakdown cover the operational detail this profile compresses.
Best for: anyone on any OS who wants the most capable general desktop agent shipping today, and organizations that need agent capability with real admin controls behind it.
5. ChatGPT After Atlas: Agent Mode, ChatGPT Work, and the Desktop Superapp
Execution locus: remote. Everything agentic OpenAI ships runs in a cloud virtual computer, and as of this month, OpenAI's desktop story is consolidating rather than fragmenting: Atlas shuts down August 9, 2026, the Windows Atlas build is cancelled, and Atlas's capabilities merge into the unified desktop application OpenAI announced in March 2026, which combines the ChatGPT app, the browser, and Codex into one surface - Wikipedia. If you are an Atlas user, the migration is undramatic but real: your agentic browsing moves into the ChatGPT desktop app and Agent Mode, so export anything you keep in Atlas-specific storage before the deadline and expect the in-browser agent experience to feel more like "agent in an app with a browser tool" than "agent in your browser." We ranked Atlas as a top-two pillar in July; we were wrong about its lifespan within thirty days, and that correction is on us.
The product that replaces Atlas in this ranking's #2 slot is much stronger: ChatGPT Work, launched July 9, 2026 on the GPT-5.6 family. Work takes a brief and "works in the background for minutes or even hours," then hands you a finished file: a spreadsheet, a slide deck, a report, or a working web app, connecting through a unified plugin directory that covers Slack, Microsoft Teams, Google Drive, SharePoint, email, calendars, CRMs, and project trackers - Fello AI. Distribution is classic OpenAI: the desktop app is available on every plan including Free, web and mobile rolled out to paid plans, Free and Go accounts are limited to the Terra model, and there is no separate price tag, though usage is metered like Codex rather than flat-rate. That metering is the fine print to respect: "included in your plan" and "unlimited" are not the same sentence, and heavy Work users will discover the difference.
The model layer under all of this is the GPT-5.6 family: Sol at $5/$30 per million tokens, Terra at $2.50/$15, and Luna at $1/$6, with OpenAI claiming Sol tops the Artificial Analysis Coding Agent Index at 80 while using less than half the output tokens of its closest rival - TechCrunch. The subscription lineup as of this week runs seven tiers: Free, Go at $8/mo, Plus at $20/mo, Pro at $100 and $200/mo, Business at $25/seat/mo ($20 annually), and custom Enterprise, with the heaviest agent allowances concentrated at the $200 tier - AI Pricing Guru. Our GPT-5.6 benchmark and pricing guide unpacks the family in detail.
Where OpenAI trails is precisely this guide's primary axis: nothing executes on your machine. Agent Mode and Work operate a remote computer, which makes them superb for research, document production, and web tasks, and structurally unable to touch the files on your desktop the way Cowork, Manus, or the local generation can. On the hardest verified desktop benchmark OpenAI also sits behind: GPT-5.4 holds 75.0% on classic OSWorld against Anthropic's 83-85% cluster, and GPT-5.5 lands around 14% on OSWorld 2.0 against Opus 4.8's 20.6% - OSWorld 2.0. What keeps it at #2 anyway is the strongest distribution machine in software plus a genuinely new capability shape: Work's hours-long, finished-file background jobs are something no other consumer product does as cleanly. The landscape of products chasing exactly that capability is mapped in our ChatGPT Work alternatives roundup.
Best for: people already inside the ChatGPT ecosystem who want agentic research and hours-long document production, and anyone whose automation targets live on the web rather than in local files.
6. Manus: Independent, Local-Capable, and Still Credit-Shaped
Execution locus: hybrid, and the most explicitly gated hybrid in the list. Manus runs tasks in cloud virtual machines you can watch like a screen recording, and since the March 2026 desktop app, its "My Computer" mode can also execute terminal commands and touch local files on your actual machine, with each command requiring explicit approval - Codersera. That combination (cloud scale by default, gated local reach on request) remains genuinely differentiated, and among the consumer generalists only Cowork offers comparable local access.
The corporate story is now stable enough to summarize without drama. Meta announced a $2-3 billion acquisition of Manus in late December 2025; China's National Development and Reform Commission prohibited it on April 27, 2026 before closing, and Manus continues as an independent company that had already crossed $100M annualized revenue in December 2025, eight months after launch - Codersera. The episode left a permanent mark on the market: it established that agentic AI is treated as strategically sensitive technology by at least one major power, and it left every enterprise procurement team with a new diligence question about technology provenance that no relocation to Singapore fully answers. Individuals and startups have visibly priced that risk and bought anyway; regulated enterprises largely have not.
Pricing (verified August 2026) - NoCode MBA:
| Plan | Monthly Cost | Credits |
|---|---|---|
| Free | $0 | 300 daily refresh credits |
| Standard | $20/mo | 4,000/mo plus daily refresh |
| Customizable | $40/mo | 8,000/mo plus daily refresh |
| Extended | $200/mo | 40,000/mo plus daily refresh |
Two pricing details changed since July and both favor the buyer: a new $40 middle tier fills the previously awkward gap between $20 and $200, and all paid tiers now support up to 20 concurrent tasks, with annual billing saving 17%. The persistent friction is the credit economy itself: simple tasks burn 10-50 credits while deep research or app-building runs can consume 500-1,000+, which means your effective monthly capacity depends heavily on task mix and is hard to predict in advance - NoCode MBA. In our experience running long agent tasks in production, this unpredictability is not a billing quirk but a planning problem: workflows you want to run nightly need a cost you can forecast, and credit-shaped pricing resists that.
Best for: individual power users and small teams who want one agent that can both run big cloud jobs and, with explicit approvals, work on local files, and who value watching the agent's session replay to build trust.
7. Google's Gemini Stack: Auto Browse Goes Mainstream, Spark Goes Global
Execution locus: remote-in-your-browser. Google still ships no native desktop agent; it ships agency inside Chrome and behind the Gemini app, and in the past six months that strategy quietly went from gated preview to broad deployment. Chrome auto browse launched to Google AI Pro and Ultra subscribers on desktop on January 28, 2026, handling multi-step chores (scheduling, forms, quotes, document collection) while showing its step-by-step actions in the Gemini side panel, with Google Password Manager integration and mandatory confirmation before purchases and social posts - 9to5Google. Our July version still described this as a "US-only preview" novelty; that undersold it.
The bigger July development is Gemini Spark, Google's background agent, which expanded from Ultra exclusivity to AI Pro subscribers in over 160 additional countries on July 30, 2026, and gained the ability to itself drive Chrome auto browse: Spark can now use your logged-in Chrome and saved passwords to execute web errands end to end, with user approval required for sensitive actions like payments and stated protections against prompt injection - 9to5Google. Read that architecture carefully, because it is a preview of where all the platform players are heading: a background agent that commands a browser agent, both living inside subscriptions you may already pay for. At $19.99/mo for AI Pro, this is the cheapest way for a mainstream user to get a working background-plus-browser agent combination today.
The limitations define the buyer. Everything is Chrome-bound and account-bound: there is no native file agent, no terminal, nothing that touches your desktop applications, so "desktop automation" here means "web automation from your desktop." There is a daily cap on agentic actions - 9to5Google. And Google's agent surface still carries the company's trademark product-churn risk: Mariner is the second dead Google agent brand in two years, which is a real planning consideration for anyone building durable workflows on Google's agent APIs rather than just using the consumer features. The underlying model economics across Google's lineup are tracked in our model benchmarks and pricing roundup.
Best for: people who live in Chrome and Google Workspace and want ambient, cheap agentic help with zero new tools, now including a background agent in most of the world.
8. Lapu AI: The Local-First Native Desktop Agent
Execution locus: fully local, and that is the entire thesis. Lapu is a native desktop application for Mac and Windows whose agent has direct access to your files, terminal, and desktop, with an explicit architectural promise: files and data remain on your machine, with "no Lapu AI cloud storage, no remote workspace, and no ambient data collection" - Lapu AI. Where the cloud generalists ask you to trust a vendor's data handling, Lapu's answer is structural: there is nothing to trust because nothing leaves.
The accountability design is the best-in-class shape in this ranking, and it is why an unbenchmarked newcomer earns a top-five slot. Every sensitive action requires explicit approval before proceeding: file writes, shell commands, and desktop automation are gated by default, and every agent action (tool invocations, file operations, model calls) is logged and visible in real time with timestamps - Lapu AI. This is the exact permission-and-audit pattern that section 17 argues production agent systems converge on after they get burned; Lapu ships it as the default posture rather than an enterprise add-on. The feature set around it is pragmatic: cross-app workflows, reusable skills with scheduled automation, and a floating mini chat reachable from any app.
Pricing - Lapu AI:
| Plan | Monthly Cost | What You Get |
|---|---|---|
| Free | $0, no card | Core desktop agent, file and terminal tools |
| Premium | $20/mo | Workflow scheduling, email support |
| Teams | Custom | Unlimited seats, central billing |
The honest counterweight is the capability column of our table. Lapu publishes no benchmark results, has no OSWorld entry, and has not yet accumulated the public production track record of the platforms above it, which is why it scores a 5 where Cowork scores a 10. Local-first also means the agent's competence is bounded by which models it can use and how well its planning layer drives them, and young products iterate fast in ways that occasionally break workflows. Our judgment call, stated as such: the architecture is right for where this market is going, the proof burden is still open, and the free tier makes verifying it against your own workload cheap.
Best for: privacy-conscious individuals and teams on Mac or Windows who want a real local agent with per-action approvals and a full audit trail, and who are willing to trade some frontier capability for full data custody.
9. O-mega: An AI Workforce Instead of a Single Agent
Full disclosure first, as always: O-mega is our product, scored with the same rubric as everything else, and its execution locus is remote by design. Every other consumer entry in this ranking is one agent you supervise. O-mega's unit is an AI workforce: multiple persistent agents with names, roles, credentials, and schedules, each running browser sessions and computer sessions in the cloud, coordinated by orchestration rather than by you relaying context between chat windows.
The distinction earns its keep at a specific scale, and we are precise about where. For a single recurring task, a single agent (Cowork, Agent Mode, Manus) is the right shape and the cheaper deal. The workforce shape wins when you are offloading a function: all of outbound research, all of listing management, all of weekly reporting. At that point you want separate agents with separate logins, separate process memories, and a delegation chain between them, so the unit you manage is the process rather than the prompt. Because execution happens in cloud sessions, O-mega is OS-agnostic (direct it from a Mac, a Windows machine, or a phone) and structurally incapable of touching your local files, which is the honest trade-off against everything in sections 4, 6, and 8: local-custody buyers should choose a local product. The wider landscape of agents that operate computers on your behalf, including ours, is mapped in our computer-use agents review.
Two operational notes from running this in production, because they generalize to every remote agent you might buy. First, credential and session management is the real product surface: agents doing recurring web work need persistent, isolated browser identities, and most of the reliability gap between demo and production lives there. Second, audit trails per agent are not an enterprise vanity feature; they are how you debug an autonomous system at all, which is why our accountability score is the rubric's highest alongside Microsoft's OS-level containment. On cost, workforce pricing sits above single-agent consumer subscriptions, and the economic frame that justifies it (or does not, for single-task buyers) is the one we built in our true cost of agentic AI report.
Best for: founders and operators who want to delegate entire recurring business processes to a coordinated set of agents with roles, credentials, and audit trails, rather than supervising one assistant task by task.
10. goose (Block): The Open-Source Local Agent That Grew Up
Execution locus: fully local, fully yours. goose is Block's open-source AI agent, shipped as a native desktop app for macOS, Linux, and Windows, a full CLI, and an embeddable API, licensed Apache 2.0, and it runs on your machine against whichever model you connect: it supports 15+ providers including Anthropic, OpenAI, Google, Ollama, OpenRouter, Azure, and Bedrock - goose. Connectivity runs through the Model Context Protocol with an ecosystem of 70+ extensions, which in practice means the same connector standard the commercial platforms adopted is the native tongue here.
The governance development is what moves goose from "interesting repo" to "infrastructure you can bet on": the project moved under the Agentic AI Foundation at the Linux Foundation, making it vendor-neutral and community-governed for the long term - goose. Recall section 1's structural pattern: standalone agent products keep getting absorbed or shut down by their platform owners. An Apache-licensed, foundation-governed agent is the one shape that pattern cannot kill, which matters if you are automating processes you intend to still be running in 2028. The price shape follows from the license: the software is free and you pay only the model tokens you consume, with a local-model path through Ollama for the full-custody configuration.
The trade-offs are the standard open-source contract, stated without romance. Setup means bringing your own API keys and configuring providers and extensions, which filters out most non-technical users; capability is a function of the model you connect rather than a tuned first-party stack, so goose with a budget model is a very different product than goose with a frontier one; and guardrails are yours to configure, with no vendor holding a safety layer above you. For technical users that last point is a feature. The surrounding self-hosted ecosystem, including how goose compares to building your own stack, is covered in our open-source personal AI guide and our open-source AI coders roundup.
Best for: technical users and teams who want a free, foundation-governed local agent with a real desktop app, full model choice including local models, and no vendor who can sunset it.
11. Microsoft Copilot Actions, Agent Workspace, and Fara-7B
Execution locus: local, at the OS level, which is a position no one else on this list can occupy. Windows 11 ships experimental agentic features in which Copilot takes on automated tasks around file organization, scheduling, and email, running inside a contained workspace under an agent-specific low-privilege account, with everything disabled until you explicitly turn it on under the AI components area of Settings - PCWorld. Our July version implied default-on agentic Windows was imminent; the verified reality is more modest. Windows 11 26H2 arrives late September to early October 2026 as an enablement package (a small switch-flip update, not a reinstallation) that unlocks features already staged through monthly 25H2 updates, and the agentic features remain opt-in preview - PCWorld.
Judged as an agent today, this is the weakest capability entry in the top ten: a bounded set of file and app chores, off by default, well behind the generalists. Judged as infrastructure, it remains the most consequential, because Microsoft is answering the question every IT department asks about every product above: who is accountable when an agent acts? Giving agents an identity the OS understands, a workspace the OS contains, and an audit surface the OS produces is something only the OS vendor can do, and once those primitives are default-on, third-party agents will be pushed to run inside them the way apps were pushed into sandboxes a decade ago.
The second half of Microsoft's entry is Fara-7B, its MIT-licensed on-device computer-use model, which runs quantized on Copilot+ hardware so screen pixels never leave the machine; it remains the most credible fully-local model path on consumer Windows, with the significant caveat that small on-device models sit far below frontier cloud capability on every hard benchmark. That capability gap is the price of custody, and for bounded repeated workflows on sensitive data it is often a price worth paying. Microsoft's parallel Cowork-adjacent productivity agents are a separate track we analyze in our Copilot Cowork breakdown; the OS primitives, not those apps, are the durable story here.
Best for: Windows-first organizations planning for OS-enforced agent identity and containment, and users who want to experiment with agentic Windows today by opting in rather than adopting third-party software.
12. Simular: Research Frontier With a Consumer Beta
Execution locus: split personality, and the profile has to hold both halves honestly. The research half is Agent S3, Simular's open-source compositional framework, released October 2, 2025, which wraps frontier models with planning and grounding and its signature Behavior Best-of-N technique: run several complete attempts, convert each trajectory into a behavior narrative, and have a judge pick the best. That lifted OSWorld from 62.6% single-run to 69.9%, within a few points of the 72% human baseline, capping a lineage that ran 20.6% (Agent S) to 48.8% (Agent S2) in about a year - Simular. S3 drives real machines across macOS, Windows, and Linux, and remains the cheapest way for an engineering team to own a near-frontier desktop agent outright: framework free, pay only model tokens.
The consumer half is newer and less settled than we would like for a ranked product. Simular's consumer offering, centered on the Sai agent, is in public beta: it runs "in a private remote environment that is always available," checks with you "before critical actions," and advertises approvals built into every important step, but ships no published pricing beyond mentions of monthly and yearly plans with a free trial - Simular. Note the architecture honestly: despite the macOS packaging, the current Sai product executes in a remote environment, so buyers seeking the local-first custody of section 8 or 10 should read the fine print rather than the app icon.
That split explains the score. Capability is top-tier and independently benchmarked, which almost nothing else in the bottom half of this table can say; friction is the worst in the ranking, because the capable half needs engineers and the packaged half is a beta with unpublished pricing. The trajectory is the reason it stays ranked: Simular is the only entrant moving along the research-to-consumer path with published, verifiable results at every step, and its bBoN idea has already been quietly absorbed into commercial stacks, which is the usual afterlife of good open agent research.
Best for: engineering teams that want a state-of-the-art self-hosted desktop agent without per-seat licensing, and early adopters willing to run a beta from the team with the best published capability record outside the big labs.
13. Perplexity Computer: A 19-Model Cloud Agent Behind a $200 Wall
Execution locus: entirely remote, which is why the most talked-about agent launch of the spring sits last in a desktop-automation ranking, and we want to be precise about both halves of that sentence. Launched February 25, 2026, Perplexity Computer is a cloud agent that coordinates 19 different models (spanning Claude, GPT, Gemini, and Grok families) to run complex multi-step workflows autonomously, and by the standard of cloud orchestration it is genuinely ambitious - SentiSight. Nothing about it, however, runs on your machine or touches your files: it is a research-and-execution service you visit, not a desktop agent you install.
The access economics are the steepest in the category. Computer is not sold standalone: it requires Perplexity Max at $200/month ($2,000 annually, or $325/seat for Enterprise Max), which includes 10,000 monthly credits that expire without rollover, with a default additional spending cap of $200 that users can raise to $2,000 - SentiSight. One warning for cross-shoppers: at least one widely-read competitor listicle claims Computer is available on the $20 Perplexity Pro tier; the sources we checked this week are unambiguous that it requires Max, so price your evaluation at $200, not $20. Credit-consumption on long workflows is substantial, and the spending-cap design tells you Perplexity itself expects overage to be a normal event.
What you get for that money is real: multi-model orchestration means the system routes work to whichever model suits each step, and observed strength on long research-and-produce workflows is exactly what you would predict from that design. But the buyer this guide serves should apply the section 2 test first. If your automation needs live in files, apps, and logged-in sessions on your machine, Computer cannot reach them at any price; if your needs are web-shaped research and production at volume, ChatGPT Work delivers a comparable capability class inside a $20 Plus plan, metered rather than walled at $200. Ranking last here is a statement about fit for desktop intent, not about engineering quality.
Best for: heavy research users already justifying Perplexity Max who want long autonomous web workflows, and buyers comparing multi-model orchestration approaches rather than desktop reach.
14. The New Desktop-Native Generation: OpenOwl, Bytebot, UI-TARS, Fazm
Beneath the ranked ten, a generation of desktop-native tools emerged in the last year that this guide previously did not cover at all, and the omission mattered: these products are where the local-execution thesis of section 2 is being explored fastest. None of them yet has the track record for a top-ten slot; all of them are worth knowing, and one or two will likely be ranked entries in the next refresh. What unites them is that every one runs on or against your actual machine, and most are open source, which per section 10's argument makes them structurally shutdown-proof in a market where platform vendors keep killing agent surfaces.
OpenOwl is desktop automation as an MCP building block: install via npm i -g openowl, connect it to Claude, Codex, or any MCP-compatible assistant, and it sees your screen, clicks, and types across apps and browsers, with screenshots, files, and keystrokes "processed entirely on your machine" and signed binaries for verification; pricing runs free, Pro at $19.99/mo, and a $499 done-for-you setup - OpenOwl. Bytebot attacks containment instead: each agent gets a full containerized Linux desktop (Firefox, Thunderbird, VS Code preinstalled) running in an isolated Docker container on your own infrastructure, Apache 2.0 licensed and free, with costs limited to model API fees; it supports password managers including 1Password and Bitwarden for authenticated workflows, and records action history with before-and-after screenshots - Bytebot. UI-TARS Desktop is ByteDance's open-source native GUI agent (Apache 2.0, 38.4k GitHub stars), driving Windows, macOS, and browsers through the UI-TARS and Seed-VL vision-language models with fully local processing - GitHub. And Fazm is a free, open-source native macOS shell for the CLI-agent era: it wraps Claude Code, Codex, and Gemini CLI agents in a Mac-native UI with up to 40 parallel sessions, hold-to-talk voice input, and reach into Chrome and native Mac apps via accessibility APIs, on your existing Claude subscription - Fazm.
The pattern across the four is worth stating because it predicts the category's next two years. The new generation does not compete with Cowork on model quality; it competes on architecture: local processing, open licenses, MCP as the universal connector, and containment designs (Bytebot's throwaway desktops, Lapu's approval gates) that treat the agent as an untrusted process rather than a trusted assistant. That is the same conclusion production operators reach independently (section 17), arriving bottom-up from open source rather than top-down from enterprise requirements. When frontier-model quality becomes cheap and interchangeable (and the API price war among GPT-5.6, Opus 5, and their successors is doing exactly that), architecture is what remains to compete on.
15. What Left the Top 10: Nova Act, Skyvern, and Context
Three products from our previous ranking are no longer in the top ten, and none of them got worse. They were removed because this list's job is to match the desktop-automation intent, and all three are, by their own current positioning, not desktop agents. Keeping them ranked against Cowork and Lapu was a category error we inherited from the article's older structure; moving them here with updated facts is the honest correction.
Amazon Nova Act remains the cleanest labor-priced automation meter in the market: $4.75 per agent-hour of elapsed working time, with parallel agents billed separately and human-in-the-loop wait time explicitly excluded from the meter - AWS. It is a browser-automation SDK and AWS service for engineering teams, not a desktop product, and as the invisible engine inside internal tools it is excellent; that story began with Amazon's original Nova launch. Skyvern still does one unglamorous thing well (vision-driven automation of hostile web forms and portals), but its public pricing page now shows only "5,000 free credits every month" and an enterprise book-a-demo path; the Hobby and Pro tiers we quoted in July are no longer displayed, so treat any specific paid pricing you read elsewhere, including our own older coverage, as unverified until you see it in-app - Skyvern. It is web-only by design, and industrial web automation at that tier raises its own infrastructure questions, the kind we cover in our stealth browser guide. Context has completed its move upmarket into a full enterprise platform (Workspace, Engine with 800+ connectors, Unify, Evals) with hosted, VPC, air-gapped, and on-prem appliance deployments, and its flagship Qualcomm deployment now advertises 1,600 production workflows at 98% accuracy, up from the "roughly 100 workflows" figure we cited in July - Context.
The Context number deserves one paragraph on its way out of the ranking, because it carries the most important enterprise lesson in this guide. The 98% did not come from a better model; per Context's own positioning it comes from evals, rubrics, and golden sets applied to the same base model at every stage. Verification infrastructure, not model selection, is what separated a failed agent program from sixteen hundred production workflows. That lesson transfers to every product in the top ten, and it is the entire subject of section 17.
16. Verified Pricing Across the Board (August 2026)
Pricing rotted faster than any other fact class in this article's history, so this section is a single reference table with every number re-checked against the vendor's official page or a current primary source during the first week of August 2026. The sources sit next to each row's product profile above; the table exists so you can screenshot one thing.
| Product | Free tier | Entry paid | Mid tier | Top tier |
|---|---|---|---|---|
| ChatGPT (Agent Mode + Work) | $0, desktop app included | Go $8/mo; Plus $20/mo | Pro $100/mo | Pro $200/mo; Business $25/seat |
| Claude (Cowork) | Limited free chat | Pro $17/mo annual | Max 5x $100/mo | Max 20x $200/mo; Team $20/seat |
| Google AI (auto browse + Spark) | Gemini free tier | AI Pro $19.99/mo | - | AI Ultra tiers above |
| Manus | 300 credits/day | $20/mo (4,000 credits) | $40/mo (8,000) | Extended $200/mo (40,000) |
| Lapu AI | Free, no card | Premium $20/mo | - | Teams custom |
| goose | Free, Apache 2.0 | Model tokens only | - | - |
| OpenOwl | Free tier | Pro $19.99/mo | - | $499 done-for-you |
| Perplexity Computer | - | - | - | Max $200/mo (10,000 credits) |
| Nova Act | Playground | $4.75/agent-hour metered | - | Volume via AWS |
| Skyvern | 5,000 credits/mo | Not published | - | Enterprise, book a demo |
The distribution of entry prices tells the market's whole competitive story in one chart: everything that wants mainstream adoption has converged on a $17-20 monthly band, and the one product priced an order of magnitude above it is betting on a premium-bundle strategy rather than volume.
Two buying rules fall out of the numbers. First, start in the $20 band and instrument: every major capability class (hybrid desktop agent, background work agent, browser agent, local agent) now has a competent entry at or under $20/month, so the rational path is to measure how often you hit a limit before paying for a $100-200 tier. Second, respect the difference between flat, metered, and credit pricing: flat tiers (Claude, Lapu, Google) are predictable; metered products (ChatGPT Work's Codex-style metering, Nova Act's agent-hours) scale honestly with usage but need monitoring; credit systems (Manus, Perplexity) are the hardest to forecast, and if your workflow is recurring and nightly, forecastability is worth paying for. The deeper economics of agent pricing models are the subject of our true cost of agentic AI report.
17. What Actually Breaks in Production: An Operator's View
This section exists because it is the one thing an aggregator cannot write. We run browser-session and computer-session agents in production at O-mega every day, and the gap between "the demo worked" and "the workflow ran unattended for a month" is where all the real product differences live. What follows is qualitative and first-hand; none of it is a timed lab study, and we present it as operator experience, not benchmark data.
The first thing that breaks is authentication, not intelligence. The hardest step in most real web workflows is not reasoning about the page; it is being allowed onto the page at all: login walls, two-factor prompts, session expiry mid-task, and anti-bot systems that treat a competent agent exactly like the automation it is. This is why per-product details that look boring on a pricing page (Bytebot's password-manager support, Chrome auto browse's Password Manager integration, persistent agent browser identities) predict production success better than five points of OSWorld. The second thing that breaks is long-horizon drift, and OSWorld 2.0's 20% scores are the benchmark community catching up to what operators already knew: per-step reliability compounds, so a step that succeeds 95% of the time fails somewhere in almost every forty-step run. Production systems survive this with checkpoints and verification between stages, never with a single heroic end-to-end attempt, which is why the eval-shaped machinery of section 15 keeps winning enterprises. The third is the instruction/content boundary: prompt injection, where page or document content gets treated as instructions, remains unsolved in the general case, and every serious vendor's mitigation (confirmation gates at Google, approval-gated commands at Manus and Lapu, contained agent accounts at Microsoft) is a containment strategy rather than a cure.
The practical checklist we would give any team adopting anything from this guide is short. Grant minimal scopes and treat every credential an agent holds as inventory to track. Prefer agents that show their work: session replays, action logs, and timestamps are debugging surfaces, not marketing. Keep confirmation gates on for anything touching money, credentials, or outbound communication. Verify between stages of any workflow longer than a few steps. And treat "the agent did it" as an incident class your team will eventually file, because at production volume, you will.
The reason to be optimistic anyway is that the market is converging on exactly these patterns from every direction at once: enterprise platforms via evals, open source via containment-first architecture, OS vendors via agent identity, and consumer products via approval gates. The direction of travel in agent infrastructure has been visible since the earliest autonomous LLM agents: capability arrives first, accountability arrives second, and the products that survive are the ones that ship both.
18. Which Agent for Which Person: The Decision Framework
Rankings compress too much, so here is the routing logic we actually use when people ask, laid out as the three questions from section 2 plus one about scale. First: where does the work live? Web-shaped tasks (forms, research, portals, purchases) never need local file access, and granting it anyway is pure risk; file-and-app-shaped work needs a hybrid or local agent. Second: where may the data go? If the answer is "nowhere," the local generation (Lapu, goose, Fara-7B, OpenOwl) is your entire menu, and you accept the capability discount knowingly. Third: who verifies the output? You personally, a confirmation gate, or eval infrastructure; the longer the workflow, the further right you need to be on that spectrum. Fourth: what is the unit of scale? One assistant for one person, a workforce for a business function, or OS-governed fleets for an enterprise.
A few concrete personas make it tangible. A consultant on a Mac drowning in documents and decks: Claude Cowork on Pro at $17/month, full stop. A Windows analyst whose company forbids cloud file access: Lapu's free tier this week, goose with a local model if engineering support exists, and watch the 26H2 agent features land as the OS-level path this autumn. A marketer already on ChatGPT Plus: flip on Agent Mode, give ChatGPT Work a real brief, and buy nothing new until the metering bites. A two-person e-commerce brand with recurring supplier-portal drudgery: start with the $20 band, and graduate to a workforce platform like O-mega when the processes multiply past what one supervised agent can hold. An engineering team automating internal workflows: Bytebot's containerized desktops or Agent S3 for control, Nova Act's metering when the flows harden into production.
The two buying mistakes we see most have not changed, only their price tags: over-purchasing autonomy (a $200 tier for a task the $20 band handles) and under-purchasing accountability (any agent, at any price, with no plan for verifying output). Start cheap, instrument your limits, upgrade on evidence, and spend the saved money on verification.
19. Future Outlook
Three near-term developments look highly probable from the visible curves, and one longer arc matters more than all of them. First, OSWorld 2.0 scores will climb the way classic OSWorld's did: the 20% frontier of today is the 38% Operator moment of the last cycle, and Anthropic's decision to lead its Opus 5 announcement with OSWorld 2.0 cost-performance signals where the labs are pointing their training effort - Anthropic. Expect professional-length workflows to move from "unreliable" to "supervised-viable" within a small number of model generations. Second, the desktop superapp consolidation completes: OpenAI folds Atlas into one desktop application this month, Google now runs a background agent that drives its own browser agent, and Anthropic's Cowork-plus-plugins increasingly resembles an operating layer; the standalone agent app is a dying form factor at the platform tier, which structurally raises the value of the open-source escape hatch. Third, the local generation gets its benchmark moment: products like Lapu and OpenOwl currently compete on architecture without published capability numbers, and whichever of them first posts credible independent benchmarks will convert the privacy-first audience fastest.
The longer arc is the one we flagged in July and stand by with more evidence now: the unit of purchase is shifting from tool to outcome. ChatGPT Work hands you a finished file, not a chat log. Nova Act meters agent labor by the hour. Workforce platforms sell processes rather than seats. Enterprises measure workflows at 98% accuracy rather than models at 85% benchmarks. Every one of those is the same underlying transition: paying for completed work instead of paying for access to capability. When it completes, guides like this one will rank standing services with reliability records the way businesses today compare contractors, and the durable human role consolidates around specification and verification: deciding what is worth doing, and confirming it was done.
20. Conclusion
Thirty days invalidated three pillars of our previous version, and that fact is the guide. This market now updates faster than any publishing cadence, which is why this refresh leads with its changelog, verifies every number against a source we opened this week, and tells you plainly where our July judgment was wrong: we bet on Atlas weeks before its shutdown, cited a flagship model one day before its replacement, and under-covered the desktop-native generation that buyers are actually searching for.
The August 2026 shape of the market is legible, though. Claude Cowork is the default answer for general desktop delegation, now running the model that leads the hardest benchmark that exists. ChatGPT Work made hours-long, finished-file background work a $20-tier commodity even as Atlas died. Manus and the Gemini stack hold the hybrid and browser lanes. The local-first generation (Lapu, goose, and the section 14 bench) is the fastest-moving front, trading frontier capability for custody and accountability, and the workforce and enterprise tiers (O-mega, the Context-class platforms) win where the unit of delegation is a process rather than a task. Route yourself with the three questions: where does the work live, where may the data go, and who verifies the output.
We will refresh this guide again when the facts rot, which the last thirty days suggest will be soon. If this update proves anything, it is that in this category the honest ranking is not the one with the boldest claims but the one with the most recent verification dates.
This guide reflects the AI desktop automation landscape as of August 5, 2026. Every price and benchmark was checked against the linked primary sources during the first week of August 2026. This market changes monthly: verify current details on official pages before purchasing.