The evidence-based answer to "can AI do my job," read straight from OpenAI's GDPval benchmark and the 2026 labor data.
On the hardest test yet built for real work, the best AI model produced a deliverable that professional graders rated as good as or better than a human expert's 47.6% of the time - arXiv. That number belongs to OpenAI's GDPval, a benchmark of 1,320 real economic tasks across 44 occupations in the 9 largest sectors of U.S. GDP, and it is the closest thing we have to a direct answer to the question every knowledge worker is now asking in private.
But 47.6% is not "AI can do your job," and it is not "your job is safe" either. It is a deliverable rated in a blind comparison, on a one-shot task, by a grader who never had to live with the consequences. The distance between "produced a document an expert preferred" and "did the job, and answered for it on Monday" is the whole story, and it is a distance most headlines skip.
This guide reads GDPval carefully and then goes past it. It breaks down exactly what the benchmark measures, the real per-model results, the "100x cheaper" claim and why it collapses to roughly 1.4x once a human stays in the loop, the live 2026 scoreboard, and the independent labor data from METR, Anthropic, and Stanford that tells you which jobs are actually moving. It is written for the person whose job is on the table, not the person selling the tool.
Contents
- What GDPval Actually Measures (and Why It Broke the Mold)
- The Headline Number: 47.6% and the Year It Took to Get There
- The Machinery: How 1,320 Tasks Get Graded Blind
- The 100x Illusion: Cost, Speed, and the Human Who Never Left
- GDPval in 2026: The Live Scoreboard and a Human Baseline of 1000
- The Occupation Map: 44 Jobs, 9 Sectors, and Where the Losses Cluster
- Benchmark Versus Job: The Five Things GDPval Cannot See
- What the Independent Data Says About Your Job
- Replacement in the Wild: Klarna, Salesforce, Amazon, IBM
- The Deployment Divide: Why 95% of Pilots Show No Profit
- How to Actually Put AI on Your Work: Platforms and Pricing
- First Principles: Task, Role, and the Price of Expertise
- The Outlook: The End-of-2026 Bet and What 2027 Holds
- Conclusion: So, Can AI Do Your Job?
Before the analysis, the table below scores the practical places you can actually put an AI on your work, on the five things that decide whether it does the job or just hands you a first draft. Each cell carries the score and the evidence behind it, and the whole table is sorted by final score, highest first. The criteria and weights are explained directly beneath it. This is the "how to apply" bottom line; the "can it" analysis follows.
| # | Platform | What It Does | Task Breadth (25%) | Autonomy & Reliability (25%) | Cost to Deploy (20%) | Governance & Fit (15%) | Setup Effort (15%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Anthropic Claude | General agent that topped GDPval; Cowork + Code | 9 - led GDPval 2025 and 2026, broad deliverables | 9 - computer + browser use GA, autonomous Cowork/Code | 8 - Team $20/seat/mo, Code included | 8 - RBAC, spend limits since Cowork GA | 7 - powerful, developer-leaning | 8.4 |
| 2 | OpenAI ChatGPT + AgentKit | General agent + drag-drop agent builder | 9 - widest knowledge-work coverage, GPT-6 Astra | 8 - Agent mode + AgentKit, still needs review | 8 - Business $20/user/mo, AgentKit at API rates | 8 - Enterprise controls, Connector Registry | 8 - low for chat, moderate for AgentKit | 8.3 |
| 3 | Google Gemini Enterprise | Workspace-wide agents on one platform | 8 - broad, Workspace + custom agents | 7 - Agent Platform launched Apr 2026, maturing | 8 - Business $21/seat/mo | 8 - deep Google Cloud + Workspace fit | 7 - platform config required | 7.6 |
| 4 | Microsoft Copilot Studio | Office-native agents + computer use | 8 - broad across M365, computer-using agents GA | 7 - governance-first, still largely assistive | 7 - M365 Copilot $30/user/mo | 9 - deepest enterprise stack (Entra, Purview) | 7 - studio build effort | 7.6 |
| 5 | O-mega | Builds and runs a whole company from one chat | 8 - end-to-end business, narrower per single task | 7 - high autonomy, shorter public track record | 7 - one conversation, no code | 6 - smaller vendor, less governance history | 9 - non-technical, lowest setup | 7.4 |
| 6 | Lindy | Low-code builder for custom agent workflows | 7 - many workflows, you assemble them | 7 - solid automation, credit-metered | 8 - $29.99 to $199.99/mo, transparent | 6 - SMB-oriented controls | 7 - low-code builder | 7.1 |
| 7 | Salesforce Agentforce | Production agents bound to CRM and support | 6 - powerful but CRM/customer-process bound | 8 - real deployments, $1.2B ARR | 6 - $2/conversation adds up for general work | 9 - deep CRM integration, trust layer | 6 - needs Salesforce + config | 7.0 |
The criteria are chosen from first principles about the actual question, "can this do a knowledge worker's job," not "is this a good product." Task Breadth (25%) asks how wide a range of real job tasks it can take on rather than one narrow lane. Autonomy & Reliability (25%) asks whether it runs multi-step work at acceptable reliability or needs constant supervision. Cost to Deploy (20%) is the real price to put it on work. Governance & Fit (15%) covers security, data, and how it slots into an organization. Setup Effort (15%) rewards getting to value without a technical team, which matters most to the non-technical reader. Salesforce Agentforce lands last here despite being a commercial success, because on the specific axis of "do a general worker's job" it is deliberately narrow and priced per conversation; its 9 for governance and its $1.2B ARR show where it genuinely wins - Salesforce. The scores are directional, and every one is defended in the cell so you can disagree with the weighting rather than the data.
1. What GDPval Actually Measures (and Why It Broke the Mold)
For most of the last decade, the benchmarks that ranked AI models measured things no employer ever pays for. MMLU tested multiple-choice trivia. GPQA tested graduate science questions. These are proxies for intelligence, and they were useful when models were bad, but they have a fatal flaw for the question in front of us: nobody has ever been fired because a coworker could not answer a physics multiple-choice question, and nobody has ever been promoted because they could. The benchmarks measured recall and reasoning in the abstract, and the abstract is not where jobs live. A job is a stream of messy, specified, consequential deliverables, and until late 2025 there was no serious public measurement of whether a model could produce them.
GDPval is OpenAI's attempt to close exactly that gap, and its design choices are the most interesting thing about it. Rather than asking a model a question, GDPval hands it a real work request drawn from a real occupation, complete with reference files, and asks for the actual deliverable a professional would produce: a financial model, a legal brief, a nursing care plan, a marketing deck, a set of engineering drawings - arXiv. The tasks were not written by AI researchers guessing at what lawyers do. They were written and vetted by professionals averaging 14 years of experience in each field, with 30 tasks per occupation, which is how you arrive at the 1,320-task full set - arXiv. The benchmark covers the top 9 sectors of U.S. GDP, from Professional and Scientific Services to Health Care, Manufacturing, Finance, Government, and Retail, and the occupations inside them collectively account for roughly $3 trillion in annual wages - Marketing AI Institute.
What makes those deliverables gradeable is also what makes GDPval hard to game. The tasks span the real output formats of professional work: spreadsheets and financial models, slide decks, legal briefs and contracts, engineering diagrams and CAD files, clinical documents, and long-form written analysis. Each is a format where a trained professional can look at two versions and say which is better, and why, which is precisely the property a knowledge-recall benchmark lacks. A multiple-choice question has one defensible answer and no craft. A quarterly board deck has a hundred small judgment calls about emphasis, structure, and what to leave out, and it is the accumulation of those calls that separates competent work from expert work. By insisting on the deliverable rather than the answer, GDPval measures craft under specification, which is a far better proxy for economic value than anything that came before it, and a much harder thing to fake with a confident tone.
The image below is one real GDPval deliverable, an engineering design task, and it makes the point better than any description: this is not a chatbot prompt, it is a piece of professional work with a right and a wrong answer that a trained eye can judge.
Why this matters is straightforward: a benchmark that measures deliverables is the first one that can be argued about at a dinner table. When a model scores well on GDPval, the claim is no longer "it is smart," it is "it produced work a professional in your field could not reliably tell apart from a human's." That is a claim about your job, stated in the currency of your job, which is why GDPval landed with more force than any leaderboard before it. How to read the rest of this guide follows from that framing: treat every number as a statement about a deliverable, then ask what the deliverable leaves out, because the gap between the deliverable and the job is where your answer actually lives. For a broader tour of how these capabilities get benchmarked, our guide to AI computer-use benchmarks in 2026 maps the wider testing landscape.
2. The Headline Number: 47.6% and the Year It Took to Get There
The single most quoted GDPval result is that Claude Opus 4.1 reached a 47.6% "win-or-tie" rate against human experts on the open-sourced gold subset, meaning that in blind pairwise judging, its deliverable was rated better than or as good as the human's on 47.6% of tasks - arXiv. Read one way, that is astonishing: a model matched or beat a 14-year professional on nearly half of real work samples. Read another way, it is a loss: the human deliverable was still preferred on the majority, roughly 52%, and "as good as" is doing quiet work inside that 47.6% because a tie is not a win. Both readings are true, and holding them at once is the discipline this entire topic demands.
The word buried inside the metric deserves scrutiny, because a win and a tie are not the same event. A tie means a professional could not confidently pick the human's work over the model's, which is a real and unsettling result, but it is not the model being better; it is the model being indistinguishable on that sample. Strip the ties out and the pure win rate, the share of tasks where the expert actively preferred the machine, is meaningfully lower than 47.6%. This is why careful readers of GDPval landed on the phrase near parity rather than superiority: on the balance of the evidence, the best 2025 model reached the zone where its output was often hard to tell from an expert's, while still losing outright more often than it won. Near parity on a one-shot deliverable is an extraordinary technical achievement and a much narrower claim than the headlines made it.
The field behind the leader tells you more than the leader does. GPT-5 scored 39.0%, o3 35.2%, and o4-mini 29.1%, while the older GPT-4o managed just 12.5% - arXiv. Google's Gemini 2.5 Pro and xAI's Grok 4 landed lower still, around 25% and 24% respectively in the paper's reported figures. The models the paper tested were a specific 2025 vintage: GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4 - The Decoder. The chart below shows the full field, and the shape of it is the point.
The trajectory hidden in that chart is the reason the benchmark scared people. From GPT-4o to GPT-5, GDPval performance more than tripled in about a year, from 12.5% to 39.0% - arXiv. That is not the flat line of a maturing technology approaching a ceiling; it is a steep climb caught mid-ascent, and the paper's own summary is that frontier performance is "improving roughly linearly over time" with the best models "approaching industry experts in deliverable quality" - arXiv. How to apply this: the headline number matters less than the slope. A snapshot of 47.6% tells you where one model stood in mid-2025; the tripling tells you that whatever number is current when you read this is already stale, and probably higher. That is the uncomfortable core of the "can AI do my job" question, and it is why the rest of this guide spends as much time on the benchmark's blind spots as on its scores. For how the frontier models themselves now rank on agent work, our best LLM for AI agents, September 2026 ranking tracks the current order.
3. The Machinery: How 1,320 Tasks Get Graded Blind
A benchmark is only as trustworthy as its grader, and this is where a lot of AI evaluation quietly falls apart. If you let a model grade its own homework, or if you grade "quality" with a rubric a model wrote, you get numbers that flatter the machine. GDPval's designers understood this, and their grading method is the reason the 47.6% figure carries weight rather than marketing gloss. Each deliverable is judged by an experienced professional in the relevant occupation, shown the original request and reference files, and asked to rank two or more unlabeled submissions without knowing which came from a human and which from a model - arXiv. The comparison is blind and pairwise, producing a clean verdict of win, tie, or loss against the human's own work on the same task.
That human-graded core is expensive and slow, so GDPval also ships an automated grader, and its calibration is more reassuring than you might expect. The automated grader agrees with human expert graders 66% of the time, which sounds middling until you learn that human experts agree with each other only 71% of the time - arXiv. In other words, the machine grader is within five points of the irreducible disagreement between two qualified humans looking at the same professional deliverable. Real work does not have a single correct answer, and that inter-rater gap is a feature of professional judgment, not a flaw in the benchmark. The multi-stage review pipeline that produced and validated these tasks is shown below.
The reason this machinery matters to your job is that it removes the easiest objection. You cannot dismiss GDPval as "AI grading AI" or "a rigged rubric," because the load-bearing results come from professionals in the field comparing work blind, and the automated grader is validated against them. When someone tells you the benchmark is hype, the honest rebuttal is not that the grading is wrong; it is that the tasks themselves are one-shot and stripped of context, which is a real and serious limitation covered in section 7. First, though, the most misunderstood claim in the entire GDPval story deserves its own section, because it is the one that turns a benchmark result into a business case, and the business case is weaker than the number implies.
4. The 100x Illusion: Cost, Speed, and the Human Who Never Left
The statistic that traveled fastest from GDPval was not the win rate. It was the claim that frontier models complete these tasks "about 100 times faster and 100 times cheaper" than human experts - The Decoder. That single line launched a thousand LinkedIn posts about the end of professional work, and it is technically true, which is what makes it so misleading. The 100x figure counts only inference time and API billing. It measures the moment the model spends producing a draft, priced at pennies of compute, against the hours a human spends producing the same thing at a professional wage. On that narrow comparison, a model that drafts a consulting deck in four minutes for a few dollars really is roughly two orders of magnitude faster and cheaper than an analyst who takes most of a day.
The problem is that the narrow comparison is not a job. The paper's own numbers make this concrete. A human expert took an average of 404 minutes and cost $361 per task; the gold subset's tasks ran even longer, averaging 9.49 hours of expert time. But the moment you insert the step every real workplace requires, a human reviewing and fixing the model's output, the advantage collapses. Under the paper's realistic "let the model try, then a human corrects it" workflow, GPT-5 came out roughly 1.1x to 1.4x faster and 1.2x to 1.6x cheaper than the unaided expert, not 100x - arXiv. The expert review alone still took about 109 minutes and $86. The gap between 100x and 1.4x is not a rounding error. It is the entire cost of the human who checks the work, catches the hallucinated figure, and signs their name to the result.
Put concrete numbers on it. Suppose the task is a market-analysis deck that a human analyst delivers in the benchmark's average of about seven hours. The model drafts a version in perhaps four minutes of inference for a few dollars of API cost, which is the 100x story in a sentence. But that draft is not shippable: a senior analyst still spends the better part of two hours reading it against the source data, correcting a mislabeled chart, rewriting the recommendation the model hedged, and confirming the numbers are real before their name goes on it. The paper prices that review at roughly $86 and 109 minutes, and it is not optional, because the one figure the model invented is the one that gets quoted in a meeting. The net result is a genuinely faster, cheaper deck, but the saving came from compressing the analyst's day, not eliminating the analyst, and the analyst's scarcest contribution, being accountable for the number, was never automatable in the first place.
This is the first-principles heart of the "can AI do my job" question, so it is worth stating plainly. The model can produce the deliverable for roughly 1% of the cost. The job is the deliverable plus the oversight, and the oversight is now most of the remaining cost and all of the accountability. GDPval measures the cheap part and leaves the expensive part to you. That does not make the technology unimportant; a durable 1.4x productivity gain across knowledge work would be one of the largest economic events in living memory. But it is a very different claim from "your job costs 100x too much," and the people making hiring decisions have started to notice the difference, which is exactly why so many deployments stall. We traced that stall in depth in our analysis of why most AI agent pilots never scale, and it rhymes with everything in this section.
5. GDPval in 2026: The Live Scoreboard and a Human Baseline of 1000
The original GDPval paper is a snapshot from mid-2025, and in a field that tripled a benchmark score in a year, a snapshot ages fast. The most useful 2026 artifact is the GDPval-AA v2 leaderboard run by Artificial Analysis, which takes the 220-task gold subset, gives each model shell access and a web browser inside an agentic harness, and scores the resulting deliverables through blind pairwise comparison - Artificial Analysis. The crucial design choice is the scale: rather than a raw win rate, it uses Elo ratings with the human expert anchored at 1000, so you can read at a glance how far above or below a human baseline any model sits.
As of mid-September 2026, the top of that board sits far above the human anchor. Claude Fable 5.1 leads at roughly 1745 Elo, with Claude Opus 5 close behind near 1735, and the leaderboard updates continuously as new models arrive - Artificial Analysis. One reason the 2026 board sits so far above the 2025 paper is not just better models but a better harness. The original paper tested models largely one-shot, and it explicitly found that more reasoning effort, more task context, and more scaffolding all raise GDPval scores - arXiv. The Artificial Analysis run leans into exactly that, giving each model shell access, a browser, and an agentic loop to research, draft, check, and revise before submitting. That is a fairer picture of how the tools are actually deployed, since nobody ships a raw first draft to a client, but it also means the board measures the model plus its scaffolding, not the model alone. An Elo gap of that size against a 1000 baseline implies the leading model's deliverable wins the blind comparison the large majority of the time. Taken at face value, the frontier did not just approach the human expert between 2025 and 2026; it passed the anchor and kept climbing. Anthropic's lead is a genuine throughline here, since Claude Opus 4.1 topped the original human-graded paper and Claude Fable 5.1 tops the 2026 board.
There is a caveat that matters as much as the number, and skipping it would be dishonest. The 2026 board is scored by an LLM judge, not by the human expert panel that graded the original paper. An automated judge is fast and consistent, but LLM judges are known to favor the fluent, well-formatted, confident output that frontier models are specifically tuned to produce, which can inflate scores relative to a skeptical human professional weighing whether they would actually ship the work. The right way to hold both facts is this: on the original, human-graded benchmark the best 2025 model matched experts under half the time; on the 2026, machine-judged board the best models sit well above the human anchor. The truth about your job is somewhere in the tension between those two methods, not at either extreme. For the current standing of these specific models on agent workloads, see our head-to-head on GPT-6 Astra versus Claude Fable 5.1 for agents, and for the wider agent-capability picture our agentic computer use guide goes deeper on what "shell access and a browser" actually buys.
6. The Occupation Map: 44 Jobs, 9 Sectors, and Where the Losses Cluster
Averages hide the thing you actually care about, which is not "can AI do a job" but "can AI do mine." GDPval is granular enough to answer that better than any survey, because it reports performance by occupation, by sector, and by task type. The 44 occupations span the recognizable middle of the professional economy: software developers, lawyers, financial analysts, registered nurses, mechanical engineers, editors, project managers, and more, chosen because they sit in the highest-GDP sectors and because their core deliverables can be produced and judged on a computer. The full map of occupations is below, and it is worth studying to locate yourself on it.
Two patterns in the results tell you where exposure is highest. First, models came closest to expert parity in Government, Retail Trade, and Wholesale Trade, and lagged furthest in the sectors where deliverables depend on dense, proprietary, or safety-critical context - arXiv. Second, and more useful for self-assessment, win rates were highest on shorter tasks of zero to two hours and declined steadily as tasks got longer. A task that takes a professional twenty minutes is far more exposed than one that takes a professional two days, because the long task carries accumulated context, judgment, and coordination that a one-shot prompt cannot hold. There is also a deliverable-type signal: Claude led on nearly every format including slides and spreadsheets, while GPT-5 led specifically on pure text, and each model's edge traced to its strengths, with Claude stronger on aesthetics and formatting and GPT-5 stronger on factual accuracy.
It helps to decompose a few named occupations the way GDPval implicitly does. A software developer's week contains highly exposed tasks, writing a standard function, fixing a well-described bug, generating a test suite, and far less exposed ones, deciding what to build, negotiating a schema change across teams, and owning an outage at 2 a.m. A lawyer's exposed tasks include first-draft clauses, document review, and case summaries, while the protected core is client counsel under uncertainty, courtroom judgment, and bearing professional liability for the advice. A registered nurse sits almost entirely outside GDPval's reach, because the deliverable is care given to a body in a room, and the benchmark explicitly excludes physical and in-person work. The pattern that emerges is that occupations are not uniformly exposed; they are exposed in their specified, documentary middle and protected at their two ends, the strategic front where the work is defined and the accountable back where it is owned. Two people with the same title can sit on opposite sides of that line depending on how much of their week is middle.
What to do with this map depends on the shape of your work, not its title. The exposed core of almost any knowledge job is the set of short, well-specified, self-contained deliverables: the standard memo, the routine analysis, the first-draft deck, the boilerplate contract clause, the status report. The protected core is the opposite: long-horizon work laced with context, stakes, and other people. Almost every real job is a blend, which is why "will AI take my job" is the wrong question and "which parts of my job is AI about to absorb, and what is left when it does" is the right one. Financial services is a useful worked example of that blend in practice, and we mapped it in how the financial sector automates with AI agents. The framework for deciding which side of the line a given task falls on comes in section 12.
7. Benchmark Versus Job: The Five Things GDPval Cannot See
GDPval is the best measurement we have, and it still measures far less than a job. Its designers are refreshingly honest about this, and their disclaimers are more valuable than most of the coverage the benchmark received. The single most important line in the paper is that "tasks are precisely-specified and one-shot, not interactive" - arXiv. Real work is almost never precisely specified and almost never one-shot. It arrives ambiguous, changes halfway through, and gets built across a dozen exchanges with people who themselves do not fully know what they want. A benchmark that hands the model a clean, complete brief has already done the hardest and most human part of the job before the clock starts.
The exclusions compound from there. GDPval's initial version covers only knowledge work performable on a computer, and explicitly excludes manual and physical labor, tasks requiring extensive tacit knowledge, access to personally identifiable information, proprietary internal tools, and person-to-person communication - arXiv. Each exclusion is a category of work the benchmark is silent about, and several of them are where the actual value of many jobs lives. The deepest gap is philosophical: GDPval scores a deliverable a grader prefers, not an accountable outcome in the world. Professional work has an iceberg quality, where the visible artifact is a fraction of the value, and the rest is negotiation, judgment under uncertainty, relationship, and the willingness to be wrong and answer for it - Snorkel AI.
Accountability is the exclusion that matters most, because it is invisible until something breaks. Imagine the model builds a flawless-looking financial model with one transposed assumption buried in a hidden cell. On GDPval, a grader comparing it against a human version might well prefer it, because it looks cleaner. In a real firm, that model goes into a board deck, the board approves a number based on it, and three weeks later the error surfaces. The question the workplace asks next has no benchmark equivalent: who is responsible. A model cannot be fired, sued, or held to a fiduciary duty, so the accountability does not vanish when the deliverable is automated; it concentrates onto whichever human signed off. That is why so much automated work still routes through a person at the end, and why the reviewer's role, the one that looks like a bottleneck, is often the part that was never really about producing the deliverable at all.
The diagram below draws the gap explicitly.
The practical lesson is not that the benchmark is wrong but that it is narrow in a knowable way, and the narrowness maps directly onto job security. The more your work looks like a clean one-shot deliverable, the more GDPval is measuring you. The more it looks like ambiguous, iterative, accountable, relational work, the more GDPval is measuring a fraction of you and calling it the whole. This is why two people with the same job title can face wildly different exposure: one spends the day turning specified inputs into standard outputs, and the other spends it absorbing ambiguity and owning consequences. The benchmark can only see the first person. Understanding that difference is the beginning of a real answer, and it is reinforced by what the independent, non-OpenAI data shows about the labor market, which is the subject of the next section.
8. What the Independent Data Says About Your Job
A benchmark from a model vendor, however well built, is an interested witness. The stronger evidence about your job comes from independent researchers measuring different things, and three sources matter most. The first is METR, which measures not deliverable quality but task length: the duration of task an AI can complete at a given reliability. METR's original finding was that the length of task a model can handle at 50% reliability doubled roughly every seven months from 2019 through 2025 - METR. Its January 2026 update recomputed and accelerated that trend, putting the doubling time at about 196 days overall and 88 days when measured only since 2024, with Claude Opus 4.5 reaching a 50% time horizon of 320 minutes - METR. Task length is the exact axis GDPval flagged as the boundary of exposure, and it is the axis that is moving fastest.
It is worth translating METR's numbers into job terms, because a time horizon sounds abstract until you map it onto a workday. A 50% time horizon of 320 minutes means a leading model can now complete, at even odds, a task that would take a skilled human more than five hours of focused work, roughly a full morning, before it loses the thread. Two years earlier that figure was measured in minutes. The reliability caveat matters, since 50% is a coin flip and no one ships coin-flip work unreviewed, but the direction is the whole point: the length of coherent, multi-step work a model can hold together is doubling on the order of months, and long-horizon work was precisely the protected end of the occupation map in section 6. If that doubling continues, the boundary between exposed and protected tasks does not hold still; it migrates upward into work that felt safe a year ago.
The second source measures how people actually use these systems. Anthropic's Economic Index, drawn from real Claude traffic, found that consumer use of Claude.ai splits 52% augmentation and 45% automation, while first-party enterprise API use runs closer to 75% automation - Anthropic. The single most common task across the board is "modifying software to correct errors," and computing and mathematical tasks dominate usage. The consumer split is the important number, because it says that when individuals reach for AI, they mostly use it to augment their own work rather than to replace someone. The chart below shows that split.
The third source is the one that should get your attention, because it measures jobs, not usage. Stanford's "Canaries in the Coal Mine" study, using ADP payroll data, found a 13% relative decline in employment for workers aged 22 to 25 in the most AI-exposed occupations while older and less-exposed workers stayed stable, and its August 2026 update widened that gap to about 19% - Stanford Digital Economy Lab. The adjustment showed up as reduced hiring rather than layoffs or wage cuts, which is how labor markets usually absorb this kind of shock. This does not mean older workers are safe; it means the first-order effect landed on the entry-level rungs, where the work is most specified and most learnable, and where GDPval scores highest.
Against all three, there is a serious counter-signal that has to be reported. A Yale Budget Lab analysis found "no AI jobs apocalypse" through late 2025: the share of workers in high, medium, and low AI-exposure jobs stayed remarkably steady, with no clear break tied to AI-exposure measures at the aggregate level - Brookings. Both things are true at once. The aggregate labor market has not cracked, and a specific, young, exposed cohort is already feeling it. The honest reading is that the effect is real, concentrated, and early, not broad and finished, and the New York Fed's data agrees, with recent-graduate unemployment elevated near 5.6% and underemployment at 42% in 2026 - New York Fed. We collected the wider workforce picture in the honest truth about the impact of AI on the workforce, which sits alongside this section.
9. Replacement in the Wild: Klarna, Salesforce, Amazon, IBM
Studies describe the shape of the change; corporate actions show its edge. The most cited early case is Klarna, whose OpenAI-built assistant, by the company's own account, handled 2.3 million conversations in its first month, did the work equivalent of 700 full-time agents, and covered two-thirds of customer service chats - Klarna. That is a genuine automation result, and it is also a cautionary tale, because Klarna later rehired human agents after service quality complaints, a detail the triumphant version of the story usually omits. The lesson is not that the automation failed; it is that the last mile of customer trust turned out to be part of the job, and the benchmark deliverable was not the whole job.
The larger and less reversible signal comes from companies restructuring around the technology. Salesforce cut roughly 4,000 customer support roles, reducing that headcount from about 9,000 to 5,000, with CEO Marc Benioff saying plainly that AI agents now handle about half of customer interactions and that "I need less heads" - Fortune. Amazon's Andy Jassy told staff in a memo that AI "will reduce our total corporate workforce" over the coming years as efficiency gains compound - Amazon. These are not pilots; they are permanent shifts in how large employers size their teams, and they cluster exactly where GDPval predicts, in specified, high-volume, computer-based work.
The counter-case in the same set of companies is the most instructive part. IBM automated 94% of routine HR tasks with its AskHR agent and replaced several hundred HR workers, yet its total employment went up, because the savings were redeployed into roles the automation could not touch - Entrepreneur. Duolingo's "AI-first" memo similarly promised to stop using contractors for work AI can handle while continuing to hire where humans add value - The Register. The pattern across all four is consistent with the data in section 8: AI is absorbing tasks and reshaping headcount at the exposed edge, sometimes shrinking teams and sometimes just changing what the team does, while total demand for labor has not collapsed.
The fork between IBM and Salesforce is the decision every organization now faces, and it is a choice, not a law. Faced with the same automated efficiency, one firm banks it as smaller headcount and the other redeploys it into work the automation cannot do, and both are rational under different strategies. The redeploy path bets that freed capacity converts into more output, more products, more customers served, which is the historical pattern of every prior automation wave, from the spreadsheet that did not end accounting to the ATM that did not end bank tellers for decades. The cut path bets that demand is fixed and the only win is lower cost. Which bet is right depends on whether the organization has more valuable work its people could be doing, and most do, which is why the aggregate labor market has bent rather than broken even as specific functions shrink.
The chart below shows how the exposure concentrated on the youngest workers.
How to apply this if it is your team: watch the boundary between "handled" and "owned." Klarna automated the handling and got burned on the ownership. The roles most exposed are the ones defined by handling volume; the roles most durable are the ones defined by owning outcomes. That distinction is also the difference between a mid-market company that automates a function and one that merely adds a chatbot, which we broke down in automation in the mid-market: hire or automate.
10. The Deployment Divide: Why 95% of Pilots Show No Profit
If frontier models can match experts on half of one-shot tasks, why has the corporate world not been transformed overnight? The answer is the single most important number in the applied AI economy, and it is not a capability number. Roughly 95% of enterprise generative-AI pilots delivered no measurable profit-and-loss impact, according to MIT's Project NANDA study on the state of AI in business, measured against an estimated $30 to $40 billion in enterprise spending - Fortune. The reason is not that the models are dumb. NANDA's own diagnosis is organizational: most deployed systems do not retain feedback, adapt to context, or improve over time, so they stall in the gap between an impressive demo and a process that actually moves the numbers.
The shape of that stall is worth seeing up close, because it repeats. A team buys a support agent, and in the demo it resolves a dozen sample tickets flawlessly, which is real: those tickets sit inside the frontier. It goes to production, and within a month the numbers disappoint, not because the model got worse but because production is where the jagged edges live. The agent confidently quotes a wrong refund policy it half-remembered, cannot see the account history that lives in a system nobody connected, and has no memory of the fix a human made to the same problem yesterday, so it repeats the mistake. None of those are intelligence failures; they are integration, context, and memory failures, and every one of them is organizational work the subscription did not include. The 5% of pilots that cross the divide are almost always the ones that treated the model as the easy 20% and budgeted for the unglamorous 80%: the data plumbing, the guardrails, the feedback loop, and the slow mapping of exactly which tasks are safe to hand over.
That gap has a technical name in the research: the jagged frontier. A Harvard and BCG study of 758 consultants found that on tasks inside the AI frontier, access to a strong model boosted quality by around 40%; on tasks outside it, the same consultants using AI were roughly 19 percentage points less likely to reach the correct answer, because the model produced confident, plausible, wrong output they trusted - Harvard AI Institute. The frontier is jagged because it does not follow human intuition about what is hard; a model can ace a task a person finds difficult and fail one a person finds trivial, with no warning label. Deploying AI into a job means learning the exact shape of that jaggedness for that job, which is slow, expensive organizational work, not a subscription.
The more optimistic and equally sourced finding is that where AI does land, it tends to augment rather than replace. A study of 5,179 customer-support agents found AI access raised productivity by 14% on average, with gains up to about 34% for the least-experienced workers and near zero for the most experienced, and the authors' explicit conclusion was that you usually benefit more by augmenting workers than by trying to replace them - Stanford SIEPR. The chart below shows that skew, and it carries a subtle warning for the labor market: if AI most helps novices reach expert-level output, it compresses the value of junior experience, which is precisely the cohort the Stanford employment data found under pressure.
There is even a cooling signal worth respecting. Gartner predicts more than 40% of agentic-AI projects will be canceled by the end of 2027 on cost, unclear value, and inadequate controls, and it warns of "agent washing," where ordinary automation gets rebranded as autonomous agents - MarTech. Independent Census data analyzed by Apollo's Torsten Slok even showed large-firm AI adoption dipping from a mid-2025 peak near 15% back toward 9% as the first wave of enthusiasm met the deployment divide - Sherwood. None of this says the capability is fake. It says the distance from "can" to "does" is measured in quarters of organizational change, which is why the platforms that shorten that distance are where the real leverage sits.
11. How to Actually Put AI on Your Work: Platforms and Pricing
Capability without a place to run it is a demo. The practical question, once you accept that a model can do the exposed parts of your work, is where you deploy it, at what price, and how much of the deployment divide the platform absorbs for you. The market has settled into a few pricing shapes, and understanding them is the difference between a controllable bill and a runaway one. The per-seat model, favored by the general assistants, charges a flat monthly price per user and folds most usage in. The consumption model, favored by the enterprise agent platforms, charges per conversation, per action, or per credit, which aligns cost to value but makes budgeting harder. And the autonomous-builder model charges for outcomes rather than seats. Each fits a different answer to "can AI do my job."
Among general assistants, the pricing is now remarkably uniform. OpenAI's ChatGPT Business is $20 per user per month billed annually - Coworker. Its AgentKit agent-builder, launched at OpenAI's DevDay in October 2025, lets teams assemble multi-step agents that bill at standard API rates rather than a separate fee - TechCrunch. Anthropic's Claude Team runs $20 per seat per month, with Claude Code included in paid plans and Claude Cowork, its autonomous desktop agent, generally available since April 2026 - Anthropic. Google's Gemini Enterprise starts at $21 per seat per month for the Business tier - TechCrunch, and Microsoft 365 Copilot is $30 per user per month with Copilot Studio adding prepaid credit packs at $200 per month for 25,000 credits - Microsoft. The chart below compares entry seat prices, though the real cost depends on usage.
The consumption platforms tell a different story. Salesforce Agentforce charges $2 per conversation, or Flex Credits at $500 per 100,000 credits with a standard action costing 20 credits, about $0.10 - Salesforce. That is efficient for a well-scoped support flow and expensive if you point it at open-ended knowledge work, which is exactly why it scored narrow-but-strong in the table at the top.
It is worth pausing on what each shape optimizes for, because the choice is really a bet about where your job's value sits. The general assistants optimize for breadth and low friction: a single seat that will draft, analyze, code, and summarize across every part of your week, which is why they top the scoring table for the generic knowledge worker. Anthropic's Claude earns its lead on two counts, having topped GDPval in both the 2025 human-graded paper and the 2026 board, and having shipped Claude Cowork, the autonomous desktop agent noted above, with enterprise controls like role-based access and spend limits. Microsoft's advantage is the opposite kind: it wins on governance and fit precisely because Copilot lives inside the Office and identity stack an enterprise already runs, which is worth more to a regulated buyer than raw model capability. Salesforce, meanwhile, is the clearest illustration that commercial success and general-purpose reach are different axes, since Agentforce passed $1.2 billion in ARR while staying deliberately bound to customer-facing processes rather than open-ended knowledge work - Salesforce.
On the builder end sit the autonomous-company platforms, where the unit is not a seat but a business function. Lindy runs $29.99 to $199.99 per month on a credit model - Lindy, and O-mega takes the idea furthest, building and running an entire company, its website, app, billing, content, and admin, from a single conversation, which is a different bet on "can AI do the job" than renting a smarter assistant by the seat. That autonomous-builder category is where the "hire a workforce rather than a tool" thesis lives, and we ranked the field in hire an AI workforce to run your company in 2026.
Which shape you choose should follow the shape of the work, not the marketing. If your exposed tasks are the specified, short deliverables from section 6, a $20 general assistant captures most of the value with the least setup, and the honest advice is to start there before anything heavier. If the work is a repeatable, high-volume process, a consumption platform that meters to value is worth the integration cost. And if the goal is to stand up a function or a whole business with no team behind it, the autonomous builders are the only category that even attempts it. For the enterprise end of that decision, our best enterprise AI agent platforms of 2026, ranked compares governance and cost in more depth, and for the assistant end, ChatGPT Work versus Claude Cowork settles the day-to-day choice. The broader menu of what OpenAI's own agent stack now offers is covered in our OpenAI workspace agents guide.
12. First Principles: Task, Role, and the Price of Expertise
Step back from the numbers and reason from the ground up, because the consensus framing of this question is wrong in a way that matters. The consensus asks "which jobs will AI replace," which treats a job as an atomic unit that either survives or dies. But a job is not atomic. A job is a bundle of tasks, and a task is a mapping from inputs to a deliverable. GDPval measures the mapping. When intelligence becomes cheap, the price of performing any single mapping falls toward the cost of the compute, which is roughly what the 100x figure captures. That is a real and large change. But the value of a job was never the mapping alone; it was the bundling, the judgment about which mapping to run, the accountability for the result, and the relationships that decide whether the result is even wanted.
This reframes the entire question. The correct unit of analysis is not the job and not the worker but the task, and every job is a portfolio of tasks with different exposure. When the cheap-intelligence input arrives, it does not delete jobs; it reprices tasks, driving the specified, self-contained, low-stakes tasks toward zero and leaving the ambiguous, contextual, accountable tasks roughly where they were, or even raising their value because they are now the scarce complement to abundant machine output. This is why the Stanford data found the effect at the entry level: junior roles are, by design, weighted toward the specified tasks that senior workers have already mastered, which is how juniors learn. Cheap intelligence does not just threaten the junior worker's output; it threatens the training ladder that turns juniors into seniors, and that second-order effect is the one worth losing sleep over.
The apprenticeship economics deserve spelling out, because they are how the entry-level effect becomes a mid-career problem. Professions have always subsidized training: firms overpay juniors relative to their immediate output because the junior work, the research memo, the first-pass model, the document review, is also how someone becomes a senior who can do the accountable work. When cheap intelligence absorbs exactly that junior work, the immediate cost saving is obvious, but the pipeline that produced next decade's experts quietly stops running. A firm that automates all its analyst tasks in 2026 may discover in 2032 that it has no partners, because the path from analyst to partner ran directly through the tasks it deleted. This is the strongest argument against treating AI purely as a cost-cutting tool, and it is an argument from self-interest, not sentiment: the organizations that keep humans in the loop on exposed tasks are not being soft, they are protecting the mechanism that turns juniors into the seniors the machine still cannot replace.
The diagram below sketches how to locate any single task on the exposure map.
The practical conclusion is neither the doomer's nor the booster's. The worker who treats AI as a threat to defend against will lose the exposed tasks and keep nothing. The worker who treats AI as a tool to absorb the exposed tasks reclaims the time for the accountable, relational, judgment-heavy work that reprices upward. The same logic runs at the level of the firm: the company that uses cheap intelligence to deliver more outcomes wins over the company that uses it to cut its way to the same outcomes, because the market pays for outcomes, not for the intelligence input. We argued the macro version of this in the agent economy and the economics of digital labor, and it is the frame that makes the outlook in the next section legible rather than frightening.
13. The Outlook: The End-of-2026 Bet and What 2027 Holds
Forecasts in this field age badly, so it is worth being precise about who is predicting what. The most quoted forecast, that "at least one model will match the performance of human experts across many industries before the end of 2026," is often misattributed to OpenAI. It is actually Julian Schrittwieser's, an Anthropic researcher, from his essay on underestimating exponential progress, where he also projected models working autonomously for full eight-hour stretches by mid-2026 and cited both GDPval and METR as his evidence - julian.ac. OpenAI's own public posture around GDPval was notably more cautious, describing the top models as "getting close to expert-level performance" while stressing that current AI is not replacing whole jobs and that real workflows still need human review, iteration, and integration - The Decoder. The gap between the researcher's forecast and the vendor's caution is itself information: the people closest to the frontier are more bullish than the people who have to ship it.
The trend line under the forecast is real regardless of who drew it. The paper's chart of frontier performance over time, reproduced below, shows the roughly linear improvement that makes the end-of-2026 bet plausible even if the exact date slips. Whether the line bends is the whole question, and there are honest reasons to expect it might, from the deployment divide of section 10 to the possibility that the remaining tasks are harder in kind, not just in degree.
The harder-in-kind possibility is the one the trend charts cannot rule out. GDPval's easy wins came on short, specified tasks; what remains are the long, ambiguous, accountable tasks, and it is not obvious that the same scaling that conquered the first conquers the second, because they may demand capabilities, persistent memory, genuine world-grounding, reliable calibration, that do not simply fall out of a bigger model. The counterargument is equally live: every previous ceiling that looked like a difference in kind, coding, competition math, long-context reasoning, turned out to be a difference in degree that scale eventually crossed. Nobody honest knows which pattern holds for the accountable core of professional work. What the evidence does support is a directional bet rather than a dated prophecy: the exposed tasks keep getting absorbed, the frontier keeps migrating into what felt safe, and the safe harbor keeps shrinking even if it never fully closes.
Two forces will shape 2027 more than model quality. The first is cost, which keeps falling: the entire market moved from models to agent runtimes in 2026, with OpenAI shipping full-duplex voice in the API at $0.05 per minute - Testing Catalog, Anthropic moving computer use, browser use, and its Skills and Files APIs to general availability - Anthropic, and Google consolidating everything into its Gemini Enterprise Agent Platform - Google Cloud. Falling cost widens the set of tasks where automation pencils out. The second force is capital and consolidation: Mistral raised a Samsung-led 3 billion euro round, Europe's largest, to stay in the race - Mistral, and SpaceX's record $60 billion acquisition of the maker of Cursor showed how quickly the agent-tooling layer is being bought up - Yahoo Finance. The direction of travel is clear even when the dates are not: agents are getting cheaper, more autonomous, and more embedded in the tools where work already happens. Our view of where that leads is in the future of business as an autonomous agent workforce.
14. Conclusion: So, Can AI Do Your Job?
The honest answer is a decision framework, not a yes or no, because the question was always underspecified. Can AI do the specified, short, self-contained, low-stakes, computer-doable tasks inside your job? On the evidence of GDPval, increasingly yes, and the trend line is steep. The best model already matched or beat human experts on 47.6% of exactly those tasks in 2025 - arXiv, and the 2026 boards put the frontier well above the human anchor. If your job is mostly that kind of task, the exposure is real and it is here now, and the productive response is to become the person who wields the tool rather than the person it routes around.
Can AI do the ambiguous, iterative, accountable, relational, context-heavy work that most senior roles are actually made of? Not yet, and not soon on the current evidence. The 100x cheaper claim collapses to roughly 1.4x once a human has to review and own the output - arXiv, 95% of enterprise pilots still show no profit impact - Fortune, and the jagged frontier means a model that aces one task will confidently botch a neighboring one. The deliverable is cheap; the judgment and the accountability are not, and those are what the market pays a professional for.
The synthesis is that AI is not coming for jobs so much as repricing the tasks inside them, driving the specified work toward zero and making the human parts scarcer and, if you position for it, more valuable. The people at real risk are the ones whose roles are almost entirely exposed tasks, and the cohort feeling it first is the entry level, where the training ladder itself is under strain. The people with the most to gain are the ones who use cheap intelligence to do more of the accountable work, not to do the same work with fewer people. If you want to run that experiment at the level of a whole function or company rather than a single task, autonomous builders like O-mega exist to test exactly that thesis, alongside the general assistants and enterprise platforms scored at the top of this guide. Whichever path you take, the move is the same: find the exposed tasks in your own week, hand them over deliberately, and reinvest the time in the work no benchmark can grade.
This guide reflects the AI landscape and the GDPval evidence as of September 2026. Benchmark results, model names, pricing, and labor data change quickly, so verify current details against the primary sources linked throughout before making decisions.