The insider guide to Salesforce's named Agentforce agents: what Casey, Paige, Carter, Hunter, Marshall, Piper and Fin actually do, what they cost, and where they win or fail.
On September 11, 2026, Salesforce stopped shipping features and started shipping coworkers. The company gave seven of its AI agents human first names, a job title each, and a runtime that lets them work a goal for weeks instead of answering a single question - Salesforce Newsroom. It was timed four days before Dreamforce 2026 (September 15-17 in San Francisco), and it reframed the entire pitch: the product is no longer "AI that answers," it is "AI that does the job."
Here is the problem that framing hides. Salesforce's own benchmark shows leading agents completing only 58% of single-turn CRM tasks and about 35% of multi-turn tasks - arXiv. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027 - Gartner. And a Bloomberg investigation found several flagship Agentforce demos were largely aspirational, with little of the shown functionality live - Bloomberg Law. So the interesting question is not "did Salesforce name its agents." It is whether named, long-horizon agents are a genuine capability shift or a naming exercise stapled to the same reliability ceiling every agent vendor is fighting.
This guide answers that from first principles. It breaks down exactly what each of the seven agents does, the long-horizon runtime that is the real news, the three overlapping pricing models you will actually pay, how Agentforce works under the hood (the Atlas Reasoning Engine, the Trust Layer, MCP and A2A), the $3.6 billion Fin acquisition that reshaped its customer-service story, the nine rival platforms competing for the same budget, and the security and ROI gaps that separate the marketing from the deployment. Every number here is sourced to a primary document.
Contents
- What Salesforce actually launched: from "answer" to "work"
- Meet the seven: Casey, Paige, Carter, Hunter, Marshall, Piper, Fin
- The long-horizon runtime: memory, durable execution, dynamic steering
- Under the hood: the Atlas Reasoning Engine and the trust stack
- What it costs: the three-headed pricing model
- Does it actually work? Proof points, ROI, and the reality check
- The Fin acquisition and the shift to outcome pricing
- The competitive field: nine platforms fighting for the same budget
- Security and governance: ForcedLeak, the Trust Layer, and the control plane
- The market and the money: financials, forecasts, and the model backdrop
- How to decide: a first-principles buyer's framework
The enterprise AI agent platforms, scored
Before the detail, here is the whole field on one scorecard. Any buyer evaluating Agentforce is really choosing among ten platforms that all promise autonomous agents, so this table ranks them on the four or five things that actually decide value in production. The scores are weighted, and each cell carries the evidence behind it, not a bare number. Agentforce leads, but by a narrow margin, and its lowest marks (pricing clarity, raw task accuracy) are exactly where the rest of this guide spends its time.
The criteria are chosen from what an enterprise buyer pays for, not from a generic feature checklist. Agent capability and autonomy (30%) is the range of genuinely autonomous, production-grade work the platform can do. Data and workflow integration (25%) is how deeply the agent is grounded in the systems of record where the work already lives. Pricing transparency and outcome alignment (20%) is whether you can predict the bill and whether it tracks value delivered. Governance, security and observability (15%) is the control plane, testing, and audit surface. Openness and model flexibility (10%) is whether you are locked to one vendor's stack.
| # | Platform | What it is | Capability (30%) | Data integration (25%) | Pricing (20%) | Governance (15%) | Openness (10%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Salesforce Agentforce | 7 named job-ready agents on a CRM system of record | 8 - broadest role set (service to supply chain) + long-horizon runtime; capped by 35% multi-turn accuracy | 9 - Data 360 + native CRM grounding, deepest for its own customers | 6 - three overlapping models, now with $2/resolution outcome pricing and a free tier | 8 - Trust Layer, Command Center, Testing Center, AI Harness; ForcedLeak (9.4) happened | 7 - native MCP + A2A + multi-model Atlas, but heavy data gravity | 7.8 |
| 2 | Microsoft Copilot / Agent 365 | Agents across M365, Dynamics, Copilot Studio | 7 - vast breadth but more DIY, fewer prebuilt named agents | 9 - Microsoft Graph and M365 data gravity | 6 - $30/user Copilot + $0.01/credit Studio + $15 Agent 365 stack up | 8 - Agent 365 registry + Entra Agent ID + Purview | 7 - MCP support, Azure model choice, but stack gravity | 7.5 |
| 3 | Sierra | Outcome-priced CX agents, Bret Taylor's startup | 8 - elite resolution quality, Sierra Horizon long-horizon, 40% of Fortune 50 | 7 - integrates but is not a system of record | 8 - transparent outcome pricing near $1.50/resolution | 7 - strong but less of an enterprise control plane | 7 - model-agnostic, proprietary platform | 7.5 |
| 4 | Intercom / Fin | Packaged CX agent, now Salesforce-owned | 8 - claims 76% end-to-end resolution, Apex model | 6 - customer-support scoped, not broad | 9 - cleanest model at $0.99 per outcome | 6 - solid but support-scoped | 6 - proprietary, now inside Salesforce | 7.2 |
| 5 | Google Gemini Enterprise | The "front door" for agents on Google Cloud | 7 - strong models, agent gallery, newer as named product | 7 - Workspace and connectors, weaker CRM/ERP gravity | 7 - clear $21-$60/user tiers | 7 - Google Cloud security, control plane maturing | 8 - A2A originator, open protocols | 7.1 |
| 6 | SAP Joule | 40+ agents embedded in ERP | 7 - 40+ agents, 2,400 skills, ERP-centric | 8 - deep ERP system-of-record gravity | 6 - AI Units at roughly $0.08-$0.18/action | 7 - enterprise-grade, ERP-scoped | 6 - SAP stack gravity | 7.0 |
| 7 | ServiceNow AI Agents | AI Agent Orchestrator on the Now Platform | 8 - strong in IT and employee workflows | 8 - deep service system of record | 4 - no public dollar pricing, opaque "assists" | 8 - mature workflow governance | 6 - some MCP/A2A, Now gravity | 7.0 |
| 8 | O-mega | Horizontal autonomous-company builder | 7 - broad multi-agent workforce, computer and browser use | 6 - connects across your stack, not a system of record | 7 - usage-based and transparent, smaller scale | 6 - controls exist, not an incumbent control plane | 9 - model-agnostic, not locked to any SaaS suite | 6.8 |
| 9 | Decagon | CX pure-play, $4.5B valuation | 8 - strong autonomous resolution | 6 - support-scoped integration | 6 - per-conversation and per-resolution, less public | 6 - support-scoped | 6 - proprietary | 6.6 |
| 10 | Workday Illuminate | HR and finance agents | 6 - HR and finance agents, newer | 8 - HR and finance system of record | 5 - Flex Credits, hybrid, less transparent | 8 - Agent System of Record + Agent Passport | 6 - Workday stack gravity | 6.6 |
Read the table as a map of trade-offs, not a verdict. The incumbents (Agentforce, Microsoft, ServiceNow, SAP, Workday) win on data gravity: their agents sit on top of the system of record where the work already happens, which is why an Agentforce service agent can close a case and a Joule agent can post to the general ledger. The pure-plays (Sierra, Decagon, Fin) win on pricing honesty and resolution quality, because they charge only when the agent finishes the job and they have nothing to protect but that number. The horizontal builders like O-mega win on openness, running agents across your whole stack rather than inside one vendor's walls. No single row is right for everyone, and the closeness of the top scores (7.8 down to 7.0 covers seven platforms) is the real story: this is a field of narrow leads, not a runaway.
1. What Salesforce actually launched: from "answer" to "work"
The cleanest way to understand September 11 is to notice what changed in the verb. For two years, enterprise AI was sold as an assistant: it summarized, drafted, and answered. Salesforce's new portfolio is sold as a worker: it is framed as "job-ready agents built to take on high-value work," and the company is explicit that the shift is from getting AI to answer questions toward getting AI to do work across business functions - Salesforce Newsroom. That is not a marketing nicety. It is a claim about autonomy over time, and it is the axis on which the whole launch should be judged.
Start from first principles about why a vendor would give an agent a name at all. A name is not a feature; it is a contract of expectation. When you call a piece of software "Einstein" or "Copilot," you signal a tool. When you call it "Hunter, your outbound sales rep," you signal an employee, and employees are judged on outcomes, not clicks. Salesforce leaned all the way in: customers can tailor each agent and even give it their own name, so it "becomes an extension of their brand and workforce" - Salesforce Newsroom. The naming is a deliberate move up the abstraction ladder, from capability (what the software can do) to role (what job it holds). It also raises the stakes, because a named worker that fails 35% of the time reads very differently from a tool that occasionally gives a wrong answer.
The launch bundled three things that are worth separating, because they are not equally new:
- Seven named agents - most are repackaged and hardened versions of existing Service, Sales, Commerce and Employee agents, given roles and defaults.
- The long-horizon runtime - the genuinely new engineering, letting an agent pursue a goal across days and weeks.
- A governance layer - the Trusted Enterprise AI Harness and an AI Control Plane to manage agents across vendors.
The proof-of-scale number Salesforce attached to all of this is its new usage metric, the Agentic Work Unit (one discrete task an agent completes). The company says it has delivered 7 billion Agentic Work Units across Agentforce and Slack, including 3.2 billion in the most recent quarter - Salesforce Newsroom. That framing matters because it moves the scoreboard away from seats sold and toward work performed, which is the same move the naming makes: stop counting licenses, start counting jobs done. Whether the jobs were done well is a separate question this guide returns to in section 6, but the strategic intent is unambiguous, and it is the clearest signal yet of where the entire category is heading.
It is also worth being precise about the platform name, because the market keeps garbling it. "Agentforce 360" is not "Agentforce 3.0." Agentforce 360 is the umbrella brand Salesforce launched at Dreamforce 2025 (October 13, 2025), rebranding its whole stack (Data Cloud became Data 360, Sales Cloud became Agentforce Sales, and so on), and it succeeded the numbered releases that ran through Agentforce 3 in June 2025 - Salesforce. The seven named agents are the commercial expression of that platform, not a new version number. If you read this article's predecessor, our breakdown of Agentforce 2.0, the through-line is visible: each release moved the product from a configurable toolkit toward a set of ready-made workers, and the named agents are the endpoint of that trajectory.
2. Meet the seven: Casey, Paige, Carter, Hunter, Marshall, Piper, Fin
The seven agents are not seven equal products. Six are generally available today, one is in pilot, and each maps to a specific Salesforce cloud and a specific job. Understanding them individually matters because the differences in maturity and design are larger than the shared branding suggests, and a buyer who treats them as a uniform "portfolio" will misprice the risk. Below, each agent is described by the job it holds, the channels it works, its availability, and the customer proof point Salesforce chose to attach to it. Treat the proof points as vendor claims (they are Salesforce-selected and mostly self-reported), useful as a signal of intended use, not as audited results.
The announcement image Salesforce led with makes the framing literal: a roster of purpose-built workers rather than a feature list.
Before the profiles, the shared architecture is worth stating once: every one of these agents runs on the same Atlas Reasoning Engine, grounds in the same Data 360 layer, and is governed by the same Einstein Trust Layer (all covered in section 4). What differs is the topic library, the default actions, and the channels each is wired into. That is why a customer can deploy Casey in weeks: the reasoning and safety machinery is shared, and only the job-specific configuration changes.
Casey, the service agent
Casey is the customer-service agent, resolving issues across voice, SMS, WhatsApp and web chat, including returns and escalations - Salesforce Newsroom. It is generally available now and is the most battle-tested of the seven, because it is essentially the productized version of the Agentforce Help Agent that already runs Salesforce's own support portal. Salesforce cites the travel-tech firm Engine fully resolving 50% of chat inquiries with Casey without a human - Unite.AI. Casey is also the agent that carries Salesforce's boldest pricing experiment, pay-per-resolution, which we unpack in section 5. If you are evaluating one agent to prove the concept, Casey is the lowest-risk starting point precisely because customer service is the domain with the clearest success signal (did the ticket close) and the most training data.
Paige, the employee-experience agent
Paige handles IT and HR requests, resolving them across Slack, portals and the tools employees already use - Salesforce Newsroom. It is generally available and points Agentforce squarely at the internal-service market that ServiceNow has owned for a decade. Salesforce cites Autism Queensland resolving 70% of administrative requests through Paige - Salesforce Newsroom. The strategic logic is that Slack (which Salesforce owns) is where employee questions already land, so an HR or IT agent that lives in Slack has a distribution advantage no standalone tool can match. Paige is the agent most likely to trigger a direct competitive fight, because employee service is exactly where Microsoft (via Copilot in Teams) and ServiceNow (via Now Assist) are strongest.
Carter, the commerce agent
Carter is the shopper agent: it helps customers discover and compare products, answer questions, and convert with in-chat checkout - Salesforce Newsroom. It is generally available and is the clearest example of an agent designed to move revenue rather than deflect cost. Salesforce says retailer Hibbett handled 90% of core shopper journeys with Carter within six weeks of going live - Unite.AI. Commerce is a revealing test case because the agent is not just answering, it is transacting, which means a wrong answer is not a bad experience, it is a lost sale or a fraudulent one. Carter is where the reliability ceiling has the most direct financial consequence, and where the in-chat checkout capability is the genuine differentiator versus a plain product-search bot.
Hunter, the outbound sales agent (the one that matters)
Hunter is the outbound sales agent, and it is the only one of the seven still in pilot, with general availability planned for November 2026 - Salesforce Newsroom. It is also the most important, because it is the first agent to run on the new long-horizon runtime. Hunter works a pipeline from research to outreach, collaborating with human sellers over weeks and months rather than answering a single query - Salesforce Newsroom. Salesforce cites Perk attributing 60% of its sales pipeline to Hunter - Futurum Group. The reason Hunter is the showcase is structural: outbound sales is the canonical long-horizon job. It requires holding context across a multi-week campaign, resuming after interruptions, and adapting to a prospect's responses, which is exactly what the new runtime is built to do. If Hunter works at GA, the runtime is real. If it slips or underperforms, the "long-horizon" claim is the first thing to discount.
Marshall, the supply-chain agent
Marshall orchestrates end-to-end back-office processes, automating manual work with deterministic execution and an audit record - Salesforce Newsroom. It is generally available and is the odd one out, because supply chain and back-office work reward determinism over conversation. Where Casey and Fin are judged on natural-language resolution, Marshall is judged on whether a process ran correctly every time, which is why Salesforce stresses deterministic execution and auditability rather than resolution rates. This is the agent most aligned with classic process automation, and it is where the comparison to traditional robotic process automation and the new agentic BPA is most direct: Marshall is Salesforce's answer to the RPA incumbents, with reasoning bolted onto deterministic steps.
Piper, the inbound pipeline agent
Piper is the inbound pipeline-generation agent, working across websites and inboxes to engage, qualify, and convert inbound leads into pipeline - Salesforce Newsroom. It is generally available, and Salesforce cites Asana driving 4x conversation volume with Piper, at an average deployment time of 45 days - Unite.AI. Piper and Hunter are the two halves of the revenue engine (inbound and outbound), and the fact that Piper is GA while Hunter is still in pilot tells you something real: qualifying an inbound lead is a shorter-horizon task than running a multi-month outbound campaign, so it was easier to ship. The 45-day deployment figure is a useful reality check on the "job-ready" branding: ready does not mean instant, it means weeks of configuration rather than months of custom build.
Fin, the customer agent from a $3.6 billion acquisition
Fin resolves complex customer-experience workflows across every channel, and it is not homegrown: it arrived through Salesforce's $3.6 billion acquisition of Fin (formerly Intercom), which closed on September 10, 2026 - Salesforce Newsroom. Fin is powered by a proprietary model Salesforce and Fin call Apex, purpose-built for customer support, and Salesforce claims it resolves on average 76% of support volume end-to-end - Salesforce Newsroom. The proof point is pointed: Anthropic resolves 79% of its conversations autonomously through Fin - Unite.AI. Fin's presence alongside Casey is initially confusing (two customer-service agents), and section 7 explains why Salesforce keeps both. The short version: Fin is the fast, packaged option; Casey and the Agentforce platform are the deeply customizable one.
Taken together, the seven agents cover the revenue and service surface of a modern enterprise: two on service (Casey, Fin), one on employee experience (Paige), one on commerce (Carter), two on sales (Piper inbound, Hunter outbound), and one on the back office (Marshall). The deliberate gap is anything requiring heavy domain judgment (legal, clinical, financial-advisory), which is where the accuracy ceiling bites hardest and where Salesforce has wisely not planted a named flag. For a buyer, the practical takeaway is that the maturity gradient runs from Casey and Fin (proven service) down to Hunter (unproven long-horizon), and your pilot should start at the proven end.
3. The long-horizon runtime: memory, durable execution, dynamic steering
If you strip the marketing away, the long-horizon runtime is the only genuinely new engineering in the launch, and it is worth understanding precisely because it is also the hardest thing to get right. The problem it solves is the defining weakness of every chatbot-descended agent: they are stateless and short-lived. They answer a question and forget it. They cannot hold a goal across a Tuesday and a Thursday, cannot resume a plan after a system restart, and cannot absorb a correction halfway through a multi-step job. A "sales rep" that forgets the deal between conversations is not a rep. So Salesforce built a runtime that adds persistence, and it is built on three named mechanisms - Salesforce Newsroom.
The first is memory, which preserves context and progress across sessions so that work does not stop when an interaction ends. The second is durable execution, which keeps plans running over time and lets an agent resume or course-correct as conditions change. The third is dynamic steering, which adapts the agent's behavior based on an individual user's feedback and direction - Salesforce Newsroom. Put those together and the claim is that an agent can now pursue goals across days and weeks rather than completing a single task, while also learning new skills, collaborating with other agents, and improving over time.
Reason about why this is genuinely hard, because the difficulty is the whole story. A short-horizon agent fails cheaply: it gives one wrong answer, the user rephrases, and the cost is a moment of friction. A long-horizon agent compounds error. If it holds a wrong assumption in memory, it carries that error across every step for a week. If durable execution resumes a plan that was already off track, it accelerates in the wrong direction. This is why the industry's autonomy is capped not by whether an agent can take an action, but by whether it can take hundreds of actions in sequence without drifting. Carnegie Mellon's TheAgentCompany benchmark found the best model completing only 30.3% of realistic multi-step office tasks autonomously - arXiv. Long-horizon runtimes are the industry's bet that better memory and steering can push that number up, and Hunter is Salesforce's first live test of the thesis.
The dependency worth flagging is that long-horizon autonomy makes grounding and memory hygiene mission-critical. An agent that remembers everything also remembers stale facts, contradictions, and poisoned inputs. This is a large enough design problem that memory has become its own architecture discipline, which we covered in our guide to AI agent memory in 2026. Salesforce's answer is to anchor memory in Data 360 (the governed data layer) rather than in the model's context window, so that what the agent "remembers" is auditable and can be corrected. Whether that is enough to keep a multi-week campaign on the rails is the single most important open question about the entire launch, and it will not be answered until Hunter reaches general availability in November.
Make the failure mode concrete, because it is easy to underrate. Imagine Hunter three weeks into an outbound campaign, having recorded in memory that a prospect's budget cycle closes in Q1. If that fact was wrong (the prospect misspoke, or a note was mis-parsed), durable execution does not just repeat the error, it plans around it: it schedules follow-ups, tailors messaging, and prioritizes the account on a false premise, and every downstream step inherits the mistake. A stateless chatbot cannot make this class of error because it has no plan to corrupt. This is why long-horizon autonomy raises the value of dynamic steering so sharply: the human seller's mid-campaign correction is not a nicety, it is the mechanism that stops a small memory error from compounding into a wasted month. The runtimes that win will be the ones that make their memory legible enough for a human to spot and fix the wrong fact before it propagates, which is a transparency problem as much as an intelligence one.
For a buyer, the practical implication is a change in how you scope a pilot. A short-horizon agent can be evaluated in a day: send it a hundred tickets and count resolutions. A long-horizon agent has to be evaluated over the horizon it claims to work, which for Hunter is weeks. That means the honest evaluation of the runtime cannot happen at launch, cannot happen in a demo, and cannot be inferred from a resolution rate. It requires running the agent on a real multi-week goal and auditing not just the outcome but the trajectory. The vendors that win the long-horizon race will be the ones whose agents fail safely and visibly when they drift, and that is a governance property as much as a capability one, which is why the control plane in section 9 matters as much as the runtime itself.
4. Under the hood: the Atlas Reasoning Engine and the trust stack
To evaluate any of these agents you have to understand the machine they run on, because the reasoning engine and the safety stack are where the real reliability and risk live, not in the agent's name. Agentforce's core is the Atlas Reasoning Engine, which Salesforce calls the brain of Agentforce and which was incubated at Salesforce AI Research - Salesforce. Atlas runs a ReAct-style loop (reason, act, and observe, repeated until a user goal is fulfilled), it grounds every step in enterprise data through retrieval, and it is deliberately prompted to expose its reasoning to reduce hallucination - Salesforce. If that pattern sounds familiar, it is the same reason-act-observe architecture that underpins most serious agent frameworks, which we explain in our primer on multi-agent orchestration.
The mechanics are worth walking through, because they explain both the strengths and the failure modes. Atlas builds a dynamic plan from a grounded prompt, evaluates data, and refines its actions in a loop until the desired outcome is reached. Crucially, it grounds itself using retrieval-augmented generation over structured and unstructured data in Data 360, so the agent's answers are anchored to the customer's own records rather than the model's training data - Salesforce. User inputs are classified and routed to specialized subagents (renamed from "topics" in April 2026), each holding a relevant set of actions and instructions - Salesforce. This topics-actions-instructions model is the configuration surface: an admin defines when an agent should use a topic, what it is allowed to do, and how it should behave, all in natural language.
The following diagram shows how a single request moves through the stack, from the user's message to a governed action and back.
Safety runs through the Einstein Trust Layer, and its request-response journey is a fixed sequence: grounding, then data masking of PII, then prompt defense, then a secure LLM gateway, then zero data retention with third-party model providers, then toxicity detection across five harm categories, then demasking, then an audit trail - Salesforce Trailhead. The zero-retention guarantee rests on contractual agreements with the model providers so that no customer prompt or response is retained by them. This is the layer enterprise buyers most need to interrogate, because it is the difference between an agent that can safely touch CRM data and one that leaks it, a distinction that section 9's discussion of the ForcedLeak vulnerability makes vivid.
Two integration points complete the picture, and both are genuinely open rather than proprietary. First, Agentforce ships a native MCP client, so agents can connect to any Model Context Protocol server without custom code, governed by an admin-curated allowlist to prevent context bloat - Salesforce. Second, Salesforce co-supports the open Agent2Agent (A2A) protocol for cross-vendor agent collaboration, alongside Google and a large group of vendors. If you are weighing which protocol matters for your own stack, we compared them directly in MCP vs A2A in 2026. The short version: MCP lets an agent use tools, and A2A lets agents talk to each other, and Agentforce now speaks both, which is a meaningful hedge against lock-in even though the data still lives in Salesforce.
The observability and lifecycle tooling is what separates a demo from a production deployment, and it is where Agentforce is genuinely ahead of the pure-plays. The Agentforce Testing Center auto-generates hundreds of synthetic test interactions and runs them in parallel, so you can regression-test an agent before it faces a customer - Salesforce. The Command Center, introduced with Agentforce 3, is a unified observability pane built on OpenTelemetry that integrates Datadog, Splunk and Wayfound - Salesforce. On the build side, Agentforce Builder is a conversational studio for designing agents, Agent Script is a human-readable, portable language for the if-then logic and handoffs, and Agentforce Vibes is the enterprise "vibe coding" IDE, re-engineered onto Salesforce's own Coding Agent Platform running the Claude Agent SDK for Anthropic models with Claude Sonnet 4.5 as a default - Mastra. The presence of this tooling is the strongest argument that Agentforce is built for scaled deployment rather than pilots, and it is exactly the surface that the outcome-based pricing in the next section is designed to monetize.
5. What it costs: the three-headed pricing model
Agentforce pricing is the single most confusing thing about the platform, and the confusion is not accidental: Salesforce runs three pricing models at once, and which one you pay depends on the agent, the edition, and when you signed. Getting this right is worth real money, because the difference between the models can be an order of magnitude at scale. The three models, all confirmed on Salesforce's own pricing page, are legacy per-conversation, consumption via Flex Credits, and per-user editions, now joined by a fourth, outcome-based, model for the Help Agent - Salesforce.
Start with the two consumption metrics, because they are the foundation. The legacy per-conversation price is $2 per conversation (a 24-hour session), and it is still listed - Salesforce. The newer and now-default model is Flex Credits, priced at $500 per 100,000 credits (so $0.005 per credit), where a standard agent action consumes 20 credits (about $0.10) and a voice action consumes 30 credits (about $0.15) - Salesforce. Flex Credits were introduced in May 2025 specifically to replace the per-conversation-only model, because per-conversation pricing punished customers whose agents did more work per session - Salesforce.
Layered on top are the per-user editions, and in September 2026 Salesforce simplified a messy tier structure into three bundled editions. Here is the current shape of the per-seat pricing, with the credit allocations that make the comparison meaningful.
| Edition | Price (per user/month) | Included Flex Credits | Replaces |
|---|---|---|---|
| Foundations | $0 | 200,000 (via Enterprise Edition) | Free entry tier |
| Agentforce User License | $5 | none (buy credits separately) | Add-on seat |
| Core | $195 | 500,000 | Enterprise (~$175) |
| Advanced | $395 | 1,000,000 | Unlimited (~$350) |
| Max | $550 | 2,750,000 | Agentforce 1 |
The edition prices and credit bundles come from Salesforce's 2026 simplification announcement - Salesforce, with the Core and Advanced list prices rising roughly 11-13% versus the editions they replace while Max held its price and nearly tripled its included credits - CIO. The strategic read is clear: Salesforce is discounting the credits (the consumption) while holding or raising the seat price, which is a bet that once agents are embedded, consumption grows faster than seat count. That is the same land-and-expand logic that made Data Cloud consumption its fastest-growing line.
The genuinely new idea, and the one that best aligns price with value, is outcome-based pricing for the Casey Help Agent: $2 per resolution, charged only when the agent resolves an issue autonomously from start to finish, with no charge if the customer asks for a human or walks away unhappy - Constellation Research. A resolution is defined rigorously (at least two turns of messages, not abandoned, with acceptable final feedback), and $2 equals 400 Flex Credits, so the outcome model is a repackaging of consumption into a value-aligned unit - Salesforce Ben. This is the most defensible pricing in the entire enterprise agent market, because it inverts the risk: you pay for jobs done, not tokens burned. It is also the model the pure-plays pioneered, which is why Salesforce's move here reads as a competitive response to Sierra and Fin as much as a customer-friendly gesture.
The chart below shows how the outcome-based prices compare across the vendors that publish them, and the spread is instructive.
For a non-technical buyer, the practical guidance is to model your cost on the metric that matches your workload, not on the seat price. If you run high-volume customer service, the outcome model at $2 per resolution is predictable and value-aligned. If you run complex multi-action agents (Marshall orchestrating a back-office process), Flex Credits at $0.10 per action can add up fast, because a single "job" might be dozens of actions. If you have a small pilot, start on Foundations for free and buy credits as you go. The one thing you should not do is assume a per-seat edition price is your total cost, because the consumption meter runs on top of it. For the fuller history of how this pricing evolved from the original flat per-conversation model, our earlier breakdown of Agentforce pricing traces the shift, and it is a useful companion for anyone modelling a budget.
6. Does it actually work? Proof points, ROI, and the reality check
This is the section that decides everything, and it is where you have to hold two true things at once: Agentforce has real, measurable wins, and it also has a documented gap between what is demoed and what is deployed. Reasoning from first principles, the reason both are true is that agent performance is bimodal. On narrow, well-bounded, high-volume tasks (customer-service deflection, FAQ resolution), agents are genuinely good. On open-ended, multi-turn, judgment-heavy tasks, they are not yet reliable. The proof points cluster in the first category, and the failures cluster in the second, so which story you hear depends entirely on which tasks the storyteller picked.
The strongest, best-sourced win is Salesforce's own help portal, because it is the one deployment Salesforce cannot fake (its customers use it daily). Agentforce resolves about 76% of customer queries without a human there, has handled more than 1.7 million conversations (later passing 2 million), and escalates only 5% of inquiries to a human engineer - Salesforce. The site takes more than 60 million visits a year and the agent cut response time by 65% for 90% of users, operating in seven languages that cover 94% of case volume. That is a real, large-scale, self-consumed deployment, and it is the single most credible data point in the entire Agentforce story precisely because Salesforce eats its own cooking here.
The named-customer proof points are weaker evidence (mostly Salesforce-selected and self-reported) but consistent, and their consistency is itself a signal. OpenTable reported 73% case resolution within three weeks, a 40% improvement on its previous bot, across 11,000-plus conversations a week - Salesforce. 1-800Accountant autonomously resolved 70% of chat engagements during the 2025 tax week - Salesforce. Wiley reported a 213% ROI and a 40%-plus lift in self-service efficiency - Salesforce. Pandora's agent handles 45,000 conversations a month at 60% deflection. Read as a set, these say the same thing: in customer service, at high volume, on bounded queries, Agentforce reliably deflects the majority of contacts. If your job is the customer-support agent category, the deflection economics are real.
Now the reality check, because it is essential and it comes from credible, independent sources. Salesforce's own research benchmark, CRMArena-Pro, found leading LLM agents completing only 58% of single-turn CRM tasks and about 35% of multi-turn tasks, with "near-zero inherent confidentiality awareness" by default - arXiv. That is Salesforce's own scientists reporting that agents drop below coin-flip reliability the moment a task requires several dependent steps. Layer on the widely cited MIT report finding 95% of generative-AI pilots failing to reach production (a figure whose narrow methodology critics have challenged) - Fortune, and Gartner's forecast that over 40% of agentic AI projects will be cancelled by end of 2027 due to cost, unclear value, and weak controls - Gartner. The pattern is not that agents do not work; it is that the demo-to-production gap is enormous, a dynamic we dissected in why most AI agent pilots never scale.
The sharpest critique is a Bloomberg investigation published May 22, 2026, "Salesforce Touts AI Promise Over Reality in SaaSpocalypse Fight," which found that several flagship Agentforce demos were largely aspirational. A University of Chicago Medicine promotional video showed patients using Agentforce to refill prescriptions and book appointments, but Bloomberg reported "little of that AI functionality is live," with callers instead getting keypad menus routed to human schedulers - Bloomberg Law. The investigation named Williams-Sonoma and Finnair demos in the same vein. Marc Benioff's defense was that he does not want a customer thinking they bought something that is not yet real, but the reporting is a documented instance of the marketing running ahead of the deployment, and it is exactly why a buyer should discount any demo and insist on a reference customer running the same workload in production.
The most consequential real-world data point is what Salesforce did to its own workforce. Benioff said he cut support headcount from about 9,000 to about 5,000 (roughly 4,000 roles), stating plainly "because I need less heads," with customer interactions now split "50% with agents, 50% with humans" and support costs down about 17% - Fortune. Whatever you make of the framing (it drew significant backlash), it is the most honest signal in the entire story: Salesforce is betting its own operating model on this working. The counterweight is adoption breadth. Analysis of Q1 FY27 suggested that only around 6% of Salesforce's roughly 150,000 customers were on a paid Agentforce plan, and a TD Cowen partner survey characterized adoption as "subdued," with a large share of partners reporting little client interest - Yahoo Finance. The honest synthesis: the technology deflects real volume in customer service today, and it is nowhere near the autonomous-worker reliability its naming implies for complex, multi-step jobs.
7. The Fin acquisition and the shift to outcome pricing
The $3.6 billion Fin acquisition is the clearest evidence of where Salesforce thinks the customer-service market is going, and understanding it explains the otherwise-baffling decision to ship two customer-service agents (Casey and Fin) at the same launch. Salesforce announced the definitive agreement to acquire Fin (formerly Intercom) on June 15, 2026, and closed it on September 10, 2026, months ahead of the guided timeline - Salesforce. The deal brought in a customer base of more than 30,000 companies and Fin's AI team, with founder Eoghan McCabe remaining as CEO of the unit - Salesforce.
Reason about why a company with its own customer-service agent (Casey) would pay $3.6 billion for another one. The answer is not the technology; it is the pricing model and the deployment speed. Fin's commercial model is $0.99 per outcome (a resolution, a procedure handoff, or a disqualification), with a 50-outcome monthly minimum and $9.99 for qualifications, all published transparently on Fin's own page - Fin. That transparent, self-serve, outcome-only model is exactly what Salesforce's enterprise sales motion cannot easily replicate, and it is exactly what the SMB and mid-market want. Salesforce said it directly: Fin's packaged, fast-to-deploy offerings complement Agentforce's customizable enterprise platform - Salesforce. In other words, Fin is the on-ramp and Agentforce is the destination.
There is a hype-filter caveat worth stating clearly, because it applies to both agents. Salesforce claims Fin resolves 76% of support volume end-to-end, but independent analysis of real customer deployments puts typical resolution rates closer to 42-50% - Gleap. Both numbers can be true: the 76% is an average across favorable configurations, and the 42-50% is what a typical customer sees. The lesson for a buyer is the same one that runs through this entire guide: the headline resolution rate is a ceiling, not an expectation, and your own rate will depend on your knowledge base, your query mix, and how aggressively you tune the escalation threshold.
The strategic significance is that Salesforce is now hedged across the two pricing philosophies that are fighting for the category. Through Casey and the Agentforce platform, it sells the customizable, deeply-integrated, per-user-plus-consumption model that enterprises with complex CRM data want. Through Fin, it sells the packaged, pure-outcome, self-serve model that smaller and faster-moving buyers want. That is not redundancy; it is a portfolio that covers the whole market, and it is a direct response to the pure-plays (Sierra, Decagon) whose entire pitch was that the incumbents could not offer clean outcome pricing. Salesforce just bought its way to a credible answer, and the $3.6 billion price tag is a measure of how threatening it considered that gap.
8. The competitive field: nine platforms fighting for the same budget
Agentforce is not being evaluated in a vacuum. Every enterprise buyer weighing it is also hearing from Microsoft, ServiceNow, Google, SAP, Workday, and a set of well-funded pure-plays, and the structural insight is that they are not all the same kind of competitor. The market splits cleanly into two camps, and the camp a vendor belongs to predicts almost everything about its pricing, its strengths, and its lock-in. The diagram below lays out the split.
Microsoft is the most serious threat because it competes on the same data-gravity logic as Salesforce, only anchored in the productivity stack instead of CRM. Microsoft 365 Copilot is $30 per user per month, Copilot Studio sells consumption at $200 per 25,000-credit pack (about $0.01 per credit), and Agent 365, launched at Ignite on November 18, 2025, adds a $15 add-on to govern agents as first-class identities - Microsoft. Where Salesforce grounds agents in Data 360, Microsoft grounds them in the Microsoft Graph, and for a company that lives in Teams, Outlook, and Excel, that gravity is at least as strong as Salesforce's is for a company that lives in the CRM. The Agent 365 control plane and Entra Agent ID also make Microsoft the leader in agent identity, a governance frontier we examined in Okta vs Entra for AI agent identity.
ServiceNow and the other suite incumbents each own a different system of record, and each is racing to put agents on top of it. ServiceNow overhauled its pricing into three AI-native tiers (Foundation, Advanced, Prime) in April 2026 with Now Assist bundled and consumption metered in "assists," though it notably declines to publish dollar pricing - TechTarget. SAP Joule now spans 40-plus agents and 2,400 skills, priced through consumption "AI Units" at roughly $0.08 to $0.18 per action - Redress Compliance. Google renamed Agentspace to Gemini Enterprise and priced it at $21 to $60 per user per month across editions - Workagent. Each of these wins in its home territory (ServiceNow in IT service, SAP in ERP, Workday in HR and finance) and each is weakest outside it, which means the real competitive question for a buyer is rarely "which agent platform is best" but "which system of record does this workload live in."
Workday deserves a closer look because it competes directly with Paige on the employee-experience surface and it has taken the most distinctive governance stance. Its Illuminate agents span HR and finance, priced through a Flex Credits consumption model, and its differentiator is the Agent System of Record, a registry that governs both first- and third-party agents, complete with an Agent Passport for independent testing and continuous monitoring - Workday. That is a pointed bet: where Salesforce and Microsoft race to own the agents, Workday is racing to own the system of record for agents themselves, on the theory that HR should manage AI workers the way it manages human ones. For a company whose most valuable data is its people and payroll, that framing lands, and it is a reminder that the incumbents are not converging on one design, they are each extending their existing gravity into agents.
The pure-plays compete on a completely different axis, and they are winning the pricing-honesty argument. Sierra, Bret Taylor's startup, raised a $950 million Series E at roughly $15.8 billion in May 2026, counts more than 40% of the Fortune 50 as customers, and prices on outcomes at about $1.50 per resolution - TechCrunch. Decagon reached a $4.5 billion valuation in a January 2026 Series D on roughly $100 million of annualized revenue, with a median customer contract around $432,750 - Sacra. These companies have no system of record to defend, so they can charge only when the agent resolves the issue, and their entire pitch is that they resolve at a higher rate than the incumbents' bolt-on agents. That pitch is exactly what forced Salesforce to buy Fin and to launch $2-per-resolution pricing.
There is a third position that neither camp occupies, and it is where the horizontal builders sit. Instead of grounding agents in one vendor's system of record or selling a single outcome-priced skill, platforms like O-mega run a model-agnostic autonomous workforce that spans your whole stack, wiring agents across CRM, email, browser, and internal tools rather than living inside any one suite. The trade-off is real and cuts both ways: you give up the deep native grounding that makes an Agentforce service agent so good inside Salesforce data, and you gain openness and freedom from lock-in, which matters most to buyers who refuse to concentrate their agent strategy inside a single vendor's walls. For an enterprise whose work does not live primarily in one system of record, the horizontal approach (or a deliberate mix of a suite agent plus a horizontal one) is a legitimate third answer, and it is the natural home for anyone who wants to hire an AI workforce that runs across the whole company rather than one department.
The practical decision framework that falls out of this is simple. If most of your work and data already live in one suite, the incumbent that owns that suite has a near-insurmountable grounding advantage, and you should default to it unless its pricing or accuracy disqualifies it. If your primary job is high-volume customer service and you want predictable economics, an outcome-priced pure-play (or Salesforce's Fin) is often the better buy. And if you refuse to concentrate risk in one vendor, a horizontal, model-agnostic platform is the hedge. The one strategy that consistently fails is buying the platform with the best demo, because as section 6 showed, the demo is the least reliable input you have.
9. Security and governance: ForcedLeak, the Trust Layer, and the control plane
Autonomy and access are the same thing pointed in two directions, and that is why security is not a footnote to the agent story, it is the story's hardest constraint. An agent that can resolve a customer's issue can, if manipulated, exfiltrate that customer's data, because both actions use the same permissions. The canonical proof is ForcedLeak, a critical Agentforce vulnerability rated CVSS 9.4, disclosed by Noma Security. It let an attacker exfiltrate CRM data through an indirect prompt injection in the Web-to-Lead flow, by planting malicious instructions in a lead's Description field and routing the leaked data to an expired Salesforce-allowlisted domain the attacker bought for about $5 - The Hacker News. Salesforce patched it by enforcing trusted URLs, but the lesson generalizes: an agent that reads untrusted input and holds real permissions is a new and expanded attack surface, and the defense has to be architectural, not a filter bolted on afterward. We wrote a full treatment of this class of attack in AI agent security and prompt-injection defense.
This is where the Einstein Trust Layer from section 4 earns its place, because it is Salesforce's architectural answer. By masking PII before the prompt reaches the model, defending against injection, enforcing zero data retention with providers, and writing an audit trail, the Trust Layer is designed to make the agent's power auditable and bounded - Salesforce Trailhead. ForcedLeak is proof that the layer is not infallible (the injection reached the agent through an unguarded input path), but the presence of a coherent, inspectable safety stack is a genuine advantage over cobbling agents together from raw model APIs. The right way to read ForcedLeak is not "Agentforce is insecure" but "any agent with CRM access is a target, and you must evaluate the vendor's entire input-validation and permission model, not just its model quality."
The governance story got a significant addition on September 10, 2026, the day before the agent launch, with the Trusted Enterprise AI Harness and an AI Control Plane. The Harness is a six-part governance framework (Trusted Context, Agency, Action, Governance, Security, and Models), and the Control Plane is a single place to see, manage, and control agents and AI across the enterprise, including agents from other vendors, rolling out from early fiscal 2028 - Salesforce. The reason this matters is a structural fact about how enterprises actually run agents: VentureBeat reported that 85% of organizations run two or more agent orchestration platforms simultaneously, and 53% expect their primary control plane by end of 2026 to be hybrid - VentureBeat. Salesforce is betting that whoever governs the multi-vendor agent sprawl owns the most valuable position, even if it does not own every agent.
Reason about why the control plane may be the smartest move in the whole launch. If the future is many agents from many vendors (which the 85% figure suggests), then the scarce resource is not another agent, it is a trustworthy place to govern all of them. Owning the control plane is a way to stay central even as the agents themselves commoditize, which is the same strategic logic Microsoft is pursuing with Agent 365. For a buyer, the implication is that your agent-governance decision may end up more consequential than your agent-vendor decision, because the governance layer is where identity, audit, and kill-switches live. Choosing an agent is choosing a worker; choosing a control plane is choosing the management system for your entire agent workforce, and the second choice is much harder to reverse.
10. The market and the money: financials, forecasts, and the model backdrop
The financial numbers behind Agentforce are large enough to explain why Salesforce is betting the company on it, and specific enough to test against skepticism. In the most recent quarter (Q2 of fiscal 2027, reported August 26, 2026), Salesforce posted $11.3 billion in revenue, up 11%, with Agentforce ARR exceeding $1.5 billion, up more than 240% year over year, and combined Agentforce plus Data 360 ARR nearing $3.9 billion - Salesforce. Benioff framed it directly: "AI is delivering value across every layer of our platform ... with ARR about to cross $4 billion" - Salesforce. The growth trajectory is the real evidence, so it is worth seeing it quarter by quarter.
The adoption curve behind that revenue is equally steep, and it is the number that best captures how fast enterprises moved from curiosity to contract. Salesforce reported cumulative Agentforce deals closed rising from more than 12,500 in mid-2025 to more than 18,500 by late 2025 to more than 29,000 by January 2026 - Salesforce. One honest caveat belongs on these figures: the combined ARR includes roughly $1.1 billion of Informatica Cloud ARR from that acquisition, so part of the jump is inorganic, and the "deals closed" number counts closed deals, not necessarily fully deployed ones. Still, the direction is unambiguous, and it explains the market's reaction: the stock jumped about 22.5% the day after the Q2 earnings beat, though the September 11 agent launch itself moved it only about 3% - 247 Wall St.
Zoom out and the analyst forecasts explain why every enterprise vendor is racing into this category at once. Gartner projects agentic AI software spend reaching $985 billion by 2030 at a 62.7% compound growth rate - Gartner, and separately warns that $234 billion of enterprise application spend is at risk from agentic disruption by 2030 - CIO Dive. IDC forecasts agentic AI exceeding 26% of worldwide IT spending, about $1.3 trillion, by 2029 - IDC, and Bloomberg Intelligence sizes the broader generative-AI market at $2.3 trillion by 2032 - Bloomberg. The chart below places the headline agentic-AI forecasts side by side; note the differing scopes and target years in the subtitle, because these are not measuring the same thing.
The model backdrop matters here too, because agent quality is downstream of model quality, and September 2026 was an unusually busy month for frontier releases. OpenAI shipped GPT-6 Astra on September 3 at premium pricing of $10 and $50 per million tokens - CNBC. Google released Gemini 3.8 Flash around September 2 at an introductory $0.75 and $3.75 per million (doubling on January 1, 2027), with a verified Terminal-Bench 2.1 score of 89.4% (not the 90.8% that circulated on aggregator blogs) - Google DeepMind. Anthropic's current lineup includes Claude Fable 5.1 at $10 and $50 per million with a 75% cache-read cut, alongside Claude Opus 5 - Anthropic. For a buyer, the relevance is indirect but real: Agentforce's Atlas engine can call OpenAI, Anthropic, and Gemini models, so the agent you deploy inherits whichever frontier model you route it to, and the economics of that choice shift monthly. If you want the current cost-per-task picture, our best LLM for AI agents ranking tracks it, and the GPT-6 Astra cost breakdown drills into the premium tier specifically.
The first-principles synthesis of all this money is that the market is pricing agentic AI as a platform shift, and Salesforce is spending like it believes that. It cut 4,000 of its own support roles, paid $3.6 billion for Fin, and rebuilt its pricing around consumption and outcomes, all bets that only pay off if agents move from deflecting FAQs to doing genuine work. The forecasts say the prize is measured in the hundreds of billions to trillions. The benchmarks say the technology is not there yet for complex tasks. Both are true at once, and the gap between them is precisely the risk (and the opportunity) that a buyer is underwriting when they sign an Agentforce contract in 2026.
11. How to decide: a first-principles buyer's framework
Strip away the names and the forecasts, and the buying decision reduces to three structural questions, because everything else is downstream of them. The first is where your work and data actually live, because grounding beats everything. An Agentforce service agent is excellent inside Salesforce data and mediocre outside it, exactly as a Joule agent is excellent inside SAP and a Now Assist agent is excellent inside ServiceNow. If 80% of the work you want to automate already runs on Salesforce, the grounding advantage is so large that Agentforce is the default, and the burden of proof is on any alternative to overcome it. If your work is spread across many systems, no single suite agent has that advantage, and a horizontal or best-of-breed approach becomes rational.
The second question is what kind of task you are automating, because the reliability ceiling is task-shaped. For high-volume, bounded, well-documented tasks (customer-service deflection, FAQ resolution, lead qualification), agents work today, and the evidence in section 6 is strong: 70-plus percent resolution is a realistic, repeatable outcome. For open-ended, multi-step, judgment-heavy tasks, the 35% multi-turn success rate is your warning, and the honest move is to keep a human in the loop and treat the agent as an accelerant, not a replacement. The mistake to avoid is buying the "autonomous worker" story for a task that is actually a "multi-step judgment" job, because that is where pilots die.
The third question is how you will govern and measure it, and this is the one buyers most often skip. Before you deploy, decide what a "resolution" means for you, how you will audit the agent's trajectory (not just its outcome), how you will contain a prompt-injection attempt like ForcedLeak, and which control plane will hold your kill-switch. Salesforce's Testing Center, Command Center, and Trust Layer exist precisely because these are the questions that separate a demo from a deployment, and a vendor that cannot answer them for your workload is selling you the demo. The practical test is simple: ask for a reference customer running your exact workload in production, at your volume, and ask what their real resolution rate is, not the headline one.
Put the three questions together and a decision framework falls out. If your data lives in one suite and your task is bounded, buy the incumbent that owns that suite and start with its most proven agent (for Salesforce, that is Casey or Fin, not Hunter). If your task is high-volume service and you want predictable economics, weigh an outcome-priced option (Fin at $0.99, Sierra near $1.50, or Casey at $2) against the incumbent's consumption model. If you refuse to concentrate your agent strategy in one vendor, add a horizontal, model-agnostic platform like O-mega to the evaluation, and treat openness as a first-class criterion rather than an afterthought. And regardless of which you choose, start narrow, measure the real resolution rate, and expand only where the number holds, because the entire history of this category, from the 95% pilot-failure statistic to the aspirational demos, says that the winners are the teams that scaled what worked rather than the teams that believed what was shown.
The deepest truth about September 11 is that Salesforce did something clever with abstraction. By naming its agents and giving them jobs, it moved the conversation from "can this software do a task" (where the honest answer is "sometimes") to "which job should this worker hold" (where the buyer's imagination fills the gap). That is a powerful sales move and a genuine product improvement at the same time, because the long-horizon runtime is real engineering aimed at the real bottleneck. But the naming does not change the physics: an agent is exactly as reliable as its worst multi-step failure, and no first name fixes that. The enterprises that win with Agentforce in 2026 will be the ones that hire Casey to deflect tickets today, keep a human next to Hunter until the runtime proves itself, and never mistake a well-cast demo for a colleague who has actually done the job.
This guide reflects the enterprise AI agent landscape as of September 15, 2026, the opening day of Dreamforce 2026. Pricing, model availability, and agent capabilities in this category change frequently, so verify current details against primary sources before purchasing.