The insider's guide to the two most valuable AI support startups, the field around them, and how to actually pick one.
In May 2026, Sierra raised $950 million at a $15.8 billion valuation. Three months earlier, in January, its closest rival Decagon had tripled its own valuation to $4.5 billion in under six months - Bloomberg. Two companies, barely three years old, selling the same promise from opposite playbooks, have become the most heavily funded pure plays in enterprise AI outside the foundation model labs themselves.
The promise is deceptively simple: an AI agent that talks to your customers, understands what they want, takes real action on your systems, and resolves the issue without a human ever touching it. Customer support turned out to be the first enterprise job that autonomous AI could actually do at scale, and the money has followed with a speed that looks reckless until you understand why support was always going to be the wedge.
But here is the problem. The category is a fog of self-reported resolution rates, opaque negotiated pricing, and marketing numbers that collapse the moment an independent tester runs the same workload. A vendor that advertises 76% resolution may deliver closer to 40% in your account. A platform priced at "$0.99 per resolution" can quietly cost six figures a year. And the two most-funded names, Sierra and Decagon, are not even the highest-scoring options for most buyers once you weigh what actually matters.
This guide breaks down exactly how a modern support agent works, ranks every major platform on weighted, sourced criteria, dissects Sierra and Decagon in depth, maps the wider field of a dozen serious competitors, decodes the pricing wars, and shows where these agents genuinely succeed, where they fail badly enough to end up in court, and how to decide between building, buying, or running support as one part of a fully autonomous company. Where it helps, we point to deeper dives in our own coverage of the economics of digital labor and what AI is really doing to the workforce.
Before the details, it helps to hear the thesis from the person who set the category's terms. Sierra CEO Bret Taylor, former co-CEO of Salesforce and current chairman of OpenAI's board, laid out the whole argument (agents as the product, outcome-based pricing, support as the entry point) at Sequoia's AI Ascent.
Contents
- The 2026 scorecard: every major platform ranked
- Why customer support became AI's first real job
- The anatomy of a modern support agent
- Sierra: the outcome-priced enterprise agent
- Decagon: the concierge built on operating procedures
- Sierra versus Decagon, head to head
- The wider field: everyone else building support agents
- Pricing models and the economics of a resolution
- Where it works: real deployments and hard numbers
- Where it fails: limits, backlash, and liability
- Build, buy, or run it inside an autonomous company
- Measuring what matters: evals, benchmarks, and the model layer
- The 2026 outlook and a decision framework
1. The 2026 scorecard: every major platform ranked
Before the deep dives, here is the whole field in one view. The table below scores the ten most relevant customer support agent platforms on the criteria that actually determine outcomes for a buyer, each weighted by how much it moves the decision. This is the summary. The rest of the guide is the depth behind each cell, because a number without its reasoning is worthless, and a ranking that hides its assumptions is just an opinion in a costume.
The criteria are built from first principles, not from a vendor feature checklist. What a buyer is really purchasing is a resolved conversation at an acceptable cost, delivered on the channels their customers use, without a multi-quarter implementation and without a billing model that punishes them for succeeding. That reduces to five things worth weighing: how well the agent actually resolves issues, whether the pricing model aligns the vendor's incentive with yours, whether it survives enterprise reality (integrations, security, scale), how far it reaches across voice and other channels, and how fast it goes live.
| Rank | Platform | What it does | Resolution (30%) | Pricing align (20%) | Enterprise (20%) | Voice/omni (15%) | Time-to-value (15%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Intercom Fin | Support agent with transparent per-outcome pricing on any helpdesk | 7.5 - 67% verified avg, 38-53% independent | 9.5 - $0.99 per outcome, published rate card | 8.5 - 6,000+ customers, being bought by Salesforce | 7.0 - chat, email, Fin Voice add-on | 9.5 - live in under an hour, self-serve | 8.3 |
| 2 | Sierra | Enterprise Agent OS with outcome pricing and white-glove build | 8.5 - case studies 70-77%, up to 94%; authored TAU-bench | 8.0 - pure outcome-based, but opaque and negotiated | 9.0 - claims 40%+ of Fortune 50, $15.8B | 8.5 - chat, voice, email, SMS as first-class | 6.0 - vendor-owned build, 3-7 months | 8.1 |
| 3 | Decagon | AI concierge configured through natural-language operating procedures | 8.5 - claims 70-90% autonomous, 93% quality (SR) | 6.5 - per-conversation or per-resolution, opaque | 8.0 - 100+ enterprise adds in 2025, $4.5B | 7.5 - chat, email, voice (ElevenLabs), SMS | 6.0 - heavy upfront engineering | 7.5 |
| 4 | Zendesk | Incumbent helpdesk turned AI Resolution Platform | 6.5 - built on Ultimate and Forethought acquisitions | 7.0 - ~$1.50 per resolution, 72-hour rule | 8.5 - massive install base, deep ecosystem | 7.5 - full omnichannel helpdesk plus voice | 7.0 - instant for existing customers | 7.2 |
| 5 | Parloa | Voice-first agentic platform for enterprise contact centers | 7.0 - agentic voice, blue-chip references | 5.0 - negotiated, $300K+/yr, opaque | 8.0 - $3B valuation, Microsoft and Booking.com | 9.5 - the category's leading voice specialist | 5.5 - enterprise voice integration cycle | 7.0 |
| 6 | Cognigy (NICE) | Enterprise conversational and agentic AI, now inside NICE | 6.5 - mature contact-center automation | 5.0 - enterprise custom, now NICE-bundled | 8.5 - bought by NICE for $955M | 9.0 - strong voice plus omnichannel | 5.5 - enterprise deployment | 6.8 |
| 7 | Salesforce Agentforce | CRM-native agent layered on Service Cloud and Data Cloud | 6.5 - broad agent, no strong published rate | 5.5 - three models, ~$2 per conversation | 9.5 - the incumbent giant, 6,000 paying deals | 7.5 - omnichannel, voice actions | 4.0 - 5-11 months, Data Cloud often required | 6.7 |
| 8 | Ada | Automation-first agent for very high-volume support | 6.5 - claims 70-84%, independents see 30-50% | 5.5 - quote-based, moved off per-resolution | 7.0 - 300K+ conversations, six-figure deals | 6.5 - chat plus voice | 6.0 - $10K-$30K implementation | 6.3 |
| 9 | O-mega | Autonomous company builder; support is one function it runs | 6.0 - not a benchmarked CS specialist | 6.0 - subscription for a whole operation | 5.5 - built for founders and operators | 4.5 - not a dedicated contact-center product | 8.5 - conversational, self-serve setup | 6.1 |
| 10 | Forethought | SupportGPT agents (Solve, Assist, Triage), now Zendesk-owned | 6.5 - per-customer fine-tuned models | 5.0 - opaque, ticket-volume driven | 6.5 - Zendesk-owned, sold standalone | 6.0 - multichannel, voice add-on | 5.5 - 20,000-ticket minimum | 6.0 |
The weights encode a point of view worth stating plainly. Resolution quality carries the most weight at 30% because an agent that cannot solve the problem is worthless at any price, and because this is exactly where vendor marketing and independent measurement diverge most. Pricing alignment and enterprise readiness each take 20% because the second most common way these projects fail is not the AI, it is a bill that balloons or an integration that never ships. Voice and omnichannel and time-to-value round out the last 30%, weighted lower only because they are table stakes rather than differentiators for most text-first support operations.
Two results deserve a flag up front. Intercom Fin ranks first not because it has the smartest agent, but because it is the only major platform with a fully published, self-serve, per-outcome rate card, and that transparency compounds across every other criterion. And Sierra and Decagon, the two valuation darlings, land second and third: excellent agents wrapped in opaque pricing and long, vendor-owned deployments that cost them on the exact axes a mid-market buyer cares about most. The market is not pricing these companies on today's support revenue. It is pricing them on the belief that support is the wedge into every other enterprise agent, which is the argument we unpack next.
2. Why customer support became AI's first real job
To understand why capital is pouring into support agents specifically, start with the structural question rather than the surface one. The surface question is "which support vendor wins." The structural question is: what changes when a competent customer conversation drops from roughly $6 to under $1? Support was always the enterprise function most exposed to cheap intelligence, for three reasons that have nothing to do with any particular startup. It is enormously high volume, it operates in a bounded domain where the correct answer usually exists in a knowledge base, and its output is discretely measurable. A support interaction either resolves the issue or it does not, and that binary is the seed of an entire business model.
That measurability is the key. Most white-collar work produces outputs that are hard to price by the unit, which is why software has historically been sold by the seat. Support breaks that pattern, because a resolved ticket is a countable, valuable event. The moment intelligence became cheap enough to resolve tickets reliably, it became possible to sell the software by the outcome rather than by the seat, and outcome-based pricing is not a marketing gimmick here, it is the natural consequence of a job whose value is finally legible to a billing system. We explore this shift in depth in our look at the economics of digital labor, and it explains why the pricing innovation in this category appeared before it did anywhere else.
The numbers underneath are large enough to justify the frenzy without any hype. The market for AI in customer service was worth about $13.0 billion in 2024 and is forecast to reach roughly $83.9 billion by 2033, a compound growth rate of 23.2% - Grand View Research. Independent forecasters across the broader conversational-AI space cluster in a similar band, with most 2030 estimates landing between $41 billion and $50 billion and a few outliers higher, so the exact figure is contestable but the trajectory is not.
The analyst class has committed hard to the direction of travel. Gartner projects that agentic AI will autonomously resolve 80% of common customer service issues without human intervention by 2029, driving a 30% cut in operational costs - CX Today. The same firm expects that by 2028, 70% of customer service journeys will begin and end with a third-party conversational assistant on mobile. Whether those exact percentages land is less important than the fact that the people advising the Fortune 500 have decided this is where service is going.
Adoption is following the forecasts, though unevenly. Surveys aggregated from vendor and analyst data show AI agent adoption in support organizations rising from roughly 39% in 2025 to 66% in 2026, and about 91% of customer service leaders report pressure to implement AI this year - fin.ai benchmarks. But there is a large and telling gap between piloting and shipping: one 2026 read of enterprise CX teams found 64% had run an agentic pilot while only 27% had a channel fully in production. That two-to-one gap between experiments and deployments is the honest state of the market, and it is where most of the vendor competition is actually being fought. The labor consequences of that shift are real and contested, which is why we treat them separately in our analysis of AI's true impact on the workforce.
The first-principles takeaway is that support is not merely an early adopter, it is the proving ground. Whoever demonstrates reliable, measurable resolution in support earns the reference customers, the real-world evaluation data, and the enterprise distribution to expand into sales, IT service, and every adjacent agentic function. That is why a company doing perhaps $150 million in support revenue can be valued at nearly $16 billion. The bet is not on support. The bet is on support as the door into everything else.
3. The anatomy of a modern support agent
It is tempting to think of a support agent as a chatbot with a better brain, but that mental model will lead you to buy the wrong thing. A 2026-grade agent is a small system of components, and the intelligence of the language model is only one of them. The difference between a demo that dazzles and a deployment that survives contact with real customers lives almost entirely in the parts that are not the model: the retrieval layer that grounds it in your actual policies, the action layer that lets it do things rather than just talk, the guardrails that stop it from inventing a refund policy, and the escalation logic that hands off cleanly when it is out of its depth.
At the center sits a large language model doing the reasoning, and here the news is that the raw capability is no longer the bottleneck. The current flagships (OpenAI's GPT-5.6 family of Sol, Terra, and Luna released in July 2026, Anthropic's Claude Opus 5 and Claude Sonnet 5, and Google's Gemini 3.6 Flash) are all more than capable of holding a coherent, policy-aware support conversation - TechCrunch. Because capability has commoditized, the platforms differentiate on everything wrapped around the model, which is precisely why the incumbent labs now loom as a competitive threat rather than just a supplier.
The grounding layer is retrieval-augmented generation, and it is the single most important determinant of whether an agent tells the truth. Rather than relying on what the model memorized, the agent retrieves the relevant passages from your help center, past tickets, and policy documents at query time, then answers from that evidence. Modern implementations use hybrid retrieval (keyword search fused with dense vector embeddings, then a cross-encoder reranker to sort the candidates), and getting this right typically moves real deflection by double digits. If you want the mechanics rather than the summary, our introduction to RAG covers the retrieval stack that every serious support agent depends on.
The diagram below shows how the pieces fit together, from the customer's first message through retrieval, reasoning, action, and either resolution or a clean handoff to a human.
Two components separate a real agent from a fancy FAQ bot. The first is the action layer, the set of tools the agent can call to actually do something: issue a refund, look up an order, change a subscription, update a record in the CRM. An agent that can only talk deflects questions; an agent that can act resolves problems, and the resolution rate difference between the two is the difference between a cost center and a working product. The second is the guardrail system, layered checks that validate the agent's inputs and outputs, catch prompt-injection attempts, block personally identifiable information, and fact-check the response against the knowledge base before it reaches the customer. Guardrails are why a serious deployment does not repeat the mistakes we catalog in our guide to prompt-injection defense for AI agents.
Around all of this sits the connective tissue that rarely makes the sales deck but decides the deployment. Conversation state lets an agent pause, resume, and remember context across channels so a customer who starts in chat and calls an hour later is not starting over. Escalation logic decides when to hand off, ideally with a confidence gate so the agent routes to a human before it guesses. And observability (traces, evaluations, and increasingly OpenTelemetry-based instrumentation) is what lets a team see why the agent did what it did and improve it, which matters enormously once you are running thousands of conversations a day. Coordinating multiple specialized sub-agents for retrieval, action, and escalation is its own discipline, one we cover in our piece on multi-agent orchestration.
Voice adds a final layer of difficulty that is easy to underestimate. A voice agent must run speech-to-text, then the language model, then text-to-speech, and the entire round trip has to complete fast enough to feel human, which means a latency budget in the low hundreds of milliseconds against a human tolerance of roughly 200 to 300 milliseconds before a pause feels awkward. The 2026 shift is from cascaded pipelines toward native speech-to-speech models like OpenAI's GPT-Realtime-2 and Google's Gemini 3.1 Flash Live, which collapse the steps and cut the delay. This is exactly why voice-native specialists like Parloa command their own valuations: doing voice well is genuinely hard, and the underlying architecture is different enough that it does not come free with a good chat agent. For teams weighing whether to assemble this stack themselves, our insider guide to building AI agents walks through the same components from the builder's side.
4. Sierra: the outcome-priced enterprise agent
Sierra is the company that set the category's frame, and it did so partly by force of pedigree. It was founded in 2023 by Bret Taylor and Clay Bavor, and emerged from stealth on February 13, 2024, alongside a $110 million raise - Fortune. Taylor is one of the most credentialed operators in software: co-creator of Google Maps, former CTO of Facebook, former co-CEO of Salesforce, and the current chairman of OpenAI's board. Bavor spent nearly two decades at Google running Workspace and the AR and VR labs. The pair even chose the name Sierra, a mountain range, deliberately to avoid the cold, metallic robot imagery that clings to AI, a small signal of how carefully the company manages perception.
The funding trajectory since then has been close to vertical. Sierra reached a $4.5 billion valuation in October 2024, jumped to $10 billion in September 2025, and then raised $950 million at $15.8 billion in May 2026, co-led by Tiger Global and GV, leaving it with more than a billion dollars in the bank - TechCrunch. The company says it topped $150 million in annualized revenue within eight quarters of launch and counts more than 40% of the Fortune 50 as customers, though those revenue figures are self-reported and unaudited, a caveat worth keeping in mind throughout this category.
The product is Agent OS, refreshed to version 2.0 at the company's November 2025 Sierra Summit. It is a platform for building, deploying, and improving conversational agents across chat, voice, email, and SMS, and the 2.0 release added an Agent Data Platform that gives agents persistent memory, an Insights layer for diagnosing performance, Live Assist for coaching human agents in real time, the ability to publish an agent directly into ChatGPT, and voice as a genuine first-class channel - Sierra. The screenshot below shows the interface where CX teams define an agent's journeys and the guardrails it "cannot cross," which is Sierra's central design metaphor.
Sierra's most consequential contribution to the category is not a feature but a pricing philosophy: outcome-based pricing. As the company puts it, "you pay only when the software achieves specific, valuable outcomes," where an outcome is a resolved conversation, a saved cancellation, an upsell, or a completed purchase, and unresolved or escalated conversations in most cases carry no charge - Sierra. The mechanic is elegant because it aligns the vendor's revenue with the buyer's value, but Sierra publishes no dollar figures, and third-party estimates put the effective rate around $1.50 per resolved interaction on platform contracts that start near $150,000 a year and reach $200,000 to $350,000 in year one once implementation is included - fin.ai. The opacity is the trade-off: the incentive alignment is real, but you cannot compare Sierra on a rate card because there is not one.
The customer roster is genuinely blue-chip, and Sierra publishes specific outcomes for many of them. SiriusXM (34 million subscribers), Rocket Mortgage, Airtable (an 80% resolution rate), Chime (70%+ resolution), and SoFi (a 33-point NPS gain) are all named on its site, alongside Sonos, Ramp, Wayfair, CLEAR, and dozens more - Sierra. Cross-study figures compiled by third parties cluster around 70% containment with CSAT above 4.5 out of 5, with the best deployments reaching into the 90s. Those numbers are strong, and they are also carefully selected case studies, so treat them as the top of the range rather than the expectation.
Perhaps the most revealing thing Sierra ever published was a benchmark. In June 2024 the company released TAU-bench, an evaluation of how well AI agents handle realistic, multi-turn support tasks under domain policies, and its headline finding was humbling: the best model of the day resolved fewer than half of the tasks on average, and consistency collapsed further when the same task was run eight times - Sierra. Its 2026 successor, tau2-bench, added a telecom domain and dual-control scenarios where both the user and the agent take actions. Publishing a benchmark that exposes how far agents are from perfect reliability was a shrewd move, because it reframed the problem: capability is largely solved, but reliability is not, and reliability is what enterprises pay for. If you want the fuller picture of how agents are scored, our guide to AI agent evals and benchmarks goes deep on exactly this.
The criticisms of Sierra are the natural shadow of its strengths. Skeptics call it an "LLM wrapper" because it builds on OpenAI and Anthropic models rather than owning proprietary model IP, a critique that is technically accurate and strategically overblown, since the value is in the integration and reliability layer, not the raw model. More substantively, deployments are long and expensive, often three to seven months of white-glove, vendor-owned implementation, which puts Sierra out of reach for the mid-market and makes it overkill for a business with simple support needs. And the negotiated pricing that aligns incentives also removes transparency, which is exactly why a fully published rate card, as we will see, is the single feature that separates the field's number one from its number two.
5. Decagon: the concierge built on operating procedures
Decagon is the challenger that made the establishment nervous, and it did so with a genuinely different design philosophy rather than a cheaper clone. It was founded in August 2023 by Jesse Zhang and Ashwin Sreenivas, who met at an Andreessen Horowitz retreat, and came out of stealth in June 2024 - Wikipedia. Zhang studied computer science at Harvard, previously founded the gaming startup Lowkey (acquired by Niantic), and worked at Citadel Securities and Google. Sreenivas holds computer science degrees from Stanford, founded Helia (sold to Scale AI), and came from Palantir. The mission they set is unusually clear: "empower every brand to deliver concierge customer experiences," and that word, concierge, does real work in explaining the product.
The funding story is arguably even more dramatic than Sierra's on a percentage basis. Decagon raised $35 million across seed and Series A in mid-2024, a $65 million Series B at $650 million later that year, a $131 million Series C at $1.5 billion in June 2025, and then a $250 million Series D at $4.5 billion in January 2026, led by Coatue and Index Ventures - CMSWire. That Series D tripled the company's valuation in under six months, and a March 2026 employee tender offer confirmed the $4.5 billion mark - Sacra. Revenue is harder to pin down because Decagon declined to quantify it at the Series D, but third-party estimates put annualized revenue around $35 million as of late 2025, up from roughly $10 million a year earlier, with more than 100 enterprise customers added in 2025.
The technical differentiator is a concept Decagon calls Agent Operating Procedures, or AOPs. Instead of an engineer hard-coding an agent's logic, a CX team writes instructions in natural language, which Decagon compiles into validated workflows and code, complete with Git-based version tracking and guardrails that execute validation steps in code rather than in prose - Decagon. The design intent is a clean division of labor: non-technical support leaders shape what the agent does and how it behaves, while engineers retain control over the systems, integrations, and hard rules it must obey. It is a thoughtful answer to a real organizational problem, because in most companies the people who understand support policy are not the people who can write code.
Around the AOP core sits a full platform: a core agent that integrates with Salesforce, Zendesk, Snowflake, Stripe, and internal APIs; routing and agent-assist for human handoffs; an analytics dashboard; and a QA interface called Watchtower that runs A/B tests and simulations before changes ship - Contrary Research. Decagon runs omnichannel across chat, email, voice, and SMS with cross-channel memory, and it launched voice in February 2025 through a partnership with ElevenLabs, though voice remains the less mature side of the product relative to text. At the Series D the company leaned into proactive, outbound use cases, such as calling stranded travelers to rebook them, a hint at where support agents expand once they are trusted.
The customer outcomes Decagon publishes are aggressive, and mostly self-reported. Duolingo reports deflection above 80%, ClassPass a 95% lower cost per conversation, Substack more than 90% resolution while holding team size flat, and Curology support costs cut by 65% with ticket resolution rising from 5% to 80% - Decagon. The logo wall runs deep: Notion, Bilt, Vanta, Webflow, Eventbrite, Affirm, Deutsche Telekom, Block, and Hertz among them. Company-wide, Decagon claims more than 10 million customers served, 80% average deflection, and a 93% agent-quality score. These are impressive if accurate, and the honest reader should treat them as vendor figures until an independent test confirms them, a discipline this whole category rewards.
Decagon's weaknesses are the mirror image of Sierra's, and they are instructive. The most common complaint is heavy upfront engineering, the effort required to stand up the AOPs and integrations before value appears. There are integration gaps that matter for the mid-market: Decagon supports Zendesk, Salesforce, and Kustomer but not Freshdesk or HubSpot, which cuts out a large slice of smaller companies. Reviewers also describe an AOP complexity ceiling, where the natural-language procedure model strains once every ticket needs conditional logic branching ten different ways - eesel. And like everyone else, Decagon faces the horizontal threat of Salesforce and other platforms bundling "good enough" agents into software companies already own.
6. Sierra versus Decagon, head to head
Put the two side by side and the first thing that stands out is how much they have in common. Both are pure-play customer experience agent startups founded in 2023, both are led by unusually credentialed founders, both build on frontier models rather than their own, both sell primarily to enterprises, and both, at different moments, carried a $4.5 billion valuation. They are the two names that come up first in every enterprise RFP for a support agent, and they are competing for many of the same logos. The interesting comparison, then, is not who is "better" in the abstract, but which of two coherent philosophies fits a given buyer.
The sharpest distinction is about control and deployment. Sierra runs a managed, white-glove model where the vendor owns implementation and tuning, which is why its deployments take months and its pricing is negotiated rather than published. Decagon, through AOPs, hands the CX team direct control over agent logic after an initial engineering setup, betting that the people who own support policy should own the agent. Neither is universally right. A large enterprise that wants a partner to own the outcome and has the budget for a multi-quarter build will find Sierra's model reassuring. A company with a capable in-house CX operations team that wants to iterate quickly without filing a vendor ticket for every change will prefer Decagon's self-service posture.
The positioning diverges from there. Sierra emphasizes hallucination reduction and hyper-realistic voice with distinct brand personality, leaning on its benchmark research to claim reliability leadership. Decagon emphasizes autonomous task execution and explainability, with the no-code guardrail customization of AOPs as its wedge, and claims higher headline resolution (up to roughly 90% versus Sierra's roughly 70% containment) across a customer base that spans both B2B and B2C - Contrary Research. Read those claims with the skepticism the whole category demands, because "resolution" and "deflection" are measured differently by different vendors, a definitional gap we return to in the failures chapter.
Where the two genuinely part ways is scale and market valuation, and the chart below tells that story more clearly than prose. As of mid-2026, Sierra sits at a $15.8 billion valuation on roughly $150 million of self-reported revenue, while Decagon sits at $4.5 billion on an estimated $35 million. Sierra is roughly three and a half times the valuation and around four times the revenue scale, which means the market currently views Sierra as the category's front-runner rather than treating the two as equals.
The first-principles read on the gap is that investors are pricing distribution and reference density, not product quality. Sierra's founder network, its Fortune 50 penetration, and its early ownership of the outcome-pricing narrative give it a distribution advantage that compounds, and in a market where the models are commoditizing, distribution is the durable moat. That does not make Sierra the right choice for any specific buyer. For a mid-market company with an internal CX team, Decagon's control model and lower entry point may well be the better fit, and for a company that just wants a support agent live this quarter, neither may be the answer at all. The valuation tells you who Wall Street thinks wins the platform war. It does not tell you who resolves your tickets best.
7. The wider field: everyone else building support agents
Fixating on Sierra and Decagon would be a mistake, because the most consequential moves in 2026 are being made by incumbents and specialists that most buyers will actually shortlist. The category is consolidating fast, capital is flooding into voice, and the largest software company in customer service is quietly assembling a position that could reshape the whole market. To see the shape of the field, it helps to look at where the money went, because valuations and acquisition prices reveal the market's revealed preferences better than any feature comparison.
The incumbent that matters most is Salesforce Agentforce. Its third generation launched in June 2025 with agent observability through a Command Center, native support for the Model Context Protocol, and an ecosystem of MCP servers from partners including AWS, Box, Google Cloud, and Stripe - Salesforce. Adoption has been real but bumpier than the marketing suggests: Salesforce grew from a few hundred paid deals to around 6,000 paying customers within a year, yet CEO Marc Benioff had to publicly address "low Agentforce adoption" concerns at Dreamforce, and implementations routinely run five to eleven months and often require the expensive Data Cloud add-on. Agentforce's advantage is not the agent, it is that millions of companies already live inside Salesforce, and we cover its mechanics in our Agentforce explainer and the practical deployment guide.
The most important single event in the category, though, was Salesforce agreeing to acquire Intercom's Fin business for roughly $3.6 billion in June 2026 - CNBC. Fin had built the category's benchmark for pricing transparency, a flat $0.99 per resolution on any helpdesk with no platform fee, and had reached an average resolution rate of about 67% across 7,000 customers with its third-generation agent while crossing $100 million in annualized revenue. With that acquisition, Salesforce will own both the transparent, self-serve, per-outcome leader at the low end and the heavyweight Agentforce platform at the enterprise end, covering the pricing spectrum from both directions. That is a strategic position no pure-play startup can easily counter.
The other incumbents are all pivoting hard toward autonomous resolution, and the shape of their moves is worth understanding before any of the names blur together. Zendesk, taken private in 2022 and rebuilt as an AI Resolution Platform, announced outcome-based pricing where customers "only pay for issues resolved autonomously," reportedly around $1.50 per committed resolution, and bought both Ultimate and Forethought to fill out its agent stack - Futurum. The consolidation extends beyond Zendesk. Consider the three biggest moves in one place:
- NICE acquired Cognigy for roughly $955 million in 2025, at what analysts called a 25x premium, to anchor its enterprise CX-AI platform - NICE.
- Parloa, the German voice specialist, raised a $350 million Series D that tripled its valuation to $3 billion in under a year, on the strength of logos like Microsoft and Booking.com - Sifted.
- Salesforce absorbed Fin, as above, in the deal that reframed the whole competitive map.
Those three transactions tell a consistent story: the category is consolidating into a small number of platforms, and voice is where fresh capital is concentrating because it is the hardest channel to get right. Beyond them sits a rich second tier worth knowing by name. PolyAI raised $86 million at a $750 million valuation for enterprise voice - SiliconANGLE. Cresta and ASAPP attack the real-time contact center, each valued around $1.6 billion. Crescendo took the contrarian path of buying the BPO PartnerHero to fuse AI with 3,000 human agents. And Ada, Gorgias, Gladly, Kore.ai, Kustomer, and Lorikeet each own a niche, from high-volume automation to ecommerce to regulated verticals.
The strategic lesson from the wider field is that the pure-play startups are being squeezed from two sides. Above them, incumbents with existing distribution are bundling agents into software customers already pay for. Below and beside them, specialists are winning the channels and verticals that require genuine depth, particularly voice. Sierra and Decagon command the headlines and the valuations, but a real shortlist for most buyers includes at least one incumbent (because the integration is already there) and at least one specialist (because the hard channel or the regulated vertical demands it). The question of who to buy from is inseparable from the question of how you will be billed, which is where the real fight is happening.
8. Pricing models and the economics of a resolution
Nowhere is this category more confusing, or more consequential, than pricing, and the confusion is not accidental. Because a support agent's value is finally measurable, vendors have invented at least five distinct ways to charge for it, and each model encodes a different theory of who bears the risk when the agent fails. Understanding the models is not a finance exercise, it is how you avoid signing a contract that looks cheap in the demo and costs six figures in production. The starting point is the raw economics: a human-handled ticket costs roughly $6 to $12, complex cases up to $25, while an AI resolution runs $0.10 to $2.00, a headline gap of about 12x - eesel.
That gap is real but it is not the whole story, and the honest version includes the labor market it disrupts. Fully loaded human agent costs in 2026 range from $6 to $14 an hour offshore to $28 to $45 onshore, and an in-house agent fully loaded can exceed $65,000 a year - Callforce. Against those numbers, even an AI resolution at the top of its range is dramatically cheaper, which is why vendors advertise ROI figures around $3.50 returned per dollar spent, rising to 8x for leaders. The economics are compelling enough that the interesting question is no longer whether AI is cheaper, but how the vendor captures a share of that savings, and that is exactly what the pricing model decides. For the fuller cost picture across the agent stack, see our report on the true cost of agentic AI.
The five models, and the incentive each creates, are the heart of the matter:
- Per-resolution or outcome (Fin at $0.99, Zendesk near $1.50, Gorgias around $0.90, Sierra estimated near $1.50) aligns vendor revenue with buyer value, so the vendor only earns when it succeeds.
- Per-conversation (Agentforce at ~$2, Decagon's primary model, Parloa) is predictable but bills you even when the agent fails to resolve.
- Consumption or credits (Agentforce Flex at $0.10 per action, Kore.ai in 15-minute sessions) is granular but genuinely hard to forecast.
- Per-seat (Agentforce from $125, Cresta $100 to $200) is familiar SaaS but breaks down for agents that replace rather than assist humans.
- Per-minute (PolyAI, Parloa voice) is the standard for voice and scales with call length.
The pattern in that list is a shift from paying for access toward paying for outcomes, and the diagram below traces that evolution. The direction is not arbitrary: as the value of a support interaction became measurable, the market moved toward charging for the measurable thing, which is the resolution itself.
Outcome-based pricing sounds like a clean win for buyers, but it introduces a subtle and important battleground: what counts as resolved. Because money now changes hands on that definition, vendors define it carefully. Zendesk and others use a 72-hour quiet period, meaning a conversation counts as resolved only if the customer does not come back within three days, and some platforms will double-bill a ticket if a human is not looped in within that window - eesel. The "resolved" definition is the new contract negotiation, and a buyer who does not scrutinize it can end up paying for reopened tickets, escalated cases, and conversations the customer abandoned in frustration. Salesforce running three pricing models at once is not indecision, it reflects how genuinely hard it is to translate agentic usage into a bill both sides trust.
There is one more variable that changes the economics from the vendor's and the builder's side: the cost of the underlying model. Every resolution consumes tokens, and the difference between routing a simple FAQ to a cheap model like Claude Haiku 4.5 or Gemini 3.5 Flash-Lite and sending everything to a flagship is enormous at scale. Teams that implement intelligent model routing can cut inference costs substantially without hurting quality, which is why the platforms with the best margins are quietly the best at model orchestration, a topic we cover in cutting agent costs with model routing. For a buyer, the practical takeaway is to match the pricing model to your volume predictability: choose per-resolution if your volume is spiky and you want to pay only for wins, and consider per-conversation or per-seat only if your volume is stable and you have negotiated a rate that reflects your actual resolution rate.
9. Where it works: real deployments and hard numbers
The clearest evidence that support agents are real, and not a demo trick, is the set of large-scale deployments that have been running long enough to produce audited-adjacent numbers. The canonical example is Klarna, whose OpenAI-powered assistant, launched globally in February 2024, handled 2.3 million conversations in its first month, the equivalent workload of about 700 full-time agents, and was forecast to drive a $40 million profit improvement that year - OpenAI. It resolved issues in about two minutes versus eleven for humans, matched human CSAT, and drove a 25% drop in repeat inquiries across 23 markets and more than 35 languages. Crucially, Klarna's cost per transaction fell from $0.32 to $0.19, roughly 40%, over two years.
What makes Klarna the most useful case study is not the launch but the follow-through, because it demonstrates the mature shape of a working deployment. In its Q3 2025 update, Klarna reported that the AI now does the work of 853 employees (up from 700), saving around $60 million, with response times improved and no drop in customer satisfaction, and an NPS of 73 - Customer Experience Dive. The number grew rather than shrank, which is the opposite of a failed experiment. Notably, Klarna built this on OpenAI's own tooling rather than buying a CX vendor, a fact we will return to when we weigh building versus buying, and a case study we also examine in our look at how the financial sector automates with AI agents.
Beyond Klarna, the vendor case studies converge on a consistent finding: deeply integrated, action-taking agents on well-scoped use cases genuinely reach 70% resolution and beyond. Sierra reports Airtable at 80% and Chime above 70%; Decagon reports Substack above 90% and Curology rising from 5% to 80% ticket resolution. The pattern across these wins is not the vendor, it is the setup: a clean knowledge base, real API access so the agent can act, a narrow initial scope, and rigorous measurement. The image below shows the kind of diagnostic tooling that makes this possible, in this case Sierra's Insights layer analyzing exactly why an agent transfers a conversation to a human.
But the honest version of "where it works" requires separating two metrics that vendors routinely blur, and the chart below makes the gap visible. Deflection means no human touched the conversation. Resolution means the customer's problem was actually solved. They are not the same, and the difference is often large - fin.ai. A vendor can report an 80% deflection rate while true self-service resolution sits far lower, because a customer who gives up and closes the chat counts as deflected but not resolved. The realistic industry ranges, from independent observers rather than marketing decks, are 30% to 50% for early deployments, 50% to 70% as workflows mature, and 70% to 85% for deeply integrated agents on well-scoped problems.
The practical guidance that falls out of this is to trust the setup, not the logo. The deployments that work share a recipe that any buyer can insist on: scope the agent to a bounded set of high-volume, low-risk intents first, wire it into real systems so it can act rather than just answer, ground it in a curated knowledge base, keep a human escalation path that is fast and obvious, and measure resolution rather than deflection from day one. Do those five things and a 70% resolution rate is achievable with several of the platforms in this guide. Skip them and no vendor, at any valuation, will save the project.
10. Where it fails: limits, backlash, and liability
Every serious buyer should study the failures as carefully as the successes, because the failure modes in this category are specific, repeatable, and occasionally end up in a courtroom. The most important legal precedent is Air Canada, whose chatbot gave a passenger false information about bereavement fares. When the airline argued that the chatbot was "a separate legal entity" responsible for its own statements, a Canadian tribunal rejected the argument outright, found the airline liable for negligent misrepresentation, and ordered it to pay - CBC. The ruling established a principle every deploying company now lives under: you own what your agent says, and "the AI made it up" is not a defense.
The failure that ended that principle's theoretical status was Cursor. In April 2025, the AI code editor's support agent, named "Sam," invented a nonexistent policy limiting accounts to one device, told customers this was the rule, and triggered a wave of cancellations before anyone at the company noticed - The Register. The company's co-founder apologized, blamed a combination of model hallucination and a session-management bug, and started clearly labeling AI responses. The lesson is not that Cursor was careless, it is that a hallucinated policy delivered with total confidence is one of the most dangerous things a support agent can do, because customers reasonably assume the company stands behind its own support channel. This is exactly the class of risk that guardrails and grounding are meant to contain, and why our guide to prompt-injection defense treats it as a first-order concern rather than an afterthought.
The escalation and handoff design is where most of these failures are prevented or created, which is why it deserves its own diagram. An agent that guesses when it should escalate produces Air Canada and Cursor outcomes; an agent that routes to a human the moment its confidence drops or a guardrail fires produces trust. The flow below shows the decision points that separate a safe deployment from a liability.
Beyond individual incidents, there is a genuine consumer backlash that buyers ignore at their peril. A February 2026 SurveyMonkey study found that 79% of Americans strongly prefer humans over an AI agent, only about 2% want to interact exclusively with chatbots, and 46% say AI service rarely or never leads to a successful outcome - SurveyMonkey. And yet, revealingly, 80% of people who actually used an AI chatbot reported a positive experience, and 92% of businesses report improved CSAT after deploying AI when it is paired with seamless human escalation. The contradiction resolves into a clear design rule: customers do not hate AI support, they hate being trapped in it with no way to reach a person.
Even the flagship success story contains a cautionary chapter. Klarna's CEO acknowledged in 2025 that the company had cut too deep on human agents and was rehiring, moving to a hybrid model, and stressing that "it's so critical to be clear to your customer that there will always be a human if you want" - Forbes. This is frequently misreported as Klarna abandoning AI. It did not. It walked back AI-only, not AI, and its later numbers grew. The nuance matters because it defines the mature end state: AI handles the volume, humans handle the exceptions and the emotionally loaded cases, and the two are stitched together so tightly that the customer never hits a wall. The other failure mode is quieter and more industrial, and it is playing out in the stock prices of the outsourcing giants, where the human contact-center business is being repriced in real time as AI absorbs the routine volume.
11. Build, buy, or run it inside an autonomous company
Once a buyer understands the field, the real decision is architectural rather than a simple vendor beauty contest: do you buy a standalone platform, build the agent yourself, or run support as one function inside a broader autonomous operation? Each path has a distinct cost structure and a distinct failure mode, and the right answer depends far more on your engineering capacity and your ambitions beyond support than on which vendor scored highest in a table. The frame that clarifies it is to ask what you are actually trying to own: the fastest possible time-to-value, the deepest possible control, or the ability to run many functions from one system.
Buying a standalone platform is the default, and for good reason. For a company that wants a support agent live this quarter with minimal engineering, a platform like Fin or a Zendesk agent delivers value fast, and the per-resolution pricing means you pay in proportion to what you get. The cost is money and lock-in: enterprise contracts run from $50,000 to well past $600,000 a year once implementation is included, and the opaque, negotiated pricing at the high end (Sierra, Decagon, Parloa) means you are trusting a sales process rather than a rate card. Buying is right when support is a discrete problem you want solved by specialists and you are comfortable being a tenant in someone else's platform.
Building in-house has become dramatically more viable, and Klarna is the proof: it built its own support agent directly on OpenAI's tooling rather than buying a CX vendor, and the result now does the work of over 800 agents. The reason building is newly practical is that the hard components are now available as robust primitives. Consider what a build actually requires:
- A model and an agent framework, such as the OpenAI Agents SDK or the Claude Agent SDK, which handle the reasoning loop and tool calling.
- A retrieval stack to ground the agent in your knowledge base, plus a guardrail layer for safety.
- Conversation state, observability, and evaluations, the connective tissue that makes it improvable.
- A helpdesk of record to manage tickets, escalations, and the human side.
The catch, and it is a real one, is that assembling and maintaining that stack is a serious engineering commitment, and the total cost of ownership (including the separate helpdesk seats at $55 to $175 per agent per month and the ongoing model bill) often rivals buying once you account for the team. Building is right when support is strategic enough to be a differentiator, when you have the engineering depth, and when the flexibility of owning the stack outweighs the maintenance burden. Our insider guide to building AI agents lays out that stack in detail.
The AI labs themselves have made building even more tempting, which is the disintermediation threat hanging over every pure-play vendor. OpenAI's AgentKit and Anthropic's Claude Agent SDK both target agent construction directly, and Anthropic now markets Claude for support use cases and offers a managed agent runtime. When a large enterprise can assemble a competent agent on the raw SDK, the question "why pay a wrapper?" becomes sharper, and it is the single biggest long-term risk to companies whose value is primarily integration and orchestration rather than proprietary model IP.
There is a third path that reframes the question entirely, and it is where a platform like O-mega fits as one option among the alternatives above. Instead of treating support as a standalone system to buy or a stack to build, O-mega runs support as one function inside a fully autonomous company: a single conversation builds and operates the website, the app, the billing, the content, and the admin, with customer support handled as part of that whole rather than bolted on. In the assessment table it scores highest on time-to-value and lowest on dedicated contact-center depth, which is exactly the trade-off it represents: it is not the choice for a Fortune 500 contact center with millions of voice calls, but it is a coherent option for a founder or operator who wants support to exist and improve without standing up a separate CX platform at all. The same logic underlies our argument for hiring an AI workforce to run your company, where support is one job among many rather than the whole project.
The decision, then, is not "which vendor" but "which altitude." Buy if support is a bounded problem and speed matters most. Build if support is strategic and you have the team. Run it inside an autonomous operation if support is one of many functions you would rather not manage as separate systems. The wrong move is to treat these as the same decision, because a mid-market company that buys a Fortune 50 platform overpays, and an enterprise that builds without the engineering depth ships a Cursor.
12. Measuring what matters: evals, benchmarks, and the model layer
If there is one discipline that separates teams who succeed with support agents from teams who get burned, it is measurement, and the reason is structural: this is a category where the marketing metrics and the true metrics diverge by design. The foundational skill is knowing which benchmark and which internal metric to trust, because a leaderboard score does not translate to your account, and a vendor's resolution rate does not survive contact with your tickets. Sierra's own TAU-bench and its successor tau2-bench are the most credible public benchmarks, measuring whether an agent adheres to domain policies across multi-turn, tool-using conversations, and the sobering headline from that research is that even frontier models leave substantial gaps on realistic tasks - GitHub.
The trap with public benchmarks is treating them as a ranking of products, when they are really a ranking of models under a specific test harness. Leaderboard results are not comparable across harnesses, and aggregators frequently disagree about who leads, so the right way to read them is as ranges and directional signals rather than a single winner. What actually predicts your outcome is an internal evaluation built on your own policies and your own historical tickets, run continuously against the agent before and after every change. That is what the leading platforms do with tools like Watchtower and Insights, and it is what a serious in-house build must replicate. Our full guide to AI agent evals and benchmarks covers how to construct those internal evaluations properly.
Underneath the evals sits the model layer, and here staying current genuinely matters because the models turn over every few months and the wrong choice quietly caps your ceiling. As of mid-2026, the flagships worth evaluating for support are OpenAI's GPT-5.6 family, Anthropic's Claude Opus 5 for the hardest reasoning and Claude Sonnet 5 for the speed-to-quality balance, and Google's Gemini 3.6 Flash as a cost-efficient workhorse - Anthropic. For the high-volume, low-complexity tail of support traffic, cheaper models like Claude Haiku 4.5 or Gemini 3.5 Flash-Lite resolve simple intents at a fraction of the cost, which is why model routing is a margin lever as much as a quality one. For voice, the native speech-to-speech models, OpenAI's GPT-Realtime-2 and Google's Gemini 3.1 Flash Live, are what make sub-300-millisecond conversations feasible.
The practical measurement framework that ties this together is short and worth memorizing. Measure resolution, not deflection, so you are counting solved problems rather than abandoned chats. Benchmark on your own policies, not a public leaderboard, because your edge cases are what break agents. Track cost per resolution including the model bill, not just the vendor invoice, so model routing decisions are visible. And instrument escalation quality, because how gracefully the agent hands off is what customers actually remember. A team that holds itself to those four measurements will make better vendor decisions, negotiate better contracts, and catch the failure modes from the previous chapter before they reach a customer or a courtroom.
13. The 2026 outlook and a decision framework
Step back from the individual companies and the structural picture is clear. Customer support is the first enterprise function where autonomous AI works well enough, cheaply enough, and measurably enough to change the underlying economics, and the market has responded by pouring capital into the companies best positioned to own the wedge. The forward trajectory follows from first principles rather than hype: as the models commoditize, the durable advantages shift to distribution, reliability, and integration depth, which is why an incumbent like Salesforce (now owning both Agentforce and Fin) and a distribution-rich pure play like Sierra are structurally better positioned than the quality of any single agent would suggest. Expect more consolidation, continued capital concentration in voice, and a widening gap between vendors that measure resolution honestly and those that sell deflection.
The disintermediation question is the one to watch over the next 24 months. If the AI labs make building a competent agent trivial through tools like AgentKit and the Claude Agent SDK, the pure-play platforms will have to justify their margins with something more durable than a model wrapper, and the answer will be reliability, workflow depth, and the accumulated evaluation data from millions of real conversations. That is a real moat, but it is a narrower one than a $15.8 billion valuation implies, which is why the honest outlook includes the possibility that today's valuations compress even as the underlying category keeps growing. Both things can be true: the market is real and durable, and some of the prices paid for a piece of it are not.
For the buyer, the decision framework reduces to a handful of questions asked in order. First, match the pricing model to your volume: choose per-resolution if your volume is spiky and you want to pay only for wins, per-conversation or per-seat only if your volume is stable and the negotiated rate reflects your real resolution rate. Second, weigh time-to-value against control: buy a self-serve platform like Fin if speed matters most, choose a white-glove partner like Sierra if you want the vendor to own the outcome, or pick Decagon's AOP model if your CX team wants direct control. Third, insist on the setup that actually works: bounded initial scope, real system integration, a curated knowledge base, and a fast, obvious human escalation path. Fourth, measure resolution, not deflection, and benchmark on your own policies before you sign anything.
The last question is the one most buyers skip, and it is the most important: what altitude do you want to operate at? If support is a discrete problem, buy a specialist. If it is strategic, build. And if support is simply one of many operational functions you would rather have run for you, consider running it inside a broader autonomous operation where it is one job among many. This is the same shift toward autonomous operation that runs through much of our coverage, from the agentification of business to the autonomous business guide for 2026. The companies that win with support agents in 2026 will not be the ones that bought the highest-valued vendor. They will be the ones that matched the right architecture to their real needs and measured the truth instead of the marketing.
This guide was written by Yuma Heymans (@yumahey), founder and CEO of O-mega and co-founder of the AI recruitment platform HeroHunt.ai. Having built one of the first autonomous AI recruiters in production, he has a hands-on view of how autonomous agents take over an entire business function, the same pattern now reshaping customer support, and he writes regularly on the autonomous AI workforce and the shifting AI job market.
This guide reflects the AI customer support landscape as of August 2026. Valuations, pricing, resolution rates, and available models change frequently in this category, and many revenue and resolution figures cited here are self-reported by vendors. Verify current details, and independently test any platform on your own tickets, before purchasing.