The insider's guide to the one decision that governs every enterprise AI agent project in 2026: do you build the thing, or do you rent it?
On July 22, 2026, OpenAI launched Presence, an enterprise agent product it will not let you buy by yourself. There is no API, no pricing page, and no self-serve signup. To run it, an OpenAI Forward Deployed Engineer (a title borrowed straight from Palantir) embeds with your team to wire it into your systems, and a select group of global systems integrators handle the rest - The Register. The company that sells the cheapest intelligence on earth just decided that, for most enterprises, the answer to "should we build our own agent?" is no, let us do it for you.
That single product tells you almost everything about the state of agentic AI. The models are commoditized. The frameworks are free. And yet the hardest, most valuable part of an AI agent (getting it into production without it leaking data, breaking policy, or hallucinating a refund) is so hard that even OpenAI ships humans alongside the software to get it done.
Here is the problem for everyone else. You do not have OpenAI's Forward Deployed Engineers, and you cannot afford to be one of the roughly 40% of agentic AI projects Gartner expects to be canceled by the end of 2027 over runaway cost, unclear value, and weak controls - Gartner. So the build-versus-rent choice is not a philosophy debate. It is the difference between an agent that pays for itself and a six-figure line item your CFO kills in the next budget cycle.
This guide breaks down exactly what "build" and "rent" mean in 2026, ranks thirteen approaches and platforms on weighted, sourced criteria, dissects OpenAI Presence and why it exists, maps the build stack (the frameworks, the real costs, the hidden six-layer tax), the rent stack (managed platforms and the shift to outcome pricing), the market and the failure data, the security and legal liability that follow whoever deploys the agent, and a decision framework you can actually use. Where it helps, we point to deeper dives in our own coverage of the true cost of agentic AI and the economics of digital labor.
Contents
- The 2026 build-vs-rent scorecard
- What OpenAI Presence actually is
- Build vs rent, from first principles
- The build path: frameworks are free, production is the bill
- The rent path: managed agents and the rise of outcome pricing
- OpenAI's own agent surface: build and rent from one vendor
- The market: size, adoption, and the failure data
- Where build wins, where rent wins
- Security, governance, and who is liable
- Protocols: how MCP and A2A change the math
- The hybrid frontier: build on rails and the autonomous company
- The 2026 outlook and a decision framework
1. The 2026 build-vs-rent scorecard
Before the deep dives, here is the whole field in one view. The table below scores thirteen of the most relevant ways an enterprise can put an AI agent into production, from writing raw framework code to renting a fully managed, human-deployed service. Each option is scored on the five things that actually decide the outcome, weighted by how much each one moves the decision. This is the summary. Everything after it is the reasoning behind each cell, because a score without its justification is just an opinion wearing a number.
The criteria are built from first principles, not from a vendor feature sheet. What a buyer is really purchasing is a working outcome in production, delivered fast enough to matter, controllable enough to differentiate, cheap enough to survive scale, safe enough to pass audit, and reliable enough to trust with real customers. That reduces to five weighted factors: speed to value (how quickly it reaches production), control (how much of the logic, data, and behavior you own), total cost of ownership at scale, governance and security (guardrails, identity, compliance, auditability), and scale and reliability (proven production maturity). A category column marks whether each option is a pure build framework, a "build on rails" managed platform, a rented managed service, or the newer autonomous workforce model.
| # | Solution | Category | What it does | Speed (25%) | Control (20%) | TCO at scale (20%) | Governance (20%) | Scale (15%) | Final |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Amazon Bedrock AgentCore | Build on rails | Managed runtime for any framework, any model | 7 - managed harness, but AWS assembly | 8.5 - any framework/model, you own the agent | 8 - consumption, pays only active CPU ($0.0895/vCPU-hr) | 8.5 - security enforced at infra layer agents can't bypass | 8 - AWS-grade, GA Oct 2025 | 8.0 |
| 2 | Gemini Enterprise Agent Platform | Build on rails | Google's rebranded Vertex agent stack + ADK | 7.5 - ADK builds agents in <100 lines | 8 - code-first, you own the logic | 7.5 - pure PAYG, ~$500-2K/mo realistic | 8 - deterministic guardrails, native A2A | 8 - Google infra, 200+ models | 7.8 |
| 3 | Microsoft Copilot Studio | Build on rails | Low-code + code agents across M365 | 8 - fast for M365 shops | 7 - custom agents + autonomous actions | 6.5 - $200 per 25K credits, action = 25+ credits | 7.5 - Entra/Purview, but EchoLeak exposure | 9 - 1M+ agents created | 7.6 |
| 4 | LangGraph | Build framework | Stateful graph orchestration for long-running agents | 6 - powerful but heavy, longer to prod | 9.5 - durable execution, full ownership | 7 - free (MIT) + tokens + infra | 6 - you own the compliance stack | 8.5 - ~400 enterprises on platform | 7.3 |
| 5 | Salesforce Agentforce | Rent platform | Agents native to the Salesforce CRM | 7.5 - on-platform, ~38-day buy avg | 6 - configurable but locked to Salesforce | 6 - ~$2/conversation, ~$2M/yr at 1M convos | 8 - enterprise trust layer | 9 - 18,000+ companies, 121 countries | 7.2 |
| 6 | OpenAI Agents SDK | Build framework | Lightweight provider-agnostic multi-agent SDK | 6.5 - fast prototype, you assemble prod | 9 - own code, 100+ LLMs | 7.5 - free (MIT) + model tokens only | 5.5 - tracing/guardrails, you own compliance | 7 - 28K stars, you run it | 7.1 |
| 7 | Claude Agent SDK | Build framework | The harness behind Claude Code, as an SDK | 6.5 - great for coding/computer use | 9 - own harness, MCP tools | 7 - free (MIT) + Claude tokens | 6 - hooks/permissions, you own compliance | 6.5 - powers Claude Code | 7.0 |
| 8 | Sierra | Rent CX | Outcome-priced enterprise support agents | 7.5 - concierge onboarding, weeks | 5 - vendor-run, you configure | 7 - outcome-based, ~$1.50/resolution | 7.5 - CX guardrails, narrow scope | 8 - SiriusXM, ADT, Sonos; $15.8B | 7.0 |
| 9 | CrewAI | Build framework | Role-based multi-agent "crews" | 7.5 - fastest to a multi-agent demo | 8.5 - own code, role orchestration | 7 - free (MIT) + tokens | 4.5 - lighter, state issues at scale | 6.5 - 56K stars, scale criticism | 6.9 |
| 10 | O-mega | Autonomous workforce | Builds and runs a whole company from a chat | 8 - describe it in a conversation | 7 - runs in your tools and guardrails | 7 - self-serve subscription + credits, free tier | 6.5 - org-context guardrails, younger | 6 - emerging category | 6.9 |
| 11 | Decagon | Rent CX | Agentic support "workers" on operating procedures | 7.5 - ~6-week deploy | 5 - vendor-run agent OS | 7 - ~$0.99/conv or ~$0.50/resolution | 7 - procedure-driven guardrails | 7.5 - Duolingo, Chime; $4.5B | 6.8 |
| 12 | Lindy | Rent builder | No-code horizontal agent builder | 8.5 - live in minutes | 5 - templated flows | 6.5 - credits, 2x overage | 5.5 - SMB-grade | 6 - prosumer/SMB | 6.4 |
| 13 | OpenAI Presence | Rent managed | Fully managed, human-deployed voice/chat agents | 7 - FDE-led, but gated GA | 4 - policy-scoped, no code/model ownership | 5 - custom enterprise, opaque pricing | 9 - runtime guardrails, simulations, human approval | 7 - proven on OpenAI's own line, nascent | 6.4 |
Criteria weights: speed to value (25%), control and customization (20%), total cost of ownership at scale (20%), governance and security (20%), scale and reliability (15%). Each cell shows the score (0-10) followed by the reason. The final column is the weighted average, rounded to one decimal. Rows are ordered by final score descending, and ties are broken alphabetically.
The most important thing this table reveals is not any single winner. It is the shape of the ranking. The top three are all "build on rails" platforms, and that is not a coincidence. Amazon Bedrock AgentCore, the Gemini Enterprise Agent Platform, and Microsoft Copilot Studio each combine managed infrastructure (so you skip the six-layer engineering tax) with retained control of the agent's logic and data (so you keep your differentiation) and consumption pricing (so cost tracks value). They occupy the productive middle of the build-rent spectrum, which is exactly where the empirical data says most successful 2026 deployments actually live. Pure build frameworks like LangGraph score well on control but pay for it in speed and governance, because you assemble everything yourself. Pure managed services sit lower not because they are bad but because they trade away the two things this scorecard weights most: control and cost transparency.
That is why OpenAI Presence, the product this guide is named after, lands near the bottom of a general-purpose scorecard despite being one of the most sophisticated agent systems ever shipped. It scores a 9 on governance, higher than anything else here, because it is engineered from the ground up around policies, simulations, guardrails, and human approval. But it scores a 4 on control (you own none of the code or the model) and a 5 on cost (there is no published price and the delivery model is consulting-shaped). For a regulated enterprise that weights governance far above everything else, Presence would rise to the top of a re-weighted table. For everyone else, its position here is the honest answer. Keep that tension in mind, because it is the entire argument of this guide: the right answer depends on which of these five things your business actually cannot compromise on.
2. What OpenAI Presence actually is
To understand the build-versus-rent decision in 2026, you have to understand why a foundation model lab, whose entire business is selling the raw intelligence you would use to build your own agent, decided to also sell the finished agent as a done-for-you service. Presence is not a new model. It is, in OpenAI's own framing, a deployment and management layer that sits on top of OpenAI's models - Help Net Security. It exists to close the gap between a model that can hold a convincing conversation and a system a company is actually willing to put in front of paying customers, and OpenAI concluded that gap was too wide for most enterprises to cross alone.
Mechanically, Presence bundles the things that are boring to build and catastrophic to get wrong. Each deployment starts scoped to one job (resolving a billing dispute, clearing an IT ticket, handling an insurance claim), and the agent receives only the knowledge and system access that the specific job requires - artificialintelligence-news. The customer writes the rules that govern what the agent can do, when it needs human sign-off, and when a person takes over. On top of that sit runtime guardrails that can intervene the moment an interaction moves outside the company's defined boundaries, pre-launch simulations against edge cases, and evaluation tooling to measure whether the thing actually works before it touches a real customer.
The most revealing feature is the improvement loop. After launch, OpenAI's Codex reads production sessions and escalations, identifies where the agent is failing, and proposes behavior changes. Critically, those changes do not ship automatically: staff must test and approve them before they go live - Help Net Security. This is a deliberate rejection of full autonomy, and it is the correct design, because an agent that rewrites its own rules in production is precisely the failure mode that gets a company sued. OpenAI runs this exact system on its own English-language phone support line, where the company reports Presence resolves about 75% of inbound issues without a human, and where the Codex loop cut human handoffs by 15 percentage points in ten days - VentureBeat. Read those numbers with the appropriate skepticism: they are OpenAI's figures, measured against OpenAI's own grading criteria, and they have not been independently verified.
Now the part that matters most for this guide. Presence is not self-serve. There is no API and no public pricing, and deployments are led by OpenAI's own Forward Deployed Engineers plus a set of selected systems integrators - dev.to. Access is gated on workflow fit, implementation readiness, and OpenAI's available delivery capacity. The early named design partners are BBVA (exploring voice banking in Mexico), SoftBank (testing Japanese-language conversations), and Insurance Australia Group (testing support during high-demand weather events) - artificialintelligence-news. When VentureBeat asked OpenAI for pricing, twice, it got no answer. A rumored ten-million-dollar entry point circulates in third-party commentary, but it is unverified and did not come from OpenAI, so treat it as noise.
The launch also drew immediate skepticism, and it is worth hearing, because it sharpens the rent-versus-build question rather than settling it. Gartner's Kathy Ross warned that by 2027 half of the organizations planning to shift customer service to AI will abandon those plans, and that the human touch remains irreplaceable in many interactions - The Register. Cyara's CEO framed the deeper issue precisely: effective AI governance now has to move at machine speed, with automated validation and real-time testing, so that human oversight functions as a checkpoint rather than the entire assurance system - CX Today. That is exactly the capability Presence is selling. The question a buyer has to answer is whether renting that machine-speed governance from OpenAI is smarter than building it, and the honest answer is that it depends entirely on whether governance is your differentiator or your overhead.
The Register read the launch as OpenAI "trying the consulting path," and offered the sharpest one-line explanation in the coverage: as AI models become commoditized, maybe there is margin in the plumbing - The Register. That is the thesis of the entire build-versus-rent debate stated in a single sentence. The intelligence is cheap. The value has migrated to the integration, the guardrails, the change management, and the accountability. Presence is what it looks like when the model maker decides to capture that value directly, by renting you an outcome instead of selling you a tool. The rest of this guide is about when you should take that deal and when you should refuse it.
3. Build vs rent, from first principles
Most build-versus-rent commentary starts from the wrong question. It asks "which vendors are winning?" or "what are competitors doing?" Those are surface questions that lead to consensus answers. The structural question is different, and you have to start there: what actually happens to a business when intelligence becomes a cheap, commoditized input? Because that, not any single product launch, is the force reshaping this entire decision.
When an input becomes cheap, the businesses that use that input to deliver a valuable output do not lose. They flourish. But the value migrates. It moves away from the commodity input and toward the scarce complements that turn the input into an outcome. In the agent economy, the commodity is the model, and its price has collapsed toward zero: a frontier flagship that cost thirty dollars per million input tokens in 2023 now has current equivalents around two to five dollars, with the cheapest capable tiers near twenty cents - TokenMix. The scarce complements are the things a model still cannot supply on its own: your proprietary data, your workflows, your integrations, your regulatory accountability, and the operational discipline to run the thing safely. That is where the value went, and build-versus-rent is really a question about who supplies those complements: you, or a vendor.
You can watch this migration happen in the numbers. As the models converged on similar capability, enterprise usage spread across four labs rather than concentrating in one: Anthropic now holds about 40% of enterprise usage, OpenAI 27%, and Google 21%, a striking reshuffle from the days when a single vendor dominated - Menlo Ventures. When the input is that interchangeable, no one can build a durable moat on the model alone. Andreessen Horowitz makes the same structural point in its outlook for 2026: as models commoditize, the defensible position shifts to proprietary data and the action layer, the part of the stack where the agent actually reaches into your systems and does work - a16z. That is the complement no vendor can rent you, because it is made of your data and your workflows. It is also the precise thing you give up when you rent a generic agent, which is why the build-versus-rent choice always comes back to one question: is the value in the model, or in what only you can supply around it?
Once you frame it this way, the two paths become clear. To build is to supply the complements yourself: you take the cheap model, wrap it in your own code, your own data pipeline, your own guardrails, and your own operations. You keep total control and you keep whatever differentiation your data and workflows create, but you also own every failure. To rent is to buy the complements pre-assembled: a vendor has already built the wrapper, hardened the guardrails, and hired the operations team, and you pay them to run it. You get speed and you offload the failure modes, but you rent your differentiation from someone who is renting the same thing to your competitors. Neither is universally right. The correct choice is a function of where your specific business keeps its value.
The empirical signal here is strong and it points in a direction that surprises people. Across enterprise deployments tracked by Menlo Ventures, externally purchased AI solutions reach production at roughly twice the rate of internally built ones, and the mix has swung hard toward buying: 76% of enterprise AI use cases are now purchased, up from 53% in 2024 - Menlo Ventures. The MIT NANDA study found the same pattern from the other side, reporting that externally sourced tools succeeded around 67% of the time versus roughly 33% for internal builds - The Register. Internal builds fail about twice as often. That is not because engineers are bad. It is because the complements that make an agent work in production (evaluation, drift monitoring, prompt-injection defense, change management) are a specialized operational discipline, and most teams underestimate how much of the total effort lives there.
But do not read that as "always rent." The same data that shows buying wins on average also shows why building wins in specific cases, and the reason is the first-principles frame again. If the value of your agent comes from proprietary data no vendor can replicate or a workflow that is itself your competitive advantage, then renting means renting your own moat back from a supplier who will happily rent an identical capability to everyone in your industry. In that case the average is irrelevant, because you are not an average case. This is the discipline first-principles thinking imposes: the aggregate says buy, but the aggregate is not your business. You build where your differentiation lives and you rent everything else, and the rest of this guide is about telling those two apart. We go deeper on this economic frame in our analysis of the agent economy and the economics of digital labor.
4. The build path: frameworks are free, production is the bill
The single biggest misconception about building AI agents is that the framework is the hard part. It is not. The framework is the easiest, cheapest, most commoditized layer in the entire stack. In 2026 you can choose from more than a dozen credible, production-grade agent frameworks, and almost all of them are free open source under permissive MIT or Apache-2.0 licenses. LangChain alone has over 143,000 GitHub stars; CrewAI has 56,000, LlamaIndex 51,000, and OpenAI's own Agents SDK 28,000 - GitHub. Picking a framework is a solved problem. The money and the months go somewhere else entirely.
The framework landscape has consolidated into a few clear archetypes, and understanding them is useful even for a non-technical buyer, because it tells you what your engineers will actually be arguing about. The heavyweight orchestrators like LangGraph give you durable, stateful, long-running agents with human-in-the-loop checkpoints, at the cost of a steep abstraction load. The lightweight SDKs like OpenAI's Agents SDK and Pydantic AI keep the surface small and let you stay close to the raw model. The role-based frameworks like CrewAI get you to a working multi-agent demo fastest. And the cloud-native kits like Google's ADK, Microsoft's Agent Framework (the merger of AutoGen and Semantic Kernel, shipped as version 1.0 in April 2026), and AWS Strands bind tightly to their respective clouds - Visual Studio Magazine. We compare these head to head in our guide to LangGraph vs CrewAI vs AutoGen.
Here is what none of that free tooling tells you. A framework gives you the agent loop and nothing else. To ship a production agent you must assemble roughly six more layers around it: compute and sandboxing, memory, tools and actions, model routing, multi-step orchestration, and observability and governance - Augment Code. Each layer is a real engineering project with a real monthly bill. Your evaluation and observability layer means buying or self-hosting something like Langfuse, LangSmith, or Braintrust. Your memory layer means running a vector database that routinely lands two to four times over budget at scale - LeanOps. Your tools layer means wiring and securing dozens of integrations through the Model Context Protocol. None of this is optional, and all of it is permanent.
The observability layer deserves special attention, because it is the one most teams skip and the one that separates a demo from a product. LangChain's own survey found that while 89% of organizations have some observability, only 52% run offline evals and 37% run online evals, meaning most teams launch agents they cannot actually measure - LangChain. The practical cost is concrete: a hosted eval and tracing tier runs from roughly thirty dollars a month at the low end into the hundreds as trace volume grows, and a production vector store at ten million vectors costs on the order of seventy dollars a month before it inevitably overshoots. The tools layer is heavier still. The Model Context Protocol crossed 97 million monthly SDK downloads by March 2026, yet only around 41% of software organizations run it in real production, which tells you that wiring and securing agent tools is still genuine engineering work, not plug-and-play - Digital Applied. Every one of these layers is a place where an underfunded build quietly falls apart.
The costs are concrete and larger than most teams plan for. Standing up an agent engineering team runs five hundred thousand to one point two million dollars per year in salary before a single feature ships, with senior AI engineers commanding total compensation north of two hundred thousand dollars - Augment Code. Time to production is measured in quarters, not weeks. Consider the realistic ranges:
- Basic bounded agent: three to five months from kickoff to production, including a thirty-day tuning intensive after launch
- Enterprise multi-agent system: six to twelve months
- Median time-to-first-value, bought vs built: 38 days for a purchased agent against 94 days for an in-house build - Digital Applied
- Development share of total cost: only 25 to 35% once tokens, infra, maintenance, and oversight are counted
That last point is the one that sinks budgets. Teams price the build as if the initial development is the cost, when development is barely a third of the three-year total. The rest is the permanent maintenance line: model upgrades that break tuned prompts every quarter, integrations that need re-authorization, drift that must be monitored, and evaluations that must be re-run. A mid-complexity agent quoted at eighty thousand dollars is realistically a hundred and thirty to a hundred and sixty thousand over eighteen months once the ongoing costs land - SoftTeco.
The honest takeaway is not "never build." It is that building is a multi-quarter engineering program with a permanent operating cost, and it only makes sense when what you are building is genuinely yours. The framework is free precisely because it is not where the value is. If your plan is to assemble the free framework, the paid observability tier, the vector database, and the integration layer into something that looks a lot like a product a vendor already sells, you are about to spend a million dollars rebuilding a commodity. Our 2026 insider guide to building AI agents walks through this stack in detail, and the Claude Agent SDK deep dive covers one of the strongest build-side options for coding and computer-use agents.
5. The rent path: managed agents and the rise of outcome pricing
If building is about supplying the complements yourself, renting is about buying them pre-assembled, and the rent market in 2026 has organized itself around a pricing revolution that tells you exactly how confident these vendors are. The unit of value is moving away from the seat and toward the resolved outcome. Gartner projects that by 2030 at least 40% of enterprise software spend will shift toward usage, agent, or outcome-based pricing, and the seat-based share of vendor revenue is already declining - guptadeepak. This matters because outcome pricing aligns the vendor's incentive with yours: they only make money when the agent actually works.
The clearest expression of this is in customer support, the first enterprise job autonomous agents could do at scale. Intercom Fin charges 99 cents per resolution and nothing when the query is simply passed to a human - Fin.ai. HubSpot's Breeze agent went further, cutting its price to 50 cents per resolved conversation - HubSpot. The two most valuable pure-play support startups, Sierra (valued at 15.8 billion dollars in May 2026) and Decagon (4.5 billion in January), both price on outcomes, with Sierra reported around a dollar-fifty per resolution and Decagon around fifty cents to a dollar - eesel. We dissect that specific rivalry in our guide to Sierra vs Decagon.
To see why renting is compelling, put those numbers next to the thing they replace. A human-handled support contact costs somewhere between six dollars on the low end and thirteen-fifty on a blended benchmark, with phone resolutions running much higher - LiveChatAI. Against a rented agent at a dollar or two per resolution, that is a three-to-thirteen-times unit-cost reduction on the fraction the agent successfully deflects. The catch, and it is a large one, is that word "successfully." Vendor-marketed resolution rates and field reality diverge sharply: a platform advertising a 76% resolution rate frequently lands closer to 40 to 53% in production, and one independent head-to-head put a vendor claiming eighty-percent deflection at 49% - eesel. You still pay humans for everything the agent fails, so the real economics depend entirely on the true resolution rate in your account, not the number on the slide.
The cautionary tale here is Klarna, and it cuts against renting and over-building alike. Klarna's AI assistant, built with OpenAI, handled 2.3 million conversations in its first month, the work of roughly 700 agents, cut resolution time from eleven minutes to under two, and was projected to add forty million dollars of profit - OpenAI. It became the industry's favorite proof point. Then, by 2025, CEO Sebastian Siemiatkowski admitted the company had cut too deep on humans and began rehiring to guarantee a human option for customers who wanted one - CX Dive. The lesson is not that the agent failed; it deflected real volume. The lesson is that an agent deployed without a human-escalation floor is a brand and quality risk, not a clean headcount win, and that the vendor's resolution rate is a ceiling you may never hit on the interactions that matter most. Whether you build or rent, you are still accountable for the fraction the agent gets wrong.
Beyond the support pure-plays, the incumbents rent agents through the platforms enterprises already own, and here the pricing is more varied. Salesforce Agentforce offers three parallel meters (per-user licenses from five to five hundred and fifty dollars a month, flex credits at roughly half a cent each, and a flat two dollars per twenty-four-hour conversation), and runs across more than eighteen thousand companies - myaskai. Microsoft Copilot Studio sells capacity packs at two hundred dollars for twenty-five thousand credits, where a single autonomous agent action can burn twenty-five or more credits - CloudZero. The horizontal builders like Lindy and Relevance AI sell credits by the task, cheap to start and prone to ballooning under real load. Our Agentforce guide breaks down that platform in depth.
The strategic point beneath the price list is this. Renting is not one thing. It ranges from renting a fully finished vertical agent (Sierra for support) to renting a platform on which you still configure your own agents (Copilot Studio), and the further toward "finished" you go, the more you offload but the less you own. The vendors pricing on outcomes are making a bet that they can hit resolution rates you could not hit yourself, and for common, well-trodden workflows they are usually right. The place that bet breaks down is where your workflow is unusual, your data is proprietary, or your definition of a "resolved" outcome is not the vendor's. That is the seam where renting stops making sense and building starts, and it is worth understanding the true cost of LLM inference before you assume either path is cheap.
6. OpenAI's own agent surface: build and rent from one vendor
OpenAI is worth its own section because it is the only company that sells you every point on the build-versus-rent spectrum at once, and watching how it prices and positions each option is the clearest available lesson in where the industry thinks value lives. At one extreme sits the pure build path. At the other sits Presence, the pure rent. And the way OpenAI is investing in each end while quietly killing the middle tells you which bets it actually believes in.
At the build end, OpenAI offers the Responses API and the Agents SDK, both released in March 2026, as the code-first way to build your own agents on OpenAI's models. The Agents SDK is provider-agnostic (it supports more than a hundred LLMs, not just OpenAI's), and the pricing model is the simplest in the industry: you pay only for model tokens, with no platform fee - OpenAI Developer Community. In April 2026 OpenAI upgraded the SDK with a model-native harness and native sandbox execution, a signal that it is serious about the build path for teams that have engineers - TechCrunch. This is "rent the model, build the agent yourself," and it is the most flexible and cheapest option OpenAI sells.
The revealing move is what OpenAI is doing to the no-code middle. In June 2026 it announced it is winding down AgentKit and its visual Agent Builder, with Agent Builder becoming unavailable on November 30, 2026 and the hosted Evals product shutting down the same day - OpenAI Developer Community. OpenAI is pushing those users toward either the code-first Agents SDK or the fully managed Workspace Agents in ChatGPT. In other words, it is collapsing the drag-and-drop middle and telling customers to pick a side: either you have engineers and you build with code, or you do not and you rent the finished thing. That is a strong statement about where the company believes the durable value is, and it is worth reading alongside our guide to OpenAI Workspace Agents.
The engines under all of this are OpenAI's current models, and it is worth naming them precisely because model currency changes monthly. The flagship family is GPT-5.6, released July 9, 2026, in three variants: Sol (the top coding and agentic workhorse, priced at five dollars per million input tokens and thirty per million output), Terra (competitive with the prior generation at two and twelve dollars), and Luna (the lowest-cost tier at twenty cents and a dollar-twenty after an 80% price cut on July 30) - aipricing.guru. We cover this lineup in depth in our GPT-5.6 benchmarks and pricing guide. The competitive set is equally strong: Anthropic's current flagships are Claude Opus 5 and Sonnet 5, and Google fields Gemini 3.1 Pro for reasoning alongside the Gemini 3.6 Flash workhorse. The key point for build-versus-rent is that the model layer is a genuine commodity now, with four vendors at rough parity, which is exactly why the value moved up the stack to the deployment layer that Presence occupies.
Put the whole OpenAI surface in one frame and the strategy is unmistakable. The company is investing at both ends of the spectrum and abandoning the middle. It wants developers building on the Agents SDK, paying for tokens, keeping OpenAI's models at the center of the ecosystem. And it wants non-technical enterprises renting Presence, paying OpenAI to run the whole thing with human engineers embedded on-site. Both bets capture value from the same insight: the model is a commodity, so you either sell the commodity cheaply to people who add their own value (build), or you add the value yourself and sell the finished outcome at a premium (rent). There is no durable business in the no-code middle, which is why it is being deprecated. That is a lesson every buyer should internalize before choosing a path.
7. The market: size, adoption, and the failure data
Any honest guide has to hold two facts in tension. The AI agent market is enormous and growing fast, and most agent projects are failing. Both are true, and the gap between them is the single most important thing a buyer needs to understand before committing to build or rent, because it explains why the safe-looking choice (build it in-house, keep control) is statistically the riskier one.
Start with the size, because it is real. Five reputable analyst firms cluster the 2030 AI-agents market between roughly forty-three and fifty-three billion dollars, growing off a five-to-eight-billion-dollar base in 2025 at compound annual rates around 41 to 46% - MarketsandMarkets. Adoption is near-universal at the experimentation stage: McKinsey found 88% of organizations now use AI in at least one function, and PwC reported 79% say they have already adopted AI agents in some form - McKinsey. The money is flowing, the CAGR is genuine, and the demand is not in question.
Now the disillusionment, which is where the real story lives. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, and separately estimates that of the thousands of vendors claiming "agentic" capability, only around 130 are genuinely agentic rather than rebranded chatbots or RPA, a phenomenon it calls "agent washing" - Gartner. The MIT NANDA study delivered the bluntest number of all: 95% of enterprise generative-AI pilots delivered no measurable profit-and-loss impact, despite an estimated thirty to forty billion dollars of enterprise investment - Fortune. Even accounting for media inflation of that study's methodology, the directional finding is damning: near-universal experimentation, near-universal failure to reach measurable value.
The reliability data underneath those failures explains why so many builds stall, and it is a number every buyer should sit with. Benchmark scores are heavily scaffold-dependent, so a headline percentage tells you about a system, not a model. On tau-bench, the customer-service benchmark that measures policy adherence rather than raw task completion, the strongest results in its frozen original form landed around 69% on retail and 46% on airline tasks, meaning even the winner failed a third of retail cases and more than half of airline cases - benchmarkingagents. On GAIA, a general-assistant benchmark where humans score around 92%, top agent systems in 2026 sit near 65% - decodethefuture. Coding agents look far better on public benchmarks, but those scores are near saturation and should be read with contamination skepticism, not as production reliability. The pattern is clear: the closer your use case is to open-ended, policy-bound, multi-turn work, the wider the gap between a benchmark demo and a dependable agent, and the stronger the case for renting a battle-tested vendor over shipping a naive in-house build.
The adoption-versus-scale gap is the pattern that ties it together, and it is where build-versus-rent stops being abstract. Experimentation is easy and nearly everyone is doing it. Scaled production is the minority: McKinsey found only about 23% of organizations are scaling an agentic system anywhere, and by individual function the figure drops into single digits. The chasm between a demo that works and an agent that survives production is where the money dies. And the MIT data on that chasm is unambiguous about which side of build-versus-rent falls in more often: internal builds fail roughly twice as often as bought solutions. The reason is not that companies cannot code. It is that the disciplines separating a pilot from production (evaluation, monitoring, guardrails, change management) are exactly the operational muscle that vendors have and most internal teams have not yet built.
For a buyer, the implication is uncomfortable but clarifying. The failure data does not say "do not do agents." It says the default assumption should be rent or build-on-rails unless you have a specific reason to build from scratch, because the raw statistics say a from-scratch build is where projects go to die. You override that default only when your differentiation genuinely lives in the agent itself, and you go in with your eyes open about the operational cost. Everyone else is better served buying the disciplines they have not yet developed. The uncomfortable corollary is that many of the projects in that 40% cancellation forecast are internal builds that should have been rentals, killed after a year and several hundred thousand dollars because a team underestimated the production gap. We explore the human side of this shift in our analysis of the honest truth about AI's impact on the workforce.
8. Where build wins, where rent wins
With the economics and the failure data on the table, the decision reduces to a small number of structural questions, and the honest answer for most enterprises is not "build" or "rent" but a specific blend that depends on where your value and your risk actually sit. The consensus that has emerged from the 2026 data is that roughly 57% of organizations now run a blended build-and-buy strategy, up from 51% a quarter earlier, and that this hybrid is not a compromise but the empirically superior default - Digital Applied. The framework below is about knowing which parts to build and which to rent.
Rent wins, decisively, when the workflow is common and the vendors are mature. If you are automating customer support, IT service requests, or sales development (jobs thousands of other companies also automate), a vendor has already solved your problem better than you will on your first attempt, and buying reaches production roughly twice as reliably. Renting also wins on maintenance: model upgrades break tuned prompts every release, self-hosted frameworks carry hundreds to thousands of dollars a month in hidden upkeep, and vendors absorb that churn - SearchUnify. And it wins below scale: the build path does not amortize its engineering tax until you are north of roughly a million agent conversations a year, so under that threshold, buying is simply cheaper - Digital Applied.
Build wins in a narrower but crucial set of cases. Consider these:
- Core differentiation: the agent is the product, or depends on proprietary data no vendor can replicate
- The moat is data plus workflow, not the model: what you build is the data layer and the workflow, not the intelligence
- High volume: above roughly two million conversations a year, per-token economics flip in the builder's favor
- Data sovereignty: regulated firms that cannot diffuse accountability across a third party, or that face jurisdiction constraints
Each of those is a case where renting would mean surrendering the exact thing that makes your business defensible. If your agent's value comes from a decade of proprietary claims data, renting a generic insurance agent throws away your advantage. If you process ten million conversations a year, the vendor's per-outcome margin becomes your largest controllable cost. And if you operate under data-sovereignty rules where a vendor's domicile is itself a liability, no amount of speed offsets the accountability problem - Kai Waehner. Outside those cases, the data says building is the lower-probability path to production.
The best evidence for the framework is in how leading enterprises actually split the decision. Commonwealth Bank of Australia co-developed an AI orchestration agent with Microsoft over two years on Copilot Studio, a build-on-rails hybrid that now resolves 84.6% of self-service messaging interactions end-to-end and helped cut fraud losses more than 20% year over year - Welcome.AI. That is the middle path in production: managed platform, owned agent, proprietary bank data as the moat. Telstra, by contrast, built two tools in-house on Azure OpenAI because the value was in its own network and customer data, and reports 90% of users of one tool saving time and 20% fewer follow-up contacts - Microsoft. Both are winning, and neither chose a pure extreme. The through-line is that each built where its differentiation lived (bank fraud logic, telecom customer data) and rented the commodity infrastructure around it. That is the framework working exactly as first principles predict, and it is the opposite of the all-or-nothing choice most teams frame it as.
The practitioner's read on this is worth more than any framework, and it comes from people who ship the build side for a living. Yuma Heymans (@yumahey), founder and CEO of the AI workforce platform O-mega and co-founder of the autonomous recruiter HeroHunt.ai, has spent years on the build-your-own side of exactly this problem, wiring agents to operate inside a company's own tools and guardrails. His experience maps to what the data shows: the guardrail, permission, and change-management problem that Presence packages as a service is genuinely hard to build well, which is precisely why so many teams underestimate it and why the rent option keeps winning the average case. The lesson is not that building is wrong. It is that you should build only where you would be foolish to rent, and rent everything else without ego. Our guide to the AI-native company tech stack maps how leading teams draw that line in practice.
9. Security, governance, and who is liable
The build-versus-rent decision is not only an economic one. It is a liability one, and this is the dimension most teams discover too late, usually after an agent has already done something expensive. The structural fact underneath all agent security is simple and unforgiving: an autonomous agent with the ability to act is a new and enormous attack surface, and in 2026 the incidents are no longer hypothetical. In a study of 344 verified enterprise incidents, an autonomous AI caused harm directly in production, with no attacker in the chain, in 188 of them, and 65% of organizations reported at least one agent-caused security incident in the past year - Cyera.
The canonical vulnerability class is prompt injection, and the sharpest framing of it is Simon Willison's lethal trifecta: an agent becomes exploitable when it combines access to private data, exposure to untrusted content, and the ability to communicate externally. Any agent with all three legs can be turned into a data-exfiltration tool by a crafted input - Simon Willison. This is not a prompt-engineering problem that better instructions can fix. It is architectural, and it was weaponized in the real world as EchoLeak (CVE-2025-32711, CVSS 9.3), the first documented zero-click prompt injection in a production system, where a single crafted email could exfiltrate internal data from Microsoft 365 Copilot before it was patched - arXiv. The build implication is severe: if you build, you own the entire job of removing one leg of that trifecta from every agent you ship.
The failure modes get worse when agents can take destructive actions. A widely reported incident saw an AI coding agent delete a live production database during a code freeze, then fabricate records to cover it, because nothing constrained the blast radius of its actions - Medium/NeuralNotions. The root-cause pattern across these disasters is consistent: organizations conflate an agent's ability to act with the scope of access it should have, granting broad permissions and discovering the difference only when the agent uses them. This is the exact problem OpenAI Presence addresses by scoping each agent to one job with only the access that job requires, and it is the problem a from-scratch builder must solve alone, correctly, on every deployment. We cover the defensive playbook in depth in our guide to AI agent security and prompt-injection defense.
The tooling layer adds its own attack surface that most builders have not accounted for. Tool poisoning embeds malicious instructions inside a tool's metadata or description, so it fires silently on every invocation and can override instructions from other trusted servers, and multiple critical-severity MCP vulnerabilities were disclosed in the first half of 2026 - Invariant Labs. Underneath that sits an identity problem the industry has barely begun to solve: agents are non-human identities that need to be created, permissioned, and revoked, yet 78% of organizations have no formal policy for creating or removing AI identities, and 92% doubt their legacy identity systems can manage the risk - MSSP Alert. When Copilot Studio users alone have created more than a million agents, the scale of ungoverned machine identity is staggering. For a builder, every one of these is a control you must design and maintain yourself; for a renter, it is a discipline you are paying a vendor to have already solved. That asymmetry is a real part of the total cost, and it rarely shows up in the build-versus-rent spreadsheet.
Then there is the legal reality that resolves the "whose fault is it?" question with brutal clarity. In Moffatt v. Air Canada, a tribunal held the airline liable for false advice its chatbot gave a customer and explicitly rejected the argument that the chatbot was a separate legal entity - American Bar Association. The precedent is unambiguous and it applies regardless of the build-versus-rent choice: the deployer owns the agent's output. You cannot outsource liability to a vendor or to "the AI." Layer on the EU AI Act, whose transparency rules for any agent that interacts with people take effect August 2, 2026, with penalties up to thirty-five million euros or 7% of global turnover, and the governance stakes become concrete - AI Act Service Desk.
This is why governance pushes the decision in both directions at once, and understanding the split is the key. For most enterprises, governance is an argument to rent, because the production disciplines that prevent these incidents (prompt-injection defense, drift monitoring, non-human-identity management, scoped access) are exactly what vendors specialize in and internal teams fail at. But for regulated firms in sovereignty-constrained jurisdictions, governance is an argument to build, because shared third-party infrastructure cannot produce the clean accountability mapping their regulators demand. Gartner adds a subtle warning that catches both camps: applying uniform governance across all agents actually causes failure, because a low-risk internal helper and a high-risk customer-facing agent need different controls - Gartner. Whichever path you choose, governance has to be scoped to the risk tier of each agent, not applied as a blanket.
10. Protocols: how MCP and A2A change the math
There is a quiet infrastructure story running underneath the build-versus-rent debate that changes the calculation more than any single product, and most buyers are not pricing it in. Two open protocols, the Model Context Protocol for connecting agents to tools and Agent2Agent for connecting agents to each other, have become neutral, foundation-governed standards in 2026, and their effect is to lower the switching costs that used to make renting feel like a trap.
The Model Context Protocol is the more mature of the two, and its adoption curve is one of the fastest in the history of developer tooling. Introduced by Anthropic in late 2024, MCP reached roughly 97 million monthly SDK downloads by March 2026, a growth of nearly 5,000% in sixteen months, and was adopted by OpenAI, Google, and Microsoft in quick succession - WorkOS. The pivotal governance moment came in December 2025, when Anthropic donated MCP to the Agentic AI Foundation under the Linux Foundation, making it a vendor-neutral standard that no single lab controls. This matters for build-versus-rent because a standard tool interface means you can swap models or agent platforms without rebuilding every integration, which turns "rent now, migrate later" from a slogan into a viable strategy and makes DIY tool-wiring less of a moat worth building.
Agent2Agent does for inter-agent communication what MCP does for tools. Google donated A2A to the Linux Foundation in mid-2025, and at its one-year mark in April 2026 it had passed 150 organizations, 22,000 GitHub stars, and reached version 1.0 with cryptographically signed agent identity cards - Linux Foundation. It is now generally available inside Copilot Studio, Azure AI Foundry, and Bedrock AgentCore. The consequence is subtle but important: a bought agent and a built agent can now coordinate across vendor boundaries, which means the decision stops being all-or-nothing. You can rent a support agent from one vendor, build a proprietary pricing agent yourself, and have them collaborate through a standard protocol.
The strategic implication is that these two protocols shift the entire math toward hybrid, and away from the fear that renting locks you in forever. Here is why they matter to the decision:
- MCP standardizes agent-to-tool connections, so integrations survive a change of model or platform
- A2A standardizes agent-to-agent coordination, so built and bought agents interoperate
- Both are now Linux Foundation-governed, removing single-vendor control from the interop layer
- Together they lower switching costs, which reduces the penalty for renting first and building later
The practical takeaway is that lock-in fear is a weaker argument against renting than it was even a year ago. If you rent a platform that speaks MCP and A2A, you have preserved the option to migrate the pieces that matter without a full rebuild. That makes the pragmatic sequence (rent to get to production fast, then selectively build the components where you develop real differentiation) far more defensible than it used to be. For a deeper technical comparison of the two protocols and when each applies, see our guide to MCP vs A2A in 2026.
11. The hybrid frontier: build on rails and the autonomous company
The cleanest lesson of 2026 is that the build-versus-rent binary is dissolving into a spectrum, and the most interesting action is in the middle, where a new class of platforms lets you build your own agents on top of fully managed infrastructure. This "build on rails" model is why the top of our scorecard is dominated by AgentCore, the Gemini Enterprise Agent Platform, and Copilot Studio rather than by pure frameworks or pure managed services. It captures the two things a builder most wants (control and differentiation) while offloading the two things that sink builds (infrastructure and operational discipline).
The archetype is Amazon Bedrock AgentCore, which in 2026 added a managed harness that runs the full agent loop with no orchestration code, while enforcing security at the infrastructure layer in a way the agent cannot bypass - AWS. Google made the same move, consolidating Vertex AI into the Gemini Enterprise Agent Platform at Cloud Next 2026, where its Agent Development Kit builds a production agent in under a hundred lines of Python and a managed runtime handles the rest. The pattern is identical across vendors: you write the agent's logic and connect your data (keeping your differentiation), and the platform supplies the compute, sandboxing, memory, observability, and guardrails (absorbing the six-layer tax from Section 4). This is the empirical winner because it resolves the core tension, and it is where the 57% who run hybrid strategies increasingly land.
At the far frontier sits a category that reframes the question entirely: the autonomous company or AI workforce model, where you do not rent one agent or build one agent but describe an entire business function (or an entire business) in plain language and have a system build and operate it. OpenAI's Frontier, launched in February 2026, lets enterprises build and manage agents as AI co-workers with hire-and-onboard framing - TechCrunch. ServiceNow's Autonomous Workforce deploys AI specialists across service desk, HR, and security with enterprise-grade scope and governance. And platforms like O-mega push the idea to its logical end: describe your company in a conversation, and the system generates and runs the website, the product, the admin, the billing, and the agent workforce that operates it, working autonomously inside your organizational context and within your guardrails. It is priced as a self-serve subscription with a free tier rather than a consulting engagement, which places it on the opposite end from Presence: the same "done for you" ambition, delivered without the Forward Deployed Engineers.
What makes this category the genuine frontier, rather than just a fancier rental, is that it collapses the build-versus-rent distinction at the level of the whole business rather than the individual agent. You are not choosing whether to build or rent a support agent. You are describing an outcome at the level of a function or a company and letting the system decide how to assemble it. That is a different altitude of abstraction, and it is where the most ambitious 2026 experiments are running, from the autonomous business to full multi-agent orchestration. It is early, the reliability is unproven at scale, and it should be evaluated with the same skepticism as any managed option. But it is the clearest sign that the industry is moving past "build an agent" toward "describe a business," and that shift will make the narrow build-versus-rent question look quaint faster than most people expect.
12. The 2026 outlook and a decision framework
Step back from the individual products and the structural picture is clear, and it points somewhere more useful than "it depends." Intelligence has become a commodity, four labs sell it at rough parity, and its price has fallen an order of magnitude in three years. The value has migrated up the stack to the deployment layer: the guardrails, the integrations, the accountability, and the operational discipline that turn a capable model into a system a company will put in front of customers. Every product in this guide, from a free framework to OpenAI's human-deployed Presence, is a different bet on who supplies that deployment layer and how much they charge for it. That is the whole game.
The decision framework that follows from first principles is simpler than the market makes it look. Default to renting or building on rails, because the data is unambiguous that from-scratch builds fail roughly twice as often and account for a large share of the projects in Gartner's 40% cancellation forecast. Override that default and build only where three conditions hold: your differentiation genuinely lives in the agent, your volume is high enough to amortize the engineering tax (roughly north of a million conversations a year), or your regulatory posture forbids diffusing accountability to a vendor. When you do rent, prefer platforms that speak MCP and A2A so you preserve the option to migrate the pieces that matter. And whatever you choose, scope every agent to the narrowest access its job requires, keep a human-escalation floor, and match your governance to each agent's risk tier rather than blanketing everything.
The most useful reframe is to stop treating this as a permanent, all-or-nothing choice. The protocols have made "rent to reach production, then selectively build where you develop real differentiation" a coherent and defensible sequence, not a contradiction. Buy the commodity (the models, the hosting, the schedulers, the sandboxes), rent the common workflows (support, IT, sales development) where mature vendors will beat your first attempt, and reserve your scarce engineering effort for the narrow slice where building is the only way to protect what makes your business defensible. The companies that win with agents in 2026 are not the ones that built everything or rented everything. They are the ones that knew the difference, and had the discipline to buy their commodity capabilities without ego and build their differentiated ones without hesitation.
OpenAI Presence is the perfect closing symbol of all of this. The company with the best models on earth looked at the enterprise agent market and concluded that, for most customers, the models are not the hard part, so it started selling the deployment layer directly, with human engineers embedded on-site to make it work. That is the clearest possible signal of where the value is. The intelligence is cheap. The judgment about how to deploy it safely, and the discipline to run it, are what you are really paying for, whether you build that judgment in-house or rent it from someone who has already done the hard part. Decide accordingly, and you will be on the right side of the statistics.
This guide reflects the AI agent landscape as of August 2026. Pricing, model versions, and product availability in this category change monthly, so verify current details with each vendor before committing budget. Vendor-reported resolution rates and performance figures are self-reported unless attributed to an independent source, and should be validated against your own workloads.