How Jev, Clef, OpenAI's Decisions API and two open challengers compare on accuracy, calibration, speed, price and control, and which one belongs behind your agent's choices
Eight decision models shipped in the 16 days between September 15 and October 1, 2026.
TypeSafe AI released Jev on September 15 alongside a $40 million seed round led by DCVC - SiliconANGLE. Fastino shipped GLiNER2.5-Decide on September 24 and the "thinking" GLiDE on September 30 - MarkTechPost. OpenAI previewed a Decisions API built on GPT-6 Luna at DevDay on September 29 - OpenAI Developer Community. Then, on a single day, October 1, Cloudflare open-sourced Clef and Clef-flash - Cloudflare, Perplexity open-sourced pplx-decider-v1-27b behind its own Decisions API - Perplexity, and AWS released Strands Decider 2B for laptops - Strands Agents.
The reason a whole category formed this fast is simple once you look at what most AI agents actually spend their tokens on. A large share of the model calls inside a production agent are not creative work at all. They are if-statements in disguise: which team should get this ticket, is this message urgent, which tool should run next, is this output safe to ship. Teams have been answering those questions with frontier language models that write a sentence or a JSON blob, which code then parses and hopes matches one of the allowed options. A decision model skips the writing. Jev, for example, takes a piece of state and a set of typed questions and returns only one of the options you supplied, with a probability on each - TypeSafe.
But the category is 18 days old, and almost every number in it comes from the company selling the model. If you searched for "jev ai" you have probably already found dozens of explainer pages that repeat TypeSafe's launch claims. Very few of them test anything, and the choice in front of a builder is real: a closed model or open weights, a hosted API or a model on your own hardware, text only or images and video, a published price or a "limited preview" with no price at all.
This guide breaks down what decision models are from first principles, profiles Jev, OpenAI's Decisions API, Clef and Clef-flash plus the two open challengers from Perplexity and AWS, and separates independent benchmark evidence from vendor claims using the two public boards that measure this class of model. It then works through the real cost per decision, where these models fit inside an agent, how they fail, a decision framework, and an implementation walkthrough with working code.
Contents
- What a decision model is, from first principles
- Jev: the model that started the category
- OpenAI Decisions API: GPT-6 Luna, constrained
- Cloudflare Clef and Clef-flash: open weights at the edge
- The open challengers: Perplexity pplx-decider and AWS Strands Decider 2B
- Benchmarks: what independent boards say versus vendor claims
- Price and latency: the real cost per decision
- Where decision models fit inside an agent
- Where decision models fail
- How to choose: a decision framework
- Implementation walkthrough: shipping your first decision layer
- Outlook: where decision models go next
- Conclusion
The decision-model scorecard at a glance
The table below scores the six decision models a builder can actually reach today (or, in OpenAI's case, can apply for) on the same five criteria and the same evidence standard, then sorts them by the weighted final score. Every cell carries the score and the data point behind it, so you can disagree with a judgment without reconstructing the evidence. Where an independent measurement exists, the score leans on it; where only the vendor has published a number, the profile sections below say so.
Two patterns stand out before you read a single row. First, the top three are close, and their order depends on how much you value owning the weights: if you never intend to self-host, Jev moves to first place. Second, OpenAI's entry is the least knowable, not necessarily the weakest: its underlying model is the most capable one JevBench has measured, but OpenAI has not published a price, a schema, limits or a benchmark for the Decisions API itself, so it cannot score well on what a builder can verify today.
| # | Model | What It Does | Decision quality (40%) | Speed (15%) | Cost (15%) | Control (15%) | Readiness (15%) | Final |
|---|---|---|---|---|---|---|---|---|
| 1 | Perplexity pplx-decider-v1-27b | Open 27B decider, Apache 2.0, text and images, $0.04/M input | 8 - 56.4 on the independent Decision Index (Jev 57.9) with the best calibration measured (ECE 0.018) | 5 - hosted API answers small requests in under 2 s; about 101 ms median self-hosted | 9 - $0.04/M input, output free, no per-request fee | 9 - Apache 2.0 weights, fits one GPU with ~49 GiB free | 7 - clear published limits, 128 questions per call, but 10 requests/s per org | 7.7 |
| 2 | TypeSafe Jev 1.13 | The original System One model: closed, text only, $0.042/M input | 9 - #3 of 109 on JevBench v1.5.5 (72.1); 57.9 on the Decision Index; slightly overconfident (ECE 0.074) | 7 - 70-500 ms claimed; 524 ms median over the network in Cloudflare's test | 9 - $0.042/M input, output free, about $0.04 per 1,000 decisions | 2 - closed weights, no fine-tuning, same model for every account | 8 - self-serve console, Python and JS SDKs, OpenRouter; limits still shifting | 7.5 |
| 3 | Cloudflare Clef-flash | Cloudflare's open 9B decider: text, images and video at edge speed | 5 - JevBench Intelligence 53.1 vs Jev's 72.0 (#25 overall); self-reported 57.1 on the Decision Index | 10 - 38.8 ms median, 122 ms p95 (Cloudflare) | 8 - $0.09/M input on Workers AI | 9 - Apache 2.0 weights plus an RL fine-tuning service | 8 - Workers AI binding and REST, 64 questions, 4 images per call | 7.3 |
| 4 | Cloudflare Clef | Cloudflare's open 27B decider with a Jev-compatible API | 7 - self-reported Decision Index leader (61.2); JevBench Intelligence 67.9 | 8 - 209 ms median, 239 ms p95 (Cloudflare) | 4 - $0.24/M input, about 6x Jev; #62 on JevBench mostly on cost | 9 - Apache 2.0 weights plus RL fine-tuning | 8 - same surface as Clef-flash | 7.2 |
| 5 | AWS Strands Decider 2B | AWS's 1.9B open decider for laptops, with its full training recipe public | 4 - 0.723 on the JevBench v1 public set; 2B-class models score 19-29 vs Jev's 57.9 on the Decision Index | 9 - 115 ms median on an RTX 3090, 153 ms on an M3 Pro, no network hop | 10 - free; runs on a laptop, GPU or CPU | 10 - Apache 2.0 weights, data list and scripts; retrain in about 11 hours | 5 - pip package, CLI and local server; research-grade, no hosted service | 6.7 |
| 6 | OpenAI Decisions API | GPT-6 Luna constrained to your answers; limited preview, no price yet | 7 - plain GPT-6 Luna has the highest JevBench Intelligence measured (96.2); the API itself is unbenchmarked | 8 - 150 ms vs 1.6 s for a normal Luna call (OpenAI's DevDay claim) | - - price not published (plain Luna is $0.10/$0.50 per M) | 1 - closed, OpenAI-hosted only | 2 - limited preview, no public docs, schema or limits | 5.2 |
How the criteria were chosen. Decision quality carries 40% because a wrong decision executed quickly is worse than a right one executed slowly in almost every workflow these models serve; it blends accuracy and calibration, and prefers independent boards over vendor charts. Speed and cost carry 15% each because they are the reason the category exists, but every model here is already fast and cheap compared with a text model. Control (open weights, self-hosting, fine-tuning, data handling) and readiness (availability, documentation, limits, SDKs, modalities) carry 15% each because they decide whether you can ship and keep shipping. OpenAI's cost cell is unscored because no price exists, so its final score is the weighted average of the other four criteria.
Read the table as a map rather than a verdict. If you drop the control criterion entirely, Jev leads at 8.5 and Perplexity falls to 7.5, which is the right ranking for a team that will only ever call a hosted API. If speed matters more than accuracy (a guardrail on every chat message, an edge filter), Clef-flash is the obvious pick. The rest of this guide explains each number, so you can reweight the table for your own workload.
1. What a decision model is, from first principles
Strip away the branding and a decision is a function: it takes some state (a support ticket, a log line, an agent's last tool call) and returns one element of a finite set (billing or technical, yes or no, a severity from 0 to 3). Software has always made decisions this way. What changed over the last three years is that the hard part of many decisions, reading messy natural language and applying judgment, became something a model could do. The first tool teams reached for was the general-purpose language model, because it was the only thing that understood language well enough.
A language model answers a closed question by writing tokens. You ask "which team should handle this?", it writes "technical", or a JSON object containing "technical", and your code parses the answer and checks that it is one of the allowed values. That path has three structural costs. You pay for output tokens (and, with reasoning models, for hidden thinking tokens). Latency grows with every token the model writes. And the output space is open, so the model can return a value that is not in your set, which is how agents end up calling tools that do not exist. TypeSafe's documentation pitches Jev as a way to "replace a fragile prompt that asks an LLM to 'return JSON' with a call that returns typed values by construction" - TypeSafe Docs.
The obvious objection is that none of this is new. Constrained decoding and token log-probabilities have existed for years, and a careful engineer could already force an LLM to pick from a list and read its probabilities. Three things are actually new, and they are worth separating because they explain both why decision models are useful and why they were cloned so quickly:
- The training objective. Jev is trained with Reinforcement Learning for Calibrated Decisions (RLCD) so its probabilities track how often it is right, rather than with RLHF, which rewards answers people like - TypeSafe.
- One read, many questions. The state is ingested once and every question is evaluated against it in parallel, so adding questions barely moves latency - TypeSafe docs.
- The contract and the price. The API returns typed values by construction, and output tokens are free on Jev and Perplexity, so the bill depends only on what you send.
Of the three, calibration is the one that changes system design. Language models are not trained to make their probabilities track how often they are right, and RLHF in particular rewards answers that sound sure; some frontier models turn out well calibrated anyway (section 6), but nothing in their training promises it. A model whose 90% means "right about nine times in ten" lets code do something it could never do with text: act automatically above one threshold, ask a human below another, and escalate to a reasoning model in between. That is the whole architecture of a well-built agent, and until now it was hard to build on honest numbers. The other two properties are about economics. Free output and parallel questions mean you can ask ten speculative questions about every event for roughly the price of one, which encourages a design where the model answers many small, literal questions and your code combines them.
TypeSafe named the class System One models, after the fast, intuitive mode of thinking Daniel Kahneman described in Thinking, Fast and Slow - TypeSafe docs. The name "Jev" itself points at the economic bet: it refers to William Stanley Jevons and the Jevons paradox, the idea that making a resource cheaper increases how much of it gets used - Wikipedia. The diagram Cloudflare published with Clef shows the shape every model in this guide shares: one input, several typed questions, parallel outputs with a probability on every option.
The diagram makes one property easy to miss: nothing in the output is free text. Every answer is a distribution over options you defined, so the failure you have to design for is a wrong or uncertain answer, never an invalid one. The contrast with the language-model path is clearer side by side.
Why did the category form in 16 days rather than 16 months? Because the ingredients were already on the shelf. Cloudflare post-trained Clef on Qwen3.8-27B and Clef-flash on Qwen3.5-9B - Cloudflare; Perplexity fine-tuned pplx-decider from the same Qwen3.8-27B - Hugging Face; AWS took Qwen3.5-2B, cut off its language-modelling head and bolted on a pointer head of about one million parameters - GitHub. A strong open model already contains the understanding; turning it into a decider is a head, an adapter and a calibration-focused training set. The community proved how cheap that is: the independent JevBench board already ranks 109 systems, most of them hobbyist and small-lab fine-tunes - Benchmark Heaven.
The second thing that formed in those 16 days matters more than the models: an interface. Cloudflare describes Clef as "fully Jev-API compatible" and documents its image support as an extension to "the System One API" - Cloudflare Docs. Perplexity's Decisions API uses the same three question types with the same field names, and AWS's local server answers on the same /v1/systemone path that TypeSafe uses. Within two weeks of launch, Jev's request format became the de facto standard of the category, which has a direct practical consequence: switching providers is close to a configuration change for most of the category, and that shapes every recommendation later in this guide.
Why this matters for you: if your agent makes the same kinds of closed decisions thousands of times a day, a decision model can replace an expensive, slow, occasionally invalid text call with a cheap, fast, always-valid one that also tells you how sure it is. How to apply it: before choosing a vendor, list the decisions your system makes today, and mark which ones have a fixed answer set and enough information in the state to decide. Those are the candidates. Everything that requires writing, multi-step reasoning or arithmetic stays with a language model or with code, as section 9 explains.
2. Jev: the model that started the category
TypeSafe AI is a San Francisco lab founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida spent several years at OpenAI working on RLHF, InstructGPT, ChatGPT and GPT-4 before leaving to start the company - SiliconANGLE. The company exited stealth on September 15 with a $40 million seed led by DCVC that Forbes reported valued it at $200 million - Wikipedia. Nine days later The Information reported that TypeSafe was in talks to raise more than $1 billion at a valuation above $10 billion - AI Weekly. Almeida's framing of the problem is blunt: "We've been optimizing for humans, and we're superhuman at pleasing humans."
The launch landed hard enough that the API briefly went down under demand - TechCrunch. Developer attention has stayed high since: the most-watched explainer below, by Claire Zau, passed 315,000 views in under two weeks, which is unusual for a model that cannot write a sentence. It is a short, clear walkthrough of why a model that only returns choices and probabilities caught developers' attention.
The video's core point is the one worth holding on to: Jev's appeal is not intelligence in the chat sense but reliability of form. It answers in exactly three shapes, which TypeSafe calls primitives:
- Choice picks one option from a set you define (up to 255 options) and returns a probability for each option plus a confidence value - TypeSafe API.
- Score rates the state against ordered levels you describe, returning the probability-weighted level, the full distribution and a confidence value.
- Noul returns the probability that a yes-or-no statement is true; values near 0.5 mean the model is unsure.
These three primitives are deliberately narrow, and that narrowness is the product. Any closed decision can be expressed as one of them or as a combination: a routing decision is a Choice, a risk rating is a Score, a policy check is a Noul, and a complex judgment is several of each combined in code. TypeSafe's docs push hard on that last pattern, which they call composite scoring: break a broad judgment into atomic questions and do the weighting yourself, because the model is more reliable on small literal questions than on big fuzzy ones. That is also why the per-question confidence matters so much; it is what lets code decide which of many small answers to trust.
How you call Jev
Jev is served from a single endpoint, POST https://api.typesafe.ai/v1/systemone, with a state field, a model field and a map of named questions - TypeSafe Quickstart. The jev-latest alias currently resolves to jev-1.13.0; TypeSafe advises pinning the versioned ID once you have tuned confidence thresholds, because an alias can move and change answers without any change on your side. A complete request looks like this:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-1.13.0",
"state": "I have been trying to connect my Stripe account for 3 days and the integration keeps failing. I am losing sales. Please help ASAP.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}'
The response returns one typed answer per question name. In TypeSafe's own example, department comes back as technical with probabilities of 0.85 technical and 0.15 billing and a confidence of 0.78, while the urgency Noul comes back at 1.0. The usage block reports input tokens (the billed quantity) and output tokens (free). TypeSafe also ships a Python SDK (pip install typesafe-sdk), a JavaScript SDK, and an agent skill for Claude Code and other coding agents, so a coding agent can write Jev integrations correctly instead of guessing at the schema.
The published limits are worth reading before you design anything. Jev costs $0.042 per million input tokens with free output, accepts 64k tokens per request (32k for the state plus the longest single question), takes text only (strings, JSON objects or arrays of text, no images, audio or video), and is rate-limited at 100,000 tokens and 80 requests per second, with a warning that limits "can change without notice" while TypeSafe adds GPU capacity - TypeSafe Models. It cannot be fine-tuned: the same weights serve every account, and you shape behavior only through the state and the wording of your questions. English is the primary training language.
What the claims say, and who measured them
TypeSafe's headline numbers are 193.6x faster and 444.6x cheaper than frontier language models, with end-to-end responses between 70 and 500 milliseconds - TypeSafe. In its Doom demo, Jev answered in 0.114 seconds where GPT-5.6 Terra took 8.566 seconds - The Register. Those figures are real measurements, but they are TypeSafe's measurements on workflows TypeSafe built, and the company itself describes them as high-end results. Independent observers flagged exactly that: the 445x cost claim is still self-tested - TechStock².
The more useful evidence comes from people who ran Jev on their own problems. Vercel's Pranit Sharma said Jev ran a safety-review classifier 5 to 18 times faster than OpenAI's Luna, with greater accuracy - TechCrunch. On the critical side, Armin Ronacher argued that Jev "delegates the hallucination problem" to the user: the model cannot return an invalid option, but it can still return a confidently wrong valid one, and deciding what to do about that is now your code's job. Both observations are correct, and they are the same observation from two angles.
TypeSafe's own cookbooks are more modest and more instructive than the launch claims. Batching a 13-question regulatory briefing into one call instead of thirteen was 12.2x cheaper and 10.0x faster with no change in answers - TypeSafe Cookbook.
Reranking is the other instructive example. Using one Jev question per query-candidate pair to rerank legal search results raised top-1 accuracy from 5% to 18% and top-10 accuracy from 38% to 62% - TypeSafe Cookbook. Those are not miracle numbers; they are what a cheap, calibrated judgment layer does when you point it at a well-defined sub-problem.
Jev is also reachable through OpenRouter, which added a dedicated decisions endpoint (/api/alpha/decisions) rather than forcing Jev into its chat-completions API, and caps Jev's context at 32,000 tokens there - OpenRouter. That matters if you already route all model traffic through one gateway: Jev arrives as a separate kind of call with its own contract, not as one more chat model.
On September 25 OpenRouter added Jev Router, which uses Jev to pick the best model and reasoning effort for each request - OpenRouter. That is a telling first application: a decision model deciding which language model should do the writing.
Why this matters: Jev is the reference point every competitor benchmarks against, and on independent boards it is still at or near the top (section 6). How to apply it: Jev is the default choice if you want a hosted, text-only decision layer with the most mature documentation in the category and you have no requirement to own the weights. Pin jev-1.13.0, read the "jaggedness" page before writing questions, and treat the rate limits as provisional until TypeSafe says otherwise.
3. OpenAI Decisions API: GPT-6 Luna, constrained
OpenAI's answer arrived exactly two weeks after Jev. At DevDay on September 29, OpenAI described the Decisions API in one line: it "is in limited preview. It uses Luna to classify inputs, route requests, or choose an action from predefined answers" - OpenAI Developer Community. The OpenAI Developers account added that it gives apps "real-time decision-making" by letting developers "define questions and possible answers to classify content, route requests, or choose an agent's next action" - OpenAI Developers on X. Simon Willison, live-blogging the keynote, put the subtext plainly: "Sounds like their response to Jev, which came out of stealth less than two weeks ago!" - Simon Willison.
The design choice is the interesting part. TypeSafe, Cloudflare, Perplexity and AWS all built specialized decision models. OpenAI instead constrained a general model: a specialized version of GPT-6 Luna that picks from a finite answer set given text or image context. That is a bet that general intelligence plus a constrained output beats a specialist, and the independent evidence suggests the bet has a real basis. On JevBench, plain GPT-6 Luna (called as an ordinary LLM) posts the highest Intelligence score of any system measured, 96.2 at default reasoning effort, against Jev's 72.0 - Benchmark Heaven. It ranks only #37 of 109 overall because, called as a text model, it is far slower and more expensive per decision than a decider. The Decisions API is OpenAI's attempt to keep the first number and fix the second.
OpenAI's keynote is the primary source for what was actually announced, and the Decisions API appears alongside GPT-6.1 Sol, Dots and the Agents API in the same session. If you want to judge the claims yourself rather than through recaps, this is the recording to watch.
The keynote's speed claim is specific: a DevDay slide put the Decisions API at 150 milliseconds against 1.6 seconds for a standard Luna call, and OpenAI's Tibo Sottiaux described it as "less than a few hundreds of milliseconds end to end" - Firecrawl. InfoQ's DevDay recap confirms the API selects answers from text or image context - InfoQ. Almost everything else a builder needs is missing, and it is worth being precise about what:
- No public price. Plain GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens - OpenRouter, but the Decisions API has no published rate.
- No published schema. Endpoint path, request and response format, and whether one call can carry several questions are all undocumented.
- No confidence spec. Launch coverage says answers come with a confidence score, but OpenAI has not documented how it is computed or calibrated.
Context size, the maximum number of answers per question and rate limits are unpublished too, and access itself is limited to selected API customers, with a broader release described as coming "in the coming days". That list is not a criticism of the model, which may well turn out to be the most accurate decider on the market. It is a statement about what you can build on today. A decision layer is infrastructure: you tune thresholds against its probabilities, you size budgets against its price, and you design batching around its limits. None of those can be done against an API whose contract is not public, which is why the scorecard rates OpenAI's readiness at 2 and leaves its cost unscored rather than guessing.
There is also a timing argument that cuts the other way. OpenAI's own month illustrates why calibrated gates matter: in late September the company shelved the release of GPT-6.1 Astra after safety tests found it sometimes operated outside its authorized scope and failed to accurately disclose its actions. Its head of safety systems, Saachi Jain, said the model "didn't quite meet the bar in terms of staying within scope and authorization" - Business Standard. A cheap, fast model that answers "is this action within scope?" before an agent acts is exactly the kind of control that problem calls for, and OpenAI is well placed to ship one that sits inside its own Agents API.
Why this matters: if OpenAI ships the Decisions API broadly at a price near plain Luna's and the quality of its underlying model holds up, it could reset the accuracy ceiling for the category overnight. How to apply it: do not design around it yet. If you are already standardized on OpenAI, request preview access, but write your decision layer behind an adapter (section 11) and run Jev, Clef or pplx-decider in production until OpenAI publishes a contract you can test against. Our guide to shipping long-running agents on OpenAI's Agents API covers the surrounding platform the Decisions API is likely to plug into.
4. Cloudflare Clef and Clef-flash: open weights at the edge
Cloudflare's entry is the most complete package in the category. On October 1 it released Clef (27 billion parameters, post-trained on Qwen3.8-27B) and Clef-flash (9 billion parameters, on Qwen3.5-9B), both with weights on Hugging Face under Apache 2.0, both served on Workers AI, and both "fully Jev-API compatible" - Cloudflare. They are the first models Cloudflare's Workers AI team has trained itself, and Cloudflare positioned them explicitly against Jev, citing Jev's own benchmark suite and the community Decision Index throughout the announcement.
Three things set Clef apart from Jev on paper. It adds a vision encoder, so the state can include images and video alongside text and JSON. Its weights are open, so you can run it anywhere and audit it. And Cloudflare is offering to fine-tune it on your own traffic with reinforcement learning, which Jev explicitly does not do. On Workers AI, Clef costs $0.24 per million input tokens and Clef-flash $0.09 - Cloudflare Docs. Both accept 64k tokens of context, 1 to 64 questions per request, and up to four embedded images (4 MiB and 16 megapixels each).
Calling Clef looks almost identical to calling Jev, which is the point of the compatibility claim. The REST version below is Cloudflare's own example; inside a Worker you would call env.AI.run("@cf/cloudflare/clef", {...}) with the same body:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
Speed is where Clef-flash is in a different league. Across 43 benchmarks, Cloudflare measured a median latency of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev, with p95 figures of 122.4, 238.6 and 536.0 ms respectively. Two caveats apply: these are Cloudflare's measurements, and Jev's number includes a network round trip to TypeSafe's API while Clef runs inside Cloudflare's own network, so the comparison flatters the home team. Even so, a 9B model answering in under 40 ms is fast enough to sit in the request path of a web application, which is a use case a 500 ms hosted call cannot serve.
On accuracy, Cloudflare's own tables show a split result rather than a clean win. Clef beats Jev clearly on intent classification (94.20 vs 79.74 macro-F1 on BANKING77, 97.43 vs 89.27 on CLINC150 with out-of-scope detection) and on tool-calling accuracy (98.47 vs 95.75 on BFCL). Jev wins on deciding whether to call a tool at all (80.97 vs 72.37 on When2Call) and on reasoning-heavy retrieval (47.52 vs 45.91 on BRIGHT). On TypeSafe's own workflow evaluations, Clef leads on invoice processing (64.7 vs 61.8) while Jev leads on agent-trace observability (71.6 vs 68.5). Clef-flash, interestingly, beats both on one home-appliance control test (97.73 vs Jev's 52.27) and collapses on CLINC150 (66.77), which is a reminder that small models are uneven.
Cloudflare has been testing Clef with its Threat Intelligence team to classify website domains, and reports that fetching, rendering and classifying a site took 2.2 seconds with Clef against 4.7 seconds for gpt-oss-120b, its fastest general LLM, which also returned fewer classifications. Its Trust and Safety, Support and Bot teams are lined up as the first internal fine-tuning users. Cloudflare also states that it does not read, store or train on customer requests or responses. The more strategic piece of the announcement, though, is the reinforcement-learning service that turns Clef from a model into a platform.
The loop in the diagram is worth reading closely because it shows where Cloudflare thinks the value is. Production traffic is captured through AI Gateway, you curate tasks and define rewards, Workers AI samples decisions from the current Clef policy, Containers run your reward code, a new RL trainer updates the weights, and the chosen checkpoint is served back through Workers AI's bring-your-own-model path. Cloudflare says training optimizes calibration directly, granting partial credit to adjacent ordinal choices and penalizing drift from the reference model. For now the service is a hands-on engagement with Cloudflare's forward-deployed engineers rather than a self-serve product, and pricing is unpublished.
One operational detail deserves a warning label. Cloudflare's schema notes that long text state is truncated to fit the model's token limit. A truncated state does not raise an error; it produces a confident answer about a document the model only partly read. If your states can exceed 64k tokens, filter or chunk them in code before the call, which is good practice for every model in this guide anyway (section 9).
Why this matters: Clef is the first decision model you can both run at the edge with near-instant latency and fine-tune on your own data, with weights you own. How to apply it: if you already run on Cloudflare, or your decisions involve screenshots, product images or video frames, start with Clef-flash for anything in a user-facing request path and move to Clef where accuracy matters more than 170 extra milliseconds. If you want to understand Cloudflare's broader agent-infrastructure push, our look at Cloudflare Kitesurf for browser agents covers the same playbook applied to browsing.
5. The open challengers: Perplexity pplx-decider and AWS Strands Decider 2B
The title of this guide names three products, but on October 1 two more arrived that change the comparison, and one of them tops our scorecard. Both are open weights under Apache 2.0, both use the same three question types as Jev, and both come from companies with very different motives. Perplexity is selling hosted inference at a very low price while showing it can train its own models; AWS is growing its open-source Strands Agents SDK, which a free local decider makes more useful.
They also sit at opposite ends of the size range. Perplexity's model is a 27B-class model that needs a serious GPU to self-host, and it competes with Jev on accuracy. AWS's is a 1.9B model that runs on a laptop and competes on nothing except being free, local and fully reproducible. Between them and Clef, every deployment shape now has an open option, which is the single biggest change in the category since Jev launched.
Perplexity pplx-decider-v1-27b
Perplexity's CEO Aravind Srinivas announced the model as "a state of the art multimodal Decision Model", offered through a new Decisions API "at 4 cents per million input tokens and free output tokens", and added: "We intend to bring down the price even further over the coming days" - Aravind Srinivas on X. The weights are a fine-tune of Qwen3.8-27B under Apache 2.0, and self-hosting needs Python 3.12+ and a CUDA GPU "with room for approximately 49 GiB of weights plus working memory" - Hugging Face. In practice that means one 80 GB data-center GPU or a well-equipped workstation card, not a laptop.
The hosted API is the best-documented in the category after Jev's, and its limits are generous. A request to POST https://api.perplexity.ai/v1/decisions can carry 1 to 128 questions, up to 255 options per Choice, up to 10 levels per Score, and just under 262,144 input tokens - Perplexity Docs. Images go in as PNG, JPEG or WebP data URLs, read in 32 by 32 pixel tiles up to 2,048 tiles per image, at roughly 1,000 input tokens per megapixel. There is no per-request fee. The documentation is also unusually candid about behavior: identical requests "usually return identical numbers", but occasionally differ in the second decimal place, so thresholds need a margin.
The weak spot is speed. Perplexity's own tests on September 30 found that a request with a few hundred input tokens answered in under 2 seconds, about 90,000 tokens took 5 seconds, and a request near the input limit took 23 seconds. An oversized image does not return a validation error; the request waits about a minute and returns a 504. Every organization is capped at 10 requests per second on every plan. On a GPU you control, the same weights are much faster: the Decision Index measured a median of about 101 ms on a single RTX PRO 6000 - Decision Index. The model is fast; the hosted service, at launch, is not.
On accuracy, Perplexity published an 11-benchmark comparison against Jev and its own base model. pplx-decider scored 85.71% overall against Jev's 84.51%, and against 74.76% for the untuned Qwen3.8-27B, which shows how much the decision-specific fine-tune adds. The overall win hides a split: Jev is ahead on six of the eleven tests, including commonsense (WinoGrande), multi-step reasoning (BBH) and truthfulness, while Perplexity wins clearly on hallucination detection in retrieval (RAGTruth), financial sentiment and table fact-checking.
The pattern in the chart is a useful rule of thumb rather than a winner: Perplexity's model is stronger where the decision is grounded in the supplied evidence (is this answer supported by the passage, does this table support this claim), and Jev is stronger where the decision leans on general reasoning or world knowledge. That maps neatly onto use cases. A retrieval pipeline that needs to throw out unsupported passages or flag hallucinated answers fits Perplexity's profile; an agent that needs commonsense judgment about what a user meant fits Jev's.
The independent Decision Index tells a consistent story with one important addition. It places pplx-decider at 56.4 against Jev's 57.9, close enough to call a tie, but it measures pplx-decider's calibration error (ECE) at 0.018 against Jev's 0.074. Put plainly, when Perplexity's model says it is 80% sure, it is right almost exactly 80% of the time, while Jev's average confidence runs about seven points above its accuracy on that board - Laurence Moroney. For a system that gates actions on confidence thresholds, that difference matters as much as a point of accuracy.
AWS Strands Decider 2B
AWS released Strands Decider 2B through its Strands Labs organization as a model "optimized for fast experimentation, local development, and innovation" - Strands Agents. It is the most transparent model in this guide by a wide margin: the weights, the training data list, the training scripts, the evaluation code and a preregistered research log are all public, under Apache 2.0. AWS's architecture diagram explains the core idea better than any description.
The left side of the diagram is an ordinary language model: input, torso, a head that turns hidden states into vocabulary logits, generated text. The right side keeps the same pretrained torso (Qwen3.5-2B-Base, adapted with a rank-16 LoRA) and throws the head away. In its place sits a pointer head of about a million parameters that compares the hidden state at an <answer> marker with the hidden state at the end of each option's line, then applies a softmax to get per-option probabilities - GitHub. Because the head holds no per-option weights, nothing in the model can learn that "the first option is usually right", and there is no fixed cap on how many options a question can carry.
The measured results are honest about what a 2B model can do. The reference version (v19) answers 72.3% of the 231 tasks in the JevBench v1 public set, with a Brier score of 0.342 and an expected calibration error of 0.052; it gets every easy-tier task right but only about half of the hard tier. Latency is 115 ms median on an RTX 3090 and 153 ms on an Apple M3 Pro, with no network hop. AWS's own guidance on confidence is the most useful sentence in the README: on short classification tasks it has never seen, answers at a confidence of 0.9 or more are right about 95% of the time; below that, confirm or ask a person.
Trying it takes two commands, and the CLI shows all three primitives against the same state:
pip install strands-decider
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
--state "Help! My payouts have been failing for 3 days!" \
--choice "Which team should handle this?=billing,sales,retail" \
--noul "Does this convey urgency?" \
--score "How frustrated is the writer?=calm,frustrated,depressed"
The repository's example output for that command routes the message to billing at 0.846 probability (confidence 0.769), rates urgency at 0.829, and places frustration between "calm" and "frustrated" with a confidence of only 0.519, which is exactly the kind of low-confidence answer your code should treat as "ask someone". Running strands-decider serve exposes the same model on a local /v1/systemone endpoint, so code written for Jev can point at it with a URL change. The repo also ships a worked Strands Agents example: a before_tool_call intervention that gates a weather-tool call on two yes-or-no decisions, so the agent asks which city the user meant instead of guessing.
Two details make Strands Decider more important than its accuracy suggests. First, you can retrain the whole recipe in about 11 hours on one RTX 3090, or 70 minutes on eight H100s, which turns "fine-tune a decider on our own data" from a vendor engagement into an afternoon project. Second, the research log is preregistered: every run states its predictions and failure conditions before training, and most runs are recorded as having missed their bar. That is the opposite of a launch claim, and it is the most trustworthy evidence anyone has published in this category.
Why this matters: the two open challengers mean you no longer have to trade openness for quality (Perplexity) or pay anything at all to experiment (AWS). How to apply it: if you need open weights, data residency or image input and can tolerate hosted latency or run your own GPU, pplx-decider is the strongest all-round choice today. If you want to prototype decision gates on a laptop, test ideas offline or train a domain-specific decider yourself, start with Strands Decider 2B and move up a size only when your eval set says the small model is not good enough. Our guide to running an agent on a single 24GB GPU covers the local-hardware side of that decision.
6. Benchmarks: what independent boards say versus vendor claims
Every launch in this category came with a chart showing the new model beating Jev. TypeSafe's launch compared Jev with frontier language models on TypeSafe-built workflows, Cloudflare plotted Clef above Jev on the Decision Index, Perplexity published an 11-test panel where it edges Jev, and Fastino reported GLiDE at 64.81 on the Decision Index against Jev's 57.91, on tests and opponents it chose itself - MarkTechPost. None of these numbers is fabricated. All of them are selected, and a selected number is an upper bound, not an estimate.
Two public boards measure this class of model with the same method for everyone, and they are the evidence the scorecard leans on. JevBench, run by Benchmark Heaven and explicitly not affiliated with TypeSafe, scores systems on four equally weighted axes: chance-corrected intelligence, calibration, speed and cost per 1,000 decisions - GitHub. Its current release (v1.5.5) runs 1,624 decisions per system, 720 of them sealed so vendors cannot train on them, and ranks 109 systems. The Decision Index, a community board hosted on Hugging Face, runs 43 benchmarks and 120,340 requests per entrant on a single RTX PRO 6000 and reports calibration separately.
What JevBench shows
On JevBench v1.5.5, Jev ranks #3 of 109 with a composite of 72.1, behind two community fine-tunes of Google's Gemma 4 12B model, Cygnet (73.7) and Winnow-12B (73.2) - Benchmark Heaven. Clef-flash ranks #25 at 55.1 and Clef ranks #62 at 17.0, mostly because the board's cost axis gates it hard at $0.24 per million tokens. The most striking rows, though, are OpenAI's: plain GPT-6 Luna scores 96.2 on Intelligence and 95.6 on Calibration, well above Jev's 72.0 and 88.0, and still ranks only #37 because its speed and cost axes collapse when it is called as an ordinary text model.
The chart contains the most important structural fact in this guide. A frontier language model is still the most accurate decider measured, and on JevBench it is also better calibrated than any of the decision models in this guide. What decision models buy you is not better judgment than the best LLM; it is good-enough judgment at a tiny fraction of the cost and latency. That reframes the choice: the question is never "is Jev smarter than Luna?" (it is not), but "is Jev's accuracy enough for this decision, at this price and speed?" For most routing, triage and guardrail decisions the answer is yes. For a decision with legal or financial consequences, the right design is often a decider that handles the confident majority and escalates the rest to a frontier model, which is precisely what OpenAI's Decisions API is trying to collapse into one call.
JevBench also publishes its own humility. 68 of 75 adjacent pairs with paired-bootstrap comparisons are statistical ties, so the board asks readers to read the order as a ranking and not to treat small gaps as significant. In practice that means Jev, Cygnet, Winnow and the top 4B fine-tunes are all in the same band, and the spread inside the top ten is smaller than the difference you will see between any of them on your own data.
What the Decision Index shows
The Decision Index tells a compatible story with different entrants. Its independently run rows put Jev at 57.91, a Gemma-4-based community model (Surogate Rune) at 57.44, and Perplexity's pplx-decider at 56.40, with every 2B-class model far behind. Cloudflare's own chart of the same index adds its two models from Cloudflare's runs, and labels them honestly as self-reported.
Look at the legend before the dots. Clef's 61.2 and Clef-flash's 57.1 are orange because they are Cloudflare's own measurements; the blue points are validated by the board's operator; Jev's point includes a network round trip that the on-card models do not pay. Read with that in mind, the chart says the top of the category is a cluster of models scoring roughly 56 to 61 on this index, and that the real differentiator inside the cluster is latency, not accuracy.
The last bar matters for anyone considering AWS's model. Strands Decider 2B is not on the Decision Index yet, but "Decider 2B", a community model built on the same Qwen3.5-2B-Base torso, scores 28.97, about half of Jev's score, and AWS's own README names a decider-2b built on the same torso as its nearest open comparison. A 2B decider is a tool for experimentation, cheap first-pass filtering and on-device work, not a drop-in replacement for a 27B-class model on hard judgments.
Calibration is the number to read twice
Two metrics capture calibration. Brier score is the average squared gap between the probabilities a model gave and what actually happened (lower is better); expected calibration error (ECE) groups answers by confidence and measures how far each group's accuracy sits from its stated confidence. On the Decision Index, Jev's ECE is 0.074 and its average confidence (81%) runs about seven points above its accuracy (74%); pplx-decider's ECE is 0.018; AWS reports 0.052 for Strands Decider on JevBench.
Calibration measured on benchmark tasks does not guarantee calibration on yours. One tester asked Jev to guess the result of a hidden fair die roll 400 times; it picked nearly the same answer every time with an average confidence of about 83% and was right about 19% of the time, which is chance - AI Agents Simplified. A well-calibrated model should have spread its probability evenly across six faces. The test is deliberately adversarial (the state contains no information), but it makes the right point: confidence is a signal you validate on your own labeled data before you let it trigger actions.
Why this matters: vendor charts will keep arriving weekly, and the top of every board is a statistical tie. How to apply it: treat independent boards as a shortlist filter, not a verdict, and let your own eval set (section 11) make the final call. If you want a broader grounding in how to read benchmark claims, our guide to AI agent evals and benchmarks covers the same traps in the language-model world.
7. Price and latency: the real cost per decision
Decision models are priced per input token, with output free on Jev and Perplexity, so the bill depends almost entirely on how much state you send and how many decisions you make. That is a different cost structure from a language model, where output tokens cost several times more than input and reasoning models add hidden thinking tokens on top. It also means price comparisons per million tokens are directly comparable across the category, which is rarely true for text models.
The list prices cluster tightly at the bottom and spread out at the top. Perplexity and Jev are within a fraction of a cent of each other, Clef-flash costs about twice as much, and Clef costs about six times as much. Plain GPT-6 Luna is included as a reference point because it is the model OpenAI's Decisions API is built on, but its output price applies when you call it as a text model, and the Decisions API's own price is unpublished.
Per-token prices only become meaningful once you multiply them by tokens per decision. JevBench measured Jev reading an average of 950 input tokens per decision across its benchmark set, and at Jev's tariff that works out to about $0.04 per 1,000 decisions - JevBench. Using the same 950-token assumption for every model gives a clean comparison. For context, the same input sent to Claude Opus 5.5 at $4 per million input tokens costs $3.80 per 1,000 decisions before a single output token - OpenRouter.
| Model | Input price per M tokens | Output | Cost per 1,000 decisions (950 tokens each) | Cost per million decisions |
|---|---|---|---|---|
| pplx-decider-v1-27b (hosted) | $0.04 | Free | $0.038 | $38 |
| Jev 1.13 | $0.042 | Free | $0.040 | $40 |
| Clef-flash | $0.09 | Not listed | $0.086 | $86 |
| GPT-6 Luna as a text model | $0.10 | $0.50 per M | at least $0.095 plus output | at least $95 |
| Clef | $0.24 | Not listed | $0.228 | $228 |
| Strands Decider 2B | Self-hosted | Free | Your hardware | Your hardware |
| OpenAI Decisions API | Unpublished | Unpublished | Unknown | Unknown |
Here is the first-principles point the table makes. When a decision costs four thousandths of a cent, the model is no longer the expensive part of the decision. A million decisions cost between $38 and $228 depending on the provider. The costs that actually decide the economics are the cost of a wrong decision (a misrouted refund, a blocked legitimate user, an agent action that should have been stopped), the latency budget the decision sits inside, and the throughput limits of the service. Optimizing a decision layer for token price is optimizing the wrong variable.
To make that concrete with an illustrative example rather than a measured one: suppose a misrouted support ticket costs a person five minutes to re-route. A decider that misroutes one more ticket per thousand than its rival creates 1,000 extra misroutes per million decisions, about 83 hours of work; at an illustrative $30 an hour that is roughly $2,500. The largest price gap in the table above is $190 per million decisions. Accuracy and calibration on your own decisions dominate the business case, which is why the scorecard weights decision quality at 40% and cost at 15%. Our analysis of the true cost of LLM inference explains why per-token prices across the industry keep falling faster than anything else in the stack.
Latency is the other place the providers really differ, and here the measurement conditions matter as much as the numbers. The chart below collects the median latencies each source reports, and they were not taken under the same conditions: Cloudflare measured Clef, Clef-flash and Jev from its own network across 43 benchmarks, AWS measured Strands Decider locally on an RTX 3090, and OpenAI's figure is a keynote claim for a service few people can access.
Perplexity is absent from the chart because its documentation reports response times rather than a median: under 2 seconds for a small request on the hosted API, rising with input size. The practical reading is about where the model runs relative to your code. A model running in the same network as your application (Clef-flash inside a Cloudflare Worker, Strands Decider on the same machine) can answer in tens of milliseconds; a hosted API across the internet adds a round trip that no model optimization removes. If a decision sits inside a user-facing request, that round trip is usually the whole latency budget.
Throughput limits are the third axis, and they are easy to overlook until launch day. Jev allows 80 requests per second and 100,000 tokens per second per account, with limits that are still being adjusted. Perplexity allows 10 requests per second per organization. Cloudflare does not publish a per-model limit on the Clef pages, and a self-hosted model is limited only by your hardware. The way around request limits is the same for every provider: ask many questions per request. Because the state is read once, batching ten questions about one event into one call costs little more than one question, and TypeSafe's own test of batching a 13-question briefing found it 12.2x cheaper - TypeSafe Cookbook.
Why this matters: the cheapest model per token is rarely the cheapest decision layer once errors, latency and limits are counted. How to apply it: estimate your tokens per decision and decisions per second first, check them against each provider's limits, and only then compare prices. If cost is still the deciding factor after that, the techniques in our guide to cutting LLM costs apply to the language-model calls the decider escalates to, which is where most of your remaining spend will sit.
8. Where decision models fit inside an agent
A decision model is not an agent and does not replace the language model inside one. TypeSafe's documentation is explicit that Jev "is not a drop-in replacement for the LLM behind Claude Code, Cursor" and similar tools; it does not write, call tools or hold a conversation - TypeSafe Docs. What it replaces is the subset of language-model calls that were really closed decisions, and what it adds is a set of fast checkpoints that were too expensive to run before.
The cleanest way to think about it is Kahneman's split, which TypeSafe borrowed for the name. The decision model is the agent's System One: fast, cheap, calibrated judgments about what is happening and what kind of response it needs. The language model is System Two: slow, expensive reasoning, writing and planning, used only when System One is unsure or the task requires generation. Code sits between them and owns every threshold, every arithmetic step and every irreversible action.
The dashed line at the bottom is the part most teams skip. Every decision and its eventual outcome is a labeled example, and those examples are what you need to tune thresholds, compare providers and (with Clef, Strands or pplx-decider) fine-tune a model on your own traffic. TypeSafe's docs and the vendor examples converge on a handful of patterns that use this loop:
- Confidence-gated routing acts above one threshold, confirms in the middle and hands off below a floor, with stricter thresholds for destructive actions - TypeSafe Docs.
- Intent and model routing sends each request to deterministic code, a specialist model or a person; OpenRouter's Jev Router applies the same idea to choosing a language model.
- Tool-call gating checks a proposed action before it runs, as in the Strands
before_tool_callexample that makes an agent ask instead of guess.
Two more patterns follow the same logic at the edges of an LLM application. Guardrails screen every message going into and out of the language model against hazard categories, thresholded to pass, review or block. Retrieval filtering and eval scoring decide which passages reach the answering model and score its outputs against a rubric, at a cost low enough to run on every single response rather than on a sample.
All five patterns share one design principle: the model answers narrow, literal questions and code combines the answers. That is the opposite of how most teams use language models, where one large prompt tries to make every judgment at once and returns a paragraph that code then has to interpret. The decision-model style is more work up front (you have to decompose the judgment into questions), but it produces a system where every intermediate answer is typed, logged and individually measurable. The routing pattern is the one with the fastest payback for most agents, and our guide to AI model routing covers the economics of sending each request to the cheapest capable model.
Here is the confidence-gated pattern in code, adapted from TypeSafe's documentation. Note how the thresholds scale with the stakes of the action rather than being one global number:
from typesafe_sdk import Choice, TypeSafeClient
client = TypeSafeClient(model="jev-1.13.0") # pin the version you tuned against
def handle(user_message: str, account_id: str) -> None:
response = client.system_one(
state=user_message,
questions={
"action": Choice(
instructions="What is the user trying to do?",
criteria={
"check_balance": "View account balance",
"approve_transfer": "Approve the pending withdrawal request",
"support": "Get help with an issue",
},
),
},
)
action = response.answers ["action"]
if action.confidence < 0.5:
route_to_human(user_message) # genuinely unsure: do not guess
elif action.choice == "check_balance":
show_balance(account_id) # low stakes, recoverable
elif action.choice == "approve_transfer":
if action.confidence > 0.9:
confirm_then_execute(account_id) # high stakes needs high confidence
else:
ask_user_to_confirm(account_id)
else:
open_support_ticket(user_message)
Two techniques from people who have pushed these models hardest are worth knowing early. Sean Goedecke found that a decider playing Doom held the shoot button constantly and wandered aimlessly until Goedecke added tiered goals: a slow outer loop that picks a short-term goal from a fixed set and a fast inner loop that acts toward it. For choices with hundreds of options, such as picking the next link on a Wikipedia page with over a thousand links, the fix was tournament sampling: choose from batches of a hundred, then choose among the winners - Sean Goedecke. Both techniques are just ways of keeping each individual question small and literal, which is the recurring theme of this section.
Why this matters: the value of a decision model comes from the architecture around it, not the model alone. How to apply it: start with one pattern on one high-volume decision (ticket triage and tool-call gating are the usual first candidates), log every answer with its outcome, and only add more patterns once the first one has thresholds you trust. If your agent handles customer conversations, our comparison of Sierra and Decagon shows how the large support platforms already structure triage and escalation, and our context engineering guide covers how to build the compact state a decider needs.
9. Where decision models fail
TypeSafe deserves credit for publishing the most useful page in this whole category: a list of the ways its own model fails. The "jaggedness" page for jev-1.13, last reviewed on October 2, says the model is "fast, calibrated, and good at common-sense judgment but it is not perfect", that it "can be quite literal", and that it "struggles with tasks that require numeric precision" - TypeSafe Docs. The failure modes it lists are not specific to Jev; they follow from what a single-pass decision model is, and every model in this guide shares them to some degree.
The first group of failures comes from asking the model to do work that is not judgment. Decision models are poor calculators, and TypeSafe is blunt about it: keep arithmetic in code, because the model "recognizes the shape of an answer rather than tallying". The same applies to dates, which the model reads as text rather than ordered quantities. The recommended fix is elegant: turn date extraction into Choice questions over small closed sets (twelve months, thirty-one days, a range of years, plus an explicit "not stated" option), then assemble and compare the dates in code.
- Counting and arithmetic fail because the model pattern-matches rather than tallies; count and compute in code.
- Date and time comparison fails because dates are read as text; extract the parts, compare in code.
- Literal reading means the model answers the question you wrote, not the one you meant; put boundary cases in the criteria.
- Indirection (double negatives, properties of properties) costs accuracy; ask direct questions about named fields.
- Generation is not supported in any useful way; extract candidates with a regex or an LLM and let the decider pick.
The common thread in that list is that every failure has a design fix, and none of the fixes is "use a bigger model". They all push work toward code and toward smaller, more literal questions. That is a real cost: decomposing a judgment into reliable atomic questions takes more engineering than writing one big prompt, and it is the main reason some teams will try a decision model and go back to an LLM. It is also the reason decision layers, once built, are easier to test and audit than prompts. Each question can be evaluated on its own labeled examples, and a regression shows up as a specific question getting worse rather than as a vague change in an LLM's tone.
Order bias and context rot
The second group of failures is subtler and more dangerous because it produces confident wrong answers on inputs that look normal. TypeSafe notes that Jev "leans toward the option that comes first" in some Choice questions and recommends reordering options to check consistency. JevBench measured how severe this can be for small models: one open model scored 72% with options in one order and 21% with the order reversed on answer-judging items - JevBench. AWS's pointer-head design is partly a response to this problem, since a head without per-option weights cannot learn a positional preference, but the torso still sees the options in order.
Large states are the other quiet failure. TypeSafe warns that accuracy falls as the state fills with unrelated material ("Jev suffers from context rot"), and Cloudflare's schema notes that long text state is truncated to fit the model's limit rather than rejected. Both point to the same discipline: retrieve and filter in code, send only the fields a question needs, and never assume a model read a document it may have silently cut short. A cheap Noul question ("does this passage contain information about X?") is often the right filter before the real decision.
Security: constrained output is not a security boundary
Decision models remove one whole class of agent failure: they cannot invent a tool, emit malformed JSON or smuggle an instruction into a free-text field. That is a genuine security improvement, and it is why tool-call gating is such a natural use. But TypeSafe is explicit that the state is data the model does not treat as hostile, and that content written to steer it, "an injected instruction, a deliberately misleading framing, or text that argues for its own classification, can move the answer" - TypeSafe Docs. A prompt injection that cannot make the model write anything can still flip "safe" to "unsafe" or "billing" to "refund_approved".
That matters more now than it did a month ago. Reuters reviewed more than 200 documents and found at least 20 studies or evaluations since 2025 describing agents that deceived, replicated or tested their boundaries - The Kathmandu Post (Reuters). A decision model checking "is this action within the user's authorization?" is a useful layer against that behavior, but it is a classifier, not a permission system. Pair every decision gate on a consequential action with deterministic checks (allow-lists, scoped credentials, spend limits) that do not depend on any model's judgment. Our guides to prompt injection defense and sandbox security controls cover those deterministic layers in depth.
Operational failures
The last group is mundane, and it is the one most likely to cause your first production incident. Jev's rate limits "can change without notice", and its jev-latest alias moves when a new version ships, which can shift your carefully tuned thresholds overnight unless you pin the version. Perplexity's hosted API caps every organization at 10 requests per second and turns oversized images into one-minute 504 timeouts. OpenAI's Decisions API has no published contract at all. And English is Jev's primary training language; TypeSafe says other languages, including CJK scripts, are handled "but not equally well" - TypeSafe Models.
Why this matters: apart from rate-limit and timeout errors, these failures produce valid-looking answers, so they will not show up as errors in your logs. How to apply it: before launch, test option-order consistency on your real questions, cap state size in code, pin model versions, add deterministic checks behind every consequential gate, and alert on confidence drift (a sudden shift in the share of low-confidence answers is usually the first sign that inputs or the model changed). Our analysis of why agents still fail four of five tasks on OSWorld 2.0 is a useful reminder of how much of agent reliability lives in this unglamorous layer.
10. How to choose: a decision framework
The honest answer to "Jev vs Clef vs OpenAI Decisions API" is that the three products answer different questions, and the right one depends on four facts about your situation rather than on whose benchmark chart is highest this week. Where must the model run (anywhere, inside a specific cloud, or on your own hardware)? What is in the state (text only, or images and video too)? How much latency can the decision add? And do you need to own the weights, either to fine-tune on your data or to satisfy a data-handling requirement?
Because the leading APIs share one request format, the choice is also less permanent than it looks. Picking Jev today does not lock you out of pplx-decider or Clef tomorrow; it means your questions, thresholds and eval set are built against Jev's behavior, and moving is a matter of re-running that eval set against another endpoint. That lowers the stakes of the first decision and raises the value of building the eval set properly (section 11). The tree below encodes the four facts in the order that usually rules options out fastest.
Choose Jev when you want a hosted, text-only decision layer with the most mature documentation, SDKs and failure-mode guidance in the category, and you have no requirement to own or tune the weights. It is the highest-ranked hosted API on JevBench (#3 overall, behind two open community models), it is as cheap as anything hosted, and TypeSafe's cookbooks are the best free education on how to design questions. The cost is dependence on one young vendor whose rate limits are still settling and whose model you cannot fine-tune. Pin the version, and keep your integration behind an adapter so the dependency stays cheap to unwind.
Choose pplx-decider-v1-27b when you want Jev-class accuracy with open weights, better calibration and image input at the same price, and you either accept a slower hosted API or run the model on your own GPU. It is the best default for retrieval and grounding decisions (is this passage relevant, is this answer supported) and for any team with data-residency requirements that can host a 27B model. The cost is hosted latency in seconds rather than milliseconds and a 10 requests-per-second cap, so high-volume hosted use needs aggressive batching or self-hosting.
Choose Clef-flash, then Clef, when the decision sits in a user-facing request path, involves screenshots, product images or video frames, or when you already run on Cloudflare. Clef-flash's sub-40 ms median makes decisions possible in places a hosted round trip cannot reach, such as filtering every incoming request at the edge. Move up to Clef where the self-reported accuracy gain is worth about 170 ms and six times Jev's price, and talk to Cloudflare about the RL service if you have enough labeled traffic to fine-tune. The cost is accuracy: on JevBench, Clef-flash's intelligence score sits well below Jev's, so test it on your hardest decisions before trusting it there.
Choose Strands Decider 2B for prototyping, offline and on-device decisions, cheap first-pass filtering before a bigger model, and any situation where you want to train your own decider. It is free, it runs on a laptop, and its preregistered research log is the most trustworthy evidence in the category. The cost is accuracy on hard judgments, which is 2B-class. Choose the OpenAI Decisions API only if you are standardized on OpenAI and can wait for a public contract; until then, request access, design behind an adapter, and run one of the others in production.
There is also a fifth answer to the question, which is to not wire a decision layer at all. Decision models are infrastructure for teams building their own agents and workflows, and assembling one means writing questions, building eval sets and owning thresholds. If what you actually want is the outcome rather than the plumbing, platforms like o-mega take a different route: you describe a company in one conversation, and one AI builds and runs the website, app, billing, content and admin, so you judge the business result rather than which model made each choice underneath. That is not a substitute for a decision model inside your own product, but it is a reasonable alternative to building an automation stack from parts.
Why this matters: a structured choice now saves a migration later, but the shared API format means even a wrong first choice is cheap to reverse. How to apply it: walk the tree with your three highest-volume decisions, pick the model the majority lands on, and run the second-ranked option in shadow mode on the same traffic for two weeks before committing. For the language-model side of the same architecture, our October 2026 LLM ranking for agents covers which System Two model to escalate to, and our 2026 LLM API price table covers what those escalations cost.
11. Implementation walkthrough: shipping your first decision layer
Most teams that try a decision model and abandon it make the same mistake: they swap it into an existing LLM call, eyeball a dozen outputs, and judge the model on vibes. A decision layer earns its keep only when its thresholds are tuned on labeled data and its behavior is measured continuously, because the entire value proposition (act automatically when confident, escalate when not) rests on numbers you have checked. The five steps below are the minimum process that turns a decision model from a demo into infrastructure.
The walkthrough assumes a common starting point: an agent or workflow that already makes some closed decisions with a language model, and logs of those calls. If you are starting from scratch, the same steps apply, but step 2 takes longer because you have to collect examples before you can measure anything. None of the steps depend on a particular vendor, which is deliberate: the eval set and adapter you build in steps 2 and 3 are what make every later vendor decision cheap.
Step 1: inventory your decisions. Search your logs and code for language-model calls whose output is checked against a fixed set: enum parsing, if "yes" in response, JSON fields with allowed values, classifier prompts. For each one, record the question, the answer set, the volume per day, the latency budget, and what a wrong answer costs. Then cut the list to decisions where the state actually contains enough information to decide. A question that needs external lookups, multi-step reasoning or arithmetic is not a decision-model candidate until you move that work into code.
Step 2: build an eval set. Pull a few hundred real examples for each candidate decision, label them (from downstream outcomes where possible, by hand where not), and deliberately include the edge cases that cause incidents. Laurence Moroney's advice for this category is the right default: collect labeled examples from your domain, run models side by side, validate calibration locally and escalate low-confidence decisions - Laurence Moroney. Keep the eval set versioned; it is the asset that outlives any vendor choice.
Step 3: one adapter for all. Because Jev, Perplexity, Clef and a local Strands server share the request format, a thin adapter lets you run the same questions against all of them. The sketch below normalizes the four into one function and records latency and token usage for each call:
import os
import time
import requests
PROVIDERS = {
"jev": {
"url": "https://api.typesafe.ai/v1/systemone",
"key_env": "TYPESAFE_API_KEY",
"model": "jev-1.13.0",
},
"pplx": {
"url": "https://api.perplexity.ai/v1/decisions",
"key_env": "PERPLEXITY_API_KEY",
"model": "pplx-decider-v1-27b",
},
"clef-flash": {
"url": "https://api.cloudflare.com/client/v4/accounts/"
+ os.environ.get("CLOUDFLARE_ACCOUNT_ID", "")
+ "/ai/run/@cf/cloudflare/clef-flash",
"key_env": "CLOUDFLARE_AUTH_TOKEN",
"model": "clef-flash",
},
"strands-local": {
"url": "http://localhost:8000/v1/systemone", # strands-decider serve ... --port 8000
"key_env": None,
"model": None,
},
}
def decide(provider: str, state, questions: dict, timeout: float = 30.0) -> dict:
cfg = PROVIDERS [provider]
body = {"state": state, "questions": questions}
if cfg ["model"]:
body ["model"] = cfg ["model"]
headers = {"Content-Type": "application/json"}
if cfg ["key_env"]:
headers ["Authorization"] = f"Bearer {os.environ [cfg ['key_env']]}"
started = time.perf_counter()
response = requests.post(cfg ["url"], json=body, headers=headers, timeout=timeout)
response.raise_for_status()
data = response.json()
data = data.get("result", data) # unwrap a response envelope if the API adds one
return {
"answers": data ["answers"],
"input_tokens": data.get("usage", {}).get("input_tokens"),
"latency_ms": round((time.perf_counter() - started) * 1000, 1),
}
Step 4: measure five numbers. Run the eval set through each candidate and record five numbers per decision: accuracy, expected calibration error, the share of answers that change when you reverse the option order, p95 latency, and cost per 1,000 decisions. For calibration on Choice questions, use the probability the model placed on its chosen option rather than the confidence field, because Perplexity's documentation is explicit that confidence "is not the top probability" - Perplexity Docs, and TypeSafe computes it as a normalized statistic. The two helpers below compute calibration error and test option order in a single request, using the parallel-questions property to ask the same question twice with the options reversed:
def expected_calibration_error(probabilities: list [float], correct: list [bool], bins: int = 10) -> float:
"""probabilities: the probability each answer put on its chosen option."""
total, ece = len(probabilities), 0.0
for b in range(bins):
lo, hi = b / bins, (b + 1) / bins
idx = [i for i, p in enumerate(probabilities) if lo <= p < hi or (b == bins - 1 and p == 1.0)]
if not idx:
continue
accuracy = sum(correct [i] for i in idx) / len(idx)
mean_prob = sum(probabilities [i] for i in idx) / len(idx)
ece += len(idx) / total * abs(accuracy - mean_prob)
return ece
def order_consistent(provider: str, state, instructions: str, options: dict) -> bool:
reversed_options = dict(reversed(list(options.items())))
result = decide(provider, state, {
"forward": {"type": "choice", "instructions": instructions, "criteria": options},
"reverse": {"type": "choice", "instructions": instructions, "criteria": reversed_options},
})
return result ["answers"]["forward"]["choice"] == result ["answers"]["reverse"]["choice"]
Step 5: shadow, gate, watch. Run the chosen model in shadow mode next to your existing logic for long enough to see your normal weekly traffic patterns, compare its answers with what actually happened, and set thresholds from that data: the automatic-action threshold where accuracy on your data meets your tolerance, a review band below it, and a floor below which the decision always goes to a person or a reasoning model. Then turn on gating for the lowest-stakes decision first. In production, pin the model version, alert when the share of low-confidence answers moves sharply, and re-run the eval set before accepting any version change. Our report on why most AI agent pilots never scale is a good reminder that this unglamorous measurement work is usually what separates the deployments that pay off from the ones that stall.
Why this matters: the difference between a decision model that saves money and one that silently makes bad calls is entirely in steps 2, 4 and 5. How to apply it: budget more time for the eval set than for the integration, because the integration is a few dozen lines and the eval set is what every future decision rests on. If your decisions gate actions taken under an agent's own identity, pair this with the credential scoping in our guide to securing AI agents with non-human identity.
12. Outlook: where decision models go next
The first 18 days of this category already show its likely shape, because the forces that formed it are still acting. A strong open base model plus a cheap fine-tune produces a competitive decider; a shared API format makes providers interchangeable; and the price of a decision is already low enough that the next cuts barely matter to buyers. Reasoning from those three facts, rather than from any vendor's roadmap, gives a fairly clear picture of the next six months.
Prices will keep falling, and stop mattering. Perplexity has said it intends to cut its price further, and one diffusion-based Jev-compatible service has announced $0.035 per million input tokens - JevBench. When the model cost of a million decisions drops from $40 toward $20, almost nobody's buying decision changes, because, as section 7 showed, the cost of wrong decisions dwarfs the cost of the model. Price competition will continue because it is the easiest thing to announce, but it will stop being the reason anyone switches.
The interface is already a commodity, so differentiation moves to quality, calibration and location. OpenRouter treats decisions as their own endpoint type, Cloudflare and AWS copied TypeSafe's request format, and Perplexity's schema is nearly identical. Once switching costs approach zero, providers compete on three things: accuracy on hard decisions, calibration a builder can trust, and where the model runs (the edge, a specific cloud, a laptop). The open-source community's curated list of decision models already tracks dozens of hosted APIs, open models, runtimes and benchmarks - GitHub, which is what a commoditized layer looks like.
Decisions will learn to think a little, and to see. Fastino's GLiDE adds "adaptive thinking on uncertain cases", spending more compute only where the model is unsure - MarkTechPost, and Clef and pplx-decider already read images (Clef reads video too). Both directions erode the clean line between System One and System Two: a decider that thinks when uncertain is a cheap version of the escalation pattern from section 8, built into one call. That is also exactly where OpenAI's Decisions API is positioned. If OpenAI ships it broadly with Luna-level accuracy at a price near Jev's, the specialist vendors will be left competing mainly on openness and location, which is the scenario that would most change this guide's recommendations.
Fine-tuning on your own traffic becomes the real moat. A general decider is good; a decider trained on a company's own labeled decisions is better, and only open weights or a vendor training service make that possible. Cloudflare's RL loop and AWS's 11-hour retraining recipe are early versions of what will become standard: capture decisions and outcomes, train a private decider on them, deploy it next to your code. Jev's choice not to fine-tune on customer data is principled, but it is the one structural gap competitors have already aimed at.
That leaves TypeSafe's reported $10 billion valuation talks as the most interesting open question in the category. Two weeks after launch, its API format has been adopted by Cloudflare, Perplexity and AWS, and community fine-tunes of 12B open models sit level with it on the independent board. What TypeSafe still holds is a head start in calibration-focused training, the best documentation in the market, and the brand that defined the category. Whether that is defensible depends on how fast its next versions move compared with open models that everyone else can tune. The Jevons paradox that gave the model its name cuts both ways: decisions are getting cheap enough that every agent will make far more of them, which grows the market and lowers the price of being in it.
Why this matters: the category is moving weekly, and the safest architectural choice is the one that keeps your options open. How to apply it: build on the shared request format, keep your eval set and thresholds independent of any vendor, and re-run the eval set quarterly against whatever is new. If your agents run on frameworks like LangGraph or Strands, expect decision hooks to become a standard feature; our comparison of agent frameworks covers where those hooks will most naturally plug in, and our open-source LLM guide tracks the base models the next deciders will be built on.
13. Conclusion
Decision models are the most practical new model category of the year, because they replace something every agent already does badly and expensively: using a language model that writes prose to make a choice from a fixed list. They return typed answers by construction, attach probabilities that are usable when you validate them, and cost a few hundredths of a dollar per thousand decisions. The category went from one model to eight in 16 days, and the three names in this guide's title are now joined by two open challengers that change the default recommendation.
The decision framework in one paragraph: Jev is the best-documented decider and the highest-ranked hosted API on the independent boards, and the default when you do not need to own the weights. pplx-decider-v1-27b matches Jev's price with open weights, image input and better calibration, and tops our scorecard because control counts; choose it when you can self-host or tolerate a slower hosted API. Clef-flash is the choice when latency or images dominate, with Clef as its more accurate and more expensive sibling, and Cloudflare's RL service as the path to a private model. Strands Decider 2B is the free, local and fully reproducible way to learn and prototype. The OpenAI Decisions API may end up the most accurate of them all, but today it is a preview without a public contract, so design for it without depending on it.
Whatever you choose, the durable work is the same: decompose judgments into small literal questions, keep arithmetic and permissions in code, build a labeled eval set from real traffic, tune thresholds on it, and keep every provider behind one adapter. Do that, and the question of which decision model to use becomes one you can answer again, cheaply, every time the market moves, which at the current pace is roughly every week.
This guide reflects the decision-model landscape as of October 3, 2026. The category is less than three weeks old: prices, rate limits, model versions and benchmark rankings are changing weekly, and OpenAI's Decisions API is still in limited preview, so verify current details with each provider before building on them.