Key takeaways
- A decision model returns a typed answer with a probability for each option, and writes no text.
- In AY Automate’s test, Jev answered 2 to 3.5 times faster than small LLMs at about the same accuracy.
- Independent results are mixed: Clef led a 77-case routing pilot, and a DeBERTa classifier beat Jev on prompt injection.
- Fine-tune a classifier when categories are stable and labels plentiful; use a decision model when they change often.
- Set confidence thresholds on your own labelled data, because some decision models were overconfident in tests.
An AI decision model reads text, and sometimes images, and answers questions you define in advance with typed values instead of generated text. Each answer is a yes/no probability, one option from a list you supply, or a score on a scale, and it comes with a probability for every option. TypeSafe AI made the category popular with Jev, launched in early access on 15 September 2026. The launch post reached 1,989 points on Hacker News, and by our count at least 20 more stories about decision models passed 100 points there over the next three weeks, Cloudflare’s Clef launch among them at 641.1
Use a decision model when software needs the same bounded judgement many times, within about a second, and the categories change too often to retrain a classifier. Use an LLM with structured output when the answer needs words, arithmetic, dates or several steps of reasoning. Fine-tune a classifier when the categories are stable, you have thousands of labelled examples and answers must arrive in tens of milliseconds.
This post explains how decision models differ from autoregressive LLMs, lists the options as of 7 October 2026 and summarises four independent tests, then gives a decision rule, patterns and a checklist. We haven’t benchmarked these models ourselves. Every figure links to its source, and vendor claims are marked as such.
How an AI decision model differs from an LLM
An autoregressive LLM writes its answer one token at a time. A JSON schema can stop it writing anything outside your labels, but the label is still decoded token by token, and a reasoning model thinks first. Output tokens also cost more: GPT-6 Luna and Claude Haiku 5.5 both charge $0.10 per million input tokens and $0.50 per million output tokens.23
A decision model doesn’t write at all. The published designs share a shape: the input and the candidate answers go through a transformer once, and a small head scores each candidate. Cloudflare says Clef runs its frozen Qwen backbone in a prefill-only pass and scores the valid choices in parallel.4 Strands Decider 2B swaps the language-model head of Qwen3.5-2B for a pointer head of just over a million parameters, which scores the hidden state at each option against the hidden state at the answer position.5
TypeSafe hasn’t published Jev’s architecture. Its launch post describes Jev as “unstructured state in, typed probabilistic decisions out”, and its docs say Jev reads the input once and evaluates every question against it in parallel.67 Since the work is mostly reading, the vendors charge for input. Jev costs $0.042 per million input tokens with free output, and OpenAI’s Decisions API bills input tokens only.8
The hosted APIs share three kinds of question. A yes/no question returns the probability that a statement is true; TypeSafe calls it a noul, short for Bernoulli, and OpenAI calls it a predicate.98 A choice question returns one of your options with a probability for each, and a score question returns a probability-weighted position on ordered levels you describe. The answer can’t fall outside your options, which TypeSafe presents as a model that never makes type errors.6 Whether the answer is right is what the tests below measure.
TypeSafe trains Jev with an unpublished method it calls Reinforcement Learning for Calibrated Decisions (RLCD). The goal is that, over many predictions, answers given 0.8 are right about 80% of the time.10 The closest published method, according to Sebastian Raschka, is RLCR, which adds a Brier score to the correctness reward; he notes there’s no known link between the two.1112 Cloudflare trained Clef with cross-entropy plus a Brier loss.4
A classic classifier differs in where the labels live. A fine-tuned ModernBERT learns a fixed label set from labelled examples, so a new category means new labels and another training run.13 A decision model reads the options and their descriptions in each request, so one model can route tickets in one call and grade search results in the next.
Decision models give up text. They can’t write a reply or explain an answer, and TypeSafe suggests extracting values by having Jev choose among candidates your code has found. Its list of known weak spots for jev-1.13, last reviewed on 2 October, adds arithmetic and counting, comparing dates, multi-hop instructions, long inputs padded with irrelevant detail, adversarial content and, in some cases, a lean towards the first option in a choice.14 The Strands team says the single parallel pass makes its model much worse than reasoning models at complex problems.5
Jev, OpenAI’s Decisions API, Cloudflare Clef and the open models
These are the decision models available as of 7 October 2026, with two small LLMs and a fine-tunable encoder for comparison. Prices are list prices.
| Option | Kind | Released | Weights | Price per million tokens |
|---|---|---|---|---|
| Jev 1.13 (TypeSafe AI)7 | Decision model | 15 September, early access | Closed | $0.042 input, output free |
| OpenAI Decisions API8 | Decision API on GPT-6 Luna | 6 October, public beta | Closed | $0.10 input, no output or cache charges |
| Clef (Cloudflare)15 | Decision model, 27B | 1 October | Apache 2.0 | $0.24 input on Workers AI |
| Clef-flash (Cloudflare) | Decision model, 9B | 1 October | Apache 2.0 | $0.09 input on Workers AI |
| Strands Decider 2B5 | Decision model, 2B | 1 October | Apache 2.0 | Run it yourself |
| Claude Haiku 5.53 | Small LLM | 7 October | Closed | $0.10 input, $0.50 output, prompts up to 100K tokens |
| GPT-6 Luna2 | Small LLM | 22 September | Closed | $0.10 input, $0.50 output |
| ModernBERT base and large13 | Encoder to fine-tune, 149M and 395M | December 2024 | Apache 2.0 | Run it yourself |
Jev. TypeSafe is admitting developers from a waitlist. The current model, jev-1.13.0, takes up to 64K tokens per request (32K for the input plus the longest question), reads text only and allows up to 255 options in a choice.76 The launch post claims end-to-end responses in 70 to 500 ms. It traces the 193.6-times speed and 444.6-times cost advantages on TypeSafe’s home page to four in-house workflow evals, which score each model against the average answer of GPT-6 Astra and Claude Fable 5.1 instead of labelled ground truth.6 Against GPT-5.6 Terra, AY Automate measured 3.6 times the speed and 40 to 49 times lower cost per decision.16
OpenAI’s Decisions API has been in public beta since 6 October.2 It runs only on gpt-6-luna, takes text and inline base64 images, and can return a refusal in place of an answer. OpenAI says it answers about 10 times faster than the Responses API, without saying how that was measured.8
Cloudflare’s Clef and Clef-flash came out on 1 October as open weights, built on Qwen3.8-27B and Qwen3.5-9B. Both take images, have a 65,536-token context and, Cloudflare says, are fully compatible with Jev’s API.174 Cloudflare’s performance figures are its own. It says Clef leads the Jev Decision Index, a community leaderboard on Hugging Face,18 and beats Jev in three of four TypeSafe workflow evals, by 0.3 to 2.9 points, while losing the fourth by 3.1. It reports median latencies of 209 ms for Clef and 39 ms for Clef-flash, against 524 ms for Jev.4
Strands Decider 2B was published on 1 October by the Strands Agents team, with Apache 2.0 weights, training data and scripts. The team reports median latencies of about 115 ms on an Nvidia RTX 3090 and about 153 ms for small tasks on an M3 MacBook, and third place of 33 models in its size class on JevBench’s public set.5
llama.cpp gained a /v1/systemone server endpoint for decision models in release v0.6.0 on 5 October. It serves Clef, with image input, and the open Jev-style models laya, julia-1, lev, openjev, kev and nimble, and build b11447 added pplx-decider on 6 October.19 Strands Decider wasn’t in llama.cpp’s release notes up to 7 October.
Jev vs LLM vs classifier in published tests
Four independent tests published by 7 October compare decision models with LLMs or classifiers on labelled data. Each group of rows is one test, so compare rows inside a group. Costs are the published cost per 1,000 decisions multiplied by 1,000.
| System | Kind | Accuracy | Median latency | Cost per million decisions | Task labels for training |
|---|---|---|---|---|---|
| Prompt injection, yes/no (Red Hat, 2 October)20 | |||||
| DeBERTa-v3-base injection classifier, 184M21 | Classifier | 89.0% | 54 ms¹ | Not reported | Trained by its publisher |
| Qwen3.6-35B as judge | LLM | 89.3% | 313 ms | Not reported | None |
| Jev 1.13 | Decision model | 86.4% | 348 ms¹ | Not reported | None |
| Content safety, yes/no (Red Hat) | |||||
| Jev 1.13 | Decision model | 86.2% | 360 ms | Not reported | None |
| Qwen3.6-35B as judge | LLM | 85.5% | 308 ms | Not reported | None |
| granite-guardian-hap-125m | Classifier | 80.3% | 33 ms¹ | Not reported | Trained by its publisher |
| 77-way intent routing, BANKING77, 231 items (AY Automate, 20 September)16 | |||||
| GPT-5.6 Terra, strict JSON, minimal reasoning | LLM | 84.0% | 1.17 s² | $1,957 | None |
| Jev 1.13 | Decision model | 78.8% | 0.33 s² | $40 | None |
| Claude Haiku 4.5, same settings | LLM | 76.2% | 1.02 s² | $1,255 | None |
| 77-way intent routing, BANKING77, 77 items (S Anand, 7 October)22 | |||||
| Clef | Decision model | 96.1% | 0.75 s | $444 | None |
| Clef-flash | Decision model | 94.8% | 0.49 s | $166 | None |
| GPT-6 Astra, low reasoning | LLM | 88.3% | 1.72 s | $6,285 | None |
| OpenAI Decisions API | Decision model | 79.2% | 0.31 s | $100 | None |
| GPT-6 Luna, prompted for JSON | LLM | 77.9% | 1.07 s | $56 | None |
| Jev | Decision model | 75.3% | 0.39 s | $72 | None |
| Sentiment, IMDb test set, 25,000 reviews (Sebastian Raschka, 29 September)11 | |||||
| ModernBERT-large, 395M, fine-tuned | Classifier | 96.5% | 55 min 41 s for the set³ | Not reported | 25,000 labelled reviews |
| Jev 1.13, choice question | Decision model | 96.5% | 22 min 24 s for the set³ | $26 | None |
- Red Hat ran the classifiers on a MacBook Pro M1 CPU. Jev’s figure includes the round trip from a machine in the UK to TypeSafe’s servers in the US, which Red Hat says adds at least about 56 ms. The article’s detailed table gives DeBERTa a median of 80.4 ms and a P95 of 276.4 ms, against 54.1 ms in its summary table.
- Median across all three tasks in the study, through OpenRouter from one laptop.
- Time to score all 25,000 reviews, which measures throughput. ModernBERT ran on a DGX Spark after 3 hours of fine-tuning, and Jev ran through TypeSafe’s API.
Red Hat’s two tests went opposite ways. The DeBERTa model, running on a laptop CPU, beat Jev by 2.7 points on prompt injection, and Jev beat the 125M-parameter granite classifier by 5.9 points on content safety.20 The authors concluded that pre-trained classifiers remain very competitive, and that apart from the compact Laya model, the decision models they tested weren’t faster, cheaper or more accurate than an LLM judge.
Within the category, models differ widely. On the same 77 routing cases, Clef and Jev were 20.8 points apart. The pilot is small, and Jev’s 95% interval runs from 64.6% to 83.6%, but Clef also got 6 cases right that GPT-6 Astra missed while Astra got none that Clef missed. GPT-6 Luna got one more case right through the Decisions API than through chat, and answered in 0.31 s instead of 1.07 s.22
Against small LLMs, Jev was about as accurate and much faster and cheaper. In AY Automate’s test it was 2.0 to 3.6 times faster than the four LLMs at the median, and 4.7 to 7.5 times cheaper per decision than the two cheapest. On 77-way routing only GPT-5.6 Terra was reliably more accurate, by 5.2 points.16
Per-token prices don’t settle the cost per decision. In S Anand’s pilot, GPT-6 Luna through chat cost less per decision than Jev or the Decisions API. The decision models were also sent label descriptions, and Jev read 1,706 input tokens per case on average against 485 for Luna through chat.22 Tokenizers differ too. AY Automate found that Jev counted 360 input tokens for a prompt GPT-5.6 Terra counted as 153.16
How the published tests were run
We didn’t run any of these tests. These are their setups, and what each leaves out.
- Red Hat (Rob Geada, Mac Misiura and Shelton Cyril; updated 5 October) used two class-balanced yes/no benchmarks from EvalHub’s NeMo Guardrails library and doesn’t give their sizes. Latency is the round trip to a NeMo Guardrails server. The classifiers ran on a laptop CPU, the LLM judges on a GPU cluster and Jev over TypeSafe’s API. The experiments were in English, and two of the article’s tables disagree on some latencies.
- AY Automate (Adel Dahani) ran 791 labelled decisions on 19 September. There were 160 for 8-way routing, 231 for 77-way routing and 400 from deepset’s prompt-injection set, all with labels and no descriptions. The LLMs ran at temperature 0 with minimal reasoning. Every call went through OpenRouter from one laptop, and the 95% intervals are 6 to 13 points wide.
- S Anand’s pilot, in the version committed on 7 October, used one case from each of BANKING77’s 77 categories, scored by exact match. Decision models got label descriptions as well as names, chat models names only. Latency is end-to-end from one machine.
- Sebastian Raschka scored the 25,000-review IMDb test set, with ModernBERT fine-tuned on the 25,000-review training set. He couldn’t check whether IMDb was in Jev’s training data, and AY Automate couldn’t check its public datasets either.
Vendor-reported figures appear only in the options section, labelled there. We found no independent classification results for Claude Haiku 5.5, which came out on 7 October.
What the confidence numbers are worth
The probability on every answer is a large part of the appeal, because code can act on confident answers and escalate the rest. The published results on how far to trust it are mixed.
In S Anand’s pilot, Jev’s mean confidence was 87.4% against 75.3% accuracy, and the Decisions API’s was 85.5% against 79.2%. Clef and Clef-flash were slightly underconfident, at 90.6% and 92.8% against 96.1% and 94.8%, and had the pilot’s best Brier scores.22 Prompted LLMs varied as much. Claude Opus 5 and Claude Fable 5.1 averaged within a point of their accuracy, while GPT-6 Luna through chat was 14 points overconfident.
A model can be overconfident on average and still rank its answers well. In AY Automate’s test, keeping only Jev’s answers at 0.90 confidence or above raised accuracy from 83.8% to 95.5% on 8-way routing, with 70% of items answered. A cascade that sent everything under 0.80 to GPT-5.6 Terra matched Terra’s accuracy at 26% to 28% of its cost and about half its mean latency.16
Confident errors never reach the fallback. At 0.90 or above, Jev was still wrong on 5 of 112 answers in the 8-way task, and all five filed a direct-debit complaint under card payments. Terra made the same five calls, because the two labels overlap.16 OpenAI’s guide says to set thresholds from labelled examples, weighing the cost of false positives against false negatives.8 TypeSafe’s docs say to pin a model version once thresholds are tuned, since the jev-latest alias moves when a new model ships.7
When to use a decision model, an LLM or a classifier
The table turns the evidence above into a starting rule. It weighs the latency budget, the cost per million decisions, the need for calibrated probabilities, the labelled data you have and how often the categories change.
| If | Start with | Evidence |
|---|---|---|
| Categories change often or differ per customer, labels are few, and the budget is under a second | A hosted decision model such as Clef-flash, Jev or OpenAI’s Decisions API | Options travel with each request; medians of 0.31 to 0.75 s and $26 to $444 per million decisions in the tests above |
| Categories stay stable for months, you have thousands of labelled examples, and volume is high or the budget is tens of milliseconds | A fine-tuned classifier such as ModernBERT or DeBERTa | ModernBERT matched Jev on IMDb after training on 25,000 labels; DeBERTa beat Jev on prompt injection running on a laptop CPU |
| The answer needs an explanation, extracted text, arithmetic, dates or several reasoning steps | An LLM with structured output, such as Claude Haiku 5.5 or GPT-6 Luna at low effort | Decision models return only typed values, and TypeSafe lists numbers, dates and multi-hop questions as weak spots |
| You want close to frontier accuracy at a fraction of frontier cost | A cascade, with a decision model first and a stronger LLM below a confidence gate | Matched GPT-5.6 Terra at 26% to 28% of its cost with a 0.80 gate |
| Inputs may be adversarial, as with prompt injection | A dedicated classifier, with rules in code around it | DeBERTa beat Jev by 2.7 points, and TypeSafe says adversarial content can move Jev’s answers |
| Data must stay on your own hardware | Clef-flash, Strands Decider 2B or a fine-tuned encoder | Apache 2.0 weights; llama.cpp serves Clef and several Jev-style models |
Start with the latency budget. A check that must finish in under 100 ms points to a classifier running next to your code. Red Hat’s pre-trained classifiers had medians of 33 to 80 ms on a laptop CPU, while a hosted model also pays for the network, at least about 56 ms in Red Hat’s setup.20 Hosted decision models fit budgets from about 300 ms to a second. An LLM for classification still makes sense when a second or more is acceptable and you want its reasons as well as its label.
Then count your labels. Whichever option you pick, you need a labelled set from your own traffic to measure accuracy and choose thresholds. At a few hundred items per task, AY Automate’s 95% intervals were 6 to 13 points wide, so treat smaller gaps between models as noise until you have more data.16
Patterns for intent routing, guardrails and rule checks
A study of 2,170 public GitHub projects that used Jev, collected on 22 September, found attribute judgement and scoring in wide use.23 The three patterns below come from the vendors’ docs and examples.
Intent routing. Classify each request once, act on confident answers and send the rest to a person or a stronger model. TypeSafe’s docs route support messages this way. Order-status questions go to plain code, product and returns questions to specialist LLMs, and unclear messages or complex complaints to a person.24 In the illustrative sketch below, decide stands in for any decision API, and the thresholds are set on labelled data.
# Illustrative sketch: `decide` stands in for any decision API.
INTENTS = {
"order_status": "Asks where an existing order is",
"refund": "Wants money back for an order or a charge",
"technical": "Reports a problem using the product",
"other": "Fits none of the above",
}
def route(message, order):
r = decide(state=message, questions={
"intent": {"type": "choice", "options": INTENTS},
"upset": {"type": "yes_no", "question": "Is the customer upset?"},
})
intent = r["intent"]
if intent["confidence"] < MIN_CONFIDENCE: # chosen on labelled data
return "human"
if intent["choice"] == "refund":
# Amounts and dates are arithmetic, so code checks them.
if order.total > REFUND_LIMIT or order.days_since_delivery > RETURN_DAYS:
return "human"
return "refund_flow"
if intent["choice"] == "technical":
upset = r["upset"]["probability"] > UPSET_CUTOFF
return "senior_support" if upset else "support_llm"
return "order_lookup" if intent["choice"] == "order_status" else "human"
The “other” option follows OpenAI’s advice to give unmatched inputs somewhere to go, and the refund limits stay in code because TypeSafe lists arithmetic and dates as weak spots.814
Guardrails. A guardrail asks a yes/no question about an input or an action before it runs. The Strands example asks whether a proposed tool call’s arguments come from something the user said, and whether it’s too early to call the tool. Code turns the two probabilities into an action, which in the demo sends the agent back to ask the user.5 Jev scored highest on content safety in Red Hat’s test. For prompt injection we’d put a dedicated classifier first, since DeBERTa did better there and TypeSafe says content written to steer the model can move Jev’s answer.2014
Rule checks. Anything code can compute exactly should stay in code, and the model should be asked only for the judgement. It’s the split we use around LLM judges, where ordinary code checks allowed values, ranges and required fields on every answer, as in our pipeline with two reviews and a judge. A decision model’s typed output removes the parsing step, but range and policy checks still run in your code.
A checklist before you ship one
- Build a labelled set from your own traffic, a few hundred items per decision, and hold part of it back.
- Run your shortlist on it with your real prompts and option descriptions. Record accuracy, p50 and p95 latency from your own region, and cost per million decisions from billed tokens.
- Compare mean confidence with accuracy, and read every confident error. Overlapping categories show up there first.
- Choose thresholds on one part of the set and confirm them on the part you held back.
- Add an “other” option, and a path to a person or a stronger model.
- Shuffle the option order in tests and check the answers hold, since Jev can lean towards the first option.
- Pin the model version once thresholds are tuned, and re-run the set before you move to a new one.
- Keep arithmetic, dates and policy limits in code.
- Log the model version, probabilities and final route for every decision, and count escalations and refusals. Our post on a pipeline that reported 0 dropped shows what goes missing when only failures are counted.
-
Hacker News, “Introducing System One Models and Jev”, 15 September 2026, https://news.ycombinator.com/item?id=49717558, and “Clef: Open-weight decision models, and new RL fine-tuning platform”, 1 October 2026, https://news.ycombinator.com/item?id=49923692. We counted later stories with Hacker News Search, https://hn.algolia.com/ ↩
-
OpenAI, “Changelog”, entries for 22 September 2026 (GPT-6 Luna) and 6 October 2026 (Decisions API), https://developers.openai.com/api/docs/changelog ↩↩↩
-
Anthropic, “Claude Haiku 5.5”, released 7 October 2026, https://platform.claude.com/docs/en/models/haiku-5-5/overview, and “Pricing”, https://platform.claude.com/docs/en/about-claude/pricing ↩↩
-
Cloudflare, “Introducing Clef: our open-source decision models, and new RL fine-tuning platform”, 1 October 2026, https://blog.cloudflare.com/clef-decision-models/ ↩↩↩↩
-
Strands Agents, “Introducing Strands Decider 2B: a small, open source, decision model”, 1 October 2026, https://strandsagents.com/blog/introducing-strands-decider/ ↩↩↩↩↩
-
TypeSafe AI, “Introducing System One Models & Jev”, 15 September 2026, https://typesafe.ai/blog/introducing-system-one-models-and-jev ↩↩↩↩
-
TypeSafe AI, “Models”, https://docs.typesafe.ai/models ↩↩↩↩
-
OpenAI, “Decisions”, https://developers.openai.com/api/docs/guides/decisions ↩↩↩↩↩↩
-
Simon Willison, “Jev introduces a new shape of LLM—System One, aka Decision Models”, 21 September 2026, https://simonwillison.net/2026/Sep/21/jev/ ↩
-
TypeSafe AI, “System One”, https://docs.typesafe.ai/concepts/system-one ↩
-
Sebastian Raschka, “Language Models for Text Classification: From Bag-of-Words to Jev”, 29 September 2026, https://magazine.sebastianraschka.com/p/classifier-history-and-jev ↩↩
-
Damani et al., “Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty”, arXiv, July 2025, https://arxiv.org/abs/2507.16806 ↩
-
Warner et al., “Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference”, arXiv, December 2024, https://arxiv.org/abs/2412.13663 ↩↩
-
TypeSafe AI, “Jev 1.13 jaggedness”, last reviewed 2 October 2026, https://docs.typesafe.ai/model-jaggedness/jev-1.13 ↩↩↩
-
Cloudflare, “Workers AI pricing”, https://developers.cloudflare.com/workers-ai/platform/pricing/ ↩
-
AY Automate, “We Tested Jev on 791 Labeled Decisions Against Four LLMs”, 20 September 2026, https://www.ayautomate.com/blog/jev-vs-llm-benchmark ↩↩↩↩↩↩↩
-
Cloudflare, Workers AI model pages for clef and clef-flash, https://developers.cloudflare.com/workers-ai/models/clef/ and https://developers.cloudflare.com/workers-ai/models/clef-flash/ ↩
-
Hugging Face, “Jev Decision Index”, https://huggingface.co/spaces/multimodalart/jev-decision-index ↩
-
llama.cpp, release v0.6.0 (5 October 2026), https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0, and build b11447 (6 October 2026), https://github.com/ggml-org/llama.cpp/releases/tag/b11447 ↩
-
Red Hat Developer, “Benchmarking AI decision models against traditional guardrails”, 2 October 2026, https://developers.redhat.com/articles/2026/10/02/benchmarking-ai-decision-models-against-traditional-guardrails ↩↩↩↩
-
Protect AI, “deberta-v3-base-prompt-injection-v2”, Hugging Face, https://huggingface.co/protectai/deberta-v3-base-prompt-injection-v2 ↩
-
S Anand, “Decision models vs frontier LLMs on BANKING77”, version of 7 October 2026, https://sanand0.github.io/llmevals/jev/, with code and data at https://github.com/sanand0/llmevals/tree/715e704f70/jev ↩↩↩↩
-
Ling, Xue and Ye, “Jev in the Wild: A Data-Driven Analysis of the Jev Model’s Functionality, Applications and Ecosystem”, arXiv, 24 September 2026, https://arxiv.org/abs/2609.30216 ↩
-
TypeSafe AI, “Intent routing”, https://docs.typesafe.ai/patterns/intent-routing ↩
Frequently asked questions
What is a decision model in AI?
An AI decision model reads text, and sometimes images, and answers questions you define in advance with typed values instead of text. Each answer is a yes/no probability, one of your options or a score, with a probability for every option. TypeSafe’s Jev, OpenAI’s Decisions API and Cloudflare’s Clef are examples.
What is Jev AI?
Jev is a decision model from TypeSafe AI, launched in early access on 15 September 2026. It takes text and typed questions and returns choices, yes/no probabilities and scores with confidence values. It costs $0.042 per million input tokens with free output, and it can’t generate free text.
Is Jev better than an LLM for classification?
It is faster and cheaper, and in independent tests about as accurate as small LLMs. In AY Automate’s 791-decision test, Jev was 3.6 times faster than GPT-5.6 Terra and 40 to 49 times cheaper per decision, but Terra was 5 to 6 points more accurate on intent routing.
How much does the OpenAI Decisions API cost?
With gpt-6-luna, the only supported model, input costs $0.10 per million tokens, and there are no charges for output tokens or caching. Regional processing premiums and long-context multipliers still apply. The API has been in public beta since 6 October 2026.
What is Cloudflare Clef?
Clef and Clef-flash are decision models Cloudflare released on 1 October 2026 as Apache 2.0 open weights. Clef has 27B parameters and Clef-flash 9B. Both accept images, follow Jev’s API format, and cost $0.24 and $0.09 per million input tokens on Workers AI.
Should I use an LLM for classification or a fine-tuned classifier?
Fine-tune a classifier such as ModernBERT when the categories are stable and you have thousands of labelled examples. Use a decision model when categories change often and you need answers within about a second. Use an LLM with structured output when the answer needs an explanation, arithmetic or several reasoning steps.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.