Buying guides
When to use Jev, an LLM or traditional machine learning
- Most AI calls inside a product are decisions about text. Send each one to the cheapest method that makes it reliably on your own labeled cases.
- Jev is fast and cheap. Independent tests found it 6 to 14 times faster than Claude Opus 5, well short of the 193.6 times TypeSafe's homepage claims against LLMs.
- A small classifier trained on your own labels is Jev's real rival, and most comparisons leave it out.
- Route cheapest first, gate each action on confidence, and keep a person on adverse decisions.
We think most AI calls inside a product are decisions about text, and each one should go to the cheapest method that makes it reliably on your own labeled cases. Decisions multiply. One support ticket can need four before anyone writes a reply: which queue, how urgent, whether the customer is angry and whether the request fits policy.
In mid-September 2026, TypeSafe AI launched Jev, which it calls the first System One model. Jev answers typed questions with probabilities and writes no text.
What this comparison covers
Jev competes on one kind of work, decisions about text, such as routing, triage, moderation and scoring. Writing text or code, and multi-step agent work, stay with large language models, or LLMs. Numeric prediction, such as fraud scores and forecasts, stays with tabular machine learning, or tabular ML. TypeSafe's own docs say Jev is "not a calculator."
The name System One comes from Daniel Kahneman's System 1, fast and automatic thinking, which his Nobel lecture says the slower System 2 monitors. Where a common default sends every decision to the strongest model, we would use that model to check what the fast methods flag.
What TypeSafe claims, and what the tests measured
Jev is credible, and TypeSafe itself qualifies the headline numbers. Jev takes text, which TypeSafe calls the state, and answers typed questions about it in parallel. It picks an option, rates on a scale you define, or gives the probability that a statement is true. TypeSafe trains it for calibrated answers, meaning that of the answers rated 0.8, about 80% should be right.
The launch post expects the homepage multipliers below to be "on the higher end of real world gains," and calls the zero-hallucination figure, a format guarantee, "not empirical."
TypeSafe's claims against independent measurements
vendor claim
193.6x
faster than LLMs, on TypeSafe's own workflow tests
measured
6 to 14x
faster per decision than Claude Opus 5, in the two tests that measured speed
vendor claim
444.6x
cheaper than LLMs, on the same workflow tests
measured
120 to 320x
cheaper per decision than Claude Opus 5, in the same two tests
Both measured ranges are large wins for Jev, and well short of the homepage multipliers.
Structured outputs from OpenAI and Anthropic now give LLM answers a fixed format too, so Jev's difference is speed, price and a probability on every answer.
The known failure modes of jev-1.13, the current version, matter more than the multipliers. Counting and date logic belong in code, Jev leans toward the option listed first, and misleading instructions in the input can move the answer. Accuracy is best in English.
What the first independent tests found
The early tests support Jev's speed and price. Accuracy and calibration depend on the task, and the tests are small and mostly self-published.
- A 72-decision test put Jev at 94.4% accuracy, Claude Haiku 4.5 at 91.7% and Claude Opus 5 at 98.6%.
- A test on 900 synthetic tickets set one label by a company rule absent from the text. Jev scored 44.7%, against 25% by chance, yet its answers carried an average probability of 0.74. The author's advice is "Calibrate per question, not per model."
- A preregistered study, which fixed its method before seeing results, found Jev's choice probabilities close to calibrated where human annotators agree, and overconfident where they split.
- A radiology paper found Jev a practical, low-cost judge of report errors, ahead of a public DeBERTa model trained to judge whether one statement supports another. A locally run LLM pipeline beat Jev on clinically significant errors.
The rival most comparisons leave out
Once you have a few hundred labeled examples, we think a small classifier trained on them is the method to beat: a fine-tuned DeBERTa-class model, SetFit, or sentence embeddings fed to a logistic regression.
In tests by Bucher and Martini, fine-tuned models such as RoBERTa and DeBERTa V3 beat GPT-3.5, GPT-4 and Claude Opus used zero-shot "in all cases," across four kinds of text classification. Fine-tuned RoBERTa beat BART, which they call arguably the most consistent zero-shot model, on every dataset above 200 labeled examples.
SetFit, from Tunstall and colleagues, can need far fewer. Its README says that with 8 labeled examples per class, on one sentiment dataset, it was competitive with RoBERTa Large trained on 3,000.
Such a classifier runs on your own compute with no per-call fee, its probabilities can be calibrated with standard tools, and the text stays in your environment, which helps with data residency. It costs labeling, and retraining when categories change. Jev needs neither to start.
We know of no matched, independent test of Jev against a classifier trained on the task's own labels, so this part of our comparison is reasoning, not evidence. It is also the most useful test a team with labeled data can run.
The five methods side by side
Five ways to make a decision about text
| Rules | Tabular ML | Trained text classifier | System One model, such as Jev | LLM | |
|---|---|---|---|---|---|
| Labeled data | Not needed, but the policy must be written down (no) | Required (yes) | Required, often a few hundred examples (yes) | Needed to test, not to start; there is no training step (partly) | Needed to test, not to start (partly) |
| Latency | Effectively instant | Milliseconds, often in process | No network call; depends on model size | 70 to 500 ms claimed; about 185 to 430 ms measured | About 1.2 to 5 seconds in the same tests |
| Cost per call | Close to zero | Close to zero; labeling and retraining cost more | Your own compute; labeling and retraining cost more | $0.042 per million input tokens, output free, per TypeSafe | $0.82 to $6.40 per 1,000 decisions in one test, against $0.02 for Jev |
| Explainability and calibration | Fully explainable, no probabilities | Per-feature explanations with SHAP, and standard calibration tools | Probabilities, calibrated with the same tools | Probabilities, no reasons; calibration varies by task | Written reasons can omit what drove the answer; token probabilities measure certainty, not correctness |
| Where it runs | Your code (yes) | Your environment (yes) | Your environment (yes) | TypeSafe's API (no) | A vendor's API, or an open model you host (partly) |
| Regulated decision | Easiest to audit | Auditable with documentation and monitoring | Same as tabular ML | Pin the version, log probabilities, keep a person on adverse outcomes | Keep a person on adverse outcomes; a written reason is not a decision record |
Rules
- Labeled data
- Not needed, but the policy must be written down (no)
- Latency
- Effectively instant
- Cost per call
- Close to zero
- Explainability and calibration
- Fully explainable, no probabilities
- Where it runs
- Your code (yes)
- Regulated decision
- Easiest to audit
Tabular ML
- Labeled data
- Required (yes)
- Latency
- Milliseconds, often in process
- Cost per call
- Close to zero; labeling and retraining cost more
- Explainability and calibration
- Per-feature explanations with SHAP, and standard calibration tools
- Where it runs
- Your environment (yes)
- Regulated decision
- Auditable with documentation and monitoring
Trained text classifier
- Labeled data
- Required, often a few hundred examples (yes)
- Latency
- No network call; depends on model size
- Cost per call
- Your own compute; labeling and retraining cost more
- Explainability and calibration
- Probabilities, calibrated with the same tools
- Where it runs
- Your environment (yes)
- Regulated decision
- Same as tabular ML
System One model, such as Jev
- Labeled data
- Needed to test, not to start; there is no training step (partly)
- Latency
- 70 to 500 ms claimed; about 185 to 430 ms measured
- Cost per call
- $0.042 per million input tokens, output free, per TypeSafe
- Explainability and calibration
- Probabilities, no reasons; calibration varies by task
- Where it runs
- TypeSafe's API (no)
- Regulated decision
- Pin the version, log probabilities, keep a person on adverse outcomes
LLM
- Labeled data
- Needed to test, not to start (partly)
- Latency
- About 1.2 to 5 seconds in the same tests
- Cost per call
- $0.82 to $6.40 per 1,000 decisions in one test, against $0.02 for Jev
- Explainability and calibration
- Written reasons can omit what drove the answer; token probabilities measure certainty, not correctness
- Where it runs
- A vendor's API, or an open model you host (partly)
- Regulated decision
- Keep a person on adverse outcomes; a written reason is not a decision record
Start from the left, and send only the cases that fail to the right. Tokens are the word fragments models bill by.
In the EU, GDPR Article 22 limits decisions based solely on automated processing that have legal or similarly significant effects. Where such a decision rests on a contract or explicit consent, the safeguards must include the right to human intervention. This is not legal advice.
A routing pattern that uses all five
We would route each decision cheapest first, then let confidence decide.
- A decision: which queue, how urgent. Leads to Rules, System One model, Trained model.
- Rules: what code can compute. Leads to Confidence gate.
- System One model: text, few labels. Leads to Confidence gate.
- Trained model: tabular ML or classifier. Leads to Confidence gate.
- Confidence gate: a threshold per action. Leads to Automatic (high), LLM (mid), Person reviews (low).
- Automatic: high confidence.
- LLM: a second look.
- Person reviews: low or adverse. Leads to Trained model (labeled examples).
The routing, step by step
Rules take what code can compute or policy can state
Amounts, dates and customer tier. A model cannot apply a policy it was never given, as the synthetic-ticket test shows.
Tabular ML takes numeric predictions
Fraud scores, for example, where tree-based models like XGBoost remain strong.
A trained text classifier takes text decisions with enough labels
And stable categories.
A System One model takes text decisions with few labels
Or new categories.
A gate sets a confidence threshold for each action
TypeSafe's routing example sends anything under 0.6 to a support agent and approves a transfer automatically only above 0.85. Set yours from your own test set and risk appetite.
An LLM takes cases that need text or multi-step reasoning
It can also give mid-confidence cases a second look.
A person reviews low-confidence cases and adverse decisions
Plus a random sample of confident ones, because the preregistered study found Jev most overconfident where people disagree.
Each human decision becomes a labeled example
It tests every method, and later trains steps 2 and 3.
This keeps the human-review rate in our cost per successful outcome formula small.
Test the claims on your own data
Vendor figures come from the vendor's setup, and TypeSafe says so. Trust your own test set over any published benchmark.
Five checks on your own data
Latency
Measure median and 95th-percentile latency from your region at peak volume, against Jev's default limit of 80 requests per second.
Cost per correct decision
Include LLM fallbacks and human review, and check that the number survives a price change.
Accuracy against real outcomes
Score it next to a keyword rule and a trained classifier, in every language you serve.
Calibration for each question type
Set thresholds per question.
A pinned model version
Use a versioned model ID, because aliases like jev-latest move with each release.
Where we land
Jev is a useful option for text decisions at volume, where the options are known in advance and labels are few.
TypeSafe's manifesto warns that assistant-style training leads to "AI that requires humans in the loop instead of running in the background." We agree that confident, low-risk decisions should run in the background, and differ on the person in the loop, whose reviews produce the labels that train cheaper methods and the accountability that rules like Article 22 expect.
The parts worth owning are the decision definitions, the labeled test set, the thresholds and the review queue. They carry over to the next model, as we argued in the agent is a compiler.
If most of your model calls are text decisions, compare the five methods on your own labeled cases and rank them by cost per correct decision. To compare notes on the method, talk to us.
Sources
TypeSafe, vendor claims:
- Launch post by Diogo Almeida. It says its evals were "generally run from our laptops on the West Coast" and, on price, "We can't prove it isn't subsidized." The homepage carries the speed and cost multipliers.
- Docs: System One, quickstart, models with rate limits, language support, data retention and one set of weights for every account, failure modes, routing
- Manifesto and system-one-adapter-python on GitHub, which runs the same typed questions through an LLM for comparison
Coverage and availability:
- The Register and VKTR, launch coverage, 16 September 2026. Both note that the zero-hallucination claim does not mean every answer is correct.
- Vercel and Netlify AI Gateway changelogs. Both gateways offer Jev, which TypeSafe describes as early access.
Independent tests:
- stern9/jev-bench, which measured median latency of about 185 milliseconds for Jev and 2.6 seconds for Claude Opus 5, ejs-5/jev-benchmark, scienthoon/jev-ood-calibration, GautamTalksDev/jevbench
- Huang et al., radiology paper, arXiv, September 2026
Trained text classifiers:
- Bucher and Martini, "Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification", arXiv, 2024
- Tunstall et al., "Efficient Few-Shot Learning Without Prompts", the SetFit paper, arXiv, 2022, and huggingface/setfit on GitHub
- microsoft/DeBERTa and sentence-transformers on GitHub
LLMs:
- OpenAI, structured outputs and logprobs cookbook; Anthropic, structured outputs. Both structured-output features hold an answer to a JSON schema you define, apart from cases such as refusals.
- OpenAI, GPT-4 technical report, Figure 8, which found that post-training "hurts calibration significantly"
- Anthropic, Reasoning models don't always say what they think
Traditional machine learning:
- scikit-learn, probability calibration; XGBoost; SHAP
- Grinsztajn, Oyallon and Varoquaux, tree-based models on tabular data, 2022
Kahneman and law:
- Daniel Kahneman, Thinking, Fast and Slow and Maps of Bounded Rationality, Nobel lecture, 2002. The lecture credits the System 1 and System 2 labels to the psychologists Keith Stanovich and Richard West.
- GDPR, Article 22
Related posts: cost per outcome and the agent is a compiler.