Buying guides

When to use Jev, an LLM or traditional machine learning

  • Most AI calls inside a product are decisions about text. Send each one to the cheapest method that makes it reliably on your own labeled cases.
  • Jev is fast and cheap. Independent tests found it 6 to 14 times faster than Claude Opus 5, well short of the 193.6 times TypeSafe's homepage claims against LLMs.
  • A small classifier trained on your own labels is Jev's real rival, and most comparisons leave it out.
  • Route cheapest first, gate each action on confidence, and keep a person on adverse decisions.

We think most AI calls inside a product are decisions about text, and each one should go to the cheapest method that makes it reliably on your own labeled cases. Decisions multiply. One support ticket can need four before anyone writes a reply: which queue, how urgent, whether the customer is angry and whether the request fits policy.

In mid-September 2026, TypeSafe AI launched Jev, which it calls the first System One model. Jev answers typed questions with probabilities and writes no text.

What this comparison covers

Jev competes on one kind of work, decisions about text, such as routing, triage, moderation and scoring. Writing text or code, and multi-step agent work, stay with large language models, or LLMs. Numeric prediction, such as fraud scores and forecasts, stays with tabular machine learning, or tabular ML. TypeSafe's own docs say Jev is "not a calculator."

The name System One comes from Daniel Kahneman's System 1, fast and automatic thinking, which his Nobel lecture says the slower System 2 monitors. Where a common default sends every decision to the strongest model, we would use that model to check what the fast methods flag.

What TypeSafe claims, and what the tests measured

Jev is credible, and TypeSafe itself qualifies the headline numbers. Jev takes text, which TypeSafe calls the state, and answers typed questions about it in parallel. It picks an option, rates on a scale you define, or gives the probability that a statement is true. TypeSafe trains it for calibrated answers, meaning that of the answers rated 0.8, about 80% should be right.

The launch post expects the homepage multipliers below to be "on the higher end of real world gains," and calls the zero-hallucination figure, a format guarantee, "not empirical."

TypeSafe's claims against independent measurements

  • vendor claim

    193.6x

    faster than LLMs, on TypeSafe's own workflow tests

    TypeSafe homepage

  • measured

    6 to 14x

    faster per decision than Claude Opus 5, in the two tests that measured speed

    jev-bench and jev-benchmark

  • vendor claim

    444.6x

    cheaper than LLMs, on the same workflow tests

    TypeSafe homepage

  • measured

    120 to 320x

    cheaper per decision than Claude Opus 5, in the same two tests

    jev-bench and jev-benchmark

Both measured ranges are large wins for Jev, and well short of the homepage multipliers.

Structured outputs from OpenAI and Anthropic now give LLM answers a fixed format too, so Jev's difference is speed, price and a probability on every answer.

The known failure modes of jev-1.13, the current version, matter more than the multipliers. Counting and date logic belong in code, Jev leans toward the option listed first, and misleading instructions in the input can move the answer. Accuracy is best in English.

What the first independent tests found

The early tests support Jev's speed and price. Accuracy and calibration depend on the task, and the tests are small and mostly self-published.

Could this change break production? Accuracy on 180 decisions

Could this change break production? Accuracy on 180 decisions
Value
Keyword rule60.6%
Jev63.9%
GPT-5.6 Terra66.1%
Claude Opus 580.0%
The risk task from ejs-5/jev-benchmark, 868 decisions in all, labeled from git history. Jev beat the rule and trailed both LLMs on all four tasks.
  • A 72-decision test put Jev at 94.4% accuracy, Claude Haiku 4.5 at 91.7% and Claude Opus 5 at 98.6%.
  • A test on 900 synthetic tickets set one label by a company rule absent from the text. Jev scored 44.7%, against 25% by chance, yet its answers carried an average probability of 0.74. The author's advice is "Calibrate per question, not per model."
  • A preregistered study, which fixed its method before seeing results, found Jev's choice probabilities close to calibrated where human annotators agree, and overconfident where they split.
  • A radiology paper found Jev a practical, low-cost judge of report errors, ahead of a public DeBERTa model trained to judge whether one statement supports another. A locally run LLM pipeline beat Jev on clinically significant errors.

The rival most comparisons leave out

Once you have a few hundred labeled examples, we think a small classifier trained on them is the method to beat: a fine-tuned DeBERTa-class model, SetFit, or sentence embeddings fed to a logistic regression.

In tests by Bucher and Martini, fine-tuned models such as RoBERTa and DeBERTa V3 beat GPT-3.5, GPT-4 and Claude Opus used zero-shot "in all cases," across four kinds of text classification. Fine-tuned RoBERTa beat BART, which they call arguably the most consistent zero-shot model, on every dataset above 200 labeled examples.

SetFit, from Tunstall and colleagues, can need far fewer. Its README says that with 8 labeled examples per class, on one sentiment dataset, it was competitive with RoBERTa Large trained on 3,000.

Such a classifier runs on your own compute with no per-call fee, its probabilities can be calibrated with standard tools, and the text stays in your environment, which helps with data residency. It costs labeling, and retraining when categories change. Jev needs neither to start.

We know of no matched, independent test of Jev against a classifier trained on the task's own labels, so this part of our comparison is reasoning, not evidence. It is also the most useful test a team with labeled data can run.

The five methods side by side

Five ways to make a decision about text

RulesTabular MLTrained text classifierSystem One model, such as JevLLM
Labeled dataNot needed, but the policy must be written down (no)Required (yes)Required, often a few hundred examples (yes)Needed to test, not to start; there is no training step (partly)Needed to test, not to start (partly)
LatencyEffectively instantMilliseconds, often in processNo network call; depends on model size70 to 500 ms claimed; about 185 to 430 ms measuredAbout 1.2 to 5 seconds in the same tests
Cost per callClose to zeroClose to zero; labeling and retraining cost moreYour own compute; labeling and retraining cost more$0.042 per million input tokens, output free, per TypeSafe$0.82 to $6.40 per 1,000 decisions in one test, against $0.02 for Jev
Explainability and calibrationFully explainable, no probabilitiesPer-feature explanations with SHAP, and standard calibration toolsProbabilities, calibrated with the same toolsProbabilities, no reasons; calibration varies by taskWritten reasons can omit what drove the answer; token probabilities measure certainty, not correctness
Where it runsYour code (yes)Your environment (yes)Your environment (yes)TypeSafe's API (no)A vendor's API, or an open model you host (partly)
Regulated decisionEasiest to auditAuditable with documentation and monitoringSame as tabular MLPin the version, log probabilities, keep a person on adverse outcomesKeep a person on adverse outcomes; a written reason is not a decision record
  • Rules

    Labeled data
    Not needed, but the policy must be written down (no)
    Latency
    Effectively instant
    Cost per call
    Close to zero
    Explainability and calibration
    Fully explainable, no probabilities
    Where it runs
    Your code (yes)
    Regulated decision
    Easiest to audit
  • Tabular ML

    Labeled data
    Required (yes)
    Latency
    Milliseconds, often in process
    Cost per call
    Close to zero; labeling and retraining cost more
    Explainability and calibration
    Per-feature explanations with SHAP, and standard calibration tools
    Where it runs
    Your environment (yes)
    Regulated decision
    Auditable with documentation and monitoring
  • Trained text classifier

    Labeled data
    Required, often a few hundred examples (yes)
    Latency
    No network call; depends on model size
    Cost per call
    Your own compute; labeling and retraining cost more
    Explainability and calibration
    Probabilities, calibrated with the same tools
    Where it runs
    Your environment (yes)
    Regulated decision
    Same as tabular ML
  • System One model, such as Jev

    Labeled data
    Needed to test, not to start; there is no training step (partly)
    Latency
    70 to 500 ms claimed; about 185 to 430 ms measured
    Cost per call
    $0.042 per million input tokens, output free, per TypeSafe
    Explainability and calibration
    Probabilities, no reasons; calibration varies by task
    Where it runs
    TypeSafe's API (no)
    Regulated decision
    Pin the version, log probabilities, keep a person on adverse outcomes
  • LLM

    Labeled data
    Needed to test, not to start (partly)
    Latency
    About 1.2 to 5 seconds in the same tests
    Cost per call
    $0.82 to $6.40 per 1,000 decisions in one test, against $0.02 for Jev
    Explainability and calibration
    Written reasons can omit what drove the answer; token probabilities measure certainty, not correctness
    Where it runs
    A vendor's API, or an open model you host (partly)
    Regulated decision
    Keep a person on adverse outcomes; a written reason is not a decision record

Start from the left, and send only the cases that fail to the right. Tokens are the word fragments models bill by.

In the EU, GDPR Article 22 limits decisions based solely on automated processing that have legal or similarly significant effects. Where such a decision rests on a contract or explicit consent, the safeguards must include the right to human intervention. This is not legal advice.

A routing pattern that uses all five

We would route each decision cheapest first, then let confidence decide.

  1. A decision: which queue, how urgent. Leads to Rules, System One model, Trained model.
  2. Rules: what code can compute. Leads to Confidence gate.
  3. System One model: text, few labels. Leads to Confidence gate.
  4. Trained model: tabular ML or classifier. Leads to Confidence gate.
  5. Confidence gate: a threshold per action. Leads to Automatic (high), LLM (mid), Person reviews (low).
  6. Automatic: high confidence.
  7. LLM: a second look.
  8. Person reviews: low or adverse. Leads to Trained model (labeled examples).
Each human decision becomes a labeled example that tests every method and later trains the cheaper ones.

The routing, step by step

  1. Rules take what code can compute or policy can state

    Amounts, dates and customer tier. A model cannot apply a policy it was never given, as the synthetic-ticket test shows.

  2. Tabular ML takes numeric predictions

    Fraud scores, for example, where tree-based models like XGBoost remain strong.

  3. A trained text classifier takes text decisions with enough labels

    And stable categories.

  4. A System One model takes text decisions with few labels

    Or new categories.

  5. A gate sets a confidence threshold for each action

    TypeSafe's routing example sends anything under 0.6 to a support agent and approves a transfer automatically only above 0.85. Set yours from your own test set and risk appetite.

  6. An LLM takes cases that need text or multi-step reasoning

    It can also give mid-confidence cases a second look.

  7. A person reviews low-confidence cases and adverse decisions

    Plus a random sample of confident ones, because the preregistered study found Jev most overconfident where people disagree.

  8. Each human decision becomes a labeled example

    It tests every method, and later trains steps 2 and 3.

This keeps the human-review rate in our cost per successful outcome formula small.

Test the claims on your own data

Vendor figures come from the vendor's setup, and TypeSafe says so. Trust your own test set over any published benchmark.

Five checks on your own data

  1. Latency

    Measure median and 95th-percentile latency from your region at peak volume, against Jev's default limit of 80 requests per second.

  2. Cost per correct decision

    Include LLM fallbacks and human review, and check that the number survives a price change.

  3. Accuracy against real outcomes

    Score it next to a keyword rule and a trained classifier, in every language you serve.

  4. Calibration for each question type

    Set thresholds per question.

  5. A pinned model version

    Use a versioned model ID, because aliases like jev-latest move with each release.

Where we land

Jev is a useful option for text decisions at volume, where the options are known in advance and labels are few.

TypeSafe's manifesto warns that assistant-style training leads to "AI that requires humans in the loop instead of running in the background." We agree that confident, low-risk decisions should run in the background, and differ on the person in the loop, whose reviews produce the labels that train cheaper methods and the accountability that rules like Article 22 expect.

The parts worth owning are the decision definitions, the labeled test set, the thresholds and the review queue. They carry over to the next model, as we argued in the agent is a compiler.

If most of your model calls are text decisions, compare the five methods on your own labeled cases and rank them by cost per correct decision. To compare notes on the method, talk to us.

Sources

TypeSafe, vendor claims:

Coverage and availability:

  • The Register and VKTR, launch coverage, 16 September 2026. Both note that the zero-hallucination claim does not mean every answer is correct.
  • Vercel and Netlify AI Gateway changelogs. Both gateways offer Jev, which TypeSafe describes as early access.

Independent tests:

Trained text classifiers:

LLMs:

Traditional machine learning:

Kahneman and law:

Related posts: cost per outcome and the agent is a compiler.

· views
Amol PatilFounderFounded extendfuture in 2019. Has shipped computer vision, voice and agentic systems into production across ten industries. Amol Patil on LinkedIn

Working on something in this territory?

Tell us what you are trying to win. We answer within one business day, from the people who build.