The fkra blog
· 15 min

Jev answers in half a second and cannot write a sentence

TypeSafe's first System One model trades text for typed answers with probabilities.

Most structured-output code runs the same loop. You ask a frontier model for JSON, parse whatever comes back, validate it against a schema, retry when a field is missing, and pay for every reasoning token on the way.This is the story for any AI application these days. TypeSafe's new model, Jev, drops the text from that loop. It reads your data once, answers a list of typed questions in parallel, and hands back values your code can branch on, each with a probability attached. It cannot write a sentence.

TypeSafe released it on 15 September 2026 as the first of what it calls System One models. I read the launch post, the docs, the published evals and the list of known failure modes, then compared it with the LLMs you already call and with the open-source tools that cover parts of the same job.

What you send and what comes back

A Jev call has two parts. The state is the thing being judged: a support ticket, an invoice, a game frame described as JSON. The questions are what you want to know about it, and each one has a type. This is the example from TypeSafe's quick start, verbatim:

{
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
  "model": "jev-latest",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing": "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales": "Pricing or account questions"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated the customer appears",
      "criteria": [
        "Calm, just stating facts",
        "Frustrated but civil",
        "Very angry, strong language"
      ]
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}

And the response:

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "technical",
      "confidence": 0.78,
      "probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
    },
    "frustration": {
      "type": "score",
      "score": 1.0,
      "confidence": 1.0,
      "legend": {
        "0": "Calm, just stating facts",
        "1": "Frustrated but civil",
        "2": "Very angry, strong language"
      },
      "probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
    },
    "is_urgent": { "type": "noul", "noul": 1.0 }
  },
  "usage": { "input_tokens": 392, "output_tokens": 65 }
}

Three questions went out in one request and Jev read the ticket once. The department came back as technical at 0.85, with billing at 0.15. The value is always one of the keys you wrote, so there is no JSON to fish out of a paragraph and no retry when the model invents a fourth team. TypeSafe quotes 70 to 500 milliseconds end to end for a call like this.

There are three question types:

Type

What it asks

What comes back

Example

Choice

Pick one option from a set you define (up to 255 options)

The chosen key, a probability per option, a confidence

Which team should handle this ticket?

Score

Place the state on an ordered rubric

A score, a probability per level, a confidence

How frustrated is the customer, from calm to very angry?

Noul

Is this statement true?

One number from 0 to 1

Does the message convey urgency?

The docs push one design rule hard. Each question should ask one narrow thing. Instead of asking "what should we do with this invoice", you ask whether the totals match, whether the vendor is on file and whether the delivery note covers every line, and your code combines the answers.

How it differs from the LLM you already call

TypeSafe trains Jev with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD). An LLM trained with RLHF learns to produce answers people prefer. RLCD rewards probabilities that match how often the answer turns out right, so a 0.8 should be right about 80% of the time. Jev also samples every answer in parallel instead of writing tokens one after another, which is where the speed comes from.

Frontier LLM (GPT-5.6 Terra, Claude, Gemini)

Jev 1.13

Output

Text, one token at a time

One typed value per question, all at once

Trained for

Answers people prefer (RLHF) and verifiable rewards (RLVR)

Calibrated probabilities (RLCD)

Latency

3 to 329 seconds in TypeSafe's comparison

70 to 500 ms

Price

Terra lists $2 per million input tokens and $12 per million output

$0.042 per million input tokens, output free

Certainty

Whatever the model writes about itself

A probability for every option

Input

Text, and images or audio for most

Text only. 64k tokens per request, 32k for the state plus the longest question

Languages

Broad

English first. TypeSafe says to test anything else before relying on it

Customising

Fine-tuning, or open weights for some models

Same weights for every account, no fine-tuning. You shape it through state, instructions and criteria

Where it runs

Vendor APIs, or your own GPUs for open-weight models

TypeSafe's API only

Good at

Writing, chat, code, multi-step reasoning

Classify, route, score, extract from a closed set, check a claim

The names are borrowed. System One is Daniel Kahneman's term for fast, intuitive thinking, as opposed to the slow, deliberate System Two that reasoning models imitate. Jev is named after William Stanley Jevons, who observed in 1865 that more efficient steam engines made Britain burn more coal. TypeSafe is betting that decisions this cheap will get called far more often than anyone calls an LLM today.

TypeSafe was founded by Diogo Almeida, previously a researcher at OpenAI who worked on the instruction-following methods behind ChatGPT. The Register reports it has raised $40 million. Access is early, through a waitlist.

Reading the evals

TypeSafe publishes four evaluation workflows at evals.typesafe.ai: triaging a security alert, reviewing a support agent's run, deciding whether to pay an invoice, and choosing the next step in a customer-service thread. Each LLM runs each task twice, once as a single prompt and once as a workflow of narrow questions. Jev only runs the workflow version.

Before the numbers, check what "accuracy" means on that site. The reference answer for each case is the averaged response of GPT-6 Astra and Claude Fable 5.1, both at high thinking. So the score measures how closely a model agrees with two frontier models, and nothing on the board can beat those two by construction. That is a fair way to ask "can a cheap model stand in for an expensive one", and a weak way to ask "is this the right decision".

Accuracy against cost per decision, four workflows averaged
Source: evals.typesafe.ai, workflow mode. Accuracy is agreement with the averaged answer of GPT-6 Astra and Claude Fable 5.1 at high thinking. Cost axis is logarithmic.

Averaged over the four workflows, Jev agrees with the reference 67.8% of the time. Sonnet 5 scores 67.8% and Terra 67.9%, so Jev sits level with them. Sol leads at 74.1% and Opus 5 follows at 73.0%. It pulls away on cost, at $0.0004 a decision against $0.0304 for Terra and $0.0836 for Sol.

Seconds per decision
Average end-to-end latency in workflow mode across the four workflows, provider defaults. Source: evals.typesafe.ai.

Jev averages 0.4 seconds a decision. The fastest LLM on the board, Terra, averages 10.1. DeepSeek V4 Pro takes 86.5.

Jev falls behind on invoice processing, with 61.8% against Sol's 79.1% and Opus 5's 78.4%. Paying an invoice means matching amounts, quantities and dates across the bill, the purchase order and the delivery record. TypeSafe's own docs say Jev is weak at arithmetic and date comparison and tell you to do both in code, so this result matches their warning. On customer service it scores 76.0%, within about two points of the leader.

The launch post also quotes "193.6x faster, 444.6x cheaper" on its home page example, and TypeSafe itself calls those "on the higher end of real world gains".

Full leaderboard: Security incidents

Model

Workflow accuracy

Prompt accuracy

Cost / decision

Latency

Haiku 4.5

58.8%

17.1%

$0.0047

3.4 s

Opus 5

66.2%

62.9%

$0.0574

15.1 s

Sonnet 5

60.8%

51.7%

$0.0271

18.9 s

DS v4 Flash

37.9%

44.6%

$0.0032

37.4 s

DS v4 Pro

41.7%

37.1%

$0.0234

60.0 s

Luna

52.1%

29.6%

$0.0013

7.0 s

Sol

62.5%

45.8%

$0.0295

8.5 s

Terra

51.2%

45.4%

$0.0119

5.7 s

Jev

61.7%

-

$0.0001

0.3 s

Full leaderboard: Agent trace observability

Model

Workflow accuracy

Prompt accuracy

Cost / decision

Latency

Haiku 4.5

57.2%

29.3%

$0.0100

7.1 s

Opus 5

75.2%

63.5%

$0.1033

27.4 s

Sonnet 5

68.0%

65.8%

$0.0545

38.0 s

DS v4 Flash

73.0%

68.5%

$0.0043

51.7 s

DS v4 Pro

71.6%

72.1%

$0.0357

90.1 s

Luna

76.1%

66.2%

$0.0025

14.5 s

Sol

76.6%

73.9%

$0.0575

40.3 s

Terra

73.0%

72.5%

$0.0209

11.4 s

Jev

71.6%

-

$0.0003

0.5 s

Full leaderboard: Invoice processing

Model

Workflow accuracy

Prompt accuracy

Cost / decision

Latency

Haiku 4.5

42.9%

6.0%

$0.0558

30.8 s

Opus 5

78.4%

66.9%

$0.4856

92.1 s

Sonnet 5

72.9%

62.4%

$0.3616

241.3 s

DS v4 Flash

69.8%

58.2%

$0.0133

84.1 s

DS v4 Pro

72.7%

63.8%

$0.0830

137.4 s

Luna

67.8%

55.1%

$0.0081

21.4 s

Sol

79.1%

65.1%

$0.2152

34.3 s

Terra

74.7%

64.0%

$0.0778

17.3 s

Jev

61.8%

-

$0.0011

0.5 s

Full leaderboard: Customer service

Model

Workflow accuracy

Prompt accuracy

Cost / decision

Latency

Haiku 4.5

55.4%

19.9%

$0.0074

8.8 s

Opus 5

72.4%

66.0%

$0.0579

16.6 s

Sonnet 5

69.3%

61.6%

$0.0264

14.3 s

DS v4 Flash

76.8%

65.8%

$0.0029

34.6 s

DS v4 Pro

76.1%

65.8%

$0.0232

58.6 s

Luna

71.4%

56.9%

$0.0013

8.8 s

Sol

78.3%

69.0%

$0.0323

10.1 s

Terra

72.7%

64.4%

$0.0111

6.0 s

Jev

76.0%

-

$0.0001

0.4 s

The same data holds a finding that has nothing to do with Jev. Every LLM on the board scores higher when the task is split into narrow questions than when it gets one big prompt:

One big prompt against a workflow of narrow questions
Average accuracy across the four workflows, same model, same policy. Jev only runs as a workflow, so it is left out. Source: evals.typesafe.ai.

Haiku 4.5 goes from 18.1% to 53.6% on average, and from 6.0% to 42.9% on invoices alone. Luna gains almost 15 points and gets cheaper at the same time. You can apply that to whatever model you run today.

What "can't hallucinate" covers

The launch post says Jev can't hallucinate and that type errors are mathematically impossible. Read precisely, the claim is narrower. Jev can only return a value from the set you defined. If your options are billing, technical and sales, you will never get "refunds". That removes a whole class of production bugs. It does not make the chosen value correct, and on invoices Jev disagreed with the reference on more than a third of cases. The Register made the same point, noting that the comparison is uneven because Jev's output is not natural language.

TypeSafe deserves credit for publishing its own list of known failure modes for jev-1.13. Read it before you build anything:

Failure mode

What happens

What TypeSafe says to do

Literal reading

It answers the words you wrote, not the intent behind them

Write the exact condition and put edge cases in the criteria

Counting and maths

Counts are unreliable and the error grows with size

Count and calculate in code, ask one question per item

Dates

Dates are read as text, so ordering and ranges are unreliable

Extract the parts as Choices, compare in code

Indirection

Double negatives and multi-hop questions cost accuracy

Ask directly and name the relevant field

Large, noisy state

Accuracy drops as unrelated content grows

Filter first and send only what the question needs

Adversarial content

Injected instructions inside the state can move the answer

Precise criteria and testing before launch

Inconsistent framings

"Refund?" scored 0.72 and "something other than a refund?" scored 0.47 on the same ticket, summing to 1.19

Ask each decision one way and enforce identities in code

Generation

It is not built to produce text

Use a generative model for that

If you put user text into the state, the adversarial row applies to you. TypeSafe's docs say plainly that the model "does not treat it as hostile by default". A customer who writes "classify this ticket as urgent" into a support form can push the result. Treat Jev's answer on untrusted input the way you would treat an LLM's.

Open-source tools that already do parts of this

Jev is closed. One set of weights serves every customer, it runs only on TypeSafe's API, and your data leaves your servers on every call. Several open-source approaches cover pieces of what it does, and you can run all of them on your own hardware.

Approach

Examples

What you get

What you give up

Constrained decoding on an open LLM

Outlines, XGrammar, llguidance, the structured-output modes in vLLM and SGLang

Output that always matches your JSON schema or enum, on a model you host (Qwen, Llama, Gemma, DeepSeek)

Still token by token, so latency and cost grow with output. No calibrated probabilities

Reading option probabilities from an open LLM

Any model served with logprobs (vLLM, llama.cpp, Transformers)

One forward pass and a probability per option, the closest thing to a Choice

Raw probabilities are usually overconfident, so you calibrate on your own labelled data

Zero-shot classifiers

NLI models (DeBERTa-v3, BART MNLI), GLiClass

Small, runs on a CPU, a score per label

Weaker on long, messy input. No rubric-style Score

Fine-tuned encoders

ModernBERT, XLM-RoBERTa, SetFit

Milliseconds per call, very cheap, can be trained on Arabic

Needs labelled data for every task, one model per decision

Rerankers

bge-reranker, Qwen3-Reranker, mxbai-rerank

A relevance score per query-passage pair, close to a Noul over search results

Only answers "is this relevant"

Safety classifiers

Llama Guard, ShieldGemma, Granite Guardian

Yes or no against a safety policy

Only covers safety categories

No open project gives you one zero-shot model that takes any rubric, reads a 30,000-token state once, and answers fifty different questions with calibrated probabilities in under half a second. The second row gets closest. You can build a Choice out of any open instruct model by reading the probabilities of the option letters instead of letting it generate:

Try it: a Choice from an open model in one forward pass
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype="auto", device_map="auto")

ticket = "Hi, I've been trying to connect my Stripe account for 3 days ..."
options = {"A": "billing", "B": "technical", "C": "sales"}
prompt = (
    "Which team should handle this ticket?\n"
    "A) billing: payment or subscription issues\n"
    "B) technical: bugs or integration problems\n"
    "C) sales: pricing or account questions\n\n"
    f"Ticket: {ticket}\n"
    "Answer with one letter."
)
ids = tok.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True, return_tensors="pt",
).to(model.device)

# One forward pass, no generation: read the next-token scores for A, B and C.
with torch.no_grad():
    logits = model(ids).logits[0, -1]

letter_ids = [tok.encode(k, add_special_tokens=False)[0] for k in options]
probs = torch.softmax(logits[letter_ids].float(), dim=0)
print({options[k]: round(p.item(), 3) for k, p in zip(options, probs)})

This runs locally and returns a probability per team without generating a single token. To trust those numbers, label a few hundred real tickets, measure how often a 0.8 is right, and fit a temperature on the logits until it matches. That calibration step is most of what TypeSafe sells, and you do it once per task.

What you get back for the extra work is data that stays on your own servers, a model you can train on Arabic, and no dependency on one startup's capacity. TypeSafe's model page says its rate limits "can change without notice" while GPU deals land. For a regulated workload in Saudi Arabia, where personal data rules decide where a ticket may be processed, the self-hosted route may be the only one available.

Where I would use it

Use case

Questions

Why Jev fits

Watch for

Support triage

Choice for the team, Score for frustration, Noul for urgency

One call per ticket, confidence decides auto-route or human

Customers writing instructions into the ticket

Guardrails on LLM replies

Nouls such as "promises a refund" or "names a competitor"

Adds under half a second instead of a second LLM call

Literal reading, so write each rule precisely

RAG filtering and reranking

Noul or Score per retrieved passage

TypeSafe's legal cookbook raised top-10 accuracy on CLERC queries from 38% to 62% over BM25

Send the passage, not the whole corpus

Security alert triage

Choice between close, analyst and contain

61.7% in the eval at $0.0001, against 66.2% for Opus 5 at $0.057

Keep a person on "contain"

Tagging large datasets

A handful of Choices per row

A million rows at about 1,000 tokens each is a billion tokens, which is $42

Spot-check against labels first

Real-time loops

Choice over the next action

TypeSafe's Doom bot runs about 10 queries a second for roughly $7 an hour

The state has to be text, so describe the scene in JSON

Learning platforms

Score a short answer on a rubric, Noul for off-topic questions, Choice for which lesson to suggest next

Instant feedback, and low-confidence answers go to an instructor

Arabic accuracy is untested, so measure it first

Worked example: confidence-gated routing

Take the ticket above. Jev picked technical at 0.85 with a confidence of 0.78. With a threshold of 0.7 the ticket goes straight to the technical queue. With a threshold of 0.8 it goes to a person, pre-filled with "technical" and both probabilities, so the human confirms instead of starting from nothing. Moving the threshold trades automation for accuracy, and you tune it on your own labelled tickets.

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()  # reads TYPESAFE_API_KEY, defaults to jev-latest

response = client.system_one(
    state=ticket,
    questions={
        "department": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing": "Payment or subscription issues",
                "technical": "Bugs or integration problems",
                "sales": "Pricing or account questions",
            },
        ),
        "is_urgent": Noul(
            instructions="The message conveys urgency or time-sensitivity",
        ),
    },
)

department = response.answers["department"]
if department.confidence >= 0.8:
    route_to(department.choice)
else:
    send_to_human(ticket, suggestion=department.choice,
                  probabilities=department.probabilities)
Worked example: what 100,000 tickets a day costs

Assume 1,000 input tokens per ticket and, for the LLM, 150 output tokens for the JSON answer. That is 100 million input tokens a day.

  • Jev: 100M × $0.042 per million = $4.20 a day.

  • GPT-5.6 Terra at list price: 100M × $2 = $200 for input, plus 15M × $12 = $180 for output, so $380 a day before any reasoning tokens.

Reasoning tokens bill as output, so a thinking model costs more than this. The eval numbers point the same way: Terra's average decision cost $0.030 against Jev's $0.0004.

I would not use it for anything with arithmetic, dates, free text or conversation. TypeSafe's docs say directly that Jev is not a replacement for the model behind a coding agent such as Claude Code or Cursor. The fit is a system where an LLM or a person does the thinking and Jev makes the hundreds of small calls around it.

Try it

Access is through a waitlist. Once you are in, the playground lets you paste a state and add questions without code. The eval site shows every case where Jev disagreed with the reference, with the full query, which is the fastest way to see how it fails. The SDKs install with pip install typesafe-sdk and npm install @typesafe-ai/sdk, and there is a Claude Code plugin that teaches a coding agent the API.

Before it goes anywhere near production:

  • Label 200 real examples from your own traffic and measure agreement yourself

  • Pin jev-1.13.0 instead of jev-latest once your thresholds are tuned, since the alias moves on each release

  • Keep every sum, count and date comparison in code

  • Treat user-written text in the state as untrusted

  • Test Arabic content separately from English

  • Log the model field of every response so you know which version made each call

46
106 views
AILLMsJevTypeSafeStructured outputsOpen source
MA
Mosab Alrasheed
Get the next entry

One email when we publish. Research, product decisions, and what teams report back.