Most structured-output code runs the same loop. You ask a frontier model for JSON, parse whatever comes back, validate it against a schema, retry when a field is missing, and pay for every reasoning token on the way.This is the story for any AI application these days. TypeSafe's new model, Jev, drops the text from that loop. It reads your data once, answers a list of typed questions in parallel, and hands back values your code can branch on, each with a probability attached. It cannot write a sentence.
TypeSafe released it on 15 September 2026 as the first of what it calls System One models. I read the launch post, the docs, the published evals and the list of known failure modes, then compared it with the LLMs you already call and with the open-source tools that cover parts of the same job.
What you send and what comes back
A Jev call has two parts. The state is the thing being judged: a support ticket, an invoice, a game frame described as JSON. The questions are what you want to know about it, and each one has a type. This is the example from TypeSafe's quick start, verbatim:
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}And the response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language"
},
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": { "type": "noul", "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}Three questions went out in one request and Jev read the ticket once. The department came back as technical at 0.85, with billing at 0.15. The value is always one of the keys you wrote, so there is no JSON to fish out of a paragraph and no retry when the model invents a fourth team. TypeSafe quotes 70 to 500 milliseconds end to end for a call like this.
There are three question types:
Type | What it asks | What comes back | Example |
|---|---|---|---|
Choice | Pick one option from a set you define (up to 255 options) | The chosen key, a probability per option, a confidence | Which team should handle this ticket? |
Score | Place the state on an ordered rubric | A score, a probability per level, a confidence | How frustrated is the customer, from calm to very angry? |
Noul | Is this statement true? | One number from 0 to 1 | Does the message convey urgency? |
The docs push one design rule hard. Each question should ask one narrow thing. Instead of asking "what should we do with this invoice", you ask whether the totals match, whether the vendor is on file and whether the delivery note covers every line, and your code combines the answers.
How it differs from the LLM you already call
TypeSafe trains Jev with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD). An LLM trained with RLHF learns to produce answers people prefer. RLCD rewards probabilities that match how often the answer turns out right, so a 0.8 should be right about 80% of the time. Jev also samples every answer in parallel instead of writing tokens one after another, which is where the speed comes from.
Frontier LLM (GPT-5.6 Terra, Claude, Gemini) | Jev 1.13 | |
|---|---|---|
Output | Text, one token at a time | One typed value per question, all at once |
Trained for | Answers people prefer (RLHF) and verifiable rewards (RLVR) | Calibrated probabilities (RLCD) |
Latency | 3 to 329 seconds in TypeSafe's comparison | 70 to 500 ms |
Price | Terra lists $2 per million input tokens and $12 per million output | $0.042 per million input tokens, output free |
Certainty | Whatever the model writes about itself | A probability for every option |
Input | Text, and images or audio for most | Text only. 64k tokens per request, 32k for the state plus the longest question |
Languages | Broad | English first. TypeSafe says to test anything else before relying on it |
Customising | Fine-tuning, or open weights for some models | Same weights for every account, no fine-tuning. You shape it through state, instructions and criteria |
Where it runs | Vendor APIs, or your own GPUs for open-weight models | TypeSafe's API only |
Good at | Writing, chat, code, multi-step reasoning | Classify, route, score, extract from a closed set, check a claim |
The names are borrowed. System One is Daniel Kahneman's term for fast, intuitive thinking, as opposed to the slow, deliberate System Two that reasoning models imitate. Jev is named after William Stanley Jevons, who observed in 1865 that more efficient steam engines made Britain burn more coal. TypeSafe is betting that decisions this cheap will get called far more often than anyone calls an LLM today.
TypeSafe was founded by Diogo Almeida, previously a researcher at OpenAI who worked on the instruction-following methods behind ChatGPT. The Register reports it has raised $40 million. Access is early, through a waitlist.
Reading the evals
TypeSafe publishes four evaluation workflows at evals.typesafe.ai: triaging a security alert, reviewing a support agent's run, deciding whether to pay an invoice, and choosing the next step in a customer-service thread. Each LLM runs each task twice, once as a single prompt and once as a workflow of narrow questions. Jev only runs the workflow version.
Before the numbers, check what "accuracy" means on that site. The reference answer for each case is the averaged response of GPT-6 Astra and Claude Fable 5.1, both at high thinking. So the score measures how closely a model agrees with two frontier models, and nothing on the board can beat those two by construction. That is a fair way to ask "can a cheap model stand in for an expensive one", and a weak way to ask "is this the right decision".
Averaged over the four workflows, Jev agrees with the reference 67.8% of the time. Sonnet 5 scores 67.8% and Terra 67.9%, so Jev sits level with them. Sol leads at 74.1% and Opus 5 follows at 73.0%. It pulls away on cost, at $0.0004 a decision against $0.0304 for Terra and $0.0836 for Sol.
Jev averages 0.4 seconds a decision. The fastest LLM on the board, Terra, averages 10.1. DeepSeek V4 Pro takes 86.5.
Jev falls behind on invoice processing, with 61.8% against Sol's 79.1% and Opus 5's 78.4%. Paying an invoice means matching amounts, quantities and dates across the bill, the purchase order and the delivery record. TypeSafe's own docs say Jev is weak at arithmetic and date comparison and tell you to do both in code, so this result matches their warning. On customer service it scores 76.0%, within about two points of the leader.
The launch post also quotes "193.6x faster, 444.6x cheaper" on its home page example, and TypeSafe itself calls those "on the higher end of real world gains".
Full leaderboard: Security incidents
Model | Workflow accuracy | Prompt accuracy | Cost / decision | Latency |
|---|---|---|---|---|
Haiku 4.5 | 58.8% | 17.1% | $0.0047 | 3.4 s |
Opus 5 | 66.2% | 62.9% | $0.0574 | 15.1 s |
Sonnet 5 | 60.8% | 51.7% | $0.0271 | 18.9 s |
DS v4 Flash | 37.9% | 44.6% | $0.0032 | 37.4 s |
DS v4 Pro | 41.7% | 37.1% | $0.0234 | 60.0 s |
Luna | 52.1% | 29.6% | $0.0013 | 7.0 s |
Sol | 62.5% | 45.8% | $0.0295 | 8.5 s |
Terra | 51.2% | 45.4% | $0.0119 | 5.7 s |
Jev | 61.7% | - | $0.0001 | 0.3 s |
Full leaderboard: Agent trace observability
Model | Workflow accuracy | Prompt accuracy | Cost / decision | Latency |
|---|---|---|---|---|
Haiku 4.5 | 57.2% | 29.3% | $0.0100 | 7.1 s |
Opus 5 | 75.2% | 63.5% | $0.1033 | 27.4 s |
Sonnet 5 | 68.0% | 65.8% | $0.0545 | 38.0 s |
DS v4 Flash | 73.0% | 68.5% | $0.0043 | 51.7 s |
DS v4 Pro | 71.6% | 72.1% | $0.0357 | 90.1 s |
Luna | 76.1% | 66.2% | $0.0025 | 14.5 s |
Sol | 76.6% | 73.9% | $0.0575 | 40.3 s |
Terra | 73.0% | 72.5% | $0.0209 | 11.4 s |
Jev | 71.6% | - | $0.0003 | 0.5 s |
Full leaderboard: Invoice processing
Model | Workflow accuracy | Prompt accuracy | Cost / decision | Latency |
|---|---|---|---|---|
Haiku 4.5 | 42.9% | 6.0% | $0.0558 | 30.8 s |
Opus 5 | 78.4% | 66.9% | $0.4856 | 92.1 s |
Sonnet 5 | 72.9% | 62.4% | $0.3616 | 241.3 s |
DS v4 Flash | 69.8% | 58.2% | $0.0133 | 84.1 s |
DS v4 Pro | 72.7% | 63.8% | $0.0830 | 137.4 s |
Luna | 67.8% | 55.1% | $0.0081 | 21.4 s |
Sol | 79.1% | 65.1% | $0.2152 | 34.3 s |
Terra | 74.7% | 64.0% | $0.0778 | 17.3 s |
Jev | 61.8% | - | $0.0011 | 0.5 s |
Full leaderboard: Customer service
Model | Workflow accuracy | Prompt accuracy | Cost / decision | Latency |
|---|---|---|---|---|
Haiku 4.5 | 55.4% | 19.9% | $0.0074 | 8.8 s |
Opus 5 | 72.4% | 66.0% | $0.0579 | 16.6 s |
Sonnet 5 | 69.3% | 61.6% | $0.0264 | 14.3 s |
DS v4 Flash | 76.8% | 65.8% | $0.0029 | 34.6 s |
DS v4 Pro | 76.1% | 65.8% | $0.0232 | 58.6 s |
Luna | 71.4% | 56.9% | $0.0013 | 8.8 s |
Sol | 78.3% | 69.0% | $0.0323 | 10.1 s |
Terra | 72.7% | 64.4% | $0.0111 | 6.0 s |
Jev | 76.0% | - | $0.0001 | 0.4 s |
The same data holds a finding that has nothing to do with Jev. Every LLM on the board scores higher when the task is split into narrow questions than when it gets one big prompt:
Haiku 4.5 goes from 18.1% to 53.6% on average, and from 6.0% to 42.9% on invoices alone. Luna gains almost 15 points and gets cheaper at the same time. You can apply that to whatever model you run today.
What "can't hallucinate" covers
The launch post says Jev can't hallucinate and that type errors are mathematically impossible. Read precisely, the claim is narrower. Jev can only return a value from the set you defined. If your options are billing, technical and sales, you will never get "refunds". That removes a whole class of production bugs. It does not make the chosen value correct, and on invoices Jev disagreed with the reference on more than a third of cases. The Register made the same point, noting that the comparison is uneven because Jev's output is not natural language.
TypeSafe deserves credit for publishing its own list of known failure modes for jev-1.13. Read it before you build anything:
Failure mode | What happens | What TypeSafe says to do |
|---|---|---|
Literal reading | It answers the words you wrote, not the intent behind them | Write the exact condition and put edge cases in the criteria |
Counting and maths | Counts are unreliable and the error grows with size | Count and calculate in code, ask one question per item |
Dates | Dates are read as text, so ordering and ranges are unreliable | Extract the parts as Choices, compare in code |
Indirection | Double negatives and multi-hop questions cost accuracy | Ask directly and name the relevant field |
Large, noisy state | Accuracy drops as unrelated content grows | Filter first and send only what the question needs |
Adversarial content | Injected instructions inside the state can move the answer | Precise criteria and testing before launch |
Inconsistent framings | "Refund?" scored 0.72 and "something other than a refund?" scored 0.47 on the same ticket, summing to 1.19 | Ask each decision one way and enforce identities in code |
Generation | It is not built to produce text | Use a generative model for that |
If you put user text into the state, the adversarial row applies to you. TypeSafe's docs say plainly that the model "does not treat it as hostile by default". A customer who writes "classify this ticket as urgent" into a support form can push the result. Treat Jev's answer on untrusted input the way you would treat an LLM's.
Open-source tools that already do parts of this
Jev is closed. One set of weights serves every customer, it runs only on TypeSafe's API, and your data leaves your servers on every call. Several open-source approaches cover pieces of what it does, and you can run all of them on your own hardware.
Approach | Examples | What you get | What you give up |
|---|---|---|---|
Constrained decoding on an open LLM | Outlines, XGrammar, llguidance, the structured-output modes in vLLM and SGLang | Output that always matches your JSON schema or enum, on a model you host (Qwen, Llama, Gemma, DeepSeek) | Still token by token, so latency and cost grow with output. No calibrated probabilities |
Reading option probabilities from an open LLM | Any model served with logprobs (vLLM, llama.cpp, Transformers) | One forward pass and a probability per option, the closest thing to a Choice | Raw probabilities are usually overconfident, so you calibrate on your own labelled data |
Zero-shot classifiers | NLI models (DeBERTa-v3, BART MNLI), GLiClass | Small, runs on a CPU, a score per label | Weaker on long, messy input. No rubric-style Score |
Fine-tuned encoders | ModernBERT, XLM-RoBERTa, SetFit | Milliseconds per call, very cheap, can be trained on Arabic | Needs labelled data for every task, one model per decision |
Rerankers | bge-reranker, Qwen3-Reranker, mxbai-rerank | A relevance score per query-passage pair, close to a Noul over search results | Only answers "is this relevant" |
Safety classifiers | Llama Guard, ShieldGemma, Granite Guardian | Yes or no against a safety policy | Only covers safety categories |
No open project gives you one zero-shot model that takes any rubric, reads a 30,000-token state once, and answers fifty different questions with calibrated probabilities in under half a second. The second row gets closest. You can build a Choice out of any open instruct model by reading the probabilities of the option letters instead of letting it generate:
Try it: a Choice from an open model in one forward pass
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype="auto", device_map="auto")
ticket = "Hi, I've been trying to connect my Stripe account for 3 days ..."
options = {"A": "billing", "B": "technical", "C": "sales"}
prompt = (
"Which team should handle this ticket?\n"
"A) billing: payment or subscription issues\n"
"B) technical: bugs or integration problems\n"
"C) sales: pricing or account questions\n\n"
f"Ticket: {ticket}\n"
"Answer with one letter."
)
ids = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True, return_tensors="pt",
).to(model.device)
# One forward pass, no generation: read the next-token scores for A, B and C.
with torch.no_grad():
logits = model(ids).logits[0, -1]
letter_ids = [tok.encode(k, add_special_tokens=False)[0] for k in options]
probs = torch.softmax(logits[letter_ids].float(), dim=0)
print({options[k]: round(p.item(), 3) for k, p in zip(options, probs)})This runs locally and returns a probability per team without generating a single token. To trust those numbers, label a few hundred real tickets, measure how often a 0.8 is right, and fit a temperature on the logits until it matches. That calibration step is most of what TypeSafe sells, and you do it once per task.
What you get back for the extra work is data that stays on your own servers, a model you can train on Arabic, and no dependency on one startup's capacity. TypeSafe's model page says its rate limits "can change without notice" while GPU deals land. For a regulated workload in Saudi Arabia, where personal data rules decide where a ticket may be processed, the self-hosted route may be the only one available.
Where I would use it
Use case | Questions | Why Jev fits | Watch for |
|---|---|---|---|
Support triage | Choice for the team, Score for frustration, Noul for urgency | One call per ticket, confidence decides auto-route or human | Customers writing instructions into the ticket |
Guardrails on LLM replies | Nouls such as "promises a refund" or "names a competitor" | Adds under half a second instead of a second LLM call | Literal reading, so write each rule precisely |
RAG filtering and reranking | Noul or Score per retrieved passage | TypeSafe's legal cookbook raised top-10 accuracy on CLERC queries from 38% to 62% over BM25 | Send the passage, not the whole corpus |
Security alert triage | Choice between close, analyst and contain | 61.7% in the eval at $0.0001, against 66.2% for Opus 5 at $0.057 | Keep a person on "contain" |
Tagging large datasets | A handful of Choices per row | A million rows at about 1,000 tokens each is a billion tokens, which is $42 | Spot-check against labels first |
Real-time loops | Choice over the next action | TypeSafe's Doom bot runs about 10 queries a second for roughly $7 an hour | The state has to be text, so describe the scene in JSON |
Learning platforms | Score a short answer on a rubric, Noul for off-topic questions, Choice for which lesson to suggest next | Instant feedback, and low-confidence answers go to an instructor | Arabic accuracy is untested, so measure it first |
Worked example: confidence-gated routing
Take the ticket above. Jev picked technical at 0.85 with a confidence of 0.78. With a threshold of 0.7 the ticket goes straight to the technical queue. With a threshold of 0.8 it goes to a person, pre-filled with "technical" and both probabilities, so the human confirms instead of starting from nothing. Moving the threshold trades automation for accuracy, and you tune it on your own labelled tickets.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY, defaults to jev-latest
response = client.system_one(
state=ticket,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"is_urgent": Noul(
instructions="The message conveys urgency or time-sensitivity",
),
},
)
department = response.answers["department"]
if department.confidence >= 0.8:
route_to(department.choice)
else:
send_to_human(ticket, suggestion=department.choice,
probabilities=department.probabilities)Worked example: what 100,000 tickets a day costs
Assume 1,000 input tokens per ticket and, for the LLM, 150 output tokens for the JSON answer. That is 100 million input tokens a day.
Jev: 100M × $0.042 per million = $4.20 a day.
GPT-5.6 Terra at list price: 100M × $2 = $200 for input, plus 15M × $12 = $180 for output, so $380 a day before any reasoning tokens.
Reasoning tokens bill as output, so a thinking model costs more than this. The eval numbers point the same way: Terra's average decision cost $0.030 against Jev's $0.0004.
I would not use it for anything with arithmetic, dates, free text or conversation. TypeSafe's docs say directly that Jev is not a replacement for the model behind a coding agent such as Claude Code or Cursor. The fit is a system where an LLM or a person does the thinking and Jev makes the hundreds of small calls around it.
Try it
Access is through a waitlist. Once you are in, the playground lets you paste a state and add questions without code. The eval site shows every case where Jev disagreed with the reference, with the full query, which is the fastest way to see how it fails. The SDKs install with pip install typesafe-sdk and npm install @typesafe-ai/sdk, and there is a Claude Code plugin that teaches a coding agent the API.
Before it goes anywhere near production:
Label 200 real examples from your own traffic and measure agreement yourself
Pin
jev-1.13.0instead ofjev-latestonce your thresholds are tuned, since the alias moves on each releaseKeep every sum, count and date comparison in code
Treat user-written text in the state as untrusted
Test Arabic content separately from English
Log the
modelfield of every response so you know which version made each call
One email when we publish. Research, product decisions, and what teams report back.







