Overview / The concept
01 · How it worksThe concept, mechanically
Strip away the branding and a System One model is a narrow, well-defined thing: a function that maps state + declared questions to typed answers with probabilities, with no text generation anywhere in the path. Here is exactly what that means.
Where the name comes from
"System One" is a reference to Daniel Kahneman's Thinking, Fast and Slow, where System 1 is fast, intuitive judgment and System 2 is slow, deliberate reasoning. TypeSafe places these models firmly in the first bucket: fast, focused judgments rather than multi-step deliberation.
The company is aware that Kahneman's System 1 is also associated with being error-prone. Their stated position is that "System One Models can be made more reliable than its alternatives" — an argument they say they will develop later.
"Jev" is named after William Stanley Jevons, the 19th-century economist behind the Jevons paradox: as steam engines got more efficient, coal consumption rose rather than fell, because cheap power found new uses. TypeSafe's bet is that intelligence behaves the same way. Make a decision cost a fraction of a cent and you will put decisions places you would never have paid an LLM to make one.
Why the naming matters practically
The label is not decoration — it predicts the failure modes. Because these are System 1-shaped models, they are documented as weak at exactly the things that need System 2: multiple hops of reasoning, layered indirection, arithmetic, and date comparison. TypeSafe's own guidance is to keep all of that in ordinary code and reserve the model for the part that is genuinely a judgment.
Simon Willison has suggested a plainer name for the category — "decision models" — on the grounds that it describes the job rather than the metaphor.
Anatomy of a request
Two parts. State is the context — the thing you want judged. Questions is a named map of typed questions. You can send as many questions as fit the context window, and they are all evaluated against the same state.
{
"model": "jev-latest",
"state": "Shoes arrived two weeks late and in the
wrong size. Also I see two charges
on my card.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, refunds, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems"
}
},
"escalate": {
"type": "noul",
"instructions": "Does this need urgent human attention?"
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]
}
}
}
{
"answers": {
"department": {
"type": "choice",
"choice": "returns",
"confidence": 0.21,
"probabilities": {
"returns": 0.47,
"shipping": 0.28,
"billing": 0.25
}
},
"escalate": {
"type": "noul",
"noul": 0.93
},
"frustration": {
"type": "score",
"score": 1.44,
"confidence": 0.78,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.00, "1": 0.56, "2": 0.44 }
}
}
}
Example adapted from the Kev repository README, which mirrors TypeSafe's request and
response shape. Note that department illustrates the design's own honesty:
the top option wins at 0.47 with a confidence of 0.21 — three departments all plausibly
apply, and the numbers say so.
What counts as state
TypeSafe's own documentation states plainly that state content "written to adversarially steer the model — whether that is an injected instruction, a deliberately misleading framing, or text that argues for its own classification — can move the answer." If your state comes from untrusted users and your decision is a guardrail, that is the threat model you most care about. Test it before you ship it.
The three primitives
The model can only answer inside a space you define. Everything it is allowed to say is enumerated in the request. There are exactly three shapes.
| Primitive | Question it answers | Example answer space | Returns |
|---|---|---|---|
| Choice | Which one of these options? | billing · technical · account | Selected option, a probability for every option, plus an overall confidence |
| Score | Where on this ordered scale? | 0 = calm · 1 = frustrated · 2 = very frustrated | A continuous score, the underlying distribution across levels, plus confidence |
| Noul | Is this statement true? | true / false | A single float between 0 and 1 |
Choice — relative
A Choice settles which option, so its probabilities are relative: the distribution is forced to sum to something. This is the primitive for routing, triage, label assignment and picking among candidates.
Noul — absolute
Each Noul is an independent absolute judgment, which means every Noul in a set can come back low. If you want to know "is this any of these categories at all?", that is a set of Nouls, not a Choice.
Score — ordered
Levels are given lowest-first with descriptions, and you get back a value that can land between levels. Useful for severity, priority and quality ranking — but TypeSafe warns the score scale is weak at numerical calibration.
Jev supports up to 255 options in a single Choice question. Above that, TypeSafe uses a two-stage approach — score the candidates independently, then make one explicit Choice — which they acknowledge is slower. The open models in this space tend to be far more constrained: several encode answers as single letters or tokens, capping them at 16 options per pass.
Parallel vs sequential sampling
This is the single most important architectural difference, and it is what produces the cost and latency claims.
🗣️ Autoregressive (an LLM)
The model produces a sequence. To return
{"category": "billing", "urgent": true} it generates every token in order,
each one conditioned on the last. A JSON schema can constrain the shape, but generation
can still fail, stop early, or produce something your parser rejects — the ecosystem
literally has a name for this failure class ("structured output errors").
Cost scales with output length. Output tokens are typically billed at several times the input rate — TypeSafe's comparison uses roughly 5×.
⚡ Parallel (a System One model)
Every field of every answer is produced in a single forward pass over the input. There is no sequence to get wrong, because nothing is being written. TypeSafe describes the sampler as "incredibly efficient and hardware-aware."
The practical consequence: adding questions barely changes latency. You pay only the extra input tokens. Ten questions in one request cost roughly the same wall time as one.
Output is free in TypeSafe's pricing model, so the marginal cost of the ninth and tenth question is a few input tokens and essentially no time.
Safety classification, jailbreak detection, tool-call gating and output verification all have to run on every request, which makes cost and latency the binding constraint rather than raw capability. A decision layer that costs a fraction of a cent and answers in tens of milliseconds changes what is worth doing.
Why it cannot produce a type error
This is a structural property, not a trained behaviour. That distinction is the reason TypeSafe can make the claim at all.
Because the answer space is enumerated in the request and the model emits values from that space rather than text, there is nothing to parse and nothing to repair. TypeSafe states that schema matching is guaranteed, so they plot a flat 0% type-error rate — and note that the claim is not empirical, because a single counter-example would falsify it and they assert one is "mathematically impossible."
Their framing is that this matters more than it sounds. A hallucinated tool call is merely inconvenient inside an agent; it is a deal-breaker inside a system with latency guarantees, or buried several layers deep in a dependency chain where nobody is watching the output.
"Cannot hallucinate" and "cannot be wrong" are different claims, and summaries routinely conflate them. A model constrained to three allowed categories can still confidently choose the wrong one. What has been eliminated is the malformed answer, not the mistaken judgment. Constraining the shape of an answer has never guaranteed its substance — the same is true of JSON-schema-constrained LLM output. Calibration is the proposed answer to mistaken judgment, and calibration is precisely the claim that lacks independent testing.
What the probability actually means
LLMs are trained to satisfy human raters, and are documented as overconfident and inconsistent when asked for a confidence estimate. A model that can do a task 95% of the time but does not say which 5% it will fail cannot be automated around.
System One models optimise the probabilities themselves — that is the entire point of RLCD, the training method covered on the Training page. TypeSafe's stated properties are that higher confidence means higher accuracy, and that similar inputs get similar answers.
The workflow this enables is mundane but transformative: pick a threshold, act above it, and route everything below it to a human or a larger model.
From TypeSafe's own documentation: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." A well-calibrated model that says 0.9 is right about 90% of the time across many such calls. It tells you nothing about the particular call in front of you. Every system that means "90% confident" is also wrong one time in ten.
"At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it."
System One vs LLM, axis by axis
Adapted from the comparison table in TypeSafe's launch post, with the vendor's own caveats retained. Latency and price figures in the LLM column are the vendors' benchmark sources, not a controlled comparison.
| Axis | LLM | System One / Jev |
|---|---|---|
| Optimised with | RLHF (human preference) and/or RLVR (programmatically verifiable rewards) | RLCD — Reinforcement Learning for Calibrated Decisions, optimising probabilities against outcomes |
| Optimises for | Writeups and chat responses that human raters prefer; outputs that can be checked programmatically | Calibrated decisions — "epistemically honest probabilities" on decision-shaped tasks |
| Input emphasis | Unstructured data, weighted toward sequential messages | Unstructured data, weighted toward structured program state |
| Output | Strings. Flexible and can be anything: answers, code, hallucinations, refusals — or type-safe values, if you parse and validate them. Always some risk of going off the rails. | Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. Every answer carries calibrated probabilities and confidence. |
| Sampling | Sequential. One token at a time, each conditioned on the last. | Parallel. All outputs in a single query — described as "incredibly efficient and hardware-aware." |
| Price | $0.20–$10 per MTok input; output roughly 5× the input rate | $0.042 per MTok input ($42 per billion tokens); output free — "too cheap to meter" |
| Latency | 3–329 s end-to-end for frontier models, per the third-party benchmark TypeSafe cites | 70–500 ms, TypeSafe-reported. Claimed 40–200× faster for decision-shaped queries. |
| Confidence | Overconfident and inconsistent even when prompted for an estimate | Calibrated and always present; claimed to be more consistent across similar inputs |
| Failure mode | Can hallucinate, refuse, or emit an unparsable shape | Cannot leave the schema. Can still pick the wrong allowed value. |
| Best fit | Drafting, summarising, explanation, tool-using agents, reasoning, human-in-the-loop work, prototypes | Classify, route, score, extract, verify, branch — decisions inside running code |
Structural limits of the design
These are not bugs to be patched — they follow from what the model is. The per-model weaknesses (arithmetic, dates, indirection) are catalogued on the Jev page and in the open models' own model cards.
🚫 No text generation, ever
It cannot draft an email, write code, explain its reasoning, or summarise. It is not a chat model and cannot be made into one — chaining choices to spell out text is documented as working badly and being very slow. If the answer is prose, this is the wrong tool.
🎚️ Answers only where you allow
Freedom is traded for dependability. If the correct answer is not in your option list, the model cannot discover it — it will pick the nearest allowed value. "None of the above" has to be an option you thought to include.
🕳️ Opaque by design
You get a number, not an explanation. For a spam flag, which signals tipped it off? There is no reasoning trace to read, and no way to ask. This is a step backwards from LLMs on interpretability, and it makes bias auditing harder, not easier.