Overview / The concept

01 · How it works

The concept, mechanically

Strip away the branding and a System One model is a narrow, well-defined thing: a function that maps state + declared questions to typed answers with probabilities, with no text generation anywhere in the path. Here is exactly what that means.

Naming

Where the name comes from

"System One" is a reference to Daniel Kahneman's Thinking, Fast and Slow, where System 1 is fast, intuitive judgment and System 2 is slow, deliberate reasoning. TypeSafe places these models firmly in the first bucket: fast, focused judgments rather than multi-step deliberation.

The company is aware that Kahneman's System 1 is also associated with being error-prone. Their stated position is that "System One Models can be made more reliable than its alternatives" — an argument they say they will develop later.

"Jev" is named after William Stanley Jevons, the 19th-century economist behind the Jevons paradox: as steam engines got more efficient, coal consumption rose rather than fell, because cheap power found new uses. TypeSafe's bet is that intelligence behaves the same way. Make a decision cost a fraction of a cent and you will put decisions places you would never have paid an LLM to make one.

Why the naming matters practically

The label is not decoration — it predicts the failure modes. Because these are System 1-shaped models, they are documented as weak at exactly the things that need System 2: multiple hops of reasoning, layered indirection, arithmetic, and date comparison. TypeSafe's own guidance is to keep all of that in ordinary code and reserve the model for the part that is genuinely a judgment.

Simon Willison has suggested a plainer name for the category — "decision models" — on the grounds that it describes the job rather than the metaphor.

The interface

Anatomy of a request

Two parts. State is the context — the thing you want judged. Questions is a named map of typed questions. You can send as many questions as fit the context window, and they are all evaluated against the same state.

POST /v1/systemone
{
  "model": "jev-latest",
  "state": "Shoes arrived two weeks late and in the
            wrong size. Also I see two charges
            on my card.",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns":  "Exchanges, refunds, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing":  "Charges, invoices, payment problems"
      }
    },
    "escalate": {
      "type": "noul",
      "instructions": "Does this need urgent human attention?"
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Frustrated", "Very angry"]
    }
  }
}
Response — every field, with probabilities
{
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 0.21,
      "probabilities": {
        "returns":  0.47,
        "shipping": 0.28,
        "billing":  0.25
      }
    },
    "escalate": {
      "type": "noul",
      "noul": 0.93
    },
    "frustration": {
      "type": "score",
      "score": 1.44,
      "confidence": 0.78,
      "legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
      "probabilities": { "0": 0.00, "1": 0.56, "2": 0.44 }
    }
  }
}

Example adapted from the Kev repository README, which mirrors TypeSafe's request and response shape. Note that department illustrates the design's own honesty: the top option wins at 0.47 with a confidence of 0.21 — three departments all plausibly apply, and the numbers say so.

What counts as state

String
Free text — a ticket, an email, a log line, a policy paragraph, a diff.
JSON object
Structured program state — name/value pairs. TypeSafe's framing is that LLMs emphasise sequential messages while System One models emphasise structured program state.
Array of text
A list of documents or candidates — useful for reranking, or for making one decision per item.
Multi-part
Questions can be asked of the same state independently. They share the input text but cannot read each other's answers.
⚠️ State is data, and it is not treated as hostile

TypeSafe's own documentation states plainly that state content "written to adversarially steer the model — whether that is an injected instruction, a deliberately misleading framing, or text that argues for its own classification — can move the answer." If your state comes from untrusted users and your decision is a guardrail, that is the threat model you most care about. Test it before you ship it.

The answer space

The three primitives

The model can only answer inside a space you define. Everything it is allowed to say is enumerated in the request. There are exactly three shapes.

Primitive semantics per TypeSafe's documentation. "Noul" is short for Bernoulli — the CEO confirmed this on Hacker News.
Primitive Question it answers Example answer space Returns
Choice Which one of these options? billing · technical · account Selected option, a probability for every option, plus an overall confidence
Score Where on this ordered scale? 0 = calm · 1 = frustrated · 2 = very frustrated A continuous score, the underlying distribution across levels, plus confidence
Noul Is this statement true? true / false A single float between 0 and 1

Choice — relative

A Choice settles which option, so its probabilities are relative: the distribution is forced to sum to something. This is the primitive for routing, triage, label assignment and picking among candidates.

Noul — absolute

Each Noul is an independent absolute judgment, which means every Noul in a set can come back low. If you want to know "is this any of these categories at all?", that is a set of Nouls, not a Choice.

Score — ordered

Levels are given lowest-first with descriptions, and you get back a value that can land between levels. Useful for severity, priority and quality ranking — but TypeSafe warns the score scale is weak at numerical calibration.

Cardinality: 255 options, then two stages

Jev supports up to 255 options in a single Choice question. Above that, TypeSafe uses a two-stage approach — score the candidates independently, then make one explicit Choice — which they acknowledge is slower. The open models in this space tend to be far more constrained: several encode answers as single letters or tokens, capping them at 16 options per pass.

The core mechanic

Parallel vs sequential sampling

This is the single most important architectural difference, and it is what produces the cost and latency claims.

🗣️ Autoregressive (an LLM)

The model produces a sequence. To return {"category": "billing", "urgent": true} it generates every token in order, each one conditioned on the last. A JSON schema can constrain the shape, but generation can still fail, stop early, or produce something your parser rejects — the ecosystem literally has a name for this failure class ("structured output errors").

Cost scales with output length. Output tokens are typically billed at several times the input rate — TypeSafe's comparison uses roughly 5×.

⚡ Parallel (a System One model)

Every field of every answer is produced in a single forward pass over the input. There is no sequence to get wrong, because nothing is being written. TypeSafe describes the sampler as "incredibly efficient and hardware-aware."

The practical consequence: adding questions barely changes latency. You pay only the extra input tokens. Ten questions in one request cost roughly the same wall time as one.

SequentialTen separate LLM calls, one per decision
call 1+ call 2+ call 3+ …+ call 10→ 10× latency, 10× output tokens
ParallelOne System One call, ten decisions
single pass→ 10 typed answers + 10 probability distributions

Output is free in TypeSafe's pricing model, so the marginal cost of the ninth and tenth question is a few input tokens and essentially no time.

This is why "guardrails on every call" becomes viable

Safety classification, jailbreak detection, tool-call gating and output verification all have to run on every request, which makes cost and latency the binding constraint rather than raw capability. A decision layer that costs a fraction of a cent and answers in tens of milliseconds changes what is worth doing.

Guarantees

Why it cannot produce a type error

This is a structural property, not a trained behaviour. That distinction is the reason TypeSafe can make the claim at all.

Because the answer space is enumerated in the request and the model emits values from that space rather than text, there is nothing to parse and nothing to repair. TypeSafe states that schema matching is guaranteed, so they plot a flat 0% type-error rate — and note that the claim is not empirical, because a single counter-example would falsify it and they assert one is "mathematically impossible."

Their framing is that this matters more than it sounds. A hallucinated tool call is merely inconvenient inside an agent; it is a deal-breaker inside a system with latency guarantees, or buried several layers deep in a dependency chain where nobody is watching the output.

❌ But this is not the same as being correct

"Cannot hallucinate" and "cannot be wrong" are different claims, and summaries routinely conflate them. A model constrained to three allowed categories can still confidently choose the wrong one. What has been eliminated is the malformed answer, not the mistaken judgment. Constraining the shape of an answer has never guaranteed its substance — the same is true of JSON-schema-constrained LLM output. Calibration is the proposed answer to mistaken judgment, and calibration is precisely the claim that lacks independent testing.

Reading the output

What the probability actually means

LLMs are trained to satisfy human raters, and are documented as overconfident and inconsistent when asked for a confidence estimate. A model that can do a task 95% of the time but does not say which 5% it will fail cannot be automated around.

System One models optimise the probabilities themselves — that is the entire point of RLCD, the training method covered on the Training page. TypeSafe's stated properties are that higher confidence means higher accuracy, and that similar inputs get similar answers.

The workflow this enables is mundane but transformative: pick a threshold, act above it, and route everything below it to a human or a larger model.

⚠️ Read this sentence twice

From TypeSafe's own documentation: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." A well-calibrated model that says 0.9 is right about 90% of the time across many such calls. It tells you nothing about the particular call in front of you. Every system that means "90% confident" is also wrong one time in ten.

Armin Ronacher, on where the burden lands

"At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it."

Side by side

System One vs LLM, axis by axis

Adapted from the comparison table in TypeSafe's launch post, with the vendor's own caveats retained. Latency and price figures in the LLM column are the vendors' benchmark sources, not a controlled comparison.

Axis LLM System One / Jev
Optimised with RLHF (human preference) and/or RLVR (programmatically verifiable rewards) RLCD — Reinforcement Learning for Calibrated Decisions, optimising probabilities against outcomes
Optimises for Writeups and chat responses that human raters prefer; outputs that can be checked programmatically Calibrated decisions — "epistemically honest probabilities" on decision-shaped tasks
Input emphasis Unstructured data, weighted toward sequential messages Unstructured data, weighted toward structured program state
Output Strings. Flexible and can be anything: answers, code, hallucinations, refusals — or type-safe values, if you parse and validate them. Always some risk of going off the rails. Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. Every answer carries calibrated probabilities and confidence.
Sampling Sequential. One token at a time, each conditioned on the last. Parallel. All outputs in a single query — described as "incredibly efficient and hardware-aware."
Price $0.20–$10 per MTok input; output roughly 5× the input rate $0.042 per MTok input ($42 per billion tokens); output free — "too cheap to meter"
Latency 3–329 s end-to-end for frontier models, per the third-party benchmark TypeSafe cites 70–500 ms, TypeSafe-reported. Claimed 40–200× faster for decision-shaped queries.
Confidence Overconfident and inconsistent even when prompted for an estimate Calibrated and always present; claimed to be more consistent across similar inputs
Failure mode Can hallucinate, refuse, or emit an unparsable shape Cannot leave the schema. Can still pick the wrong allowed value.
Best fit Drafting, summarising, explanation, tool-using agents, reasoning, human-in-the-loop work, prototypes Classify, route, score, extract, verify, branch — decisions inside running code
Boundaries

Structural limits of the design

These are not bugs to be patched — they follow from what the model is. The per-model weaknesses (arithmetic, dates, indirection) are catalogued on the Jev page and in the open models' own model cards.

🚫 No text generation, ever

It cannot draft an email, write code, explain its reasoning, or summarise. It is not a chat model and cannot be made into one — chaining choices to spell out text is documented as working badly and being very slow. If the answer is prose, this is the wrong tool.

🎚️ Answers only where you allow

Freedom is traded for dependability. If the correct answer is not in your option list, the model cannot discover it — it will pick the nearest allowed value. "None of the above" has to be an option you thought to include.

🕳️ Opaque by design

You get a number, not an explanation. For a spam flag, which signals tipped it off? There is no reasoning trace to read, and no way to ask. This is a step backwards from LLMs on interpretability, and it makes bias auditing harder, not easier.