Field guide · data current to 28 Sep 2026

Jev and the System One model class

On 15 September 2026, TypeSafe AI released Jev — a model that generates no text at all. It takes a block of state plus typed questions and returns probabilities. This guide explains how these models work, how they are trained, why they behave the way they do, and what the fast-moving open-weight community has built in the fortnight since.

At a glance

The numbers that define the class

Cost and latency figures below are TypeSafe's own published claims, not independently replicated measurements. The failure-mode count is from TypeSafe's own documentation.

70–500 ms Claimed end-to-end latency
$0.042 Per million input tokens · output free
3 Question primitives: Choice · Score · Noul
9 Documented failure modes in Jev 1.13
The 60-second version

A decision layer, not a chatbot

A large language model answers a question by writing out an answer one token at a time. Even for a yes/no question, it generates the text and you parse it back into a boolean. The output is a string, and a string can be anything — an answer, a refusal, a paragraph of reasoning, or malformed JSON your parser chokes on.

A System One model removes generation from the loop entirely. You declare the shape of the answer up front — which fields, which allowed values — and the model fills every field in one parallel forward pass, attaching a probability to each. No tokens are generated. The output cannot violate the schema you supplied.

TypeSafe's own framing: "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."

LLMAutoregressive sampling
tok→ tok→ tok→ tok→ string

3–329 s end-to-end on frontier models. Output billed at roughly 5× input. Confidence is self-reported and usually overconfident.

System OneParallel field fill
state + questions→ Choice 0.47 / 0.28 / 0.25 Noul 0.93 Score 1.44

All questions answered in the same pass, so adding questions barely moves latency. Every answer carries a calibrated probability.

⚡ The point is not that it's faster

It's that it tells you how sure it is. A model that is right 95% of the time but cannot flag the other 5% cannot be automated around. One that reports honest uncertainty can: you branch on the probability, act on the confident cases, and escalate the rest to a human or a larger model. That single property is what turns a model call into something you can put inside a dependency chain.

The landscape

Two tracks, one interface

Jev is closed, hosted and waitlisted. Within roughly 24 hours of launch, the open-weight community reversed the interface — the request and response shape published in TypeSafe's API docs — and started competing on implementation. Reproducing the contract takes a weekend. Reproducing the accuracy is a research problem.

🔒 Jev — TypeSafe AI

Closed

Proprietary weights, no technical paper, hosted API only. Trained on 100% synthetic data with RLCD. Published claims: 40–200× faster and 40–400× cheaper than frontier LLMs on decision-shaped tasks.

Hosted only Waitlisted $0.042 / MTok in 70–500 ms

Full detail on Jev →

🔓 The open reproductions

Apache-2.0 / MIT

Two dozen independent projects of wildly varying quality. Most are LoRA adapters or option-scoring heads over Qwen3.5, Gemma or ModernBERT. Several implement TypeSafe's /v1/systemone contract, so the official SDK talks to them by changing a base URL.

Self-hosted 322M – 35B params CPU to B200 No per-call cost

Browse the catalog →

⚠️ The one number to remember

On the independent jabr classifier suite — 49 tasks, 869 cases, out-of-domain and zero-shot — hosted Jev scores 0.966 macro accuracy and the best open entrant scores 0.704. That is a 26-point gap. The open projects win on latency, cost and control; they do not yet win on out-of-the-box accuracy. The most-starred project is not the most accurate one — those are different projects.

Reading the numbers

Three kinds of evidence, kept separate

This field is eleven days old at time of writing and moves fast enough that a figure can go stale in a day. Throughout this guide, claims are tagged so you can tell what kind of thing you are reading.

Vendor

Self-reported

Speed, price and accuracy figures published by the vendor on benchmarks the vendor designed. Directionally useful, not independent. TypeSafe's own launch post includes "nuance" notes admitting several of these are best-case.

Independent

Third-party measured

JevBench, the Jev Decision Index, and the jabr suite. These run the same inputs against every system. They are one-person or small-team projects with published methodology and published limitations — read both.

Community

Builder reports

GitHub READMEs, model cards, and practitioner write-ups. Usually the most candid about failure modes, because the authors are the ones who hit them. Not comparable across projects.

How to read this guide in a hurry