Jev and the System One model class
On 15 September 2026, TypeSafe AI released Jev — a model that generates no text at all. It takes a block of state plus typed questions and returns probabilities. This guide explains how these models work, how they are trained, why they behave the way they do, and what the fast-moving open-weight community has built in the fortnight since.
The numbers that define the class
Cost and latency figures below are TypeSafe's own published claims, not independently replicated measurements. The failure-mode count is from TypeSafe's own documentation.
A decision layer, not a chatbot
A large language model answers a question by writing out an answer one token at a time. Even for a yes/no question, it generates the text and you parse it back into a boolean. The output is a string, and a string can be anything — an answer, a refusal, a paragraph of reasoning, or malformed JSON your parser chokes on.
A System One model removes generation from the loop entirely. You declare the shape of the answer up front — which fields, which allowed values — and the model fills every field in one parallel forward pass, attaching a probability to each. No tokens are generated. The output cannot violate the schema you supplied.
TypeSafe's own framing: "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
3–329 s end-to-end on frontier models. Output billed at roughly 5× input. Confidence is self-reported and usually overconfident.
All questions answered in the same pass, so adding questions barely moves latency. Every answer carries a calibrated probability.
It's that it tells you how sure it is. A model that is right 95% of the time but cannot flag the other 5% cannot be automated around. One that reports honest uncertainty can: you branch on the probability, act on the confident cases, and escalate the rest to a human or a larger model. That single property is what turns a model call into something you can put inside a dependency chain.
Where to go next
Seven pages, arranged so you can read straight through or jump to the part you need.
The concept, mechanically
Request anatomy, the three primitives, parallel vs sequential sampling, and why type-safety is guaranteed by construction rather than by training.
Read the concept → 02How they're trained
RLHF → RLVR → RLCD. Calibration as a training objective, 100% synthetic data, and what "can't hallucinate" does and does not mean.
Read about training → 03Jev in detail
Founders, funding, what TypeSafe published, what they withheld, the nine documented jagged edges, and the black-box bias concern.
Read the Jev deep dive → 04The open-weight field
Twenty-plus Jev-shaped models, sorted into four architectural families — encoders with option heads, LoRA adapters, frozen-logit readouts, and diffusion.
Browse the catalog → 05Benchmarks & leaderboards
JevBench, the Jev Decision Index, and the jabr classifier suite — what each measures, how they differ, and the caveats their own authors publish.
Compare the benchmarks → 06Putting it in a stack
Routing, guardrails, reranking, verification. Working code for the hosted API, the Python SDK, LangChain middleware, and a local server.
See the patterns → 07Sources
Every claim in this guide, tagged by kind: vendor material, independent measurement, or community report. Plus the caveats we could not resolve.
Open the source list →Two tracks, one interface
Jev is closed, hosted and waitlisted. Within roughly 24 hours of launch, the open-weight community reversed the interface — the request and response shape published in TypeSafe's API docs — and started competing on implementation. Reproducing the contract takes a weekend. Reproducing the accuracy is a research problem.
🔒 Jev — TypeSafe AI
ClosedProprietary weights, no technical paper, hosted API only. Trained on 100% synthetic data with RLCD. Published claims: 40–200× faster and 40–400× cheaper than frontier LLMs on decision-shaped tasks.
🔓 The open reproductions
Apache-2.0 / MITTwo dozen independent projects of wildly varying quality. Most are LoRA adapters or option-scoring heads over Qwen3.5, Gemma or ModernBERT. Several implement TypeSafe's /v1/systemone contract, so the official SDK talks to them by changing a base URL.
On the independent jabr classifier suite — 49 tasks, 869 cases, out-of-domain and zero-shot — hosted Jev scores 0.966 macro accuracy and the best open entrant scores 0.704. That is a 26-point gap. The open projects win on latency, cost and control; they do not yet win on out-of-the-box accuracy. The most-starred project is not the most accurate one — those are different projects.
Three kinds of evidence, kept separate
This field is eleven days old at time of writing and moves fast enough that a figure can go stale in a day. Throughout this guide, claims are tagged so you can tell what kind of thing you are reading.
Self-reported
Speed, price and accuracy figures published by the vendor on benchmarks the vendor designed. Directionally useful, not independent. TypeSafe's own launch post includes "nuance" notes admitting several of these are best-case.
Third-party measured
JevBench, the Jev Decision Index, and the jabr suite. These run the same inputs against every system. They are one-person or small-team projects with published methodology and published limitations — read both.
Builder reports
GitHub READMEs, model cards, and practitioner write-ups. Usually the most candid about failure modes, because the authors are the ones who hit them. Not comparable across projects.
- Five minutes: this page and The concept.
- Fifteen minutes: add Training and the jagged-edges table in Jev in detail.
- Evaluating for real: Benchmarks first, then Open models, then the anti-patterns at the end of Practice.