Overview / Training

02 · How they're trained

How these models are trained

The architecture of a decision model is mostly unremarkable — it is the objective that changed. Instead of being rewarded for writing text a human likes, the model is rewarded for stating probabilities that match reality. That single swap is what makes everything else on this site possible.

Objectives

Three training objectives

Alignment technique is the axis on which System One models diverge from LLMs. TypeSafe's own comparison puts it as a three-way table.

Objective Reward signal What it produces Known weakness
RLHF
Reinforcement Learning from Human Feedback
Human raters preferring one response over another Chat responses people like. The technique behind InstructGPT and ChatGPT. Optimises for being liked. Produces models that are fluent, agreeable and overconfident about their own certainty.
RLVR
Reinforcement Learning with Verifiable Rewards
Automatic pass/fail — a proof checks, a test passes, a kernel runs faster Genuine capability on problems whose correctness is cheap to check. Only applies where you have a checker. Says nothing about calibration on fuzzy judgment.
RLCD
Reinforcement Learning for Calibrated Decisions
Whether the stated probability matched the outcome Answers with "epistemically honest probabilities" on decision-shaped tasks. Only defined where there is a ground truth or a trustworthy reference distribution to measure against.
The founder's framing is unusually direct

Diogo Almeida helped invent RLHF at OpenAI, then concluded it was pointed at the wrong target. "We have lightning in a bottle, and yet it is not useful… It took me a while to come to the conclusion: the problem is we are optimizing for human language." He also told Forbes: "We've been optimizing for humans and we're super human at pleasing humans." RLCD is his answer — reward the model for being right with confidence rather than for being persuasive.

The objective

RLCD, and what it optimises

Under RLHF the training signal is a human preference. Under RLCD the signal is outcome agreement: the model says 0.8 and is scored on whether the world was true about 80% of the time it said 0.8.

That reframes the whole task. A conventional language model asked for a probability is extrapolating from text that looks like a probability. An RLCD-trained model is being directly optimised so that its numbers mean something.

It also explains the parallel sampler. If the training target is "produce these N probabilities accurately", there is no reason for the architecture to emit them one token at a time — and every reason not to, since sequential decoding is the expensive part.

⚠️ RLCD is TypeSafe's term, and its details are private

TypeSafe describes RLCD but has published no technical paper, no loss function, and no ablations. Because of that, independent projects use the label loosely. One roundup of the first week's clones noted a project "claims RLCD without justification", and that another's confidence is entropy-based rather than calibrated — those are materially different things. Entropy measures how spread out a distribution is; it is not a statement about accuracy. Treat "we use RLCD" in any model card as a claim requiring verification, not a technique you can look up.

Why uncertainty classification isn't enough

A model can be an excellent classifier and a useless confidence reporter. The benchmark that matters for automation is not accuracy alone but the pair — accuracy and how well stated confidence predicts it. That pair is why JevBench carries calibration as a quarter of its composite score and why the Jev Decision Index reports ECE for every system with a distribution.

Measurement

Calibration, measured properly

"Calibrated" is a specific, testable property. Two metrics carry almost all the weight in this literature.

ECE — Expected Calibration Error

Group the predictions into bins by the confidence the model stated. In each bin, compare the average stated confidence with the observed accuracy. A perfectly calibrated model lands on the diagonal: of everything it called 0.9, exactly ~90% was right.

JevBench publishes ECE on its hardest tier. Lower is better. For reference, in JevBench's first run the best calibrated system measured 0.009 and Jev 1.13 measured 0.027.

Brier score, and distribution fidelity

The Brier score is the mean squared error of probabilistic predictions — it penalises both being wrong and being confidently wrong. Kev's model cards publish it alongside accuracy so you can see accuracy and calibration move independently.

JevBench adds a second calibration term: fidelity to the exact gold distribution. That is stricter than "did the top answer win" — it asks whether the whole distribution matched. Systems that emit only a label score zero on it.

⚠️ Two numbers that must never be pooled

Benchmarks distinguish native probabilities — read directly off a model's output distribution in a single pass, which sum to 1 by construction — from verbalized probabilities, where an ordinary instruction model is asked to write a distribution under a JSON schema. JevBench labels every row native or verbalized and explicitly refuses to pool them into one calibration claim. It also found that only the verbalizing models needed their distributions renormalized — three systems in the first run produced sums slightly off 1 purely as a rounding artefact of writing three decimal places. These are different objects and pretending otherwise inflates the comparison.

And the sentence that keeps it honest

TypeSafe again: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." Calibration is a property of a population of calls, which is exactly why it is useful for automation — you can budget for a known error rate — and exactly why it does not let you skip verification.

Data

The data: 100% synthetic

Almeida's strongest claim is not about the model — it is about the data. Jev is trained exclusively on synthetic data, and he calls the decision to build that capability the best bet of his career: "better than our launch, in my opinion, better than RLHF. Half of [our company] is a lab that basically owns this entire subfield of statistically well-understood synthetic data."

The logic follows from what the model needs. Calibration can only be trained if you know, for each example, what the true probability distribution over the options actually is. Hand-labelled data gives you a label, not a distribution. Synthetic generation lets you construct tasks where the ground truth is known by construction — and construct millions of them, in shapes nobody would bother hand-annotating.

⚠️ And it is the least scrutinised part

Latent Space's roundup put it bluntly: not enough people are talking about the data side, which is acknowledged to be 100% synthetic. Training a model to make calibrated judgments from synthetic scenarios means its notion of "the right answer" was authored by whoever wrote the generator. That is not disqualifying — but it makes the calibration claim and the data-generation claim the same claim, and neither has an external audit. It is also the mechanism Simon Willison warned about when he called the bias risk serious: a floating-point number can conceal a great deal.

How the open projects get their data

  • Distillation from strong teachers — JevK5 v0.3 trained on 17,408 questions written by Qwen3.6-27B and GPT-6 Luna, plus 30,052 items replayed from public train splits.
  • Contrastive curation — Bespoke Nimble's recipe uses synthetic contrastive data curation to pick which examples carry signal.
  • Public labelled datasets — BANKING77, CLINC150 and similar intent sets, recast as Choice questions with full label sets.
  • Programmatically generated rubrics — policies, deadlines and rules turned into questions whose answers are computable.
The gap

What TypeSafe did not publish

Worth stating plainly, because it shapes everything on the open-models page.

❌ Not published

  • Exact architecture
  • Weights
  • Technical paper
  • Loss function or RLCD details
  • Training data or generator
  • Parameter count

✅ Stated

  • Transformer-based
  • Trained on 100% synthetic data
  • New architecture + parallel sampler
  • Trained with RLCD
  • Optimised for structured output
  • Cannot generate text

🔍 Independently reconstructed

Since the weights are closed, the community worked backwards from behaviour. Archer Hume's write-up "Jev's Architecture Unmasked" is the best-known reconstruction, and Kev states its architecture is explicitly based on it. Outside observers have separately suggested Jev may be built on top of an open-weight LLM. None of this is confirmed by TypeSafe. Treat every architectural claim about Jev as inference from the API's behaviour.

Open reproductions

Four ways to build one

The open projects differ less in ambition than in where they intervene. Every one of them is solving the same problem: get a probability distribution over a declared answer set without generating tokens. There are four practical strategies.

Architectural families, inferred from each project's own README and model card. Parameter counts and mechanisms are as published by the authors.
Family Mechanism Representative projects Trade-off
A · Encoder + option head A bidirectional encoder (ModernBERT, DeBERTa) reads state, question and all candidate options at once; a small decision head scores each option into a distribution. Nothing is generated; the whole thing is one encoder pass. Laya · Von · open-jev-deberta-v3-large Tiny (200M–421M), CPU-viable, milliseconds. Weakest zero-shot. Built to be fine-tuned, not used as-is.
B · Causal LLM + adapter + head Keep a normal Qwen or Gemma backbone, train a small LoRA plus a pointer/readout head that compares each option's representation against the question's decision representation. Kev · Bespoke Nimble · NanoJev · mini-Jev · JevK5 · jev-lite Best accuracy of the open families, reuses strong pretrained knowledge. Needs a GPU; model sizes 0.6B–9B.
C · Frozen model, logit readout Train nothing. Read the option probabilities straight off a frozen model's next-token logits at the answer position, skipping generation entirely. SemIf · AnyJev · JevK5's readout Zero training cost, uses a model you already host. Requires a causal model with sane logits — and you inherit its biases wholesale.
D · Diffusion Denoising diffusion rather than autoregressive or single-pass decoding, adapted to emit a distribution over declared options. djev (Maisa) · DiffusionGemma-based builds Newest and least characterised. Benchmarks competitively without following the same architectural logic.

Family A in one line

Make the whole model small enough that a single forward pass over state + question + every candidate option is cheap. Von is 395M parameters and reports sub-15 ms decisions.

Family B in one line

Keep the knowledge that pretraining bought you, and add the smallest possible piece that converts it into a distribution. Kev describes this as a rank-16 LoRA plus "a small pointer head that compares each option representation with the question's decision representation."

Family C in one line

The insight that made a dozen clones appear in a week: reading option logits off any open-weight model gives you the same answer shape for free. Reproducing the interface is easy. Reproducing the quality is the research problem.

The key trick

The readout: skipping generation

If you want to understand these models, this is the mechanism to internalise. An autoregressive model already computes, at every step, a probability distribution over its entire vocabulary. Normally you sample one token from it and loop. But the distribution is already there — it contains the probability of every possible next token, including the token for "A", "billing", or "yes".

A decision model's readout simply refuses to sample. It presents the candidate answers so that each maps to a token or a letter, then reads the probability of each straight out of that single distribution and renormalises across the candidates. One forward pass, no generation, and the probabilities are native rather than written down.

Step 1Render the options into the model's vocabulary
state + question→ "Answer with one letter: A) returns B) shipping C) billing"
Step 2Read the next-token distribution at the answer position
P("A") = 0.47 P("B") = 0.28 P("C") = 0.25 …every other token also has a probability
Step 3Renormalise over the candidates, return the distribution
returns 0.47 shipping 0.28 billing 0.25

A temperature parameter controls how sharp this distribution is — JevK5 publishes its fitted value (1.22) in the model card.

❌ And this is where the whole approach can quietly fail

A raw next-token distribution is not a calibrated probability about the world. It is a model's internal preference over strings. JevK5's model card makes this explicit about its intended comparison: its untrained base model scores 1.000 on easy items and 0.613 on hard ones from the same readout — meaning the readout alone carries a lot of signal, but the interesting part is entirely about which items it fails. Two other documented hazards follow. First, option order changes answers: JevBench recorded one open model scoring 72% with options in one order and 21% with them reversed, and Kev's README lists "changing option order can change an answer" as a known limitation that question isolation does not prevent. Second, knowledge questions stay unsolved — a Kev trained on Qwen3.5-9B scores 0.74 on MMLU against Jev's 0.90, and a version trained on a 35B mixture-of-experts base did not move that number, because the ceiling was the base model.

Expectations

What fine-tuning can't fix

The most useful thing in Kev's documentation is the list of things training did not solve. It is the clearest evidence available about the shape of this problem.

Problem Does fine-tuning fix it? Evidence
Base-model knowledge No. Set by the pretrained backbone. Kev-9B scores MMLU 0.74 vs Jev 0.90 and MMLU-Pro 0.52 vs 0.84. A version trained on a 35B-A3B MoE base did not move these numbers.
Option-order sensitivity No. Listed as an unresolved limitation. Kev: "Changing option order can change an answer. Question isolation doesn't prevent this." JevBench measured one open model at 72% → 21% on reversal.
Domain adaptation Yes — strongly. This is the whole value of fine-tuning. Kev-9B: 0.872 accuracy on held-out examples from trained sources; 0.822 on genuinely new sources. Nimble reports base Qwen going 66% → 90% on a curated eval.
Calibration Yes, measurably. Kev-9B Brier on new sources improves to 0.237 against Jev's 0.211 — close, and improving with scale. JevK5 v0.3 reported hard-tier ECE of 0.054 against an untrained baseline's 0.117.
Reasoning-shaped tasks Barely. The model is a classifier with no reasoning step. JevK5's model card records hard-tier multi-hop reasoning at 0.11 and temporal/numeric at 0.27.
Long inputs No. Truncation is a serving limit. JevK5 truncates state past 512 tokens, which affected 52 of 111 hard items (median state length 2,188 tokens); accuracy on those was 0.346 vs 0.458 on the rest. Jev's own context window is bounded too.
Speed Yes — dramatically. JevK5 built one CUDA graph per padded input length: 13 ms versus ~70 ms eager on an H100, with identical answers.
✅ The honest summary of the training picture

Fine-tuning a decision head onto an open model reliably buys you domain accuracy, calibration and speed. It does not buy you knowledge, reasoning, or robustness to input formatting. Those come from the base model or not at all. So the practical recipe is: pick the strongest base you can afford, fine-tune it on a few hundred examples of your decisions, refit the confidence temperature, then measure it against whatever you run today. Several projects' authors say roughly this, including the ones whose numbers look best.