Overview / Sources

07 · Provenance

Sources, and what we couldn't resolve

Everything on this site traces to one of the items below. Each is tagged by kind, because in a field eleven days old the difference between a vendor's own benchmark and an independent one is the whole game. This page also lists the contradictions we could not settle.

Legend

How sources are tagged

Vendor / primary

Material published by the party being described — TypeSafe's own blog, documentation and API reference. Authoritative about what they claim and what they shipped. Not independent about whether it works.

Independent

Measurement or analysis by a party with no stake in the outcome. The strongest evidence available here. Caveat: most are one-person projects, and all publish their own limits.

Community

Project repositories, model cards, directories and practitioner reports. Usually the most candid about failure modes, since the authors hit them first. Not comparable across projects.

Primary

TypeSafe AI — vendor and documentation

  • Vendor Introducing System One Models & Jev — Diogo Almeida's launch post, 15 Sep 2026. The origin of the LLM-vs-System-One comparison table, the RLCD description, the cost and latency claims, and the "nuance" notes quoted throughout. typesafe.ai/blog →
  • Vendor System One — concepts documentation. The canonical definition of the class, the three primitives table, the confidence caveat ("calibration is measured across groups of predictions"), and the POST /v1/systemone endpoint reference. docs.typesafe.ai/concepts/system-one →
  • Vendor Jev 1.13 jaggedness. The nine documented failure modes and their prescribed workarounds, plus the structural-invariance examples. Last reviewed 17 Sep 2026. The single most useful document in this ecosystem for anyone building. docs.typesafe.ai/model-jaggedness/jev-1.13 →
  • Vendor Primitives and state references. Choice, Score and Noul semantics; the shapes state can take. docs.typesafe.ai/primitives →
  • Vendor TypeSafe AI blog and manifesto. The framing that today's AI "was trained on the assumption that a human sits on the other side of it." typesafe.ai →
Journalism and analysis

Independent coverage

  • Independent Simon Willison — "Jev introduces a new shape of LLM — System One, aka Decision Models", 21 Sep 2026. Source of the "decision models" reframing, the black-box objection, the Cupertino / East Palo Alto experiment, the search-reranking use case, and the list of unconventional demos. simonwillison.net →
  • Independent TechCrunch — "A new kind of AI model from a ChatGPT inventor is thrilling developers", 18 Sep 2026. The founder interview quotes, the Vercel and Bryo AI testimonials, the API overload incident, the synthetic-data claim, and Armin Ronacher's "delegates the hallucination problem to the user". techcrunch.com →
  • Independent Wikipedia — Jev (AI model). Fact sheet cross-check: release date, version, funding, the confirmation that no architecture, weights or technical paper have been published, and the note that outside observers suspect an open-weight LLM underneath. en.wikipedia.org →
  • Independent Forbes — "This $200 Million Startup Wants To Fix AI's Overconfidence Problem", 15 Sep 2026. Original source for the $200M valuation and the "we've been optimizing for humans" quote. forbes.com →
  • Independent Latent Space / AINews — "Here are 6 Clones of Jev in 2 days." The clearest account of how fast the open ecosystem appeared, the per-project technical sketches, and the observation that the synthetic-data side went largely unexamined. latent.space →
  • Independent TrueFoundry — "TypeSafe AI's Jev and 'System One Models': What Actually Shipped." The clearest separation of verifiable claims from vendor claims, and the "constraining the shape of an answer has never guaranteed the substance of it" formulation. truefoundry.com →
  • Independent DataCamp — "Top 7 Open-Source TypeSafe Jev Alternatives." Mechanism-level descriptions and quick-start code for Laya, Nimble, Kev, SemIf, Rizzo Flow, Von and NanoJev. datacamp.com →
  • Independent Pinggy — "Best Open Source Jev Alternatives: 5 Self-Hosted Decision Models." Star counts read from the API on 23 Sep 2026, hand-measured latency on an M3 Pro, the 0.966-vs-0.704 framing, and the "stars measure attention, not accuracy" observation. pinggy.io →
  • Independent LangChain — "Building a Harness with Jev." The agent-loop framing, the three-question example, the TypeSafeClassifier API, and the model-routing and tool-risk-gating middlewares. langchain.com →
Measurement

Independent benchmarks

  • Independent JevBench harness — github.com/fstandhartinger/jevbench (MIT). Scoring code, item sets, per-task outcomes, revision log, and the "Limits, stated plainly" section that every caveat on the Benchmarks page is drawn from. github.com →
  • Independent JevBench v1.0 archive — 242 decisions, five axes, deliberately no combined score, full exclusion list with reasons. benchmarkheaven.com/jev-models/v1 →
  • Independent JevBench v1.2 leaderboard — the JevBench Score, four axes as a geometric mean, the 220-item hard tier, and the per-axis receipts. benchmarkheaven.com/jev-models →
  • Independent JevBench v1.4 methodology — the harmonic-mean composite, the 308 sealed decisions, and the 93-system cohort. Referenced by JevK5's card for its 2nd-of-76 placement. METHOD-v1.4.md →
  • Independent Jev Decision Index — @multimodalart on Hugging Face. 132,422 requests over 37 benchmarks, a 19-benchmark scored panel across five equal-weight areas, and published house rules banning prompt tuning, truncation and option filtering. huggingface.co/spaces →
  • Independent JevBench.dev — a separate effort benchmarking decision models on interactive harnesses (win/loss, task completion, latency) starting with StarCraft II, rather than on single typed answers. Note: a different project from the JevBench above, with a similar name. jevbench.dev →
Community

Project repositories

Model cards and READMEs are the source for every parameter count, licence, mechanism and self-reported score on the open models page.

Honesty

What we could not resolve

Every guide like this has seams. These are ours, stated rather than smoothed over. Where the sources conflict, both versions are given and neither is presented as settled.

The central claim — calibration — has no independent replication

Everything that makes this model class interesting rests on probabilities that mean something. As of 28 September 2026 we found no third-party study that validates TypeSafe's calibration numbers. JevBench measures ECE on Jev (0.027 in its first run) but that is a measurement of a benchmark run, not a replication of the training claim. TypeSafe's own launch post concedes the benchmarks are self-designed. Treat calibration as the working hypothesis of this field, not an established result.

Jev's architecture, parameter count and training data remain unknown

No paper, no weights, no parameter count. TypeSafe says "transformer-based", "new architecture", "parallel sampler", "100% synthetic data", "RLCD". Outside observers have suggested it may sit on top of an open-weight LLM. Archer Hume's reconstruction is inference from API behaviour. Any architectural statement about Jev in this guide — and in most coverage you will read — is inference rather than disclosure.

RLCD is a term, not a documented method

TypeSafe names the training technique but has published no loss function, no ablations and no paper. Independent projects use the label loosely — one first-week roundup noted a project "claims RLCD without justification" and that another's confidence is entropy-based rather than calibrated. Those are materially different things wearing the same label. If a model card cites RLCD, that is not yet a verifiable claim about how it was trained.

Popularity metrics contradicted each other on the same day

While compiling this guide, one repository's own page reported 2.5k stars while GitHub's topic index reported 7.2k for the same repository, on the same day. Third-party roundups published a third figure. This is why the open-models page ranks nothing by stars and why no star count appears in it. If you need one, read it from the repository page yourself on the day you need it.

JevBench has three cohorts in circulation with different formulas

v1.0 (no combined score, 242 decisions, five axes), v1.2 (geometric mean of four axes, 534 decisions, three-tier item set) and v1.4 (harmonic mean, added 308 sealed decisions) produce different rankings from partly different item sets. Figures quoted without a version are not comparable. We have labelled every number on the benchmarks page with its version, but coverage elsewhere frequently does not.

Reported adoption results are individual testimonials

Vercel's "5–18× faster with greater accuracy" and Bryo AI's cost comparison are single-company reports about their own workloads, relayed through a journalist. They are useful signals about where the tool helps. They are not controlled evaluations and we have not verified them independently. The same applies to the X/Twitter practitioner reports cited on the practice page.

Model naming across sources is inconsistent

The frontier models that decision models are compared against appear under differing names in different sources — for example, a Vercel classifier described as "ChatGPT Luna 5.6" in one report and "GPT-5.6 Terra" in TypeSafe's own demo notes. Some of these may be different models, different snapshots, or reporting errors. We have preserved each source's own naming rather than normalising it, because normalising would be guessing.

Ecosystem size figures use different counting methods

Within one week we found the open ecosystem described as 6 clones in 2 days, 146 open-source builds, 916 projects, and 1,858 repositories tagged jev. These are not contradictions — they count different things (direct model reproductions, directory entries, catalogue items, any repository using the tag) — but they are routinely quoted interchangeably. Treat all of them as order-of-magnitude signals.

One project's attribution varies between sources

mini-Jev appears attributed both to a GitHub account listed as r-ms and to a Hugging Face account listed as samatv256 in different directories, describing the same frozen-Qwen3-0.6B-plus-head design. Similarly, Laya's release is credited to "Convai Innovations" in one roundup while the repository lives under a personal account. We have described the mechanisms without asserting a single authorship where the sources disagree.

Everything here has a short half-life

Jev launched eleven days before this page was compiled. The open models move weekly; benchmarks are revised without version bumps in their public URLs; licences change; projects get abandoned or renamed. The system-one-adapter-python wrapper TypeSafe published, the Laya version numbers, and the JevBench cohorts were all in flux during compilation. Treat every figure as a dated observation, not a standing fact.

Corrections welcome

If you maintain one of the projects catalogued here and we got your base model, licence, mechanism or score wrong, that is a documentation failure on our part rather than a judgment about your work — several of these projects' own authors were more careful about their numbers than most of the coverage that followed. Send a pointer to the authoritative line in your own repository and it should be corrected here.