Overview / Jev

03 · The reference implementation

Jev in detail

The model that started the category, and the one almost every open project is measured against. This page is deliberately split between what TypeSafe can be verified to have shipped, what it claims, and what its own documentation admits is broken.

Reference

Fact sheet

Developer
TypeSafe AI — San Francisco, founded 2024, after roughly two years in stealth
Founders
Diogo Almeida (CEO — ~4 years at OpenAI on RLHF, InstructGPT, ChatGPT, GPT-4), Erik Gafni, Sasha Sheng
Released
15 September 2026, limited early access / waitlisted
Version in this guide
jev-1.13 (stable jev-1.13.0)
Funding
$40M seed led by DCVC; reported valuation ~$200M
Licence
Proprietary. Closed weights, closed architecture, no technical paper.
Training
Transformer-based, trained exclusively on synthetic data via RLCD
API
POST /v1/systemone · SDK default model jev-latest
Availability
One hosted API. Nothing to self-host, nothing to download.
Pricing
$0.042 per million input tokens ($42 per billion). Output tokens free.
Epistemics

Shipped vs claimed vs unknown

To TypeSafe's credit, the launch post includes its own "nuance" notes underneath each result — much of the middle column below is them marking their own homework down.

Verifiable

Checks out directly

  • Speed per call — you can time it yourself
  • Cost per call — pricing is published and transparent
  • The API contract — documented and publicly readable
  • Type-safety — true by construction, not by training
  • The product exists and serves traffic
Self-reported

Plausible, unaudited

  • 40–200× faster, 40–400× cheaper
  • 70–500 ms end-to-end latency
  • Peak figures of 193.6× faster / 444.6× cheaper
  • Calibration quality (the central claim)
  • Reliability of the parallel sampler at scale
Unknown

Nobody outside knows

  • Architecture, parameter count, weights
  • Training data and the synthetic generator
  • Public benchmark performance
  • Behaviour on non-decision-shaped tasks
  • Production reliability under sustained load

The caveats TypeSafe published against itself

Drawn from the "nuance" notes in TypeSafe's own launch post and their workflow-evals site.
Claim The company's own caveat
The 193.6× / 444.6× headline figures "We expect that these are on the higher end of real world gains."
The benchmark workflows "The content of these workflows were not deliberately chosen nor constructed to make our model look good… However, they were made by individuals on our model capabilities team, so some bias could exist."
The accuracy comparison Reference answers are the average of GPT-6 Astra and Fable 5.1, which "biases answers towards OpenAI and Anthropic's models. We likely underestimate the relative performance of our model and DeepSeek's models."
LLM comparison rows Numbers come from OpenRouter, so "there almost certainly is bias here: more complex queries might be routed to better models."
The 0% type-error plot "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots."
Sustainability of the price "We can't prove it isn't subsidized; we'll need the long-term to prove the sustainability of our pricing."
Where the timings came from "Our published evals are generally run from our laptops on the West Coast (this is where our service is currently based)."
The side-by-side demo The input state was "short, dense, and detailed… The relatively shorter input paints our model in an advantageous light."
This is unusually honest vendor behaviour — and it still isn't verification

Publishing your own counter-arguments is rare and worth crediting. But self-flagged bias is still bias, and a benchmark on a format the vendor invented, graded against a reference the vendor chose, is not evidence a third party can check. The correct reading is directionally credible, numerically unconfirmed.

Failure modes

The nine jagged edges

TypeSafe maintains a public "jaggedness" page listing known weaknesses of jev-1.13, last reviewed 17 September 2026, with a prescribed workaround for each. It is the single most useful document in this ecosystem for anyone actually building. Their own summary: the model is "fast, calibrated, and good at common-sense judgment but it is not perfect. It can be quite literal in its understanding. It struggles with tasks that require numeric precision."

# Failure mode What goes wrong Prescribed fix
1 Literal reading It answers the question you wrote, not the one you meant. Scoping words, negations and implied conditions are read at face value. Write the exact condition in the instructions. Put boundary cases in the criteria. "When you look at a wrong answer and find yourself explaining what you really meant, that explanation is the missing half of the instruction."
2 Math and counting "Jev is not a calculator." Counting is unreliable — characters, term occurrences, items in a list. It recognises the shape of an answer rather than tallying, and error grows with size. Do all arithmetic in code. To count matching items, iterate in code and ask one question per item, then add the answers yourself.
3 Numeric representations Semantic beats numeric. Hex/RGB colours underperform English colour names; assembly or binary underperforms high-level languages. It cannot reliably judge whether two RGB triples are near each other. Convert in code and pass either the computed number or a named bucket. Keep the model for the part that is genuinely a judgment, e.g. "does this colour read as a warning?"
4 Score-scale interpolation Score levels are weak at numerical calibration — you cannot reconstruct an exact magnitude by interpolating between the nearest two levels. Use the expectation only to test a threshold. Do not use score outputs to compute an exact number.
5 Date and time comparison Reads dates as text, not ordered quantities. Ordering, distance and window membership are unreliable, and worse with mixed formats, relative references, quarters or settlement windows. Split the work. Extraction is a judgment — use the model, as a Choice over twelve months / thirty-one days / a bounded year range, with an explicit "not stated" option. Code assembles and owns ordering, duration and offset.
6 Indirection Double negatives and multi-hop instructions ("a property of a property") cost accuracy. Write instructions as directly as possible. Identify relevant parts of the state by name.
7 Large state with irrelevant detail "Context rot." Accuracy falls as the state grows with content unrelated to the decision; unrelated detail acts as a distractor and makes a wrong answer harder to diagnose. Retrieve and filter in code first; send only the fields the question needs. Where filtering in state is unavoidable, use a Noul to filter for relevance.
8 Adversarial content State is not treated as hostile. Injected instructions, misleading framing, or text that argues for its own classification can move the answer. Be explicit in the criteria, and test thoroughly before deploying to many users. TypeSafe says they expect to improve this in future versions.
9 Contradictory criteria Instructions and option criteria that conflict with each other degrade reliability. Align the instruction and the criteria.

Plus one trap that isn't a "mode" — broken structural assumptions

TypeSafe documents that the model is extremely consistent — semantically similar inputs give quantitatively similar outputs — but that many structural invariants you might assume simply do not hold.

Concretely: on the same ticket, "I'm not happy with the fit. What are my options here?", the Noul form of "is the customer asking for a refund?" returned 0.22, while the equivalent yes/no Choice returned 1% yes / 99% no. And for a ticket reading "I was charged twice for the same order. Can someone look into this?", P(refund) = 0.72 and P(not_refund) = 0.47 — summing to 1.19.

❌ Don't build on arithmetic identities

A Choice over options and one Noul per option answer different questions: the Choice is relative (settling which option), while each Noul is absolute and can be low for all of them. TypeSafe's guidance is direct: don't rely on expected structural invariance, don't carry a threshold tuned on a Noul over to a Choice, and don't hold the model to arithmetic identities between separate questions. If your code has an assert that probabilities sum to 1 across questions, remove it.

Boundaries

Hard limits

255 Maximum options in a single Choice question. Beyond that TypeSafe switches to a two-stage score-then-choose, which is slower.
Bounded Context window. TypeSafe's docs point to the Models page for the exact token limit rather than advertising a headline number.
0 Tokens generated. Not trainable into a text generator — chaining choices to emit text is documented as working badly and being very slow.
1 Hosted API. No self-hosting, no local deployment, no on-premise option.
⚠️ Two more limits you have to design around
  • No answer outside the option list. If the correct answer isn't among your options, the model cannot surface it — it will select the nearest permitted value. An explicit "none of the above" has to be a choice you thought to include.
  • Everything depends on one vendor staying up. During launch week TypeSafe "briefly lost the ability to serve users from its API because demand was so high" (TechCrunch). Early access also means waitlisting, per-account limits, and no service-level guarantees.
The serious objection

The black-box problem

Simon Willison's write-up raises the objection that matters most, and it is not about latency or cost: "Something I've found a little uncomfortable about Jev is how it very much represents a regression even further towards black box machine learning systems."

LLMs are already opaque, but you can at least ask one to justify itself — and while that justification may be unreliable, it is something to interrogate. A decision model gives you a float. If it marks a message as spam, which content signals tipped it off? There is no answer to retrieve.

That makes bias auditing harder precisely where it matters most, and Willison's own experiment is the illustration: asked to score every city in the San Francisco Bay Area on a yes/no "Good city?" question, Jev rated Cupertino top and East Palo Alto bottom. His reaction — "Huh." — and then the warning:

❌ The sentence worth quoting to anyone proposing a use case

"I really hope nobody uses Jev to rank job applicants — that floating point number could conceal all manner of unseen bias baked into the models, and experimentally picking that bias apart is going to be a tricky business."

The counter-argument

Willison also concedes the mitigation: because these models are so cheap, you can run hundreds or thousands of experimental probes for a few cents. Eval suites and structured experiments become the substitute for introspection — you cannot ask why, but you can measure what. Every serious write-up in this space converges on the same advice.

Where that leaves you

Treat a decision model the way you would treat any statistical component with a distributional output: measure it on your own data, disaggregate the results by the groups you care about, and keep a human in the loop for high-stakes calls. The type-safety guarantee is about shape, and shape has never been the hard part of fairness.

Adoption

Adoption and reception

Third-party reports from launch week. These are individual testimonials, not controlled evaluations — useful as a signal of where the tool actually helps.

Who What they did Reported result
Vercel
Pranit Sharma
Replaced an OpenAI model used as a classifier reviewing commands for safety Five to 18 times faster, with greater accuracy
Bryo AI
Nikhil Mudholkar, CTO
Classifying business emails, versus Gemini Gemini slightly more accurate, but 10–20× more expensive. The real draw was the confidence score: "the only one that hands back a real probability which makes it ideal for automating workflows"
Vercel
AI Gateway usage data
Adoption measured across their gateway in the first day Reached ~13% of teams — 2× the GPT-5.6 family and 6× Fable 5.1. Described as the fastest adoption of any model in the gateway's history.
Box
Aaron Levie
Demo of classifying incident reports into escalation paths Cited as a compelling integration
LangChain Built TypeSafeClassifier plus middleware for model routing and tool-risk gating Ships in the open-source library; the browser/computer-use control loop was called the best application seen so far by LangChain's CEO
Skeptics Several ML practitioners pushed back on the framing The recurring substantive question, from Abhinav A. on X: many demos emphasised speed more than quality, and at that point there was no standard benchmark for the category. That gap is what JevBench and the Jev Decision Index have since tried to fill — see Benchmarks.
Trajectory

What's next

TypeSafe says more versions of the model are coming, in new modalities — the current release works on structured state as text, and the demos are explicitly noted as "not on images (yet)". Expect competitors too: Armin Ronacher's assessment is that others will spring up now that the utility is apparent, and that "we should have seen this earlier in many ways, but presumably because the LLMs are so cheap and subsidized, you often don't have to be creative yet."

Almeida's own close is a deliberate refusal of frontier-lab framing: asked whether TypeSafe is a frontier lab, he said "the main product of frontier labs is fear or hype. I would like our main product to be intelligence… [but we are] not a lab in the sense of, you know, like bet on infinite wealth, or a religion, or building God in a data center."

What would change this page

  • An independent replication of the calibration claim — the single biggest open question
  • A technical paper or an architecture disclosure
  • Public-benchmark results from a party with no stake in the format
  • Evidence the $0.042 / MTok price survives contact with real unit economics
  • Sustained-load reliability data past launch week
  • Any of the nine jagged edges being fixed — the page is versioned and they say many will be