Overview / Benchmarks
05 · MeasurementBenchmarks & leaderboards
When Jev launched there was no standard benchmark for the category at all — a gap practitioners flagged within days. Three now exist, measuring different things, and all three publish their own limitations. This page covers what each one actually tells you.
Why this category was hard to measure
Conventional model evaluation assumes a text answer. A decision model refuses to produce one, which breaks almost every standard harness: there is no completion to BLEU-score, no reasoning trace to grade, and a "wrong" answer can still be a perfectly legal value from the schema you supplied.
Worse, the thing that matters most is not accuracy — it is whether stated confidence predicts accuracy. A benchmark that only measures top-1 correctness tells you nothing about whether you can safely branch on the probability, which is the entire reason to use a decision model.
The scepticism at launch was explicit and, in hindsight, correct. One widely-shared critique: many demos emphasised speed more than quality, and there was no standard benchmark for this category. The leaderboards below are the community's answer.
All three benchmarks here are built by small teams or individuals, not standards bodies. Two are essentially one-person projects. That does not make them wrong — they publish methodology, artefacts and limitations, which is more than most leaderboards do — but it does mean you should read the stated limitations rather than the ranking. Every one of them says something equivalent to "treat this as a pilot."
Three benchmarks, three different questions
| Benchmark | Built by | Question it answers | Scale | Combined score? |
|---|---|---|---|---|
| JevBench | Benchmark Heaven (Florian Standhartinger) — a one-person project, self-funded | How do Jev-class decision models compare on accuracy, calibration, speed and cost, on the same items? | 534 public + 308 sealed decisions per system; 93 systems registered by v1.4.2 | Yes — since v1.2, a square-rooted geometric mean of four equally-weighted axes (a weak axis drags the score down hard). Earlier versions deliberately had no combined score. |
| Jev Decision Index | @multimodalart, published on Hugging Face | How broad is a decision model's competence across many task areas? | 132,422 requests across 37 benchmarks; 30+ open-weight models | Yes — 100 × the mean of five equally-weighted area scores |
| jabr classifier suite | Independent (referenced by third-party analyses) | How well does a general classifier handle unfamiliar domains, zero-shot? | 49 tasks, 869 cases | Macro accuracy across domains |
- Open vs closed, out-of-domain: Jev 0.966 vs best open 0.704 macro accuracy on jabr. A 26-point gap.
- Open vs closed, best case: JevK5 v0.2 placed 2nd of 76 systems on JevBench v1.4 and 1st among open entrants — 62.04 against Jev's 63.29.
- Within noise: openjev-sglang (Qwen3.6-35B-A3B) scored 95.5% against Jev's 96.3% on JevBench v1.0, with overlapping confidence intervals.
JevBench, version by version
JevBench is the most-cited independent leaderboard for this category. It is not affiliated with or endorsed by TypeSafe — whose model is one of the systems it measures. Its harness and 72 original public decisions are MIT licensed. Because it has been revised several times in under two weeks, older numbers in circulation are often being compared against newer ones. Version matters.
242 decisions per system — 72 published, 24 held out, 146 imported. One request at a time from a Hetzner server in Germany. The five axes were accuracy, cost, latency, validity of answers, and openness. The stated rationale for refusing a single score: "the trade-off between them is the point."
| System | Accuracy | Calibration (ECE) | Speed (median) | Cost / 1k | Type |
|---|---|---|---|---|---|
| GPT-5.6 Luna (low reasoning effort) | 97.1% | 0.016 verbalized | 0.97 s | $0.176 | Closed |
| Jev 1.13.0 (TypeSafe AI) | 96.3% | 0.027 native | 0.65 s | $0.027 | Closed |
| DeepSeek V4.1 Flash (thinking default) | 95.5% | 0.009 verbalized | 1.42 s | $0.245 | Open weights |
| openjev-sglang (Qwen3.6-35B-A3B, SGLang) | 95.5% | 0.042 native | 0.68 s | no tariff | Open |
| system-one-open (Gemma 4 E2B LoRA, L4) | 90.1% | 0.068 native | 0.65 s | no tariff | Open |
| open-jev-deberta-v3-large (local CPU) | 51.7% | 0.147 native | 1.77 s | no tariff | Open |
| Stopped early — shown with its own denominator, not ranked against complete runs | |||||
| Qwen3.8 27B (Chutes TEE) | 96.9% | 0.014 verbalized | 5.75 s | no tariff | Open weights |
What v1.0 established
Jev is not uniquely good — it is uniquely cheap
GPT-5.6 Luna led on raw accuracy and DeepSeek V4.1 Flash led on calibration. Jev's decisive advantage was price: $0.027 per 1,000 decisions, the cheapest metered route measured, while sitting inside the noise band of the leaders on accuracy.
The open field is close on accuracy, far on cost
openjev-sglang matched Jev within intervals at 95.5% using a 35B mixture-of-experts model on specialised hardware. That is a real result — and also why "just self-host it" is not automatically cheaper than a hosted API at $0.042 per million tokens.
Latency differences are mostly indistinguishable
Three systems landed within 0.03 s of each other at the top. When your cheap decision model and your hosted API are 30 milliseconds apart, latency stops being a differentiator and price and calibration decide it.
v1.2 — the combined score, and what went into it
v1.2 replaced "no combined score" with a defined one, which makes the ranking much easier to read and much easier to over-read. The formula is published and worth understanding, because the geometric mean means a single bad axis destroys a score.
100 − 20·log₁₀(s / 0.1 s). So 0.1 s scores 100, and every 10× slower costs 20 points.100 − 30·log₁₀($ per 1,000 decisions / $0.001). $0.001 scores 100; every 10× more expensive costs 30 points.| # | System | JevBench Score | What it is |
|---|---|---|---|
| 1 | classifier.dev (fast tier) | 84.8 | Jev 1.13.0 behind a third-party API. Leads on its flat pricing plan at full utilisation and on latency measured from the benchmark server. |
| 2 | Jev 1.13.0 | 75.3 | TypeSafe AI, proprietary API |
| 3 | SemIf (Qwen3.5-4B) | 74.6 | Open — frozen-backbone logit readout, no training |
| 4 | djev (Maisa) | 74.3 | Diffusion-based, served through a production API |
| 5 | Laya (421M) | 70.1 | Open — ModernBERT encoder plus decision head |
| 6 | open-alternative-jev (Qwen3.5-4B) | 69.8 | Open — see the option-order warning below |
open-alternative-jev is ranked using the author's own option order
(A. yes, B. no). With the options reversed (A. no, B. yes) the same
model scored 21% instead of 72% on answer-judging items. Both runs are
published. JevBench's authors flag one more thing about this row: it was the model whose
self-reported confidence behaviour made it the clearest example of how much a small model's
accuracy depends on formatting nobody thinks to vary.
v1.4 — sealed items and a harmonic mean
By v1.4 (25 September 2026) the suite had grown to 93 systems registered, 89 ranked, scored on 534 public decisions — 72 easy, 96 standard, 146 judge, 220 hard — plus 308 sealed decisions whose text stays private. Sealed items make up 20% of the intelligence axis precisely to make overfitting to the public set costly.
The composite also changed from a geometric to a harmonic mean of the four equally-weighted axes, which punishes a single weak axis even harder. In that run an open 4B model — decider-4b v2 — finished first on composite, 0.8 points ahead of Jev. Jev remained the most accurate system in the top ten. Both statements are true at once, and that is the point of a multi-axis score.
JevK5's placement, as its own card reports it
JevBench v1.4 ranked JevK5 v0.2 2nd of 76 systems and 1st among open entrants (62.04 against Jev's 63.29). On the 308 sealed decisions, JevK5 answered 33.1% correctly against Jev's 36.7% — the evaluator notes this sealed set is unusually difficult. The project is explicit that the ranking measures JevBench's own mix of accuracy, calibration, speed and cost, "not performance on every production workflow."
Because v1.0, v1.2 and v1.4 have different item sets, different scoring formulas and different cohorts, a number quoted without a version is not comparable to anything. The v1.0 archive is still published alongside the current board, which is good practice — but it means you will find conflicting figures for the same model. Ask for the version.
JevBench runs self-hosted and demo endpoints one request at a time with no other load, then multiplies their latency by 2 and adds 0.15 s "to approximate production load." The authors are explicit that this is an assumption, not a measurement, that it stands for infrastructure their tests lacked (authentication, load balancing, logging, billing, API gateway), and that raw p50/p95 are published so you can undo it. They also note that serving under load trades per-user speed for throughput by far more than 2× — citing an NVIDIA analysis — but concede that analysis used a 1.8T mixture-of-experts model on GPU clusters rather than a 4B model on one GPU, so it supports the direction rather than the exact factor.
Jev answered the same 242 decisions twice, about 16 minutes apart. Three answers changed — 1.2% of the suite, all in routing — and accuracy moved from 96.7% to 96.3%. Their guidance: read a gap of about one point between two rows as noise, and use the confidence intervals. This is unusually rigorous, and it immediately deflates a lot of leaderboard arguments.
The Jev Decision Index
Where JevBench asks "which is best on decision-shaped items", the Jev Decision Index asks how broad a model's competence is. It is published as a Hugging Face Space by @multimodalart.
- Tools & Automation
- Retrieval & Classification
- Language Understanding
- Knowledge & Reasoning
- Arts & Human Judgment
The house rules are the interesting part
A leaderboard is only as good as what it forbids, because those are the degrees of freedom a submitter would otherwise use to flatter their model. This one bans three:
- No prompt tuning. You cannot tune the prompt per benchmark item.
- No truncation. You cannot silently cut the input to fit your window — which is exactly the trick that made open-jev-deberta-v3-large's 51.7% in JevBench v1.0 look like a capability result rather than a context limit.
- No option filtering. You cannot discard candidates to raise your hit rate.
- Abstained and errored requests score zero, before chance normalisation.
JevK5 0.2.2 (on v0.2 weights) scores 36.31 — 15th of 49 submitted systems, with the 4th-best calibration. That combination is worth noting: a model can be mid-table on breadth while being near the top on the property that matters most for automation. Composite rank alone would hide it.
The jabr classifier suite
A much smaller and much harsher test: 49 tasks and 869 cases across compliance, triage, legal, DevOps, linguistics and safety, run out-of-domain and zero-shot. It is the source of the number that should temper every "drop-in replacement" claim.
| System | Macro accuracy | Recorded failure mode |
|---|---|---|
| Jev (hosted, TypeSafe) | 0.966 | — |
| Von (395M, Apache-2.0) | 0.704 | Collapses to one mode on unfamiliar domains |
| GLiNER2 (~300M, open) | 0.698 | Over-triggers on keywords |
| Laya (421M, open) | 0.583 | Compresses rating scales |
A 26-point gap out-of-domain, zero-shot. But note the qualifiers, because they are load bearing: in-domain, after a fine-tune on your own labels, the open models are competitive. What this benchmark measures is how much of the work you have to do yourself. With hosted Jev you are buying a model that already generalises. With an open model you are buying a base you will specialise. Both are legitimate; they are different products at different prices, and the 0.966 vs 0.704 comparison is the price of the difference.
Vendor-run evals
Both TypeSafe and several open projects publish their own evaluations. These are not worthless — they are often the only detailed numbers available — but they answer a different question: "does this model do well on the things its authors chose to test?"
| Who | Method | Their own stated limitation |
|---|---|---|
| TypeSafe — workflow evals | Assume a correct compute graph ("workflow") in code; compare each model's predictions against the average of the two strongest external models as reference probabilities, rather than a ground-truth classification | Workflows were written by their own model-capabilities team; the reference average biases toward OpenAI and Anthropic, so they "likely underestimate" themselves |
| Kev — model cards | Develop and test splits, separated into trained sources (held-out examples from training datasets) and new sources (unseen datasets and policy rule types); reports accuracy and Brier | Small models are slow on Apple Silicon; option order can change answers; training used at most 384 state tokens, so longer contexts were never trained on |
| JevK5 | Runs JevBench's own harness, reports McNemar significance tests on item-level changes across versions, and pre-registered its epoch sweep | Most gains are on held-out and index-style data; on JevBench's public items the 4B gained on the hard tier while losing one standard item. Quality on Jev's real-world workflows "has not yet been measured." |
| Bespoke Nimble | Curated evaluation against a controlled baseline — base Qwen improved from 66% to 90% | The eval is their own, on their own curated data, with the comparison point being Jev at 93% from a different measurement |
How to read a leaderboard in this category
A checklist distilled from the caveats the benchmark authors publish about themselves. Run it against any number you are about to quote.
- Which version? JevBench alone has three cohorts in circulation with different item sets and different formulas. A number without a version is not comparable.
- Native or verbalized probabilities? These are different objects and must never be pooled. Native distributions sum to 1 by construction; written ones are prose that needs renormalising.
- Was the benchmark's own training split used? One headline score that beat Jev came from a checkpoint fine-tuned on that benchmark's training split.
- Is the model a label-only system? It scores zero on calibration by definition, which drags a geometric or harmonic mean hard.
- How long was the input allowed to be? A truncation limit can masquerade as a capability limit — one system never saw the whole question 80 times out of 242.
- Was option order varied? 72% became 21% on the same model with the options reversed.
- Is the entry actually a different model? The top row of one leaderboard is Jev behind a different API — a hosting result, not a capability result.
- Were the high scores on a vendor-designed format? Speed and cost are directly checkable; accuracy on a self-invented rubric is not.
- Is the latency number a measurement or an assumption? One leaderboard multiplies self-hosted latency by 2 and adds 0.15 s, and says so.
- Are exclusions reported? A benchmark that quietly drops what it could not run cannot be checked. "No GPU available" is an availability fact, never a quality verdict.
- Was the held-out set actually unseen? Sending held-out items to the service to get predictions means "not public" is not the same as "not seen" — and this is not a contamination proof.
- Is the gap bigger than the noise floor? A re-run of the same model moved accuracy by 0.4 points and changed 1.2% of answers. One-point gaps are noise.
No single number in this page is decisive, and the leaderboards themselves keep saying so. The useful pattern is stable across every source: on raw decision accuracy, hosted Jev is at or near the top and the best open models are within a few points of it on decision-shaped items — but on genuinely unfamiliar domains the gap widens to roughly 26 points, and every open model brings a documented failure mode of its own. Choose on deployment story and cost, then measure on your own data, because the benchmarks can only tell you which models are worth testing.