Overview / Benchmarks

05 · Measurement

Benchmarks & leaderboards

When Jev launched there was no standard benchmark for the category at all — a gap practitioners flagged within days. Three now exist, measuring different things, and all three publish their own limitations. This page covers what each one actually tells you.

Context

Why this category was hard to measure

Conventional model evaluation assumes a text answer. A decision model refuses to produce one, which breaks almost every standard harness: there is no completion to BLEU-score, no reasoning trace to grade, and a "wrong" answer can still be a perfectly legal value from the schema you supplied.

Worse, the thing that matters most is not accuracy — it is whether stated confidence predicts accuracy. A benchmark that only measures top-1 correctness tells you nothing about whether you can safely branch on the probability, which is the entire reason to use a decision model.

The scepticism at launch was explicit and, in hindsight, correct. One widely-shared critique: many demos emphasised speed more than quality, and there was no standard benchmark for this category. The leaderboards below are the community's answer.

⚠️ Independent does not mean infallible

All three benchmarks here are built by small teams or individuals, not standards bodies. Two are essentially one-person projects. That does not make them wrong — they publish methodology, artefacts and limitations, which is more than most leaderboards do — but it does mean you should read the stated limitations rather than the ranking. Every one of them says something equivalent to "treat this as a pilot."

Overview

Three benchmarks, three different questions

Benchmark Built by Question it answers Scale Combined score?
JevBench Benchmark Heaven (Florian Standhartinger) — a one-person project, self-funded How do Jev-class decision models compare on accuracy, calibration, speed and cost, on the same items? 534 public + 308 sealed decisions per system; 93 systems registered by v1.4.2 Yes — since v1.2, a square-rooted geometric mean of four equally-weighted axes (a weak axis drags the score down hard). Earlier versions deliberately had no combined score.
Jev Decision Index @multimodalart, published on Hugging Face How broad is a decision model's competence across many task areas? 132,422 requests across 37 benchmarks; 30+ open-weight models Yes — 100 × the mean of five equally-weighted area scores
jabr classifier suite Independent (referenced by third-party analyses) How well does a general classifier handle unfamiliar domains, zero-shot? 49 tasks, 869 cases Macro accuracy across domains
📌 The three headline numbers, if you read nothing else
  • Open vs closed, out-of-domain: Jev 0.966 vs best open 0.704 macro accuracy on jabr. A 26-point gap.
  • Open vs closed, best case: JevK5 v0.2 placed 2nd of 76 systems on JevBench v1.4 and 1st among open entrants — 62.04 against Jev's 63.29.
  • Within noise: openjev-sglang (Qwen3.6-35B-A3B) scored 95.5% against Jev's 96.3% on JevBench v1.0, with overlapping confidence intervals.
The main event

JevBench, version by version

JevBench is the most-cited independent leaderboard for this category. It is not affiliated with or endorsed by TypeSafe — whose model is one of the systems it measures. Its harness and 72 original public decisions are MIT licensed. Because it has been revised several times in under two weeks, older numbers in circulation are often being compared against newer ones. Version matters.

v1.0Measured 19 Sept 2026 · five axes, deliberately no combined score

242 decisions per system — 72 published, 24 held out, 146 imported. One request at a time from a Hetzner server in Germany. The five axes were accuracy, cost, latency, validity of answers, and openness. The stated rationale for refusing a single score: "the trade-off between them is the point."

System Accuracy Calibration (ECE) Speed (median) Cost / 1k Type
GPT-5.6 Luna (low reasoning effort) 97.1% 0.016 verbalized 0.97 s $0.176 Closed
Jev 1.13.0 (TypeSafe AI) 96.3% 0.027 native 0.65 s $0.027 Closed
DeepSeek V4.1 Flash (thinking default) 95.5% 0.009 verbalized 1.42 s $0.245 Open weights
openjev-sglang (Qwen3.6-35B-A3B, SGLang) 95.5% 0.042 native 0.68 s no tariff Open
system-one-open (Gemma 4 E2B LoRA, L4) 90.1% 0.068 native 0.65 s no tariff Open
open-jev-deberta-v3-large (local CPU) 51.7% 0.147 native 1.77 s no tariff Open
Stopped early — shown with its own denominator, not ranked against complete runs
Qwen3.8 27B (Chutes TEE) 96.9% 0.014 verbalized 5.75 s no tariff Open weights
Accuracy figures carry 95% intervals; the top three overlap, so they should be read as a group rather than a podium. "no tariff" means there was no per-token price on the measured route — which the authors note "is not the same as free."

What v1.0 established

Jev is not uniquely good — it is uniquely cheap

GPT-5.6 Luna led on raw accuracy and DeepSeek V4.1 Flash led on calibration. Jev's decisive advantage was price: $0.027 per 1,000 decisions, the cheapest metered route measured, while sitting inside the noise band of the leaders on accuracy.

The open field is close on accuracy, far on cost

openjev-sglang matched Jev within intervals at 95.5% using a 35B mixture-of-experts model on specialised hardware. That is a real result — and also why "just self-host it" is not automatically cheaper than a hosted API at $0.042 per million tokens.

Latency differences are mostly indistinguishable

Three systems landed within 0.03 s of each other at the top. When your cheap decision model and your hosted API are 30 milliseconds apart, latency stops being a differentiator and price and calibration decide it.

v1.2 — the combined score, and what went into it

v1.2 replaced "no combined score" with a defined one, which makes the ranking much easier to read and much easier to over-read. The formula is published and worth understanding, because the geometric mean means a single bad axis destroys a score.

Intelligence
100 × weighted accuracy, with the hard tier weighted 30%, standard 28%, judge 28% and easy 14%
Calibration
On the hard tier: ECE plus fidelity to the exact gold distribution. Systems that emit only a label score zero here, because they have no distribution to be faithful with.
Speed
Mean of the p50 and p95 scores, where a score of 100 − 20·log₁₀(s / 0.1 s). So 0.1 s scores 100, and every 10× slower costs 20 points.
Cost
100 − 30·log₁₀($ per 1,000 decisions / $0.001). $0.001 scores 100; every 10× more expensive costs 30 points.
Combination
Geometric mean of the four, each weighted 25%. A weak axis is not averaged away — it caps the score.
Item set
534 decisions per system: 220 hard (111 public, 109 held out) plus easy, standard and judge tiers. The hard tier was written by Claude Opus 5 and GPT-5.6 Sol, cross-reviewed, then frozen and hashed before any system ran.
JevBench v1.2 leaderboard, as published on benchmarkheaven.com. Note the first row: classifier.dev's "fast tier" is Jev behind a different API — it matches Jev on intelligence and leads on price and measured latency, which is why it ranks first. That is a hosting result, not a model result.
# System JevBench Score What it is
1 classifier.dev (fast tier) 84.8 Jev 1.13.0 behind a third-party API. Leads on its flat pricing plan at full utilisation and on latency measured from the benchmark server.
2 Jev 1.13.0 75.3 TypeSafe AI, proprietary API
3 SemIf (Qwen3.5-4B) 74.6 Open — frozen-backbone logit readout, no training
4 djev (Maisa) 74.3 Diffusion-based, served through a production API
5 Laya (421M) 70.1 Open — ModernBERT encoder plus decision head
6 open-alternative-jev (Qwen3.5-4B) 69.8 Open — see the option-order warning below
❌ The most important caveat in the whole benchmark

open-alternative-jev is ranked using the author's own option order (A. yes, B. no). With the options reversed (A. no, B. yes) the same model scored 21% instead of 72% on answer-judging items. Both runs are published. JevBench's authors flag one more thing about this row: it was the model whose self-reported confidence behaviour made it the clearest example of how much a small model's accuracy depends on formatting nobody thinks to vary.

v1.4 — sealed items and a harmonic mean

By v1.4 (25 September 2026) the suite had grown to 93 systems registered, 89 ranked, scored on 534 public decisions — 72 easy, 96 standard, 146 judge, 220 hard — plus 308 sealed decisions whose text stays private. Sealed items make up 20% of the intelligence axis precisely to make overfitting to the public set costly.

The composite also changed from a geometric to a harmonic mean of the four equally-weighted axes, which punishes a single weak axis even harder. In that run an open 4B model — decider-4b v2 — finished first on composite, 0.8 points ahead of Jev. Jev remained the most accurate system in the top ten. Both statements are true at once, and that is the point of a multi-axis score.

JevK5's placement, as its own card reports it

JevBench v1.4 ranked JevK5 v0.2 2nd of 76 systems and 1st among open entrants (62.04 against Jev's 63.29). On the 308 sealed decisions, JevK5 answered 33.1% correctly against Jev's 36.7% — the evaluator notes this sealed set is unusually difficult. The project is explicit that the ranking measures JevBench's own mix of accuracy, calibration, speed and cost, "not performance on every production workflow."

⚠️ Three versions in circulation

Because v1.0, v1.2 and v1.4 have different item sets, different scoring formulas and different cohorts, a number quoted without a version is not comparable to anything. The v1.0 archive is still published alongside the current board, which is good practice — but it means you will find conflicting figures for the same model. Ask for the version.

⚠️ The latency adjustment — stated as an assumption, not a measurement

JevBench runs self-hosted and demo endpoints one request at a time with no other load, then multiplies their latency by 2 and adds 0.15 s "to approximate production load." The authors are explicit that this is an assumption, not a measurement, that it stands for infrastructure their tests lacked (authentication, load balancing, logging, billing, API gateway), and that raw p50/p95 are published so you can undo it. They also note that serving under load trades per-user speed for throughput by far more than 2× — citing an NVIDIA analysis — but concede that analysis used a 1.8T mixture-of-experts model on GPU clusters rather than a 4B model on one GPU, so it supports the direction rather than the exact factor.

How much of a gap is noise? They measured that too.

Jev answered the same 242 decisions twice, about 16 minutes apart. Three answers changed — 1.2% of the suite, all in routing — and accuracy moved from 96.7% to 96.3%. Their guidance: read a gap of about one point between two rows as noise, and use the confidence intervals. This is unusually rigorous, and it immediately deflates a lot of leaderboard arguments.

Breadth

The Jev Decision Index

Where JevBench asks "which is best on decision-shaped items", the Jev Decision Index asks how broad a model's competence is. It is published as a Hugging Face Space by @multimodalart.

Scale
132,422 requests across 37 benchmarks
Scored panel
19 benchmarks, grouped into five equal-weight areas
Metric
100 × the mean of the five area scores, on a 0–100 scale
Cohort
30+ open-weight decision models, plus Jev as the reference
Submission
Via pull request with a dataset link and hardware specifications; entries re-scored during review
The five areas, equally weighted
  1. Tools & Automation
  2. Retrieval & Classification
  3. Language Understanding
  4. Knowledge & Reasoning
  5. Arts & Human Judgment

The house rules are the interesting part

A leaderboard is only as good as what it forbids, because those are the degrees of freedom a submitter would otherwise use to flatter their model. This one bans three:

  • No prompt tuning. You cannot tune the prompt per benchmark item.
  • No truncation. You cannot silently cut the input to fit your window — which is exactly the trick that made open-jev-deberta-v3-large's 51.7% in JevBench v1.0 look like a capability result rather than a context limit.
  • No option filtering. You cannot discard candidates to raise your hit rate.
  • Abstained and errored requests score zero, before chance normalisation.
A worked reference point

JevK5 0.2.2 (on v0.2 weights) scores 36.31 — 15th of 49 submitted systems, with the 4th-best calibration. That combination is worth noting: a model can be mid-table on breadth while being near the top on the property that matters most for automation. Composite rank alone would hide it.

The reality check

The jabr classifier suite

A much smaller and much harsher test: 49 tasks and 869 cases across compliance, triage, legal, DevOps, linguistics and safety, run out-of-domain and zero-shot. It is the source of the number that should temper every "drop-in replacement" claim.

Macro accuracy across the suite. Note the failure mode recorded for each system — the benchmark authors published one for every entrant, which is the rare part.
System Macro accuracy Recorded failure mode
Jev (hosted, TypeSafe) 0.966 —
Von (395M, Apache-2.0) 0.704 Collapses to one mode on unfamiliar domains
GLiNER2 (~300M, open) 0.698 Over-triggers on keywords
Laya (421M, open) 0.583 Compresses rating scales
⚠️ Read the gap, then read what it means

A 26-point gap out-of-domain, zero-shot. But note the qualifiers, because they are load bearing: in-domain, after a fine-tune on your own labels, the open models are competitive. What this benchmark measures is how much of the work you have to do yourself. With hosted Jev you are buying a model that already generalises. With an open model you are buying a base you will specialise. Both are legitimate; they are different products at different prices, and the 0.966 vs 0.704 comparison is the price of the difference.

For contrast

Vendor-run evals

Both TypeSafe and several open projects publish their own evaluations. These are not worthless — they are often the only detailed numbers available — but they answer a different question: "does this model do well on the things its authors chose to test?"

Who Method Their own stated limitation
TypeSafe — workflow evals Assume a correct compute graph ("workflow") in code; compare each model's predictions against the average of the two strongest external models as reference probabilities, rather than a ground-truth classification Workflows were written by their own model-capabilities team; the reference average biases toward OpenAI and Anthropic, so they "likely underestimate" themselves
Kev — model cards Develop and test splits, separated into trained sources (held-out examples from training datasets) and new sources (unseen datasets and policy rule types); reports accuracy and Brier Small models are slow on Apple Silicon; option order can change answers; training used at most 384 state tokens, so longer contexts were never trained on
JevK5 Runs JevBench's own harness, reports McNemar significance tests on item-level changes across versions, and pre-registered its epoch sweep Most gains are on held-out and index-style data; on JevBench's public items the 4B gained on the hard tier while losing one standard item. Quality on Jev's real-world workflows "has not yet been measured."
Bespoke Nimble Curated evaluation against a controlled baseline — base Qwen improved from 66% to 90% The eval is their own, on their own curated data, with the comparison point being Jev at 93% from a different measurement
Method

How to read a leaderboard in this category

A checklist distilled from the caveats the benchmark authors publish about themselves. Run it against any number you are about to quote.

  • Which version? JevBench alone has three cohorts in circulation with different item sets and different formulas. A number without a version is not comparable.
  • Native or verbalized probabilities? These are different objects and must never be pooled. Native distributions sum to 1 by construction; written ones are prose that needs renormalising.
  • Was the benchmark's own training split used? One headline score that beat Jev came from a checkpoint fine-tuned on that benchmark's training split.
  • Is the model a label-only system? It scores zero on calibration by definition, which drags a geometric or harmonic mean hard.
  • How long was the input allowed to be? A truncation limit can masquerade as a capability limit — one system never saw the whole question 80 times out of 242.
  • Was option order varied? 72% became 21% on the same model with the options reversed.
  • Is the entry actually a different model? The top row of one leaderboard is Jev behind a different API — a hosting result, not a capability result.
  • Were the high scores on a vendor-designed format? Speed and cost are directly checkable; accuracy on a self-invented rubric is not.
  • Is the latency number a measurement or an assumption? One leaderboard multiplies self-hosted latency by 2 and adds 0.15 s, and says so.
  • Are exclusions reported? A benchmark that quietly drops what it could not run cannot be checked. "No GPU available" is an availability fact, never a quality verdict.
  • Was the held-out set actually unseen? Sending held-out items to the service to get predictions means "not public" is not the same as "not seen" — and this is not a contamination proof.
  • Is the gap bigger than the noise floor? A re-run of the same model moved accuracy by 0.4 points and changed 1.2% of answers. One-point gaps are noise.
✅ The bottom line

No single number in this page is decisive, and the leaderboards themselves keep saying so. The useful pattern is stable across every source: on raw decision accuracy, hosted Jev is at or near the top and the best open models are within a few points of it on decision-shaped items — but on genuinely unfamiliar domains the gap widens to roughly 26 points, and every open model brings a documented failure mode of its own. Choose on deployment story and cost, then measure on your own data, because the benchmarks can only tell you which models are worth testing.