Overview / Open models

04 · The open-weight field

Open-weight decision models

Within about 24 hours of Jev's launch, the open-weight community began publishing models with the same interface. Two weeks later there are dozens. This page catalogues the ones with real substance, grouped by how they work — and states plainly where they still fall short.

Before you browse

Three things that make this catalog different from a leaderboard

⚠️ Published accuracy is not comparable across projects

Each project reports numbers from its own evaluation suite, on its own item mix, against its own baseline. A model claiming 0.9 and a model claiming 0.7 may be measuring entirely different things. Only the independent benchmarks on the Benchmarks page put systems on the same items.

🚫 Nothing here is ranked by popularity

GitHub stars in this field were moving by thousands per day at time of writing, and sources disagreed with each other about the same repository on the same day. The most- starred project is not the most accurate one — on the independent jabr suite, the highest-starred project sits well below a 723-star one. Popularity tracks attention, not quality.

📄 Licence ≠ licence

A permissive repository licence is not a licence for the base weights. Several projects are Apache-2.0 or MIT while their Qwen or Gemma backbone keeps its own terms — and at least one prominent project (openjev-sglang) had no licence file at all when the benchmark authors checked. Verify both layers before commercial use.

❌ The gap that defines the state of the art

On the independent jabr classifier suite — 49 tasks, 869 cases across compliance, triage, legal, DevOps, linguistics and safety, run out-of-domain and zero-shot — hosted Jev scores 0.966 macro accuracy and the best open entrant scores 0.704. That is a 26-point gap on unfamiliar domains. The honest framing from the people who did the comparison: the open field "reproduces Jev's interface well and its accuracy only partly", and the projects "win on latency, price and control, not out-of-the-box accuracy." In-domain, after fine-tuning on your own labels, they are competitive. Out of domain, they are not close yet.

The catalog

Sixteen projects worth your attention

Laya

Apache-2.0
NandhaKishorM · released ~18 Sep 2026

The flagship open project by a wide margin, and the only one you install with a single pip install laya. A non-autoregressive decision engine on a ModernBERT-large backbone, with a separate multilingual checkpoint and a Router that picks between them by detecting the input script in under a millisecond.

Base
ModernBERT-large 421M (EN) + mmBERT-base 322M (100+ languages)
Mechanism
Single encoder pass over state + question + all options; a decision head scores each option into a distribution
Serving
pip install laya; laya-serve exposes POST /v1/systemone
Speed
32.8–39.5 ms per question on a T4; 7.2 ms with ten batched; ~66 ms median for three questions on an M3 Pro CPU
Read the caveat before adopting

Laya's README is unusually candid: its base checkpoints score 0.362 and 0.342 against a 0.318 random baseline and a 0.461 majority-class baseline. Its headline 0.766 — which beats Jev's published 0.727 — comes from laya-typed-decisions, fine-tuned on that benchmark's own training split. It is also weak on large option sets: 0.425 versus Jev's 0.870 on BANKING77's 77 labels, because options share a fixed 192–256 token budget. Treat it as a fast base to fine-tune, not a drop-in replacement.

Kev

Apache-2.0
Jared Palmer · 0.8B / 4B / 9B

The most complete open package: pretrained weights and training code and evaluation data and honest model cards. Serves TypeSafe's own /v1/systemone contract, so the official Python SDK works against it by changing a base URL — no application code changes.

Base
Qwen3.5-0.8B / 4B / 9B-Base (all Apache-2.0)
Mechanism
Rank-16 LoRA plus a small pointer head that compares each option representation against the question's decision representation
Serving
CUDA, ROCm and Apple Silicon. 4B and 9B fit a 32 GB Mac in bf16. One request at a time; repeated state text is cached.
Accuracy
Held-out from trained sources 0.872; genuinely new sources up to 0.822 / 0.852 (9B). New-source Brier 0.237 (9B) vs Jev 0.211.
Training code included HF Spaces demo Chess & playground demos

Von Best open zero-shot

Apache-2.0
Victor Panisa (wfzyx) · 395M parameters

The open field's quiet leader on independent measurement, and by far its smallest serious entry. Purpose-built System One model on a bidirectional ModernBERT backbone — it evaluates the state and candidate criteria together using OptionMarker representations, then produces logits for Choice, Noul and Score directly.

Base
395M bidirectional ModernBERT
Mechanism
Non-autoregressive. Descriptive anchors are provided per option so a bidirectional attention head can match premise evidence against option semantics regardless of domain.
Speed
Sub-15 ms. Optional OpenVINO acceleration for Intel GPUs (~4× on an Iris Xe iGPU).
Accuracy
0.704 macro accuracy on the independent jabr suite — the best open entrant. A fitted zero-shot Noul prior scores 85.1% on a held-out dev set.
Recorded failure mode

Collapses to a single mode on unfamiliar domains. Named for John von Neumann and Ludwig von Mises.

SemIf formerly OpenJev

MIT
Theodore Lee (TheoLeeCJ) · reuses a model you already run

Trains nothing. Take a frozen open-weight model you already host, declare your options, and read their probabilities straight off the next-token logits without sampling a single answer token. It reproduces the pattern rather than a model.

Base
Frozen Qwen3.5-4B (and a 35B variant) — no fine-tuning
Mechanism
Causal backbone plus a tiny three-class entailment/contradiction/neutral classifier on the last token
Speed
1.023 s versus 5.332 s for the same 21 decisions obtained via generated JSON — a ~5× win from skipping generation
Requirements
Python 3.10+, CUDA, a GPU that holds a 4B bf16 model. A browser WebGPU demo exists but is not a programmable endpoint.
JevBench v1.2 composite 74.6 JevK5 adapted its readout

Bespoke Nimble

Weights + recipe
Bespoke Labs · Qwen3.5-9B

The most reproducible project in the field. Bespoke Labs released the model, the training recipe, the evaluation setup and the data-curation pipeline — which makes it the one to study if you want to build your own rather than merely use one.

Base
Qwen3.5-9B with a rank-16 LoRA
Mechanism
Training and inference operate directly over candidate logits rather than decoding a response. Uses synthetic contrastive data curation plus constrained decoding.
Accuracy
Base Qwen improved from 66% to 90% on its curated eval, against Jev's 93%
Speed
~100 ms reported on an H100
Availability limit

Needs an NVIDIA GPU of roughly 18 GB for the 9B weights and ships no public endpoint — so the benchmark authors could not measure it at all. That is an access fact, not a quality verdict.

JevK5 1st open on JevBench v1.4

Apache-2.0
Alibi Serikbay (allebee) · 4B and 9B

The most rigorously documented open alternative, with pre-registered sweeps and McNemar significance tests published alongside the results. Combines a distilled LoRA with SemIf's logit-readout protocol and serves a TypeSafe-style endpoint.

Base
Qwen3.5-4B / 9B + a distilled LoRA, merged (weights Apache-2.0)
Mechanism
Softmax over the answer letters' next-token logits with one fitted temperature (1.22 on v0.3 4B). One CUDA graph per padded input length.
Training
Distilled from Qwen3.6-27B and GPT-6 Luna: 17,408 teacher-written questions plus 30,052 items replayed from 26 public train splits
Speed
13 ms versus ~70 ms eager on an H100, identical answers
Rank
JevBench v1.4: 2nd of 76 systems, 1st among open entrants (62.04 vs Jev's 63.29). Sealed set: 33.1% correct vs Jev's 36.7%.
Known weak spots, from its own card

Multi-hop reasoning 0.11 and temporal/numeric 0.27 on the hard tier. State is truncated past 512 tokens — which affected 52 of 111 hard items (median state 2,188 tokens), where accuracy fell to 0.346 vs 0.458 elsewhere. English only. Sixteen options per pass. Refuses inputs over 16,384 tokens.

NanoJev

MIT
TianyuCodings · 0.6B

The specialist of the group, and the only project that clearly beats Jev at something. A Qwen3-0.6B backbone with decision heads trained on four game environments, aimed squarely at the real-time control-loop case behind TypeSafe's own Doom demo.

Base
Qwen3-0.6B + decision heads
Result
Beat Jev 128/128 to 56/128 on ViZDoom Basic
Intended use
Real-time control loops only
Not a general classifier

This is a reflex, not a judgment layer. Use it where you need a decision every few milliseconds inside a tight loop — not for triage, routing or extraction.

openjev-sglang Best open, JevBench v1.0

No licence file
Eric Zhang (ekzhang) · for B200-class hardware

Pushes the frozen-logit approach to a large mixture-of-experts model served on SGLang. Delivers the highest open accuracy recorded on JevBench's first run by simply using a much larger backbone than everything else.

Base
Qwen3.6-35B-A3B
Serving
SGLang, targeted at B200-class accelerator hardware
Accuracy
95.5% (CI 92.8–97.8) — the best open rebuild in JevBench v1.0, statistically tied with DeepSeek V4.1 Flash and within the noise band of Jev itself
Speed
0.68 s median — tied for fastest among complete runs, and indistinguishable from Jev's 0.65 s
Two caveats

No licence file was present in the repository when the benchmark authors checked (2026-09-19); Qwen3.6 weights keep their own terms. And there was no per-token tariff on the measured route, so the cost axis could not be scored — which is not the same as free.

open-alternative-jev

Apache-2.0 repo
IkerMoel · on Hugging Face Spaces

A small, easily-tried reproduction running on an HF ZeroGPU Space. Included here mostly because it produced the field's clearest single piece of evidence about how fragile these models are to input formatting.

Base
Qwen3.5-4B (weights keep their own terms)
Score
JevBench v1.2 composite 69.8
⚠️ 72% → 21%

Ranked with the author's own option order (A. yes, B. no) it scored 72% on answer-judging items. With the options reversed (A. no, B. yes) the same model scored 21%. Both runs are published. This is the single most important number on this page if you are evaluating open decision models: test option order before you trust any accuracy figure, including your own.

system-one-open

MIT repo
mithalouni · Gemma-based

The strongest Gemma-based entry, and the reason it is worth noting is speed: it recorded the fastest median latency of any complete run in JevBench v1.0 — faster than Jev itself, measured from the same server.

Base
Gemma 4 E2B with a LoRA
Hardware
NVIDIA L4
Accuracy
90.1% (CI 86.3–93.7) on JevBench v1.0
Speed
0.65 s median, p95 0.77 s — tied for fastest
Licence is two-layered

Repository is MIT; the Gemma weights keep Google's own terms, which are not permissive in the same way.

djev

Diffusion
Maisa · commercially hosted

The leading diffusion-based decision model, and a genuine architectural outlier — it does not follow the encoder, LoRA or logit-readout logic at all. Available as a hosted endpoint at api.djev.dev rather than as downloadable weights.

Approach
DiffusionGemma-based denoising rather than autoregressive or single-pass decoding
Score
JevBench v1.2 composite 74.3 — 4th overall, ahead of Laya and every other open-weight entry on that run
Price
Announced at $0.035 per million input tokens with output free — not yet charged during its free preview
Measurement
Ran all 534 decisions including held-out items through its production API, one request at a time, with no latency adjustment applied

AnyJev

No training required
Nokia Applied Research

Corporate-research take on the same trick SemIf uses, packaged as a general utility: turn any LLM into a Jev-style decision model, with typed decisions and real probabilities, and no training at all. Built on transformers, with calibration and vLLM integration in its scope.

Mechanism
Converts an existing model's token distribution over declared options into typed answers
Requires
An open-weight model you can read logits from
Why it matters
Institutional backing suggests the readout approach is not just a hobbyist pattern

open-jev-deberta-v3-large

Apache-2.0 card
Kotoba Labs · CPU-only

Runs entirely on CPU, which makes it the most deployable option here — and it produced the most instructive benchmark result in the field, because its poor score was mostly not the model's fault.

Base
DeBERTa-v3-large (weights keep their own terms)
Deployment
Local CPU, no GPU required
Score
51.7% (CI 45.0–58.4) on JevBench v1.0
A context-limit result, not a judgment result

Its input window is shorter than several requests, so it failed to see the whole question 80 times out of 242. The benchmark authors explicitly flag that its 51.7% "is a context limit as much as a judgement one." Read as a ceiling on how much state you can use, not on how well it decides.

jevlike

MIT
Vincent Wang-Maścianica (vinnylarouge) · 40 KB of embeddings

The most architecturally interesting entry, and the smallest by a factor of thousands. A ~40 kilobyte embedding option-attention model: each candidate becomes a query that reads from a shared context representation and then receives a score.

Mechanism
Option-attention — candidates query a shared context representation rather than the context being scored per option
What ships
A trainer, not a pretrained general model. Bring labelled options and train an option-attention head on a laptop.
Released checkpoints
Doom and chess vision scorers only — there is no released general text-decision checkpoint, which is why the benchmark authors could not run it

mini-Jev

Smallest serving footprint
Author credit varies across directories · Qwen3-0.6B

A frozen Qwen3-0.6B with a 1.1 megabyte head attached, aimed at a single narrow job: tool selection. The extreme end of the "how little do you actually need" question.

Base
Frozen Qwen3-0.6B + a ~1.1 MB head
Task
Tool selection
Footprint
~9 GB disk for the weights, ~8.5 GB memory; Apple Silicon or CUDA
Could not be benchmarked

Over the benchmark authors' 4 GB memory bound and disk headroom.

Laya Vision

Multimodal
r33drichards · 201M parameters

Typed decisions about an image plus text, in one forward pass — the first open movement into the modality TypeSafe has said is next for Jev. At 201M parameters it is smaller than the text-only Laya.

Input
Image + text state
Output
The same Choice / Score / Noul typed decisions
Size
201M parameters
📇 The long tail

There are far more than sixteen. Community directories tracked 146 open-source builds in one catalogue and more than 1,858 repositories tagged jev on GitHub at time of writing. Many are demos, ports or one-day experiments. The table below lists the ones that recur across independent sources but did not warrant a full card.

Not independently benchmarked unless noted. Descriptions are as published by each project's own repository or directory listing.
Project What it is Family
Kev-0.5B The original tiny Kev — a LoRA adapter plus a small readout head on Qwen2.5-0.5B, small enough to run on a laptop. Superseded by the 0.8B/4B/9B family. Adapter
Tiny-Jev A 0.6B System One model: structured state in, typed decisions out. Adapter
jev-lite A QLoRA adapter that turns Gemma 4 E4B into a System One model. Adapter
open-spark-jev Local System One models on Qwen3, packaged for an NVIDIA DGX Spark. Adapter
Eikos-27B An open typed-decision model specialised for finance and trading rules. Adapter
Jeff 1 An open typed-decision model that ships with a guide to training your own. Adapter
arbiter Serve your own typed-decision model behind a Jev-shaped API. Serving
visual-jev Typed decisions over images directly, rather than over text descriptions of images. Vision
Parallel Constrained Decoding An RLCD-trained Qwen2.5-1B demo exploring open-source parallel constrained decoding as an alternative to Jev. Research
TypeLLM LLMs with type-safe generation — the types-first framing applied more broadly than decision questions. Research
open-jev (DiffusionGemma 26B-A4B) A diffusion-based build measured on an H100 through the author's own Modal account. No public endpoint. Diffusion
classifier.dev (fast tier) Top of JevBench v1.2 at 84.8 — but its fast tier is Jev behind another API, so it leads on price and measured latency rather than on being a better model. Not an open model. Hosted
Decisions

How to choose

Sorting by use case rather than by quality, because none of these wins on every axis.

If your situation is… Reach for Because
You already call hosted Jev and want to self-host with no code change Kev It implements TypeSafe's own /v1/systemone wire format, so the official SDK works by changing a base URL.
You already run an open-weight LLM you like SemIf or AnyJev Read probabilities off the logits of what you already host. Zero training cost, but you inherit that model's biases wholesale.
You need CPU-only or edge deployment, and will fine-tune Laya or Von 322M–421M parameters, milliseconds locally, no GPU in the serving path. Laya has 100+ language coverage. Fine-tune before trusting either.
You need the best open accuracy and have GPU budget JevK5-4B or openjev-sglang JevK5 is first among open entrants on JevBench v1.4; openjev-sglang matched Jev within noise on accuracy if you have B200-class hardware.
You need a decision every few milliseconds inside a control loop NanoJev The only project here with published evidence at that latency, and it beats Jev on its game environment.
Your label set is unusual enough that no pretrained checkpoint helps jevlike It ships a trainer rather than a model.
You need the most reproducible recipe to study Bespoke Nimble Weights, training recipe, evaluation setup and data-curation pipeline all published.
You need decisions about images Laya Vision or visual-jev The only open multimodal options.
✅ The advice every serious write-up converges on

Pick the model whose deployment story matches yours; fine-tune it on a few hundred of your own labelled decisions; refit the confidence temperature; test option order and input truncation explicitly; then shadow real traffic against your current solution and compare. Treat every published number — vendors' included — as a hypothesis rather than a result. These projects are barely two weeks old and move fast enough that a figure goes stale in a day.