Laya
Apache-2.0
NandhaKishorM · released ~18 Sep 2026
The flagship open project by a wide margin, and the only one you install with a single
pip install laya. A non-autoregressive decision engine on a ModernBERT-large
backbone, with a separate multilingual checkpoint and a Router that picks
between them by detecting the input script in under a millisecond.
- Base
- ModernBERT-large 421M (EN) + mmBERT-base 322M (100+ languages)
- Mechanism
- Single encoder pass over state + question + all options; a decision head scores each option into a distribution
- Serving
pip install laya; laya-serve exposes POST /v1/systemone
- Speed
- 32.8–39.5 ms per question on a T4; 7.2 ms with ten batched; ~66 ms median for three questions on an M3 Pro CPU
Read the caveat before adopting
Laya's README is unusually candid: its base checkpoints score 0.362 and
0.342 against a 0.318 random baseline and a
0.461 majority-class baseline. Its headline 0.766 — which beats Jev's
published 0.727 — comes from laya-typed-decisions, fine-tuned on that
benchmark's own training split. It is also weak on large option sets: 0.425 versus Jev's
0.870 on BANKING77's 77 labels, because options share a fixed 192–256 token budget.
Treat it as a fast base to fine-tune, not a drop-in replacement.
github.com/NandhaKishorM/laya →
Kev
Apache-2.0
Jared Palmer · 0.8B / 4B / 9B
The most complete open package: pretrained weights and training code and
evaluation data and honest model cards. Serves TypeSafe's own
/v1/systemone contract, so the official Python SDK works against it by
changing a base URL — no application code changes.
- Base
- Qwen3.5-0.8B / 4B / 9B-Base (all Apache-2.0)
- Mechanism
- Rank-16 LoRA plus a small pointer head that compares each option representation against the question's decision representation
- Serving
- CUDA, ROCm and Apple Silicon. 4B and 9B fit a 32 GB Mac in bf16. One request at a time; repeated state text is cached.
- Accuracy
- Held-out from trained sources 0.872; genuinely new sources up to 0.822 / 0.852 (9B). New-source Brier 0.237 (9B) vs Jev 0.211.
Training code included
HF Spaces demo
Chess & playground demos
github.com/jaredpalmer/kev →
Von Best open zero-shot
Apache-2.0
Victor Panisa (wfzyx) · 395M parameters
The open field's quiet leader on independent measurement, and by far its smallest serious
entry. Purpose-built System One model on a bidirectional ModernBERT backbone — it evaluates
the state and candidate criteria together using OptionMarker representations, then produces
logits for Choice, Noul and Score directly.
- Base
- 395M bidirectional ModernBERT
- Mechanism
- Non-autoregressive. Descriptive anchors are provided per option so a bidirectional attention head can match premise evidence against option semantics regardless of domain.
- Speed
- Sub-15 ms. Optional OpenVINO acceleration for Intel GPUs (~4× on an Iris Xe iGPU).
- Accuracy
- 0.704 macro accuracy on the independent jabr suite — the best open entrant. A fitted zero-shot Noul prior scores 85.1% on a held-out dev set.
Recorded failure mode
Collapses to a single mode on unfamiliar domains. Named for John von Neumann and Ludwig von Mises.
github.com/wfzyx/von →
SemIf formerly OpenJev
MIT
Theodore Lee (TheoLeeCJ) · reuses a model you already run
Trains nothing. Take a frozen open-weight model you already host, declare your
options, and read their probabilities straight off the next-token logits without sampling
a single answer token. It reproduces the pattern rather than a model.
- Base
- Frozen Qwen3.5-4B (and a 35B variant) — no fine-tuning
- Mechanism
- Causal backbone plus a tiny three-class entailment/contradiction/neutral classifier on the last token
- Speed
- 1.023 s versus 5.332 s for the same 21 decisions obtained via generated JSON — a ~5× win from skipping generation
- Requirements
- Python 3.10+, CUDA, a GPU that holds a 4B bf16 model. A browser WebGPU demo exists but is not a programmable endpoint.
JevBench v1.2 composite 74.6
JevK5 adapted its readout
github.com/TheoLeeCJ/SemIf →
Bespoke Nimble
Weights + recipe
Bespoke Labs · Qwen3.5-9B
The most reproducible project in the field. Bespoke Labs released the model, the training
recipe, the evaluation setup and the data-curation pipeline — which makes it the one to
study if you want to build your own rather than merely use one.
- Base
- Qwen3.5-9B with a rank-16 LoRA
- Mechanism
- Training and inference operate directly over candidate logits rather than decoding a response. Uses synthetic contrastive data curation plus constrained decoding.
- Accuracy
- Base Qwen improved from 66% to 90% on its curated eval, against Jev's 93%
- Speed
- ~100 ms reported on an H100
Availability limit
Needs an NVIDIA GPU of roughly 18 GB for the 9B weights and ships no public endpoint — so
the benchmark authors could not measure it at all. That is an access fact, not a quality
verdict.
github.com/bespokelabsai/nimble →
JevK5 1st open on JevBench v1.4
Apache-2.0
Alibi Serikbay (allebee) · 4B and 9B
The most rigorously documented open alternative, with pre-registered sweeps and McNemar
significance tests published alongside the results. Combines a distilled LoRA with SemIf's
logit-readout protocol and serves a TypeSafe-style endpoint.
- Base
- Qwen3.5-4B / 9B + a distilled LoRA, merged (weights Apache-2.0)
- Mechanism
- Softmax over the answer letters' next-token logits with one fitted temperature (1.22 on v0.3 4B). One CUDA graph per padded input length.
- Training
- Distilled from Qwen3.6-27B and GPT-6 Luna: 17,408 teacher-written questions plus 30,052 items replayed from 26 public train splits
- Speed
- 13 ms versus ~70 ms eager on an H100, identical answers
- Rank
- JevBench v1.4: 2nd of 76 systems, 1st among open entrants (62.04 vs Jev's 63.29). Sealed set: 33.1% correct vs Jev's 36.7%.
Known weak spots, from its own card
Multi-hop reasoning 0.11 and temporal/numeric 0.27 on the hard tier. State is truncated
past 512 tokens — which affected 52 of 111 hard items (median state 2,188 tokens), where
accuracy fell to 0.346 vs 0.458 elsewhere. English only. Sixteen options per pass. Refuses
inputs over 16,384 tokens.
github.com/allebee/jevk5 →
NanoJev
MIT
TianyuCodings · 0.6B
The specialist of the group, and the only project that clearly beats Jev at something. A
Qwen3-0.6B backbone with decision heads trained on four game environments, aimed squarely
at the real-time control-loop case behind TypeSafe's own Doom demo.
- Base
- Qwen3-0.6B + decision heads
- Result
- Beat Jev 128/128 to 56/128 on ViZDoom Basic
- Intended use
- Real-time control loops only
Not a general classifier
This is a reflex, not a judgment layer. Use it where you need a decision every few
milliseconds inside a tight loop — not for triage, routing or extraction.
github.com/TianyuCodings/NanoJev →
openjev-sglang Best open, JevBench v1.0
No licence file
Eric Zhang (ekzhang) · for B200-class hardware
Pushes the frozen-logit approach to a large mixture-of-experts model served on SGLang.
Delivers the highest open accuracy recorded on JevBench's first run by simply using a much
larger backbone than everything else.
- Base
- Qwen3.6-35B-A3B
- Serving
- SGLang, targeted at B200-class accelerator hardware
- Accuracy
- 95.5% (CI 92.8–97.8) — the best open rebuild in JevBench v1.0, statistically tied with DeepSeek V4.1 Flash and within the noise band of Jev itself
- Speed
- 0.68 s median — tied for fastest among complete runs, and indistinguishable from Jev's 0.65 s
Two caveats
No licence file was present in the repository when the benchmark authors checked
(2026-09-19); Qwen3.6 weights keep their own terms. And there was no per-token tariff on
the measured route, so the cost axis could not be scored — which is not the same as free.
github.com/ekzhang/openjev-sglang →
open-alternative-jev
Apache-2.0 repo
IkerMoel · on Hugging Face Spaces
A small, easily-tried reproduction running on an HF ZeroGPU Space. Included here mostly
because it produced the field's clearest single piece of evidence about how fragile these
models are to input formatting.
- Base
- Qwen3.5-4B (weights keep their own terms)
- Score
- JevBench v1.2 composite 69.8
⚠️ 72% → 21%
Ranked with the author's own option order (A. yes, B. no) it scored 72% on
answer-judging items. With the options reversed (A. no, B. yes) the
same model scored 21%. Both runs are published. This is the
single most important number on this page if you are evaluating open decision models: test
option order before you trust any accuracy figure, including your own.
github.com/ikermoel/open-alternative-jev →
system-one-open
MIT repo
mithalouni · Gemma-based
The strongest Gemma-based entry, and the reason it is worth noting is speed: it recorded
the fastest median latency of any complete run in JevBench v1.0 — faster than Jev itself,
measured from the same server.
- Base
- Gemma 4 E2B with a LoRA
- Hardware
- NVIDIA L4
- Accuracy
- 90.1% (CI 86.3–93.7) on JevBench v1.0
- Speed
- 0.65 s median, p95 0.77 s — tied for fastest
Licence is two-layered
Repository is MIT; the Gemma weights keep Google's own terms, which are not permissive in the same way.
github.com/mithalouni/system-one-open →
djev
Diffusion
Maisa · commercially hosted
The leading diffusion-based decision model, and a genuine architectural outlier — it does
not follow the encoder, LoRA or logit-readout logic at all. Available as a hosted endpoint
at api.djev.dev rather than as downloadable weights.
- Approach
- DiffusionGemma-based denoising rather than autoregressive or single-pass decoding
- Score
- JevBench v1.2 composite 74.3 — 4th overall, ahead of Laya and every other open-weight entry on that run
- Price
- Announced at $0.035 per million input tokens with output free — not yet charged during its free preview
- Measurement
- Ran all 534 decisions including held-out items through its production API, one request at a time, with no latency adjustment applied
JevBench adapter & results →
AnyJev
No training required
Nokia Applied Research
Corporate-research take on the same trick SemIf uses, packaged as a general utility: turn
any LLM into a Jev-style decision model, with typed decisions and real
probabilities, and no training at all. Built on transformers, with calibration and vLLM
integration in its scope.
- Mechanism
- Converts an existing model's token distribution over declared options into typed answers
- Requires
- An open-weight model you can read logits from
- Why it matters
- Institutional backing suggests the readout approach is not just a hobbyist pattern
github.com/nokia-applied-research/AnyJev →
open-jev-deberta-v3-large
Apache-2.0 card
Kotoba Labs · CPU-only
Runs entirely on CPU, which makes it the most deployable option here — and it produced the
most instructive benchmark result in the field, because its poor score was mostly not the
model's fault.
- Base
- DeBERTa-v3-large (weights keep their own terms)
- Deployment
- Local CPU, no GPU required
- Score
- 51.7% (CI 45.0–58.4) on JevBench v1.0
A context-limit result, not a judgment result
Its input window is shorter than several requests, so it failed to see the whole question
80 times out of 242. The benchmark authors explicitly flag that its
51.7% "is a context limit as much as a judgement one." Read as a ceiling on how much
state you can use, not on how well it decides.
github.com/kotoba-lang/typed-decisions →
jevlike
MIT
Vincent Wang-Maścianica (vinnylarouge) · 40 KB of embeddings
The most architecturally interesting entry, and the smallest by a factor of thousands. A
~40 kilobyte embedding option-attention model: each candidate becomes a query that reads
from a shared context representation and then receives a score.
- Mechanism
- Option-attention — candidates query a shared context representation rather than the context being scored per option
- What ships
- A trainer, not a pretrained general model. Bring labelled options and train an option-attention head on a laptop.
- Released checkpoints
- Doom and chess vision scorers only — there is no released general text-decision checkpoint, which is why the benchmark authors could not run it
github.com/vinnylarouge/jevlike →
mini-Jev
Smallest serving footprint
Author credit varies across directories · Qwen3-0.6B
A frozen Qwen3-0.6B with a 1.1 megabyte head attached, aimed at a single narrow job: tool
selection. The extreme end of the "how little do you actually need" question.
- Base
- Frozen Qwen3-0.6B + a ~1.1 MB head
- Task
- Tool selection
- Footprint
- ~9 GB disk for the weights, ~8.5 GB memory; Apple Silicon or CUDA
Could not be benchmarked
Over the benchmark authors' 4 GB memory bound and disk headroom.
Find it via the GitHub jev topic →
Laya Vision
Multimodal
r33drichards · 201M parameters
Typed decisions about an image plus text, in one forward pass — the first open
movement into the modality TypeSafe has said is next for Jev. At 201M parameters it is
smaller than the text-only Laya.
- Input
- Image + text state
- Output
- The same Choice / Score / Noul typed decisions
- Size
- 201M parameters
Directory entry →