Overview / Practice

06 · In practice

Putting it in a stack

The pattern that makes this class useful is simple and it is not "replace your LLM." It is: the LLM proposes, the decision model judges, your code executes. This page covers where that pattern pays off, with working code — and the design rules that keep it from quietly corrupting your system.

Architecture

The core loop

Two brains, deliberately different. One is slow, general and expensive and handles open-ended reasoning and generation. The other is fast, narrow and nearly free and handles the small judgments that hold the system together.

1 · ProposeThe LLM does the open-ended part
reason→ draft→ plan→ propose N candidate actions / labels / tool calls

Everything that needs generation, world knowledge or multi-step reasoning stays here. Slow and expensive, but that is what it is for.

2 · DecideThe decision model judges
state + declared options→ picked option full distribution confidence

Routing, classification, scoring, verification, safety gating. Milliseconds and fractions of a cent, so it can run on every request.

3 · ExecuteYour ordinary code branches
if confidence > threshold→ act | else→ escalate to human or bigger model

The surrounding code constrains the model's freedom, which is what makes the whole thing composable into reliable systems.

The framing worth stealing from LangChain

Agents run in a loop: a model decides what to do, a tool executes, a model evaluates the result. Tool calling and structured outputs made that loop possible — but "every decision requires another model call," and that is where the cost and latency come from. Replacing the decision step with something that costs a fraction of a cent and returns a calibrated probability is the intervention. The generation stays on the big model; the decisions move off it.

Applications

Where it pays off

Primitive choices are per TypeSafe's documentation. The practitioner examples are individually reported and not independently benchmarked.
Job Primitive Why a decision model Reported example
Classification & triage Choice You already know the label set. The task is a lookup that has been paying generation prices. Bryo AI: classifying business emails — Gemini was slightly more accurate but 10–20× more expensive
Model routing Choice Deciding "does this need the big model?" with an LLM is too expensive to do on every request, so most systems just don't bother. At this price you can. LangChain ships a ModelRouterMiddleware that picks a model per request from a declared set of criteria
Agent guardrails & tool gating Noul Risk classification has to run on every tool call, so cost and latency are the binding constraint, not capability. Vercel replaced an OpenAI classifier reviewing commands for safety: 5–18× faster with greater accuracy
Verified actions before execution Noul / Choice A hallucinated tool call is inconvenient in an agent and a deal-breaker several layers deep in a dependency chain. LangChain's AutoModeMiddleware blocks risky tool calls before the tool runs
Search & retrieval reranking Score Fetch 100 cheap candidates with an ordinary algorithm, then score all 100 in one call. The questions cost nothing extra. Simon Willison: BM25 top-100, then score each candidate for relevance against the query
Output & trace verification Noul Score, judge and jailbreak-detect LLM prompts, reasoning traces and outputs — on every call rather than by sampling. TypeSafe leads its own use-case list with verification; monitoring agents with agents is otherwise too expensive to justify
Real-time interactive applications Choice / Score Sub-100 ms means it can sit inside a user-facing interaction rather than behind a spinner. TypeSafe's Doom demo: ~10 decisions/second, about $7/hour. Browser-use agents reported "for fractions of a cent."
Map-reduce over large corpora Noul The cost floor drops far enough that tagging millions of rows becomes a question of engineering rather than budget. "Tag millions of rows without a big bill"
And the case people keep landing on first

Browser and computer-use control loops. LangChain's CEO called browser use the best application he had seen for Jev, and Browserbase was reported powering browser-use agents at a fraction of a cent. The reason is structural: an agent driving a browser makes a great many small bounded decisions — is this the right element, is this page loaded, is this action safe — and every one of them is decision-shaped.

Code

Working code

All four examples below are adapted from published repositories. The first three target a local Kev server, which implements TypeSafe's own /v1/systemone contract — so the same code works against hosted Jev by changing base_url and the API key.

Start a local server (Kev-4B)
git clone https://github.com/jaredpalmer/kev.git
cd kev

uv sync --extra serve

# Kev-4B locally; first run downloads the
# adapter and the Qwen3.5 base model.
KEV_DTYPE=bf16 uv run --extra serve \
  python -m kev.serve \
  --run jaredpalmer/kev-4b \
  --port 8009

Kev needs transformers >= 5.17. The 4B and 9B models fit a 32 GB Mac in bf16. The server handles one request at a time and caches repeated state text, but does not batch across callers — put your own queue in front of it.

Raw HTTP — the whole interface
curl -sS -X POST \
  http://127.0.0.1:8009/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "model": "kev-latest",
    "state": "Shoes arrived two weeks late and in
              the wrong size. Also I see two
              charges on my card.",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "returns":  "Exchanges, refunds, wrong items",
          "shipping": "Delivery status, delays",
          "billing":  "Charges, invoices, payments"
        }
      },
      "escalate": {
        "type": "noul",
        "instructions": "Does this need urgent human attention?"
      }
    }
  }'
Python SDK — three primitives, one call
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

# Point the official SDK at a local server.
# For hosted Jev: drop base_url, use your API key.
client = TypeSafeClient(
    api_key="local",
    base_url="http://127.0.0.1:8009",
    model="kev-latest",
)

response = client.system_one(
    state="I was charged twice. Please fix this ASAP.",
    questions={
        "billing": Noul(
            instructions="Is this ticket about billing?"
        ),
        "tone": Choice(
            instructions="What is the customer's tone?",
            criteria={
                "calm": None,
                "frustrated": None,
                "angry": None,
            },
        ),
        "urgency": Score(
            instructions="How urgent is this ticket?",
            criteria=["can wait", "this week", "today"],
        ),
    },
)

print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)
Note what is not here

No prompt string. No JSON parsing. No retry-on-malformed-output loop. No schema validation. The declaration is the prompt, and there is nothing to repair.

The payoff

The confidence gate

This is the pattern that decides whether you get any value from the confidence numbers. Without it you have an expensive classifier; with it you have a system that knows when to stop.

Gate on both confidence and the winning probability
# Two different questions:
#  - confidence: how clear-cut was the decision?
#  - probabilities[k]: how strong is the winner?

LOW_CONFIDENCE = 0.60   # tune on YOUR data
WEAK_WINNER    = 0.55   # tune on YOUR data

answer = response.choices["department"]

if answer.confidence < LOW_CONFIDENCE:
    # The model itself says this was a close call.
    escalate(ticket, reason="low confidence",
             distribution=answer.probabilities)

elif answer.probabilities[answer.choice] < WEAK_WINNER:
    # It picked something, but nothing dominates.
    escalate(ticket, reason="no dominant option",
             distribution=answer.probabilities)

else:
    dispatch(ticket, queue=answer.choice)

# Log the full distribution, not just the pick.
# This is your training data for the next iteration.
  • Don't use a single global threshold. The right cut-off depends on the cost of a mistake in that queue. A billing misroute is cheap; a medical triage misroute is not.
  • Gate on two things. A high confidence with a 0.34 winning probability means something different from a high confidence with 0.95. The distribution is the honest part of the answer.
  • Log the whole distribution, always. The pick is what your code uses today; the distribution is what lets you re-tune tomorrow without re-inferring.
  • Make escalation cheap. If the fallback is a full human review queue, the gate will get disabled the first busy week. Route to a smaller model or a rules path where you can.
  • Instrument the escalation rate. A rate that climbs over time is the earliest signal that your domain has drifted away from the data the model was tuned on.
⚠️ Remember what confidence does not mean

TypeSafe: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." A threshold of 0.9 does not mean this call is safe — it means calls like this are wrong about one time in ten. Budget for that rather than assuming it away.

Integrations

Guardrails and routing in an agent

LangChain ships first-party support, which is the easiest way to see the pattern inside a real agent loop rather than in a standalone script. Two middlewares are notable.

Model routing — spend the least that works
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
    ModelChoice,
    ModelRouterMiddleware,
)

router = ModelRouterMiddleware(
    choices={
        "fast": ModelChoice(
            model="openai:luna",
            criteria="Direct lookups, extraction, and localized changes.",
        ),
        "powerful": ModelChoice(
            model="openai:sol",
            criteria="Architecture and high-stakes decisions.",
        ),
    },
    instructions="Choose the least costly model that can complete the task.",
)

agent = create_agent(
    "openai:gpt-5.6-luna",
    middleware=[router],
)

# The router picks from the latest user message and
# uses that model for the whole run. Probabilities and
# confidence stay available in agent state.
Tool-risk gating — block before execution
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import (
    AutoModeMiddleware,
)

# Check tool calls for risky decisions before
# the tool executes, and block them if so.
guardrail = AutoModeMiddleware(tools=["bash"])

agent = create_agent(
    "openai:gpt-5.6-luna",
    middleware=[guardrail],
)
Why this one matters

Coding harnesses have shipped versions of this for a while — a classifier that checks dangerous actions before they are taken. It has historically been locked inside closed-source harness internals. Now that a cheap classifier exists, the same pattern is available to any agent you build.

Counting correctly — the documented exception
from typesafe_sdk import Noul, TypeSafeClient

client = TypeSafeClient(model="jev-1.13")

# Jev does not count reliably. So don't ask it to.
# Ask one yes/no question PER ITEM, then add up the
# answers in ordinary code.
YES = 0.5   # pick a threshold that suits your use case

items = ["typesafe", "apple", "california", "banana",
         "likes", "calibration", "orange", "vertex"]

result = client.system_one(
    {"items": items},
    {
        f"item_{i}": Noul(
            instructions=f"Is `items[{i}]` the name of a fruit?"
        )
        for i in range(len(items))
    },
)

count = sum(
    result.nouls[f"item_{i}"].noul > YES
    for i in range(len(items))
)

Adapted directly from TypeSafe's guidance. Note the shape: the model does the judgment, the code does the arithmetic. That division is the single most repeated piece of advice in this entire ecosystem.

Extraction as a Choice, not a parse
# Don't ask for a value and hope it comes back clean.
# Enumerate the possibilities and let it pick.

questions = {
    "month": {
        "type": "choice",
        "instructions": "Which month does this deadline fall in?",
        "criteria": {
            "january": None, "february": None,
            "march": None, "april": None,
            # ... twelve months, bounded set
            "not stated": "The deadline does not name a month",
        },
    },
    "day": {
        "type": "choice",
        "instructions": "Which day of the month?",
        "criteria": {str(d): None for d in range(1, 32)},
    },
}

# Every part of a date is a small closed set, which
# turns extraction into a Choice rather than free-form
# parsing -- and gives you somewhere to put an explicit
# "not stated" so a missing part is reported, not guessed.
# Code then assembles the parts and owns ordering,
# duration, offsets and weekday arithmetic.

This is TypeSafe's own recommended pattern for dates, generalised. It is also the answer to the "literal reading" failure mode: enumerate the space rather than describing it.

Discipline

Design rules and anti-patterns

TypeSafe publishes an explicit "avoid the following" list, and independent practitioners have added several more from experience. Together they form the practical constraint set for this class of model.

Rule Why Source
Don't ask the model something code can compute exactly It is "not a calculator." Counting is unreliable and error grows with the size of the thing being counted. If a regex or a parser can find it, the count belongs in code. TypeSafe
Don't hide several judgments inside one question Multiple hops of reasoning cost accuracy. Split into separate literal questions and combine the answers in code. TypeSafe
Don't send more state than the question needs "Context rot" — unrelated detail acts as a distractor and accuracy falls. Filter and retrieve in code first. TypeSafe
Don't rely on arithmetic identities between questions P(refund) 0.72 and P(not_refund) 0.47 summed to 1.19 on one ticket. A Choice and an equivalent Noul disagree by design: one is relative, the other absolute. TypeSafe
Don't assume the correct answer is in your option list The model cannot invent an answer outside the schema. If the right option isn't there, it picks the nearest permitted value. Include an explicit "none of the above" where it is possible. Inference from the design
Don't trust a probability without varying option order One open model scored 72% with options in one order and 21% reversed. Kev lists order sensitivity as an unresolved limitation that question isolation does not prevent. Independent measurement
Don't treat state as trusted Injected instructions, misleading framing, or text that argues for its own classification can move the answer. If your state comes from users and your decision is a guardrail, test that before shipping. TypeSafe
Don't assume the model is deterministic Re-running the same 242 decisions produced three answer changes. Same model, same inputs, different outputs. Independent measurement
Don't use this for explanations There is no reasoning trace and no way to ask for one. If you need to explain a decision to a regulator, an auditor or an angry customer, this model cannot help and will not try. By design
Don't deploy on high-stakes ranking without bias work A floating point number conceals whatever bias is baked in, and picking it apart is experimentally hard. The mitigation is cheap experimentation — thousands of probes for cents — not intuition. Simon Willison
✅ The positive version of the same advice
  • Enumerate, don't describe. A bounded Choice over twelve months beats a free-text date field every time.
  • Split, don't nest. One literal question per judgment, combined by your code.
  • Compute in code, judge in the model. Arithmetic, ordering, dates and offsets are yours. Meaning and relevance are the model's.
  • Filter, then ask. Narrow state to the fields the question needs before you send it.
  • Threshold, don't assume. Every decision gets a confidence gate with a cut-off tuned on your own data.
  • Measure, then trust. Cheap enough to run a thousand probes; that is the substitute for being able to ask why.
Rollout

Migrating and measuring

Two independent write-ups converge on the same honest approach, and it is not "swap and hope."

1 · Shadow first

Point a copy of production traffic at the new model and log what it would have decided, alongside what your current system actually did. Compare before you switch. One practical tip from a roundup: local servers from these projects are open by default, so set their API-key environment variable before exposing one to anything.

2 · Fine-tune, then refit

For open models, a few hundred labelled examples of your own decisions is the step that closes most of the gap. Then refit the confidence temperature — because a model fine-tuned on new labels is no longer calibrated the way its authors' card says it is.

3 · Test the boring failure modes

Reversed option order. Inputs longer than the window. Inputs with irrelevant detail. Inputs containing text that argues for a particular classification. Assertions that probabilities sum to 1. Every one of these is a documented way these models break.

4 · Instrument the escalation rate

The gated percentage of calls is your real quality metric in production — more useful than any offline benchmark, because it is measured on your traffic. Alert on it drifting upward. That, not a leaderboard position, is the number to put on a dashboard.

⚠️ And keep the model in its lane

The strongest consensus in every source on this site is not about which model to pick. It is that a decision model is a complement, not a replacement, and that the production stack of the near future holds several model types at once with a routing layer deciding which one serves a given call. One model doing everything is a prototyping pattern. If you find yourself trying to make a decision model generate text or reason in multiple hops, you have picked the wrong component.