For the last two years, the default answer to any complex heuristic problem has been the same: send it to a large language model. Support ticket routing, spam detection, tool selection, content moderation, refund approval—every one of these is a bounded decision dressed up as a generation task. We have been paying for prose we immediately throw away.
A new model category is changing that. System One models, or decision models, take unstructured state and return a typed decision: a choice from a fixed set, a score, or a probability. No tokens, no parsing, no malformed JSON, and no "I'm sorry, I can't do that." For developers, this is the difference between asking a model to write an essay and asking it to answer a multiple-choice question.
System One models mimic the fast, instinctive processing of human cognition—the "System 1" in Daniel Kahneman's Thinking, Fast and Slow. Where conventional LLMs deliberate token by token, a decision model picks immediately from defined options and returns probability weights. Jev, from TypeSafe AI, is the loudest example right now. It is also closed and in early access. What matters more is that the pattern is spreading through open-source alternatives you can run yourself. That combination—fast judgment, local execution, and predictable cost—is what makes this a structural change for software engineering rather than just another model release.
The three-layer stack
Decision models do not replace LLMs. They replace the inefficient pattern of using LLMs to do work that does not need language. The stack becomes cleaner:
| Layer | Handles | Example |
|---|---|---|
| Deterministic code | Exact rules, calculations, side effects | "If invoice > 30 days overdue, flag it" |
| Decision model | Fast semantic judgments with calibrated uncertainty | "Does this email sound like a churn risk?" |
| Generative / reasoning model | Open-ended output, planning, complex reasoning | "Draft a personalized retention email" |
Code owns control flow and risk tolerance. The decision model supplies narrow judgments. The LLM handles what actually needs words or reasoning.
Most production AI code today skips the middle layer. It asks a frontier model to produce structured JSON, then parses it, validates it, retries malformed responses, and pays for every output token. A decision model removes that overhead by construction: the caller declares the allowed answers up front, and the model returns a probability distribution over exactly those answers. In this paradigm, a malformed output is not a possibility—it is unrepresentable.
The distinction matters. Structured output tells an LLM, "Please format your answer as JSON." A typed decision model is mathematically bound to the option set you declared. It cannot produce a value outside the schema, so the failure mode shrinks from "parse error or hallucinated key" to "wrong classification"—a much easier thing to handle in code.
This also changes where judgment can live. A decision that takes 30 milliseconds on a local GPU can run inside a request handler, inside a retry loop, or on every message in a stream. A decision that takes eight seconds through a frontier API has to be batched, cached, or avoided entirely.
What this changes for developers
The practical effect is that a whole class of AI-powered features becomes production-ready. Not because the models are smarter, but because they are cheaper, faster, and easier to compose.
Useful places to start:
- Model routing — send lookups and local changes to a small model; route architecture and high-stakes decisions to a frontier model.
- Agent guardrails — approve or block bash, database, and file operations before they run.
- Input/output moderation — detect jailbreaks, policy violations, or off-topic requests in real-time.
- Ticket and email triage — route, prioritize, and flag urgency without the fragility of regex.
- Retry/continue decisions in agent loops — "Is the task done, stuck, or hallucinating?"
The common thread is that you can write the valid answers down before you deploy, and being wrong is cheap enough to catch or escalate.
There is also a clear migration path. Start with a hosted decision model to prove the judgment call is worth automating. Collect input/output pairs. Once you have labeled data and a reason to keep state local, distill the decision into a smaller, task-specific classifier. The general decision model de-risks the feature; the distilled model makes it cheap at scale.
Running this yourself
The closed option is Jev, which TypeSafe claims is two orders of magnitude faster and cheaper than frontier LLMs on classification tasks. The open options are what make this truly impactful for self-hosters.
| Model | Size | Runtime | Rough footprint | Notes |
|---|---|---|---|---|
| Laya | 322M-421M | ONNX Runtime | ~2 GB RAM / VRAM | Jev-compatible; receptron package |
| FLock this-that-model | 1.88B | vLLM / llama.cpp | ~4-6 GB VRAM | Apache 2.0, 30.9 ms on laptop GPU |
| Kev | 0.6B | ONNX | ~1-2 GB | Jev-compatible API, fine-tuned |
| NanoJev | 0.6B | Transformers | ~1-2 GB | Smallest reproducible Jev-style model |
| Von | ~395M | ONNX / Python | ~1 GB | ModernBERT + OptionMarker |
| Rizzo Flow | 1.7B-4B | llama.cpp GGUF | ~2-4 GB | Easy quantisation for CPU/GPU |
Laya is the most accessible entry point. A Node.js package loads the ONNX bundle from Hugging Face and returns Jev-compatible choice, score, and noul answers. FLock's this-that-model is the most intriguing if you already have GPU infrastructure: a 1.88B parameter Apache 2.0 model that the authors report runs in 30.9 ms on a laptop GPU, with a strictly proper scoring rule that keeps probabilities calibrated.
Deployment patterns are straightforward:
- Sidecar / same-machine: Run the decision model as a local HTTP service next to your app. Latency is low enough to call synchronously inside a request handler.
- Model server: Use vLLM or llama.cpp server for batching across multiple app instances.
- Edge / embedded: Sub-1B models can run on CPU-only machines, providing a low-latency inference point at the edge.
- Prototype with API, self-host later: Validate with Jev's hosted API, then swap to an open model once you have data and a reason to keep state local.
LocalAI does not yet carry these as first-class models, but it is OpenAI-compatible enough to work with FLock's endpoint. Laya and Kev are ONNX-native and need a thin wrapper service. The gap is small: you are running another container, not another stack.
What changes in engineering roles
Decision models move judgment from prompts and humans into a typed component. That changes the job description for several roles:
AI and platform engineers become decision-system designers. The work is decomposing workflows into decisions, defining option sets, setting confidence thresholds, and measuring calibration. It is systems engineering, not prompt engineering.
Backend engineers treat decision models as infrastructure, like a cache or a database index. Less parsing JSON from LLMs; more wiring logic around choice and noul outputs.
QA and test engineers test probabilities, not just outputs. A 0.90 confidence score should mean roughly 90% accuracy over a representative sample. Edge cases become wrong-classification risk, not malformed JSON.
Security and compliance get a smaller attack surface on the output side—no free-form text escaping into downstream systems—but the model itself joins the trust boundary. Data residency and on-prem deployment become the decisive questions, making open-source models a strategic necessity rather than a hobby.
The bottom line
Software wants decisions. Humans want explanations. LLMs are good at explanations. Decision models are good at decisions. The last few years have forced every application through the explanation layer because that was the only interface we had.
That is ending. A developer in 2027 will choose between three tools for a semantic branch: write a regex, call a decision model, or call a reasoning model. Each has a clear place. The middle one is new, and it is the one that makes agentic systems reliable enough to leave running.
For self-hosters, the timing is good. The open alternatives are small enough to run on consumer hardware, fast enough to live inside request paths, and open enough to keep your data inside your network. The smart if statement is here. It just happens to understand context.
Sources and further reading
- Thinking Fast and Slow - Daniel Kahneman
- Introducing System One Models & Jev — TypeSafe AI
- Building a Harness with Jev — LangChain
- System One models like Jev can train their own replacements — Sean Goedecke
- Jev: A System One Model, Not an LLM — innFactory
- What is an AI Decision Model? — Correlation One
- Jev AI: The Rise of System One Models — DhanushKumar
- this-that-model-1.0 paper — Cheng, Dai, Sun (FLock.io / Oxford)
- receptron/laya — Node.js ONNX runtime for Laya
- Laya open-source Jev alternative — Flowtivity
- Top open-source Jev alternatives — DataCamp
Dallum Brown
Writer and curator exploring the impact of technology on everyday life.
View All Articles