Artificial Intelligence 8 min read

The Decision Model Is the New If Statement

A new model category is emerging.

A conceptual digital render showing a chaotic stream of white and blue data particles flowing into a glowing blue pyramid. On the other side of the pyramid, the data emerges as a perfectly organized grid of blue cubes, representing the transformation of unstructured input into structured, typed decisions.
Photo: Google Gemini

For the last two years, the default answer to any complex heuristic problem has been the same: send it to a large language model. Support ticket routing, spam detection, tool selection, content moderation, refund approval—every one of these is a bounded decision dressed up as a generation task. We have been paying for prose we immediately throw away.

A new model category is changing that. System One models, or decision models, take unstructured state and return a typed decision: a choice from a fixed set, a score, or a probability. No tokens, no parsing, no malformed JSON, and no "I'm sorry, I can't do that." For developers, this is the difference between asking a model to write an essay and asking it to answer a multiple-choice question.

System One models mimic the fast, instinctive processing of human cognition—the "System 1" in Daniel Kahneman's Thinking, Fast and Slow. Where conventional LLMs deliberate token by token, a decision model picks immediately from defined options and returns probability weights. Jev, from TypeSafe AI, is the loudest example right now. It is also closed and in early access. What matters more is that the pattern is spreading through open-source alternatives you can run yourself. That combination—fast judgment, local execution, and predictable cost—is what makes this a structural change for software engineering rather than just another model release.

The three-layer stack

Decision models do not replace LLMs. They replace the inefficient pattern of using LLMs to do work that does not need language. The stack becomes cleaner:

Layer Handles Example
Deterministic code Exact rules, calculations, side effects "If invoice > 30 days overdue, flag it"
Decision model Fast semantic judgments with calibrated uncertainty "Does this email sound like a churn risk?"
Generative / reasoning model Open-ended output, planning, complex reasoning "Draft a personalized retention email"

Code owns control flow and risk tolerance. The decision model supplies narrow judgments. The LLM handles what actually needs words or reasoning.

Most production AI code today skips the middle layer. It asks a frontier model to produce structured JSON, then parses it, validates it, retries malformed responses, and pays for every output token. A decision model removes that overhead by construction: the caller declares the allowed answers up front, and the model returns a probability distribution over exactly those answers. In this paradigm, a malformed output is not a possibility—it is unrepresentable.

The distinction matters. Structured output tells an LLM, "Please format your answer as JSON." A typed decision model is mathematically bound to the option set you declared. It cannot produce a value outside the schema, so the failure mode shrinks from "parse error or hallucinated key" to "wrong classification"—a much easier thing to handle in code.

This also changes where judgment can live. A decision that takes 30 milliseconds on a local GPU can run inside a request handler, inside a retry loop, or on every message in a stream. A decision that takes eight seconds through a frontier API has to be batched, cached, or avoided entirely.

What this changes for developers

The practical effect is that a whole class of AI-powered features becomes production-ready. Not because the models are smarter, but because they are cheaper, faster, and easier to compose.

Useful places to start:

  • Model routing — send lookups and local changes to a small model; route architecture and high-stakes decisions to a frontier model.
  • Agent guardrails — approve or block bash, database, and file operations before they run.
  • Input/output moderation — detect jailbreaks, policy violations, or off-topic requests in real-time.
  • Ticket and email triage — route, prioritize, and flag urgency without the fragility of regex.
  • Retry/continue decisions in agent loops — "Is the task done, stuck, or hallucinating?"

The common thread is that you can write the valid answers down before you deploy, and being wrong is cheap enough to catch or escalate.

There is also a clear migration path. Start with a hosted decision model to prove the judgment call is worth automating. Collect input/output pairs. Once you have labeled data and a reason to keep state local, distill the decision into a smaller, task-specific classifier. The general decision model de-risks the feature; the distilled model makes it cheap at scale.

Running this yourself

The closed option is Jev, which TypeSafe claims is two orders of magnitude faster and cheaper than frontier LLMs on classification tasks. The open options are what make this truly impactful for self-hosters.

Model Size Runtime Rough footprint Notes
Laya 322M-421M ONNX Runtime ~2 GB RAM / VRAM Jev-compatible; receptron package
FLock this-that-model 1.88B vLLM / llama.cpp ~4-6 GB VRAM Apache 2.0, 30.9 ms on laptop GPU
Kev 0.6B ONNX ~1-2 GB Jev-compatible API, fine-tuned
NanoJev 0.6B Transformers ~1-2 GB Smallest reproducible Jev-style model
Von ~395M ONNX / Python ~1 GB ModernBERT + OptionMarker
Rizzo Flow 1.7B-4B llama.cpp GGUF ~2-4 GB Easy quantisation for CPU/GPU

Laya is the most accessible entry point. A Node.js package loads the ONNX bundle from Hugging Face and returns Jev-compatible choice, score, and noul answers. FLock's this-that-model is the most intriguing if you already have GPU infrastructure: a 1.88B parameter Apache 2.0 model that the authors report runs in 30.9 ms on a laptop GPU, with a strictly proper scoring rule that keeps probabilities calibrated.

Deployment patterns are straightforward:

  1. Sidecar / same-machine: Run the decision model as a local HTTP service next to your app. Latency is low enough to call synchronously inside a request handler.
  2. Model server: Use vLLM or llama.cpp server for batching across multiple app instances.
  3. Edge / embedded: Sub-1B models can run on CPU-only machines, providing a low-latency inference point at the edge.
  4. Prototype with API, self-host later: Validate with Jev's hosted API, then swap to an open model once you have data and a reason to keep state local.

LocalAI does not yet carry these as first-class models, but it is OpenAI-compatible enough to work with FLock's endpoint. Laya and Kev are ONNX-native and need a thin wrapper service. The gap is small: you are running another container, not another stack.

What changes in engineering roles

Decision models move judgment from prompts and humans into a typed component. That changes the job description for several roles:

AI and platform engineers become decision-system designers. The work is decomposing workflows into decisions, defining option sets, setting confidence thresholds, and measuring calibration. It is systems engineering, not prompt engineering.

Backend engineers treat decision models as infrastructure, like a cache or a database index. Less parsing JSON from LLMs; more wiring logic around choice and noul outputs.

QA and test engineers test probabilities, not just outputs. A 0.90 confidence score should mean roughly 90% accuracy over a representative sample. Edge cases become wrong-classification risk, not malformed JSON.

Security and compliance get a smaller attack surface on the output side—no free-form text escaping into downstream systems—but the model itself joins the trust boundary. Data residency and on-prem deployment become the decisive questions, making open-source models a strategic necessity rather than a hobby.

The bottom line

Software wants decisions. Humans want explanations. LLMs are good at explanations. Decision models are good at decisions. The last few years have forced every application through the explanation layer because that was the only interface we had.

That is ending. A developer in 2027 will choose between three tools for a semantic branch: write a regex, call a decision model, or call a reasoning model. Each has a clear place. The middle one is new, and it is the one that makes agentic systems reliable enough to leave running.

For self-hosters, the timing is good. The open alternatives are small enough to run on consumer hardware, fast enough to live inside request paths, and open enough to keep your data inside your network. The smart if statement is here. It just happens to understand context.


Sources and further reading

D

Dallum Brown

Writer and curator exploring the impact of technology on everyday life.

View All Articles

Subscribe to
The Brief

Our curated selection of tech news and other discoveries, delivered every month.

No spam. Unsubscribe anytime.

Comments (0)

Please sign in to leave a comment.

No comments yet. Be the first to share your thoughts!

Further Reading

Privacy Notice

We use essential cookies for site functionality (session management, CSRF protection) and do not track you across the web. By using this site, you acknowledge our Privacy Policy and New Zealand Privacy Act 2020 compliance.

Learn More