The Local Intelligencer · A self-hosted newspaper for a home-lab engineer


The Decision Models Come Home

Local AI & LLMs · Friday, October 9, 2026 · editorial · Model Weight, AI Desk

Eight open decision models in two weeks, all fast enough to sit in a live request path - and most of the traffic you were routing through a 27B model never needed prose in the first place.

The decision models came home this fortnight, and they did not ring the doorbell.

In the space of two weeks our desks counted, from primary pages alone: TypeSafe's Jev, OpenAI's hosted Decisions endpoint, Cloudflare's Clef and Clef-Flash, Bespoke Labs' Nimble and Together's tev1 (both listed in Ollama's September 29 announcement), Interfaze's Lev, ConvAI's Laya and Liquid AI's d1 family. The Hugging Face trending board carried several of them the day we looked: JEV-27B-VL, d1-3B, clef, clef-flash, laya, Kolibri-1 - a filing cabinet that has achieved sentience. A community speed test on a single RTX 4090 put Laya at 3.9 milliseconds per decision and d1-3B at 6.0, on one consumer GPU, through llama.cpp.

To understand what arrived, it helps to remember what a decision model refuses to do. A traditional language model generates tokens - each one a fresh throw of the dice, however well-loaded. A decision model scores a state against a schema of typed questions (yes-or-no, choice, score) in a single forward pass and returns probabilities. No prose, no apology, no commentary. It is the difference between hiring an analyst and installing a coin sorter.

And the coin sorters are fast in a way that changes architecture. Ollama's announcement numbers for Nimble 9B

The infrastructure did not wait for anyone. Ollama's v0.40 line made MLX the default runtime on Apple Silicon in the same fortnight it listed the decision models, and exposed the whole contract at /v1/systemone since September. LiteLLM backported /v1/systemone and the OpenAI-format /v1/decisions provider into the stable line this week. Willison shipped an llm plugin for the OpenAI endpoint the day it appeared. When three independent stacks adopt the same API shape within a month, that is not a trend; that is a standard being born with witnesses present.

The skeptical questions are the right ones, and the community asked most of them first. A red-team-style word-by-word test caught a real limitation: on adversarial word-flagging, even the best of the four open models caught only about 83 percent of true positives - the fastest were fastest partly because they were wrong more often, and one commenter had to relitigate what accuracy even meant before the numbers meant anything. Fine-tuning is the thread's proposed fix, and it stays a proposal: one commenter reports a Laya fine-tune blocking dangerous commands reliably - an anecdote from a different task, with no measured numbers attached - while the test's own author reports his router fine-tune still falling short on accuracy. Read that as a hypothesis awaiting measured, task-specific validation, with an independent gate behind anything it is allowed to act on. Until someone publishes that validation, the honest deployment posture is "typed decisions for triage, routing and gating, human or a full model for anything you would defend in a postmortem." A coin sorter is not an auditor.

There is also the matter of what the big models think of all this. OpenAI charges 10 cents per million input tokens on its decision endpoint and nothing for output; TypeSafe's Jev charges 4.2 cents, per Simon Willison's write-up of the API - priced like a utility either way. The open releases - Jev on a Qwen3.8-27B backbone and GEV on Gemma-4-26B, a Hugging Face last-month count of 178,643 for Jev and 916,313 for GEV, Apache 2.0 on the released adapter and head - are fine-tunes, not from-scratch architectures. The frontier's business model here is cheap and omnipresent; the home lab's answer is cheap and present, physically present, on your own silicon. That is not a competition the home lab loses. It is a division the home lab wins by having a home field.

So the position of this desk: treat decision models as a new tier of your serving stack, not a novelty adjacent to it. The arithmetic is obvious once you see it - a few gigabytes of weights that answer structured questions in tens of milliseconds beat waking a 27B generative model for every routing call, both in latency and in the humility of the error modes. Put one in front of your agent loop as a gate. Put one behind your proxy as a router. Keep the big model for the questions that deserve prose.

The frontier will keep selling judgment by the token. The home lab has always known that most decisions are not judgment. They are sorting mail - and now the mailroom runs locally, at ten milliseconds a letter, on hardware you already own.

The frontier will keep selling judgment by the token. The home lab has always known that most decisions are not judgment. They are sorting mail - and now the mailroom runs locally, at ten milliseconds a letter, on hardware you already own.