The Local Intelligencer · A self-hosted newspaper for a home-lab engineer


/v1/systemone on the LAN: a decision tier, sketched for an evening

Local AI & LLMs · Friday, October 9, 2026 · howto · Model Weight, AI Desk

One Ollama pull, a proxy config sketched against the LiteLLM docs rather than blessed, and a ten-millisecond classifier sitting next to your generative model - with your own fixture set standing between it and production traffic.

This howto is two things: an executable walkthrough of the model side, and an architecture sketch of the proxy side, labeled as such before anyone wires it into production. The design brief: a home server that answers generative questions with a real model and answers millions of small structured questions with a ten-millisecond specialist, exposed through one proxy on your LAN. Two Ollama pulls, one config file sketched but not blessed, and the satisfaction of routing support tickets faster than your router routes packets.

What you are building

One box, two tiers: llama-server (or Ollama) serving a generative model, and a decision model sitting beside it answering typed questions - yes-or-no, choice, score - in a single forward pass. LiteLLM fronts both. Applications see one OpenAI-shaped endpoint; the routing table in your config decides who handles what.

Step 1: Ollama and the decision model

Install or upgrade Ollama, then pull the decision model. Clef-Flash is the sensible first pick: a 9B multimodal decision model from Cloudflare, fine-tuned from Qwen3.5-9B, Apache 2.0, requiring Ollama 0.35.1 or later. On Apple Silicon, Ollama 0.40 runs it on MLX by default.

ollama pull clef-flash

The endpoint is /v1/systemone. Send a state and a set of typed questions; get probabilities back. The canonical smoke test, straight from the model card:

curl http://localhost:11434/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "clef-flash",
    "state": "Hello World",
    "questions": {
      "says_hello": {
        "type": "noul",
        "instructions": "Does the state text contain a greeting?",
        "criteria": {
          "true": "The state text contains a greeting.",
          "false": "The state text does not contain a greeting."
        }
      }
    }
  }'

If the JSON comes back with a probability near 1.0 under "true", your decision tier is live. A ticket-triage schema is the same shape with better adjectives: a "choice" question with one criterion per team, a "noul" for whether the customer explicitly asked for a refund, and a "score" for urgency. One request, all three answers.

Step 2: the generative tier

Pull a real model for the prose. On a Mac, Ollama's MLX default handles qwen3.8 or gemma4 out of the box; on a Linux GPU box the same pull works over llama.cpp. Nothing about this step changed this fortnight, which is precisely why it is step two and not the story.

Step 3: LiteLLM in front, sketched

LiteLLM's stable line now speaks /v1/systemone and the OpenAI-format /v1/decisions format, so the proxy can front the decision tier alongside the generative one. What follows is an architecture sketch, not a blessed config: the endpoints are real, your rack is not yet, and the details here should be verified against the LiteLLM config docs and the docker quick start before you run anything. The shape, in miniature: one config entry per tier - prose pointed at your generative model, verdict at the fast classifier - and one container fronting both, so applications keep talking to a single base URL while the config decides whether a given call is prose or a verdict. One honest caveat before you wire it up: localhost:11434 inside the LiteLLM container is the container's own loopback, not your host's Ollama, so reaching the daemon you pulled in Step 1 needs an explicit network route (host networking, for one) that you verify against the docs before running anything.

The division of labor that works in practice:

Benchmarks are other people's opinions about other people's hardware.

Step 4: verify before you trust

Benchmarks are other people's opinions about other people's hardware. Before the decision model touches production traffic, run your own fixture set: fifty real states, labels you wrote yourself, the schema you intend to deploy. Measure accuracy per question type - the community word-flagging test that caught every model's weak spot is the template, and it took one evening on one GPU. If the fast model is fast because it says "no" attractively, you want to know that from your own fixtures, not from a Reddit thread.

Step 5: the hygiene

Cost of the whole experiment: a few gigabytes of weights, one config file - sketched, not blessed - and an evening. Failure mode: discovering that half your "the model should decide" traffic never needed a language model at all - which is not a failure, it is the point.