The Local Intelligencer · A self-hosted newspaper for a home-lab engineer


The Fortnight the Whole Stack Moved

Local AI & LLMs · Friday, October 9, 2026 · feature · Model Weight, AI Desk

llama.cpp, Ollama, vLLM and LiteLLM all shipped in the same two weeks, the community ran the receipts on real hardware, and Mistral promised a trillion open weights by month's end.

Every so often the home-lab calendar produces a week where the entire stack moves at once: the runtime, the server, the proxy, the model zoo, and the benchmark that tells you whether any of it was worth the electricity. The fortnight before this issue was one of those. Here is the field report, compiled the only way this desk knows how - by fetching the primary pages and refusing, on principle, to cite anything we did not actually read.

The runtime: llama.cpp does three quiet things that matter

The b11507 release landed support for keeping a mixture-of-experts cache distributed across multiple GPUs, which sounds like a footnote until you own more than one GPU and have been pretending that a 235B-class sparse model is "basically fine" on one card. For the operators running big MoE models on mismatched secondhand hardware, this is the difference between sharding being an experiment and sharding being a configuration.

The same release window carried a CUDA top-k rework that cut a pathological 34,816-row top-k from 5,761.8 milliseconds to 941.8 - a 6x win on the exact kind of long-context generation that agentic workflows produce by default. And the b11506 server work made context checkpoints survive slot save/restore, so a restored slot can roll back to a checkpoint instead of re-processing the whole prompt. None of this makes a headline. All of it makes your evening better.

The packagers: Ollama flips its default runtime

Ollama v0.40.0 made MLX the default on Apple Silicon: architectures the MLX runtime supports now run on MLX without ceremony, covering qwen3.8, gemma4, the qwen3.5/3.6 line, the decision models (Nimble, tev1, clef, clef-flash) and an embedding model. Thirteen days later, on October 8, v0.40.2 added background upgrades for older pulls - with the original copy kept on disk and a documented jq one-liner for reclaiming the space. The Mac mini is now a first-class citizen of its own product.

The heavy iron: vLLM learns to remember its weights

vLLM v0.31.0 landed with 717 commits from 307 contributors, and the feature with the highest home-lab value-per-line is vllm preload: a daemon that keeps post-quantized weights resident in GPU memory across engine restarts, with a health endpoint and a readiness wait attached. Anyone who has cold-started a large model five times in an afternoon while iterating on configuration knows precisely what this is worth. The same release ships experimental CRIU snapshots for restoring a fully initialized engine.

The proxy: one URL to rule the decision tier

LiteLLM v1.104.2 backported /v1/systemone, the OpenAI-format /v1/decisions endpoint and the OpenAI Decisions provider into the stable line, with the same trio riding the 1.105 release candidates. The practical consequence: a home operator can route generative traffic to llama.cpp, decision traffic to clef-flash, and overflow traffic to a hosted endpoint, and present one OpenAI-shaped API to everything in the house. The proxy has quietly become the control plane of the home lab.

The community's receipts

The fortnight's best hardware thread on r/LocalLLaMA described a $2,800 rig of eight Radeon Pro V620 cards - 256 GB of VRAM from the cloud-gaming clearance bin - running Qwen3.8-Flash-Next through a custom vLLM fork at 60 to 100 tokens per second decode. The poster's own caveat is the honest part: you cannot get the cards at the quoted price anymore. The architecture lesson survives the marketplace: last-generation enterprise VRAM plus a stubborn fork beats a mortgage-sized GPU purchase for batch work.

On the training side, a LoRA-over-GGUF recipe now trains Qwen3.8-Flash-Next - a 125B-A6B sparse model - in 40 GiB of VRAM with no CPU offloading. Two years ago that sentence was a joke; today it is a configuration file.

And a community speed test of the new decision models on a single RTX 4090 timed four open decision models word-flagging through 9,534 words of centipede Wikipedia - one /v1/systemone call per word: Laya 3.9 ms, d1-3B 6.0 ms, Clef-Flash 24.4 ms, Interfaze's Lev 51 ms but with the best catch rate. The comment section then did what comment sections do, relitigating precision and recall until the numbers meant something. Treat single-community benchmarks as leads, not laws - but the shape of the result, structured decisions at single-digit or low-double-digit milliseconds on a consumer GPU, is the fortnight in miniature.

The dollar-a-token crowd got their own accounting this fortnight, of the kind this desk handles with tongs: community benchmark tables of used-market GPUs surface and sink by the week, and any specific crown in one is stale before the comments stop arguing. What survives is the caveat every such table carries: used-GPU pricing is a moving target, and the value ranking shifts with the model you actually run - a card that is worthless for one parameter count is the budget play for another. Benchmarks sort cards; your context window picks them.

The model that looms

Mistral Large 4 previewed as a 1-trillion-parameter multimodal model with 52 billion active parameters, per Mistral's announcement, trained on 3,800 of the company's own Grace Blackwell GPUs and with the weights promised by the end of the month. Simon Willison's review of the same launch reports 49 billion active and scores the model at 38 on Artificial Analysis - roughly six months behind the frontier, which by this desk's arithmetic makes it the most interesting open-weights story of the month: a 1T sparse MoE is the shape that, if the weights do arrive, a well-stocked home lab could plausibly serve at usable speeds. Nobody is going to run it on a laptop. A lot of readers are going to run it on a rack - or find out that month's promise slipped, as open-weights promises have before.

The fortnight's pattern is not any one release. It is that every layer moved in the same direction at once - toward models that fit, servers that recover, proxies that route, and communities that check the math in public. The home lab did not get a gift this fortnight. It got a raise.

The home lab did not get a gift this fortnight. It got a raise.