The Local Intelligencer · A self-hosted newspaper for a home-lab engineer


Local AI & LLMs

Friday, October 9, 2026

From the wires

Ollama makes MLX the default on Apple Silicon — ollama/ollama releases

On September 25, Ollama v0.40.0 quietly flips the default: on Apple Silicon, any architecture the MLX runtime supports now runs on MLX, no flags and no config-file archaeology. The release names qwen3.8, gemma4 and qwen3.6/3.5, adds the decision models Nimble, tev1, clef and clef-flash, and brings its first embedding model to MLX. The pull command, mercifully, stays the same.

#ollama #mlx #apple-silicon

Ollama will now upgrade your old models behind your back, politely — ollama/ollama releases

Models pulled before v0.40 are upgraded in the background the first time you run them, for speed and compatibility with the bundled llama.cpp. To keep downgrades possible, the original copy is kept on disk as a backup, and the release notes ship a jq one-liner for reclaiming that space early. Twice the models, half the free disk: an honest trade, disclosed up front.

#ollama #upgrades #disk-space

llama.cpp spreads mixture-of-experts caches across GPUs — llama.cpp releases

Support for keeping an MoE cache distributed over multiple GPUs landed this week, the sort of unglamorous plumbing that decides whether a big sparse model fits your rig at all. The same days brought CUDA radix top-k that cuts a pathological 34,816-row top-k from 5,762 ms to 942 ms, and slot save/restore that keeps its context checkpoints intact, so a restored slot can roll back instead of re-processing the whole prompt. The runtime keeps rewriting what the phrase 'runs on my hardware' means.

#llama.cpp #moe #cuda #multi-gpu

vLLM v0.31.0 ships a weight-cache daemon so restarts stop costing minutes — vllm-project/vllm releases

The flagship release packs 717 commits from 307 contributors, but the line home-lab operators should read twice is vllm preload: a daemon that keeps post-quantized weights resident in GPU memory across engine restarts, with a health endpoint and a readiness wait attached. Experimental CRIU snapshots can now restore a fully initialized engine. For everyone who has watched a giant model cold-start for the fifth time in an afternoon, this is the release notes equivalent of a cold compress.

#vllm #serving #restarts

Mistral Large 4, Le Chonk: a trillion parameters, open weights promised this month — Mistral AI

Mistral's announcement calls Large 4 a 1-trillion-parameter, 52-billion-active multimodal model - a count Simon Willison's linked review reads as 49 billion - trained on 3,800 of the company's own Grace Blackwell GPUs in Europe, with a public preview API today and the weights promised by month's end. The company also notes the model is state-of-the-art among open models on enterprise workloads like cybersecurity, finance and law. Whether the open-weights drop fits on anything short of a rack is, characteristically, next month's problem.

#mistral #open-weights #releases

Simon Willison on Mistral Large 4: back to six months behind the frontier — Simon Willison's Weblog

Willison's October 6 review scores Large 4 at 38 on Artificial Analysis - just behind DeepSeek 4.1 Flash, a 552-billion-parameter model - and calls it a huge improvement on last December's Mistral Large 3, which scored 9. His verdict: it is not a Fable-class model, but it is good to see Mistral roughly six months behind the frontier again instead of trailing out of sight. The post's other test is two pelicans on bicycles, and the high-reasoning one looks better - on fewer tokens, no less.

#mistral #benchmarks #review

Liquid AI opens the d1 decision models for the edge — Hugging Face Blog (Liquid AI)

Liquid AI released d1-3B and d1-omni-600M as open weights: decision models that answer typed questions in a single forward pass instead of generating tokens, with d1-3B scoring 48.57 on Decision Index 0.2.1, the best under 10B. The company's own timings report 16 ms per question on a Jetson AGX Thor and 8 ms on an RTX 4090. Your router no longer needs a 27-billion-parameter opinion about a support ticket.

#decision-models #liquid-ai #edge

Reflection emerges from stealth with Beam, a 501B MoE promising Apache 2.0 weights — Latent Space (AINews)

Reflection announced Beam, a 501B-total, 23B-active mixture-of-experts for coding and agentic work, trained from scratch in the US, with full Apache 2.0 weights, the model card and the technical report all promised for later this month - a preview under early-access sign-up for now. Latent Space's own tally has the current SOTA open models - GLM 5.3, Kimi K3, Qwen 3.8 Max, DeepSeek V4.1 Flash - still ahead, but notes a market segment has been waiting for exactly this. Bring a multi-GPU rig and a measure of patience; the weights are still a promise.

#open-weights #reflection #moe

LiteLLM backports the decision-model endpoints into the proxy — BerriAI/litellm releases

LiteLLM v1.104.2 backports /v1/systemone, the OpenAI-format /v1/decisions endpoint and the OpenAI Decisions provider into the stable line, with the same trio in the 1.105 release candidates. A proxy that fronts local generation can now front local and hosted decision models behind one URL. The proxy is quietly becoming the control plane of the home lab, which is either comforting or ominous depending on how your week is going.

#litellm #proxy #decision-models

The decision-model wave gets its centipede benchmark — r/LocalLLaMA

One r/LocalLLaMA poster timed four open decision models on a single RTX 4090, flagging every centipede name across 9,534 words of Wikipedia: Laya at 3.9 ms per word, d1-3B at 6.0 ms, Clef-Flash 9B at 24.4 ms, and Interfaze's Lev at 51 ms but with the best catch rate. Community numbers, to be verified before betting the farm - the poster's own comment thread immediately relitigated the precision and recall semantics. The shape of the result is the news: these models are fast enough to sit in a live request path.

#decision-models #benchmarks #community

EmbeddingGemma 2's Apache 2.0 license, and the argument for open embeddings — Simon Willison's Weblog

EmbeddingGemma 2 ships under Apache 2.0, and Simon Willison made the argument worth stealing: for embeddings, a closed hosted-only model is a bad deal, because you compute millions of vectors today and re-compute every one of them the day the vendor retires the model. Pay a provider to host it, sure - as long as the open weights exist to fall back on. Your index is a hostage; the weights are the ransom.

#embeddings #gemma #open-weights