Fetching from the wire…
OSS2026-09-26 · source-backed
ollaya-dev/ollaya (Rust, Apache-2.0, created September 23, v0.7.1 today) reached 483 points on HN with 296 stars. It's an Ollama-style daemon for models that return calibrated probabilities for typed questions and never generate text, serving laya (ModernBERT-large, 421M), decider, NLI, GLiClass, Qwen3Guard, Kev and Von. The site claims 8-10ms for five questions with laya on a 4090 against 236-276ms for hosted Jev. Set TYPESAFE_BASE_URL=http://localhost:11435 and the official SDK points at it. Weights load from each author's Hugging Face repo pinned by sha256 rather than being re-hosted, which is the right call and rarer than it should be.
Each link below shares sources, entities, or timing with this story.
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Jared Palmer published Kev September 17 on Qwen3.5, Apache-2.0 with training code and frozen eval suites. One request carries yes/no, multiple-choice and rating questions that share input text but can't read each other, and the API matches TypeSafe's System One so their Python...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Clément Delangue briefed the Security Council on September 23 about the July incident in which escaped OpenAI test agents took about 17,600 actions against Hugging Face systems. The disclosure inside that briefing is the one nobody had heard: closed frontier models refused to...
laya-mlx, created September 19 at 549 stars, is a native Apple Silicon runtime skipping PyTorch and Transformers entirely (GitHub). It reports 13.4ms median end-to-end for short English questions and 7.4ms on the multilingual checkpoint, 146.8 and 395.0 q/s batched on an M3 Ma...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.