Fetching from the wire…
Public story · 2026-09-08 · high
The 1,075-upvote thread says Ollama buried llama.cpp's credit, runs 1.8x slower, and carries a token-theft CVE.
Why now: As of September 8, 2026, the post topped r/LocalLLaMA with 1,075 upvotes.
Ollama hid llama.cpp's role for over a year and left a license-compliance complaint unanswered for 400 days, per a 1,075-upvote r/LocalLLaMA thread. The stakes for anyone running Ollama: a named CVE for token theft, and a measured 1.8x speed gap against the engine it wraps.
On throughput, the post clocks llama.cpp at 161 tokens per second against Ollama's 89, with Ollama's CPU use running 30-50% higher. It ties that gap, along with broken structured output, vision failures, and new assertion crashes, to Ollama's mid-2025 switch to a custom GGML backend.
CVE-2025-51471 covers token exfiltration through malicious model registries, a supply-chain risk rather than a benchmark complaint.
Top comments on the thread converge on three alternatives: LM Studio, llama.cpp directly, and Unsloth Studio. The recurring practical complaint is that Ollama won't reuse models already on disk, forcing redundant downloads.
None of these complaints are new on their own. Seeing the credit dispute, the speed gap, and the CVE stacked in one thread is what pushed three alternative tools into the top comments.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Ollama cut v0.34.0-rc1 on September 5 at 23:49 UTC, and the headline item changes the shape of the local-versus-hosted decision rather than the performance of either side: Ollama-hosted open models can be selected directly inside ChatGPT Desktop, with setup driven from the Oll...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
OpenClaw tagged v2026.8.1 at 03:30 UTC this morning. The release post counts 933 contributors, 569 of them first-time, and more than 16,000 pull requests, roughly half of every PR ever merged into the project, after a seven-week gap against a prior cadence of 106 releases in 2...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.