Fetching from the wire…
Public story · 2026-07-25 · high
The thread's own poster calls it a solid model that could top the 120B class if Poolside fixes the overthinking, per the 70-upvote thread.
Why now: The LocalLLaMA thread is the first sustained hands-on test of Laguna S 2.1 to surface, as of July 25.
Poolside's Laguna S 2.1 spiraled into an endless reasoning loop over a trivial prompt about walking 69 meters to a car wash, per a widely-shared LocalLLaMA thread.
That matters more than any benchmark score for builders chaining this model into agent workflows or paying per token. A loop that won't terminate costs money and time before it costs accuracy.
The thread, the model's first sustained hands-on test since launch, pulled 70 upvotes and 75 comments. Its poster wasn't dismissive. They called it a solid model that could lead the roughly 120B class outright once Poolside fixes what they described as overthinking loops.
Not every test went sideways. One practitioner used the model to solve a data-restructuring problem that had eaten days of their time. That's the kind of result a benchmark score promises but rarely proves.
Both reports describe the same model: strong on genuinely hard problems, undone by reasoning it can't switch off on easy ones.
Laguna S 2.1 runs 118B total parameters with 8B active and handles a 1M token context, scoring 70.2% on Terminal-Bench 2.1. NVFP4 weights let it fit on a single DGX Spark. That's one machine for a model this size, not a GPU rack.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Poolside told investors the license is non-exclusive, covering the system used to train its Laguna open model, plus job offers to 109 employees and a separate $1B investment at a $12B pre-money valuation. The letter insists this is neither acquisition nor acquihire, and the th...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.