Fetching from the wire…
Public story · 2026-07-25 · high
Co-CEO Eiso Kant says the Laguna S 2.1 run hit zero on-call incidents across eight weeks on a 10,000 H200 cluster.
Why now: Kant laid out the Model Factory details in the July 23 Latent Space episode.
Poolside built a 118B-parameter mixture-of-experts model in eight weeks, co-CEO Eiso Kant told Latent Space. The model has 8B active parameters per token and a 1M context window, and Poolside's team is under 70 researchers. That ratio is the story: Kant says the team runs 10,000 to 20,000 experiments a month without headcount to match.
Checkpoints are evaluable within 30 minutes of completion, per Kant. Training data streams in instead of getting pre-materialized to disk, which kills the re-staging cycles that slow most large runs.
The training run itself used a 10,000 H200 cluster and produced zero on-call incidents across the full eight weeks, Kant said. Laguna S 2.1 shipped as open weights under OpenMDW-1.1. It fits on a single DGX Spark with FP8 precision, a real constraint check on a 118B model.
Kant's thesis: model building is 90% engineering, and 95% of that reduces to improving data or compute efficiency. He doesn't say how the 30-minute eval holds up on harder capability benchmarks versus quick sanity checks. He also doesn't say what the zero-incident run cost in infrastructure work before the eight weeks started.
Each link below shares sources, entities, or timing with this story.
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
Poolside told investors the license is non-exclusive, covering the system used to train its Laguna open model, plus job offers to 109 employees and a separate $1B investment at a $12B pre-money valuation. The letter insists this is neither acquisition nor acquihire, and the th...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Per a roundup, ZAYA1-8B is an Apache-2.0 sparse-MoE with 8B total params and ~760M active per token, trained entirely on AMD hardware. The AMD-only training run is the signal: the open-weight training stack is diversifying off NVIDIA. Verify the numbers against the official mo...
NVIDIA debuted the Nemotron 3 family for agentic AI. Nano (30B params, 3B active via MoE) delivers 78% HumanEval, native 1M-token context, and 4x higher throughput than Nemotron 2 via hybrid Mamba-Transformer architecture. Critical developer angle: NVIDIA open-sources NeMo Gym...
NVIDIA claims up to 4x faster output and 30% faster agentic task completion versus comparable models, aimed explicitly at high-volume narrow work: code review, tool use, security monitoring, billing triage (NVIDIA). They also published Nemotron-RL-Agentic-Terminal-Pivot, an ag...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.