Practitioners Land on Poolside's Laguna S 2.1: Genuinely Strong in the 120B Class, Undone by Overthinking Loops
Poolside's Laguna S 2.1 — 118B total / 8B active MoE, 1M context, 70.2% on Terminal-Bench 2.1, NVFP4 weights fitting a single DGX Spark, OpenMDW-1.1 license — is now getting its first sustained hands-on scrutiny on r/LocalLLaMA, and the verdict is split in a specific way. One user reported it solved a data-restructuring problem that had taken them days (53 upvotes), while a widely-shared thread (70 upvotes, 75 comments) showed it spiraling on a trivial prompt about walking 69 meters to a car wash, with the poster explicitly clarifying it is 'a solid model' that could top the ~120B class 'if they fix the overthinking loops.' This is the practitioner signal the launch coverage on July 21 could not provide: the benchmark number is real, and the failure mode is inference-time reasoning that will not terminate — which matters far more than Terminal-Bench when the model is inside an agent loop you are paying for.
↳ Follow the thread