SourcesOpus 4.6 SWE-bench Ceiling Analysissimonwillison.net·high signalXBlueskyLinkedInCopy linkAnalysis of Claude Opus 4.6 performance ceiling at 80.8% SWE-bench and what it means for the next generation of coding agents.SourceSource pagesimonwillison.net↳ Follow the threadPolicy dependency / Stack layer"There Are No Lossless Transformations of Natural-Language Text" — Sophie Alpert's AI Writing Policy for Engineers, Endorsed by Simon WillisonSimon WillisonStack layer / Threat patternNVIDIA ships Nemotron 3.5 Lightning, a 30B MoE built for agent loops, plus NeMo Switchyard routing that cuts task cost to a third of Opus 4.8NVIDIA BlogPolicy dependency / Stack layerHinton concedes the open-weights fight at Ai4: 'I think that battle's been lost'TechCrunchStack layer / Threat patternTerminal-Bench 3.0 Lands With Frontier Models Under 40% and Claude Opus 5 Leading at 43.5%Terminal-Bench (corroborated by Scale Labs and Turing blog; r/singularity 82up/17c)Stack layer / Threat patternGitHub's playbook for the orchestrator role: agents propose, deterministic checks decideGitHub BlogStack layer / ContrastOptiver puts 30–40% of its 950 engineers on platform work — and says better AI models now get more focus than lower latencyThe Pragmatic EngineerStack layer / Contrast"Not Worth Another Token": Pruning Deep-Research Agent Context Before Retrieval Cuts Token Usage Up to 73% With Little Quality LossarXiv (via HuggingFace Daily Papers)Stack layer / Contrastcathrynlavery/diagram-design Takes #1 on GitHub Trending With +2,855 Stars in a Day — 29 Editorial Diagram Types Shipped as a Claude Code PluginGitHub Trending