VoicesWillison 5 Posts: SWE-bench Analysis GGML HuggingFace Gemini 3.1 Prosimonwillison.net·high signalXBlueskyLinkedInCopy linkWillison covered SWE-bench Feb leaderboard (Opus 4.5 leads 76.8%), GGML joins HuggingFace, Gemini 3.1 Pro, Taalas 17K tok/s custom silicon. Called GGML joining HF hard to overstate.SourceSource pagesimonwillison.net↳ Follow the threadStack layer / Update threadIBM Released Granite 4.2 as Dense Reasoning Models at 3B, 8B and 30B Under Apache 2.0 With a 512K ContextHugging Face (ibm-granite), via r/LocalLLaMAStack layer / ContrastGoogle Cloud Added Deferred-Execution Agent Pricing at Half the Inference Cost and Pooled Quota Across Business Apps and IDEsGoogle Cloud BlogPolicy dependency / Stack layerLMSM Ports the Linux Security Modules Split to LLM Serving, Cutting HarmBench ASR 39.20% to 3.32% at 98.14% ThroughputarXiv 2608.25697Threat pattern / Follow-up threadMeta's Scrapped 'Project OT' Shows Unsupervised Agents Taking Actions 'Humans Are Unlikely to Execute'Ars TechnicaStack layerEVE Online begins migrating 2.4 million lines off Stackless Python 2.7 after 16 yearsSimon Willison's Blog / CCP GamesStack layerFuzzingBrain-Bench Scores Open-Ended Bug Discovery; Claude Opus 4.8 Crashes 60 of 77 Targets but Scores 196/579arXiv 2608.25158Stack layerGemini 3.5 Transcribe Claims 2.6% Word Error Rate and Ships Two Separate APIs for Live and Recorded AudioGoogle BlogStack layerLong-Time Claude Max Subscribers Say the Chat Product Regressed Into Preamble and Double-Affirmationr/ClaudeAI