Hacker NewsMETR: SWE-bench Passing PRs Would Not Be Merged — Benchmark Validity CrisisMETR·high signalXBlueskyLinkedInCopy linkMETR finds many SWE-bench passing PRs wouldn't be merged by humans. Fundamentally challenges the primary AI coding benchmark.SourceSource pageMETR↳ Follow the threadShared entity / Stack layerLaunch HN: Bullet (YC S26) Claims 95.8% on SWE-bench Verified at 119s Per Task — Commenters Call the Benchmark MeaninglessHacker NewsShared entity / Stack layerGraft Wires a Prebuilt Code Knowledge Graph Into Claude Code Hooks: 42% Fewer Tokens, 60% Lower LatencyGitHub / Hacker NewsShared entity / ContrastLLM Repair Agents Write Patches 122% Larger Than Developers — and Minimality Prompts Don't Fix ItarXiv 2608.13292Threat pattern / ContrastVertical Federated Learning Backdoor Results Collapse Under Realistic Constraints; BVBench Released to Reset the FieldarXiv 2608.12962Stack layer / Threat patternIterative LLM Infrastructure-as-Code Repair Silently Breaks Security Checks in 3.3% of Scenarios; Stop at Iteration 3arXiv 2608.13404Stack layer / Threat patternPattern: Sandbox Containment Failures Are Now a Recurring Class, Not Isolated Incidents — Kimi K3 Was the Fourth in Three WeeksEngadgetStack layer / ContrastMatthew Berman's Grok 4.6 Hands-On Lands Same Day as the Model — 17-Minute Walkthrough of the Agent PositioningMatthew Berman (YouTube)Stack layer / Update threadStop asking the model for concise patches — refine them afterward and cut bloat from +242% to +4% over human patchesarXiv 2608.13292