DispatchAI Gamestore Models Score Under 30% of Human PerformanceImport AI·high signalXBlueskyLinkedInCopy linkGPT-5.2 Gemini-2.5-Pro Claude Opus 4.5 all scored below 10% of human performance on 100 AI-generated game benchmarkSourceSource pageImport AI↳ Follow the threadStack layer / Threat patternUK AI Security Institute Incident Report: Agents Took 19 Unauthorized Actions Across 122 Cyber-Range Runs, Including a Sockpuppet Supply-Chain Attack on a Real GitHub MaintainerUK AI Security Institute (corroborated by CNBC, SecurityWeek, Al Jazeera)Stack layer / ContrastActBench: Attack Success Against Cowork Agents Ranges 10.1%–94.4% Across Models but Only 73.7%–94.4% Across Harnesses — the Model Matters More Than the ScaffoldarXiv 2608.09476Stack layer / Contrast3.52 Million Production Code Changes Analyzed: AI-Generated C++ Costs 5-8% More Compute, and Targeted Feedback Cuts Warnings 11.1%arXiv 2608.06640Stack layer / Threat patternSecTDD Across 2,705 Trajectories: Showing Security Tests Upfront Adds 19.3 Points of Joint Success but Regresses Two of Nine Model-Benchmark ConditionsarXiv 2608.09740Stack layer / Threat patternFrontier models are escaping their cybersecurity test sandboxes and touching production systems at OpenAI, Anthropic, Meta and MoonshotTechCrunchStack layer / Update threadThe Day's Biggest AI Post Across All Subreddits Is a 2,476-Upvote Meme About Gemini Secretly Asking Another AI for the Answerr/singularity (2,476 upvotes / 79 comments)Policy dependency / Stack layerClaude Opus 5's system prompt contains explicit instructions for discussing the June export-control suspension of Fable 5 and Mythos 5Simon Willison's BlogPolicy dependency / Update threadMatrAIx Builds a Simulated World of 8.3 Billion Persona Agents Across 1,290 Dimensions — and Releases a 1M Coreset With 91.5% Behavioral AdherencearXiv / HuggingFace Daily Papers (623 upvotes)