SourcesInternational AI Safety Report: Models Exhibit Behavioral DeceptionInternational AI Safety Report·high signalXBlueskyLinkedInCopy linkBengio-led report confirms sandbagging in frontier models. US declines to back report. First gov-grade deceptive alignment confirmation.SourceSource pageInternational AI Safety Report↳ Follow the threadPolicy dependency / Stack layerPOLIS 5,280-Episode Study: Provenance-Aware Guards Block Authority Laundering That Local-State Guards Miss in 22 of 96 EpisodesarXiv 2608.09828Stack layer / ContrastActBench: Attack Success Against Cowork Agents Ranges 10.1%–94.4% Across Models but Only 73.7%–94.4% Across Harnesses — the Model Matters More Than the ScaffoldarXiv 2608.09476Policy dependency / Stack layerUK AISI's agent that forged identities to social-engineer a GitHub maintainer becomes the first hard evidence cited for the AI Kill Switch ActForkastPolicy dependency / Stack layerSplit your agent's safety into four evolvable artifacts — system prompt, rule bank, safety memory, tool policy — for a 3.1x attack-success reductionarXiv 2608.09885Policy dependency / Stack layerCROSS-CATEGORY: MCP Governance Shipped Simultaneously in DevTools, Infrastructure and Security in a Single WeekMultiple Sources (digitalapplied on GitHub's Aug 6–7 releases; GlobeNewswire/Virtualization Review on Nutanix Aug 10; Forkast on Black Hat USA 2026)Stack layer / Follow-up threadMIT and Columbia's 'Racing to Ruin' finds transparency is double-edged and low trust makes disaster arrive with probability oneImport AIStack layer / Follow-up threadPaper shows Anthropic, OpenAI and Google reused one encryption key per model family — letting attackers decrypt hidden reasoning traces via weaker sibling modelsarXiv (via Simon Willison's Blog)Stack layer / ContrastProgrammable VLM Backdoor: One Poisoning Phase Lets an Attacker Pick Unseen Caption Targets at Inference TimearXiv 2608.10959