SourcesMoltbook Safety Paper: Alignment Vanishes in Self-Evolving AI SocietiesHuggingFace Daily Papers·medium signalXBlueskyLinkedInCopy link165 HF upvotes. Safety alignment degrades when agents interact in social networks. Safe individual agents produce unsafe collective behavior.SourceSource pageHuggingFace Daily Papers↳ Follow the threadPolicy dependency / Stack layerPaul Christiano Joins the OpenAI Foundation Board and Its Safety and Security CommitteeOpenAI BlogStack layer / ContrastIBM released Granite Time Series PatchTST-FM-r2 with a commercial-friendly licenseHugging Face Blog (IBM Research)Policy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivStack layer / ContrastZvi's second Astra pass: near-perfect compliance under obvious monitoring reads as metagaming, not alignmentDon't Worry About the Vase (Zvi Mowshowitz)Stack layer / Update threadDeepSeek released V4.1-Flash: a 552B causal encoder-decoder that activates 8B params on prefill and 16B on decodeDeepSeek (Hugging Face model card)Stack layer / Update threadGander Splits a Full-Duplex Omni Agent Into a Realtime Cerebellum and a Reasoning BrainarXiv 2609.08977Stack layer / Threat pattern16% of 3,171 public agent-harness setups carry a confirmed security defect, and 3.8% ship a skill that pre-approves your shellarXiv