ResearchMOSAIC safe multi-step tool useWweb·medium signalXBlueskyLinkedInCopy linkPost-training framework: plan-check-act/refuse loop with preference RL. 50% harmful behavior reduction, 20%+ refusal on injection attacks, preserves benign performance. Tested Qwen2.5-7B, Qwen3-4B, Phi-4. arXiv 2603.03205↳ Follow the threadPolicy dependency / Stack layerAgent Framework stops forwarding headers across redirects and revalidates file skill paths before useGitHubStack layer / Threat patternDSPy adds LocalInterpreter, a persistent CPython worker that the release notes explicitly refuse to call a sandboxGitHubStack layer / ContrastControlled agentic CAD comparison: six unattended runs, 16 failures, 9 of which the tool never reportedModelRiftPolicy dependency / Follow-up threadMOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 PointsarXiv 2609.11065Stack layer / Threat patternClaude Code 2.1.269 Ships a Plugin Eval Runner and a Knob to Raise the Workflow Tool's Concurrent Agent Cap to 256Anthropic (claude-code CHANGELOG)Stack layer / ContrastTrueFoundry open-sources TrueForge, an MIT-licensed agent harness pitched against Claude Managed AgentsTrueFoundryPolicy dependency / Stack layerA replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policiesGitHubPolicy dependency / Follow-up threadAnthropic adds an 'auto' permission policy that lets the server adjudicate every Managed Agents tool call, plus a terminal attach commandClaude Platform release notes