SourcesPyVision-RL Solving Agent Interaction CollapsearXiv·medium signalXBlueskyLinkedInCopy linkAddresses RL-trained agents learning to reduce tool usage. Oversampling-filtering-ranking rollout strategy sustains interaction.SourceSource pagearXiv↳ Follow the threadStack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Policy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Stack layerA Jailbreak SoK Finds Low Final-Response Attack Success Hides Compromised Planning, Memory, and Tool StatearXiv 2609.12413Stack layer / Threat patternUnlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an AgentarXiv 2609.12808Stack layer / ContrastReflexion-Style Verbal Memory Sometimes Lowers Success Versus Plain Retry, and Replay Experiments Show WhyarXiv 2609.12404Stack layer / ContrastTwo-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gapsarXivStack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Policy dependency / Stack layerCodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the TimearXiv 2609.12464