Policy dependency / Stack layer
SkillRise Collapses Extract-Retrieve-Execute Into One RL Loop: a Single Policy Alternates Between Solving Tasks and Rewriting Its Own Skill Document
arXiv / HuggingFace Daily Papers
Stack layer / Threat pattern
Agent Harness Distillation Steals Inference-Time Harness IP From Autonomous Multi-Agent Systems Through Black-Box Queries
arXiv
Policy dependency / Stack layer
AgentSnare Traps Autonomous Pentest Agents in a Grown Decoy: Zero Real-Target Exploits Across 45 Attacker-CVE Pairs
arXiv 2607.26998
Stack layer / Contrast
Frontis-MA1 (35B) Reaches 71.21% MLE-Bench Lite Medal Average on a Single 12 GB RTX 4090, Weights Released
arXiv
Stack layer / Threat pattern
Grayscale and Color Inversion Bypass All Three Major Commercial Image-Moderation APIs
arXiv
Stack layer / Contrast
The 'Locksmith Loop' Validates Agent-Migrated COBOL-to-Java Code Against a Deterministic Oracle, Hitting 91.9% Branch Coverage
arXiv
Stack layer / Contrast
DenseOn and LateOn: Fully Open 149M Retrievers Set Size-Class SOTA at 56.20 and 57.22 nDCG@10 on BEIR
arXiv 2607.27178
Stack layer / Contrast
First post-compromise incident-response benchmark finds agents fail to proactively investigate silent intrusions
arXiv 2607.26791