Agent Harness Distillation Steals Inference-Time Harness IP From Autonomous Multi-Agent Systems Through Black-Box Queries
Agent Harness Distillation (arXiv 2607.28147, July 30) formalizes a new security problem: the inference-time harness that coordinates reasoning and action in autonomous multi-agent systems like Hermes represents substantial engineering IP, and it can be extracted through black-box interaction alone. AHD works in two stages — pre-distillation infers harness behaviors from target responses to build an initial harness, then post-distillation iteratively refines it to match the target's behavioral patterns. Experiments across real-world AMAS and multiple backbone LLMs show substantial leakage; the authors propose a deception-based defense that degrades extraction while preserving the protected agent's utility. Prior IP-leakage work assumed static, pre-configured architectures, so this extends the threat to systems whose structure emerges at inference.
Source
↳ Follow the thread