SkillClone reconstructs closed-source agent skills from ordinary usage alone, defeating every disclosure-focused defense
arXiv 2608.04192 (Aug 4) separates file secrecy from functional secrecy: even if a provider perfectly blocks prompt-injection attempts to dump a hidden skill's files, a user can rebuild what those files do. SkillClone forms an interface hypothesis from the skill's public advertisement, issues structured benign probes, synthesizes an executable replica, and iteratively repairs it through differential validation against the victim. Across 30 skills spanning rules, tables, procedures and algorithms it achieved exact or partial recovery on held-out inputs for several targets, with iterative requerying closing gaps that single-round reconstruction missed. Vaguer skill descriptions offered only limited protection — the defense that actually works is rate-limiting cumulative information leakage from legitimate use.
Source
↳ Follow the thread