Personalized Coding-Agent Skills Barely Beat No Skill at All — Generic Pooled Skills Win Across 206 Sessions From 13 Developers
Researchers extracted developer preferences from real interaction traces via rule-based bootstrapping and evidence-grounded refinement, then replayed them against a trajectory-conditioned LLM developer simulator on 206 real developer-agent sessions from 13 developers. Personalized skills gave only small and inconsistent improvements over a no-skill baseline, while generic skills pooled across all developers produced the largest and most consistent gains — with personalization helping only when a developer's history contained multiple examples relevant to the future task. The practical read for anyone building per-user agent memory: broadly transferable procedural knowledge beats individualized preference capture, and the personalization step may not pay for itself.
↳ Follow the thread