AUSO treats agent skills as a lifecycle and scores their benefit action by action
AUSO (Action-level Unified Skill Optimization) argues skills should first teach, then support capability formation, then be invoked only when they improve an individual decision, and that existing methods either keep skills external, fully internalize them, or pick between objectives using noisy task-level success rates. It starts by learning jointly from teacher guidance and environment outcomes, shifts to outcome-based policy optimization, then evaluates each sampled action under both skill-conditioned and skill-free contexts so beneficial skill-sensitive actions get stronger updates and distracting ones are suppressed. Reported gains hold on ALFWorld, WebShop and SearchQA, including out-of-distribution generalization.
Source
↳ Follow the thread