Research
Skill Habit Formation Makes Agents Reproduce Their Answer on All 456 Repeated Dispatches, Where Reasoning Arms Managed 11-26 of 42
Running 42 tasks three times each, 38-74% of agent answers disagreed across runs depending on the model, and 95.3-97.2% of generated tokens went to re-deriving known plans. The authors mine execution history for deterministic skill candidates that claim an input region and pass four gates, including a trace check against a reference's own run-to-run variance. On text-to-SQL the habit-formed variant reproduced on all 456 repeated dispatches, was non-inferior to every arm it replaced (p<0.0001), and used 14-56% fewer tokens.
Source
↳ Follow the thread