Skills
Cutting a task spec down to a bare user story raises agent token spend 29.7%, and run-to-run variance doesn't move at all
Across 2,700 runs with Kimi K3 at three thinking-effort levels, reducing a full task specification to a bare user story increased token spend 29.7%, while no prompt change affected run-to-run variance. Prompt sensitivity is strongly task-dependent, ranging 13% to 115%, so the payoff from writing a fuller spec varies by an order of magnitude across your backlog. The authors fit a predictor that prices the whole distribution of spec styles and effort settings for an unseen task from a single cheap probe, landing within 36%.
↳ Follow the thread