"Skill Use or Skill Theater": agents report using skills that provably did not change their decisions
The BACKTRACE framework pairs every skill-conditioned answer with a matched no-skill counterfactual and intervenes on skill meaning, wording, identity, content, and assignment, eliciting the agent's attribution only after the answer is committed. Across model families and domains on BACKROOMBench, stated skill use stays stable while actual causal reliance varies wildly, producing both silent uptake and purely performative use — and observational detectors (direct skill-use claims, text mentions, trace similarity, an LLM judge) fail to identify which decisions actually depended on the skill. In multi-agent settings skill influence survives communication after its source is lost, and no-skill teams still name skills and sources that were never supplied, which is a direct warning for anyone auditing agent-skill marketplaces by reading traces.
Source
↳ Follow the thread