Skills
Full-trace LLM judges of coding agents show collider bias and measure relevance, not causal contribution
This measurement framework separates three things process evaluation routinely conflates, action prediction, task uncertainty and step attribution, and instantiates step-level causal attribution with SCAE, a replay-based estimator over a structural causal model of agent execution. Across 499 file-localization episodes from 12 repositories it finds next actions are driven mainly by execution provenance rather than code-graph transitions, that uncertainty is structured at the task level rather than the step level, and that judges given the full trace exhibit systematic collider bias. If you score agent runs with an LLM judge that sees the whole trajectory, the score is measuring semantic relevance.
↳ Follow the thread