Per-Turn Credit Assignment Is the Wrong Axis When the Verifier Only Sees the Last Write
The paper defines verifier information density V_d = k/C, the fraction of an agent's C-step causal chain whose per-turn correctness the verifier actually exposes, and shows terminal-state verifiers sit deep in the low-V_d regime where targeting credit onto important turns cannot help. In shared-rollout comparisons on tau^2-bench a continuous dense reward spread uniformly beats sparse binary outcome reward, while concentrating the same advantage on progress turns is exactly as harmful as concentrating it on random turns — targeting is second-order. The mechanism is coverage: terminal-state verification collapses the signal to a single final-write turn (k=1 in 98% of rollouts) while success needs a 5-8 step chain, and measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3 against a synthetic phase boundary at V_d* ≈ 0.8.
↳ Follow the thread
No related signals yet.