Research
The Delegation Blind Spot: 4,800 Requests Show an Agent's Task Success Doesn't Reveal Which Product Change Its User Would Value
Gupta audits whether agent traces can identify user product preferences. All 36 conservative intervals stayed unresolved across two pinned model snapshots despite different execution accuracy. Recording supplied preferences resolved 3 of 9 comparisons per model, and a deterministic extractor with no model calls resolved 7 of 9, so the model-generated reports themselves added uncertainty. The paper proposes a source-labeled decision receipt and argues for keeping structured decision inputs before collecting more telemetry. The data is fully synthetic, with no human participants.
Source
↳ Follow the thread