Research
Sessions With Identical Quality Ratings Differ by Up to 70x in Interaction Cost
A productivity-oriented evaluation framework scores human-AI collaboration as outcome quality relative to interaction cost, across two datasets spanning four tasks, on the argument that completion time and outcome quality treat the collaboration itself as a black box. Sessions rated identically on quality differ by up to 70 times in interaction cost, the quality-cost relationship flips by task (some reward extended interaction, others fast convergence), and subjective user ratings are not a reliable substitute for productivity. Productive sessions are characterized by the agent probing earlier and the user spending less effort repairing the interaction.
↳ Follow the thread