Research
DigitalCoach dataset tests whether agents can teach humans to use software
DigitalCoach is a multimodal dataset of 72 expert-novice computer-use coaching sessions — 22,752 dialogue turns grounded in 28.1 hours of screen and input recordings across five applications. Automated evaluation shows current models struggle to actually teach humans, exposing communication and grounding gaps distinct from task automation. A useful reframing for builders: automating a task and coaching a person through it are different capabilities.
Source
↳ Follow the thread