Research
Adapting Agents to 30 Individuals Over 600 Tasks Lifts Solo Success 4.5-20.9% Within Tens of Tasks
Agents trained on population-scale data encode the average practitioner, but professional work lives in the departure from that average, and users cannot fully specify their criteria up front. TAHI treats cross-session human-agent interaction as the signal, folding it into agent context and weights through an evolving rubric module that crystallizes each user's training and evaluation criteria. Adapted to 30 individuals across writing and visual creation on 600 total tasks, agents improved solo task success 4.5-20.9% within only tens of tasks, and the rubric module doubles as an annotation tool producing evaluation rubrics that catch 16.0-22.3% more failures than rubrics from LMs or humans alone.
↳ Follow the thread