Controlled Crossover Experiment With 34 Participants Finds LLM Assistance Makes Human Requirements Inspection Worse, Not Faster
Requirements inspection is the practice of reading specifications for defects early, and this study put 34 participants through a crossover design where each inspected textual specifications both with and without LLM support, identifying requirements smells, classifying severity as nocuous or innocuous, and recording inspection time. Bayesian regression models per outcome variable, accounting for crossover-induced validity threats plus covariates and mediators, show LLM support negatively affected smell detection accuracy with no significant effect on classification or on task duration. So the assistant cost accuracy without buying speed, which is a useful counterweight to the assumption that adding a model to a human review loop is free.
↳ Follow the thread