AI2's TutorMoments finds tutors default to over-helping — and one prompt line moves the needle
Allen Institute for AI published TutorMoments, a benchmark that replays real decision points from 462 de-identified one-on-one math tutoring transcripts, with over 1,500 teacher-marked moments and thousands of annotations from 27 educators, and asks a model to take the tutor's turn. Across seven models tested, the default 'helpful assistant' behavior was to scaffold too much — making problems easier instead of letting students reason — and performance improved substantially once the prompt explicitly named the scaffolding-versus-rigor trade-off, though variance across models stayed wide. The transferable lesson for builders: helpfulness training actively fights any task where withholding is the correct action, and you have to state the trade-off in the prompt.
↳ Follow the thread