Research
MonitrLLM Pilot: Students Rate ChatGPT 4.19/5 While Failing 23.1% of Their Actual Goal Tasks, and Multi-Turn Conversations Fail 2.5x More Often
MonitrLLM is open-source infrastructure that links full conversation transcripts to user-reported task intent and outcome assessment, treating all three as primary evaluative signals rather than optional metadata — filling the gap between capability benchmarks, unlabeled conversation corpora, and thumbs-up feedback. A two-week pilot with 26 college students produced 206 evaluation reports with transcripts. Despite average satisfaction of 4.19/5, participants hit a 23.1% failure rate on goal tasks, and multi-turn conversations failed at 2.5 times the rate of single-turn exchanges — reframing extended interaction as a signal of difficulty rather than engagement.
↳ Follow the thread