Reasoning Gets Harder for LLMs Inside Dialogue: Multi-Turn Context Accumulation Degrades Agent Performance on Hard Tasks
arXiv·medium signal
Benchmark study shows LLMs achieving strong scores on isolated reasoning tasks exhibit significant degradation when identical tasks are embedded in multi-turn dialogue, with the performance gap widening on harder problems as context accumulates. This directly impacts agent systems that conduct iterative reasoning across tool call chains and extended conversation histories. Current agent benchmarks — which test single-shot task completion — are likely reporting inflated capability estimates relative to realistic deployed performance.