Sources
'LLMs Get Lost in Evolving User Intent' — Strong Static Benchmark Scores Don't Transfer When the User Changes Their Mind Mid-Conversation
Jihoon Tack, Philippe Laban and Jennifer Neville (arXiv 2607.20734, 22 July) introduce a framework that converts any single-turn benchmark into a multi-turn conversation where the user's objective shifts, clarifies, or reverses — turning existing benchmarks into evolving-intent testbeds with no new annotation. Their headline result: 'strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families.' This is the most directly actionable eval finding of the week for anyone shipping conversational agents, because it says leaderboard rank is close to uninformative about the failure mode users actually hit.
↳ Follow the thread