NCP-Bench: GPT-5.2 Holds Narrative Commitments Only 42% of the Time After 20 Turns, With Fact-Conflict Rates of 40–68% Across Models
arXiv 2608.08160 (submitted 8 Aug 2026) formalizes 'Narrative Commitment Preservation' and builds a benchmark of 100 narrative environments derived from movie synopses, each with a structured spec — trajectory, commitments, initial facts — that can be checked automatically throughout a player-agent/narrator-agent interaction. The headline result is a long-horizon consistency gap that is not visible in text quality: the best model, GPT-5.2, survives only 42% of runs past 20 turns, fact-conflict rates run 40–68% across models, and only isolated runs satisfy all commitments inside a 100-turn limit. The generalizable lesson for anyone building long-running agents is that fluent output is uncorrelated with commitment preservation under adversarial user intervention.
↳ Follow the thread