Fetching from the wire…
Public story · 2026-07-22 · high
Coherent input degrades attention worse than shuffled text, and models drop instructions silently instead of flagging the failure.
Why now: Both the cross-model context-rot findings and the arXiv paper's system-prompt cliff data were part of the July 22 research roundup.
Accuracy drops 30 to 50 percent before models hit their documented context limit, per new cross-model testing across GPT-4.1, the Claude 4 family, Gemini 2.5, and Qwen3.
The stakes are concrete for anyone using an LLM to monitor another LLM's long-running work. Opus 4.6 watching long agent transcripts on MonitorBench drops from 98.6% recall to 88% as context grows. The watchdog degrades right when the run gets long enough to need watching.
The degradation tracks semantic-similarity decay, not raw token count. Coherent, well-organized input degrades attention more than shuffled input, the opposite of what most builders would assume a clean document should do. Likely because coherent distractors compete for the model's attention in a way random noise can't.
A separate factorial study, per arXiv 2607.19257, ran 960 calls per model for format testing and 5,520 for scaling. Perfect-response rate fell to zero at 80 simultaneous system-prompt instructions. Recall holds near ceiling through 64-128k tokens, then falls off a cliff, one model losing 48 accuracy points at 128k.
The same paper found no reliable format advantage for markdown over plain text; winners were model-specific. Across 5,760 probes checking specifically for fabrication, models never invented an answer to cover a dropped instruction. They just quietly stop doing the thing, and nothing signals to the user what didn't run.
Each link below shares sources, entities, or timing with this story.
Cursor supports Gemini / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor supports Gemini); both cover GPT, LLM, Same, Then; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover CLAUDE, Gemini, GPT, Opus; overlapping topics (agent, drop, have).
Linked by a graph relationship (Gemini competes with Claude); both cover Claude, GPT, Opus, Qwen3; overlapping topics (agent, model, token).
Stanford benchmarked against Gemini / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Stanford benchmarked against Gemini); both cover Gemini, LLM, Opus; reported by the same outlet (arxiv.org).
Copilot uses GPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover CLAUDE, Gemini, LLM; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Claude, Gemini, GPT, Opus; overlapping topics (have, model).
Gemini built by Google / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini built by Google); both cover CLAUDE, Gemini, Same, Then; overlapping topics (agent, have).
Gemini competes with Claude / Shared entities / Same source domain
Linked by a graph relationship (Gemini competes with Claude); both cover CLAUDE, GPT, Same, Then; reported by the same outlet (arxiv.org).