Fetching from the wire…
Public story · 2026-08-24 · high
The same paper found naming a banned phrase in a prompt makes a model produce it more often.
Why now: The paper's finding surfaces alongside a Reddit thread where builders describe the identical failure without citing any research.
A production tender-response system's answer quality fell from 74% to 48%, per a new arXiv paper. The system prompt had swapped prose for nested XML tags.
That result cuts against the default most agent builders reach for, wrapping instructions in <rules> and <instructions> tags because vendor docs implied it helps. Structured markup still helped when the model was reading source documents. It only hurt once the instructions themselves got the same treatment.
The same paper found a second problem. Naming a forbidden construction in a prompt raises the odds the model produces it. Tell a model never to write a specific closing phrase, and that exact phrase turns up more often afterward.
The system still held up on the broader test. An LLM judge rated its answers as good as or better than the human-submitted version on 40 of 55 ground-truth sections. Of the 15 losses, only 6 traced to writing quality at all.
A 232-upvote r/ClaudeAI thread found the same pattern without citing any research. Correct Claude Code for adding ketchup to a coffee order. The fix gets written into the spec as 'make a coffee (without ketchup)' forever after. One reply in that thread disagrees, arguing the leftover negative constraint stops the model from re-adding what got removed. Both effects are probably real. The paper's numbers say the cost of the first outweighs the benefit of the second.
Each link below shares sources, entities, or timing with this story.
Claude Code uses WebFetch / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses WebFetch); both cover Claude Code, ClaudeAI, LLM, Reddit; reported by the same outlet (reddit.com).
Boris Cherny uses Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Boris Cherny uses Claude Code); both cover Claude Code, ClaudeAI, Reddit, There; reported by the same outlet (reddit.com).
Simon Willison released LLM / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Claude Code, ClaudeAI, There; reported by the same outlet (reddit.com).
Anthropic released Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude Code); both cover Claude Code, LLM, Nearly; overlapping topics (agent, answer, have, model).
Simon Willison released LLM / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Claude Code, ClaudeAI, Reddit; reported by the same outlet (arxiv.org, reddit.com).
Simon Willison released LLM / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Claude Code, ClaudeAI, Reddit; reported by the same outlet (reddit.com).
LLM uses OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover Claude Code, LLM; reported by the same outlet (arxiv.org).
Claude Code uses MCP / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses MCP); both cover ClaudeAI, LLM; reported by the same outlet (arxiv.org, reddit.com).