Fetching from the wire…
Public story · 2026-07-23 · high
The essay drew 614 points and 234 comments on Hacker News, and Simon Willison boosted it himself, reviving worries that fame corrupts any test this useful.
Why now: Willison boosted Castillo's essay the same week a Res Obscura piece on AI slop and a new /no-ai-slop tool were drawing similar traffic, as of July 23.
Dylan Castillo published an essay asking whether AI labs train their models to draw a good SVG pelican riding a bicycle, Simon Willison's informal test for judging new model releases. The post hit 614 points and 234 comments on Hacker News, and Willison amplified it himself.
Willison built the pelican prompt because it was obscure, a private probe models couldn't have memorized from training data already. That obscurity is the entire value of an informal benchmark: it measures whether a model can generalize to something new, not whether it has already seen the answer. Once a test goes viral enough that labs' own training pipelines are likely to ingest examples of it, that distinction collapses.
Castillo's piece is the clearest public case study of that collapse happening to a specific, widely cited eval. Every screenshot of a good pelican SVG posted online becomes a candidate training example for the next model, whether any lab intends it or not. A benchmark's shelf life runs out roughly in proportion to how popular it gets, which is a bad trade for anyone using it to compare models honestly.
The same day, a Res Obscura essay pulled 438 points arguing that quality non-fiction writing is structurally the opposite of AI slop, and a tool called /no-ai-slop shipped alongside it. Slop detection is turning into its own tooling category instead of staying a comment-section complaint. Same underlying mechanic as the pelican test: once you can name what good looks like in public, someone starts optimizing straight at the name.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Simon Willison, SVG, Willison; earlier Simon Willison coverage from 2026-06-19.
Simon Willison released Datasette / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released Datasette); both cover Simon Willison, Willison; earlier Simon Willison coverage from 2026-06-19.
Simon Willison uses Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison uses Claude Code); both cover Hacker News, Simon Willison, Willison; overlapping topics (comment, willison).
Simon Willison uses Claude Code / Shared entity: Hacker News / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison uses Claude Code); both cover Hacker News; overlapping topics (comment, point, slop).
Simon Willison uses Claude / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison uses Claude); both cover Hacker News, Simon Willison; earlier Hacker News coverage from 2026-06-10.
Simon Willison released LLM / Shared entity: Hacker News / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Hacker News; overlapping topics (arguing, comment, slop).
Simon Willison uses Claude Code / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison uses Claude Code); both cover Simon Willison, Willison; earlier Simon Willison coverage from 2026-07-21.
Simon Willison released LLM / Shared entities / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover Simon Willison, Willison; earlier Simon Willison coverage from 2026-07-19.