Fetching from the wire…
Public story · 2026-07-23 · high
The essay drew 614 points and 234 comments on Hacker News, and Simon Willison boosted it himself, reviving worries that fame corrupts any test this useful.
Why now: Willison boosted Castillo's essay the same week a Res Obscura piece on AI slop and a new /no-ai-slop tool were drawing similar traffic, as of July 23.
Dylan Castillo published an essay asking whether AI labs train their models to draw a good SVG pelican riding a bicycle, Simon Willison's informal test for judging new model releases. The post hit 614 points and 234 comments on Hacker News, and Willison amplified it himself.
Willison built the pelican prompt because it was obscure, a private probe models couldn't have memorized from training data already. That obscurity is the entire value of an informal benchmark: it measures whether a model can generalize to something new, not whether it has already seen the answer. Once a test goes viral enough that labs' own training pipelines are likely to ingest examples of it, that distinction collapses.
Castillo's piece is the clearest public case study of that collapse happening to a specific, widely cited eval. Every screenshot of a good pelican SVG posted online becomes a candidate training example for the next model, whether any lab intends it or not. A benchmark's shelf life runs out roughly in proportion to how popular it gets, which is a bad trade for anyone using it to compare models honestly.
The same day, a Res Obscura essay pulled 438 points arguing that quality non-fiction writing is structurally the opposite of AI slop, and a tool called /no-ai-slop shipped alongside it. Slop detection is turning into its own tooling category instead of staying a comment-section complaint. Same underlying mechanic as the pelican test: once you can name what good looks like in public, someone starts optimizing straight at the name.
Each link below shares sources, entities, or timing with this story.
Allen Bargi's August 15 post hit 302 points arguing that AI collaboration rewards context-sharing, examples, and feedback over precise instruction (Hacker News). The pushback holds that the piece conflates management with leadership. mikeocool calls it "the most low effort ver...
Simon Willison doesn't hand out superlatives. So when he writes that Z.ai's GLM-5.2 is "probably the most powerful text-only open weights LLM," that's worth stopping for. His June 17 evaluation walks through a 753B-parameter Mixture-of-Experts model with 40B active params, a 1...
Satya Nadella said companies routing everything through a single proprietary lab may not survive. His argument: you hand that lab your most sensitive business context, and the lab can turn it against you as a competitor. His prescription is an orchestration layer — keep the ha...
Willison launched datasette-apps (0.1a2) on June 18, hosting self-contained HTML+JS apps in a sandboxed iframe that run SQL against your data, read-only by default. He frames it as "Claude Artifacts reimagined for Datasette," artifacts backed by a JSON API to a relational data...
Simon Willison has been writing software for over 25 years. He's one of the most disciplined, transparent engineers in the Python ecosystem. And yesterday he published an essay admitting he no longer reviews every line of code that Claude Code generates for his production proj...
Promptwatch's tracking shows the share of ChatGPT search queries using site: sat at 0.3-0.5% for weeks, dipped to 0.15% on August 3-5, then jumped to 16-17% on August 8, two days after OpenAI said it was making GPT-5.6 Sol "more reliable with facts." Simon Willison Willison co...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.