Dylan Castillo investigates whether labs are training on the pelican benchmark
Dylan Castillo / Hacker News·high signal
Dylan Castillo published a deep dive testing the long-running suspicion that AI labs deliberately train models to draw SVG pelicans on bicycles — Simon Willison's informal benchmark. The post hit 614 points and 234 comments on Hacker News and was amplified by Willison himself. It's the sharpest public case study yet of what happens when an informal eval becomes popular enough to leak into training data, which is the failure mode of every viral benchmark.