Skills
How to test whether a model is overfit to your favorite eval: build a factorial grid around it and measure the cell effect
Dylan Castillo (2026-07-22) tested whether labs train specifically for the famous pelican-on-a-bicycle SVG prompt by building an 8 animals x 6 vehicles grid (48 prompts), running each three times across 7 frontier models for 1,008 SVGs, and scoring with an LLM judge. Per-lab pelican effects landed between -0.11 and +0.14 judge points and no pelican-bicycle cell effect cleared p < 0.05 — no evidence of targeted optimization. The transferable skill is the method, not the result: to check contamination on any pet benchmark, vary both axes around the famous cell and test whether the specific interaction beats the marginal effects, rather than eyeballing single outputs.
↳ Follow the thread