Logan Kilpatrick tells AI product teams to spend more than 25% of their time writing benchmarks
X/Twitter (Logan Kilpatrick)via @OfficialLoganK·low signal
Google AI Studio lead Logan Kilpatrick argued that AI-product teams should spend over a quarter of their time writing benchmarks and making sure model labs care about them, because a lab tends to improve on whatever it measures. He has made a version of this argument before ('every company building on top of AI should be making their own benchmarks'), but the 25% figure is new. On a launch day where vendor benchmark tables disagree with each other, the practical step is a private eval suite for your own tasks.