Voices
Willison's One Complaint About OpenAI's Math Release: 'I Want to See the Prompts They Used'
Willison credits OpenAI with 'a decent level of transparency' for shipping Lean formalizations, a paper, and a reconstruction of the reasoning traces, then names the two gaps that matter to practitioners: the prompts are withheld, and nothing in the release says how many $2,000 failed attempts preceded the ten successes. That second question is the whole cost story, since a 10% hit rate turns a $2,000 proof into a $20,000 proof. For builders reading capability announcements, this is the reusable check: published cost-per-success is meaningless without the denominator.
↳ Follow the thread