Hacker News
Anthropic Says Its Safety Evals Are Losing Signal Because Task Benchmarks Have Saturated
Buried in the August risk report is an admission with more downstream consequence than the headline risk ratings: Anthropic reports reduced confidence in its own safety evaluations because task-based benchmarks have saturated and no longer capture capability improvements. The same document says internal AI-assisted R&D is significantly faster but 'not yet by a factor of 2,' with acknowledged measurement difficulty — a notably deflationary number from the lab with the strongest incentive to report a larger one. HN commenters read the pair as evidence that both capability measurement and capability itself are progressing slower than the funding curve implies.
↳ Follow the thread