News
Google Research Found Frontier Models Store Almost Everything and Simply Fail to Retrieve It
A Google Research study evaluated 13 LLMs across more than 4 million responses using WikiProfile, a benchmark of 2,150 Wikipedia-derived facts tested in formats from exact context completion to multiple choice. Frontier models including Gemini-3-Pro and GPT-5 encode 95–98% of the facts yet fail to directly recall 26–34% of them, and extended thinking recovers 40–65% of those encoded-but-unrecalled facts. Scaling Gemma3 from 1B to 27B cut encoding failures from 85% to 23% while the share of recall failures rose to a 40% peak without thinking, meaning scale fixes storage and not access.
↳ Follow the thread