Encrypted Reasoning Traces Are Not Actually Private: Researchers Scraped 315,320 Blocks From Public Repos and Recovered 367 PII Artifacts and 182 Credentials
arXiv 2608.09867 (Aug 10, cs.CR) from a team including Ilia Shumailov, Jonas Geiping, and Maksym Andriushchenko shows the encrypted reasoning blocks that Anthropic, OpenAI, and Google return via API can be decoded. The core flaw: within each provider's ecosystem the encrypted blocks are fully interchangeable across sessions, users, and models, so an attacker can inject a trace produced by a capable model into a weaker, less-safeguarded variant and force plaintext decoding. From 315,320 blocks harvested out of public repositories they recovered 367 PII artifacts and 182 credentials — meaning anyone who has committed raw API responses containing encrypted reasoning to a public repo has likely leaked the model's hidden chain of thought along with whatever was in it.
↳ Follow the thread