Research
Unlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an Agent
K-Bench scores unlearning across all six channels a ReAct agent exposes, including chain-of-thought, tool calls, tool observations, and elicited summaries, and counts a query as leaked if the secret appears anywhere. When the secret lives in the prompt or the retrieval store, TOFU and MUSE report no leakage while the deployed agent leaks it on 22-86% of queries, and on structured retrieval the secret survives verbatim in the tool-observation channel with the aggregate leak rate unchanged. When the secret is in the weights, none of twenty published methods demonstrably removes it, and the top-ranked method changes depending on the base model.
↳ Follow the thread