Toketive: 38.6% of Alternative Tokenizations of the Same String Bypass Knowledge Editing and Unlearning in Open-Weight LLMs
arXiv·medium signal
arXiv 2609.29045 shows that an input string has many valid tokenizations, and non-canonical ones route around localized edits and unlearning. Across five LLMs, six datasets and six editing and unlearning methods, 38.6% of alternative tokenizations bypassed the modification. The attack needs only the released model, with no pre-edit model, training data or auxiliary classifier. Any claim that an open-weight release has 'removed' knowledge should be tested beyond canonical tokenization.