Reversible Unlearnable Examples Combine Copyright Poisoning and Watermarking, Which Interfere When Naively Stacked
This paper (2608.06211, submitted 2026-08-06) notes that the two threats to training-data copyright — unauthorized model training and malicious data leakage — resist a straightforward combination of availability attacks and watermarking because the two techniques interact negatively. The proposed mechanism generates perturbations that force a model to learn uncorrelated features by minimizing mutual information between the model's input and output, and separately uses a dual-extraction strategy with two distinct watermark extractors so the unlearnable perturbation does not destroy watermark recovery. Experiments span ImageNet, CIFAR-10, and Pets. For builders shipping datasets, this is the emerging shape of data-provenance defense: poison-plus-watermark rather than either alone.
↳ Follow the thread