为什么仅遗忘式机器遗忘需要记忆
Why Forget-Only Unlearning Needs Memorization
Luka Radi\'c · Vikrant Singhal · Amartya Sanyal
中文摘要
机器遗忘( machine unlearning)要求删除算法的输出接近于在剔除指定样本后从头重新训练得到的模型。本文研究仅遗忘式遗忘( forget-only unlearning):删除算法仅接收已训练模型和待遗忘样本,不包含任何留存数据或额外训练信息。我们追问:仅遗忘式遗忘是否总是可行?研究表明这取决于学习方法:不同数据集可生成相同的已训练模型,但在删除相同样本后所需的输出却截然不同。基于这一观察,我们推导了遗忘算法逼近重训练所需精度的下界,并针对若干标准学习算法给出具体实例。进而探讨仅遗忘式遗忘成功时必须满足的条件,并推导算法为应对任意删除请求而必须记忆的训练数据信息量下界。对于简单阈值学习器,所需信息量可达整个数据集规模,尽管常规训练仅保留一个边界点。总体而言,研究表明常规学习中丢弃的信息可能在后续删除时被需要,因此面向仅遗忘式遗忘设计的模型可能需要保留比标准训练更多的信息。
关键要点
- 01问题:仅遗忘式遗忘( forget-only unlearning)在仅有模型和待删样本、无留存数据的约束下是否可行尚不明确
- 02方法:证明不同数据集可生成相同模型但删除同一批样本后输出差异巨大,并据此推导遗忘精度与必须记忆的信息量下界
- 03结果:遗忘匹配重训练的精度存在下界,对阈值学习器等情形所需记忆量可达整个数据集规模
- 04局限:研究结论表明仅遗忘式遗忘模型需保留比标准训练更多的信息,与常规训练的信息丢弃策略相矛盾
解读
尚无解读。
原始英文摘要
arXiv:2610.10519v1 Announce Type: new Abstract: Machine unlearning asks for a deletion algorithm whose output is close to retraining from scratch without the selected forget examples. In this work, we study forget-only unlearning, where the deletion algorithm receives only the trained model and the examples to forget, with no retained data or extra training information. We ask whether forget-only unlearning is always possible. We first show that this depends on the learning method: different datasets can produce the same trained model but require very different outputs after the same examples are removed. Using this observation, we derive lower bounds on how accurately unlearning can match retraining and instantiate them for several standard learning algorithms. We then ask what must be true when forget-only unlearning succeeds. To this end, we derive lower bounds on what an algorithm must memorize about the training data to handle arbitrary deletion requests. For simple threshold learners, the required information can be as large as the entire dataset, even though ordinary training keeps only one boundary point. Overall, our results show that information discarded during ordinary learning may be needed later for deletion, so models designed for forget-only unlearning may need to retain more information than standard training does.