about
Guardrail Baselines for Unlearning in LLMs (arxiv.org)
2 points by PaulHoule on Mar 12, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of retraining a model to forget facts, you can just tell it to refuse or block the unwanted answers. These simple guardrails worked about as well as costly retraining, and they exposed flaws in how forgetting is currently scored.

Abstract

Recent work has demonstrated that finetuning is a promising approach to 'unlearn' concepts from large language models. However, finetuning can be expensive, as it requires both generating a set of examples and running iterations of finetuning to update the model. In this work, we show that simple guardrail-based approaches such as prompting and filtering can achieve unlearning results comparable to finetuning. We recommend that researchers investigate these lightweight baselines when evaluating the performance of more computationally intensive finetuning methods. While we do not claim that methods such as prompting or filtering are universal solutions to the problem of unlearning, our work suggests the need for evaluation metrics that can better separate the power of guardrails vs. finetuning, and highlights scenarios where guardrails expose possible unintended behavior in existing metrics and benchmarks.

Pratiksha Thaker, Yash Maurya, Shengyuan Hu, Zhiwei Steven Wu, Virginia Smith
arXiv:2403.03329 · cs.CL · submitted Mar 5, 2024 · updated Jun 11, 2024
abstract · pdf · html · Preliminary work, accepted to ICLR workshop SeT-LLM 2024

add comment on HN