about
The Poison of (LLMs) Alignment (arxiv.org)
17 points by throwaway888abc on Aug 30, 2023 | hide | past | pdf | 3 comments on HN

In plain words: Safety answers that teach a model to refuse harmful requests also seem to poison the rest of its training data, so the study compared tuning with and without them. Models trained with those answers scored 4-33% lower on reasoning tests than training without them.

Abstract · The Poison of Alignment

From the perspective of content safety issues, alignment has shown to limit large language models' (LLMs) harmful content generation. This intentional method of reinforcing models to not respond to certain user inputs seem to be present in many modern open-source instruction tuning datasets such as OpenAssistant or Guanaco. We introduce a novel insight to an instruction-tuned model's performance affected by the presence of alignment in supervised fine-tuning dataset. To be specific, we noticed that alignment acts as if it is poisoning the instruction dataset. Experimentally, we demonstrate that aligned answers significantly worsen the performance of the resulting fine-tuned model's on various reasoning benchmarks such as Big Bench (BBH), Massive Multitask Language Understanding (MMLU), Human Eval, and Discrete Reasoning Over Paragraphs (DROP), performing worse than the counterpart tuned without alignment by 4-33%.

Aibek Bekbayev, Sungbae Chun, Yerzat Dulat, James Yamazaki
arXiv:2308.13449 · cs.CL · submitted Aug 25, 2023
abstract · pdf · html

add comment on HN
Also discussed: Aug 2023 (2 points, 0 comments) · Aug 2023 (1 point, 1 comment)

It should not be surprising that asking an AI to not tell the truth leads to the quality of its' answers deteriorating.

I did not say they asked the AI to lie. I did they asked it not to give a correct and truthful response.

Maybe this explains much of the deterioration of ChatGPT that has been reported.

It is known that alignment tax affects smaller models but in larger models (>~100B parameters) the "tax" starts to become negative at least when trained in RLHF: https://arxiv.org/pdf/2204.05862.pdf

The largest Llama2 model has 70B parameters. They ran the experiment with 7B Llama2.

I would expect any pre-prompt or fine tuning training that is relevant to the subsequent prompts to improve the results (relative to the pre-prompting or tuning).

Likewise, any pre-prompt or fine tuning training that is NOT relevant to the subsequent prompts adds unhelpful and highly non-contextual complexity to the models job, so likely to reduce the quality of response.