In plain words: Ten language models were asked to write fake news articles from a set of false narratives, then scored on writing quality, whether they pushed the lies, and if they added safety warnings. They produced convincing articles that agreed with dangerous false claims.
Abstract
Automated disinformation generation is often listed as an important risk associated with large language models (LLMs). The theoretical ability to flood the information space with disinformation content might have dramatic consequences for societies around the world. This paper presents a comprehensive study of the disinformation capabilities of the current generation of LLMs to generate false news articles in the English language. In our study, we evaluated the capabilities of 10 LLMs using 20 disinformation narratives. We evaluated several aspects of the LLMs: how good they are at generating news articles, how strongly they tend to agree or disagree with the disinformation narratives, how often they generate safety warnings, etc. We also evaluated the abilities of detection models to detect these articles as LLM-generated. We conclude that LLMs are able to generate convincing news articles that agree with dangerous disinformation narratives.
Ivan Vykopal, Matúš Pikuliak, Ivan Srba, Robert Moro, Dominik Macko, Maria Bielikova
arXiv:2311.08838 · cs.CL · submitted Nov 15, 2023 · updated Feb 23, 2024
abstract · pdf · html
10 models, 20 disinfo narratives
A few choice ones....
Some conclusionsmeaningful differences in the willingness of various LLMs to be misused for generating disinfo
LM-based detector models seem to be able to detect machine-generated texts with high precision, providing an additional