about
NLP for Fake News Detection (2018) (arxiv.org)
7 points by painful on Jan 10, 2019 | hide | past | pdf | 1 comment on HN

In plain words: This survey compares how fake news detection is defined, what data it uses, and the language tools built to spot false stories automatically. It finds setups and datasets vary widely, and says future detectors should be more detailed, fair, and practical.

Abstract · A Survey on Natural Language Processing for Fake News Detection

Fake news detection is a critical yet challenging problem in Natural Language Processing (NLP). The rapid rise of social networking platforms has not only yielded a vast increase in information accessibility but has also accelerated the spread of fake news. Thus, the effect of fake news has been growing, sometimes extending to the offline world and threatening public safety. Given the massive amount of Web content, automatic fake news detection is a practical NLP problem useful to all online content providers, in order to reduce the human time and effort to detect and prevent the spread of fake news. In this paper, we describe the challenges involved in fake news detection and also describe related tasks. We systematically review and compare the task formulations, datasets and NLP solutions that have been developed for this task, and also discuss the potentials and limitations of them. Based on our insights, we outline promising research directions, including more fine-grained, detailed, fair, and practical detection models. We also highlight the difference between fake news detection and other related tasks, and the importance of NLP solutions for fake news detection.

Ray Oshikawa, Jing Qian, William Yang Wang
arXiv:1811.00770 · cs.CL, cs.AI · submitted Nov 2, 2018 · updated Mar 5, 2020
abstract · pdf · html · 11 pages, no figure, Accepted to LREC 2020

add comment on HN

Ouch, feeding probabilistic models training data scored with a gradient of truthfulness tags generated by humans and all their biases... surely this won't end horribly and simply serve as a method to algorithmically institute the tyranny of the majority.

If you really want to do this (You really don't, I assure you - you'll hate the end result), you've got to reach back through the AI winter and drag the granddaddy of NLP, propositional logic, into modern AI development. We'll see this employed by lawyers long before journalists.

https://en.wikipedia.org/wiki/Attempto_Controlled_English