about
Automated Fake News Detection in Social Networks (arxiv.org)
1 point by lainon on Apr 26, 2017 | hide | past | pdf | discuss on HN

In plain words: A system sorts Facebook posts into hoaxes or real news using only who liked each post, instead of reading the words. It got over 99% right even when trained on less than 1% of the posts.

Abstract · Some Like it Hoax: Automated Fake News Detection in Social Networks

In recent years, the reliability of information on the Internet has emerged as a crucial issue of modern society. Social network sites (SNSs) have revolutionized the way in which information is spread by allowing users to freely share content. As a consequence, SNSs are also increasingly used as vectors for the diffusion of misinformation and hoaxes. The amount of disseminated information and the rapidity of its diffusion make it practically impossible to assess reliability in a timely manner, highlighting the need for automatic hoax detection systems. As a contribution towards this objective, we show that Facebook posts can be classified with high accuracy as hoaxes or non-hoaxes on the basis of the users who "liked" them. We present two classification techniques, one based on logistic regression, the other on a novel adaptation of boolean crowdsourcing algorithms. On a dataset consisting of 15,500 Facebook posts and 909,236 users, we obtain classification accuracies exceeding 99% even when the training set contains less than 1% of the posts. We further show that our techniques are robust: they work even when we restrict our attention to the users who like both hoax and non-hoax posts. These results suggest that mapping the diffusion pattern of information can be a useful component of automatic hoax detection systems.

Eugenio Tacchini, Gabriele Ballarin, Marco L. Della Vedova, Stefano Moret, Luca de Alfaro
arXiv:1704.07506 · cs.LG, cs.HC, cs.SI · submitted Apr 25, 2017
abstract · pdf · html

add comment on HN