In plain words: A Dutch news dataset labels partisanship, with over 100,000 articles tagged by their publisher and 776 tagged story by story through a reader survey. The publisher-level and article-level labels let tools learn to spot partisan slant in Dutch news.
Abstract
We present a new Dutch news dataset with labeled partisanship. The dataset contains more than 100K articles that are labeled on the publisher level and 776 articles that were crowdsourced using an internal survey platform and labeled on the article level. In this paper, we document our original motivation, the collection and annotation process, limitations, and applications.
Chia-Lun Yeh, Babak Loni, Mariëlle Hendriks, Henrike Reinhardt, Anne Schuth
arXiv:1908.02322 · cs.CL · submitted Aug 6, 2019
abstract · pdf · html