In plain words: A huge collection of Reddit comments where each author marks their own sarcastic lines, with user, topic, and conversation context, to train and test sarcasm detectors. It holds 1.3 million sarcastic statements, ten times more than any earlier dataset.
Abstract
We introduce the Self-Annotated Reddit Corpus (SARC), a large corpus for sarcasm research and for training and evaluating systems for sarcasm detection. The corpus has 1.3 million sarcastic statements -- 10 times more than any previous dataset -- and many times more instances of non-sarcastic statements, allowing for learning in both balanced and unbalanced label regimes. Each statement is furthermore self-annotated -- sarcasm is labeled by the author, not an independent annotator -- and provided with user, topic, and conversation context. We evaluate the corpus for accuracy, construct benchmarks for sarcasm detection, and evaluate baseline methods.
Mikhail Khodak, Nikunj Saunshi, Kiran Vodrahalli
arXiv:1704.05579 · cs.CL, cs.AI, cs.LG · submitted Apr 19, 2017 · updated Mar 22, 2018
abstract · pdf · html · 6 pages, 4 Figures. To Appear in LREC 2018
> sarcasm is labelled by the author
They literally just searched out "/s". Clever. Though I'm guessing the "independently verified" entailed reading a lot of those comments.
Did they also read through the nonlabelled comments to catch any unlabelled sarcasm? (Guessing not since the pitch is of "self labelled sarcasm") wonder if that'll trip any usage up.