about
Extractive Adversarial Networks: Explanations for Attacks in Social Media (arxiv.org)
2 points by ironchief on Sep 6, 2018 | hide | past | pdf | discuss on HN

In plain words: The system highlights the words a classifier relies on to flag a personal attack, then a second layer checks the leftover words for missed signal. This caught cues highlight-only explanations leave out, and setting the model's default behavior by hand made its judgments more dependable.

Abstract · Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts

We introduce an adversarial method for producing high-recall explanations of neural text classifier decisions. Building on an existing architecture for extractive explanations via hard attention, we add an adversarial layer which scans the residual of the attention for remaining predictive signal. Motivated by the important domain of detecting personal attacks in social media comments, we additionally demonstrate the importance of manually setting a semantically appropriate `default' behavior for the model by explicitly manipulating its bias term. We develop a validation set of human-annotated personal attacks to evaluate the impact of these changes.

Samuel Carton, Qiaozhu Mei, Paul Resnick
arXiv:1809.01499 · cs.CL, cs.IR, cs.LG, stat.ML · submitted Sep 1, 2018 · updated Oct 19, 2018
abstract · pdf · html · Accepted to EMNLP 2018 Code and data available at https://github.com/shcarton/rcnn

add comment on HN