about
Probing Neural Network Comprehension of Natural Language Arguments (arxiv.org)
1 point by sel1 on Jul 19, 2019 | hide | past | pdf | discuss on HN

In plain words: Tested whether a language model truly understands arguments by asking which hidden assumption makes a short argument fail. Its 77% score, near the average untrained person's, came entirely from statistical shortcuts in the data; on a rebuilt version without them, all models scored at chance.

Abstract

We are surprised to find that BERT's peak performance of 77% on the Argument Reasoning Comprehension Task reaches just three points below the average untrained human baseline. However, we show that this result is entirely accounted for by exploitation of spurious statistical cues in the dataset. We analyze the nature of these cues and demonstrate that a range of models all exploit them. This analysis informs the construction of an adversarial dataset on which all models achieve random accuracy. Our adversarial dataset provides a more robust assessment of argument comprehension and should be adopted as the standard in future work.

Timothy Niven, Hung-Yu Kao
arXiv:1907.07355 · cs.CL · submitted Jul 17, 2019 · updated Sep 16, 2019
abstract · pdf · html · ACL 2019 (Updated Version)

add comment on HN
Also discussed: Jul 2019 (3 points, 0 comments) · Jul 2019 (2 points, 0 comments)