about
Glue Is Not All You Need: Discourse Based Evaluation of Language Understanding (arxiv.org)
3 points by jean-porte on Jul 23, 2019 | hide | past | pdf | discuss on HN

In plain words: A new test suite gathers 11 English datasets checking whether systems grasp implied meaning and speaker intent, not just literal wording. Models pretrained to judge whether one sentence follows from another still stumble on it, so that training does not build universal understanding.

Abstract · A Pragmatics-Centered Evaluation Framework for Natural Language Understanding

New models for natural language understanding have recently made an unparalleled amount of progress, which has led some researchers to suggest that the models induce universal text representations. However, current benchmarks are predominantly targeting semantic phenomena; we make the case that pragmatics needs to take center stage in the evaluation of natural language understanding. We introduce PragmEval, a new benchmark for the evaluation of natural language understanding, that unites 11 pragmatics-focused evaluation datasets for English. PragmEval can be used as supplementary training data in a multi-task learning setup, and is publicly available, alongside the code for gathering and preprocessing the datasets. Using our evaluation suite, we show that natural language inference, a widely used pretraining task, does not result in genuinely universal representations, which presents a new challenge for multi-task learning.

Damien Sileo, Tim Van-de-Cruys, Camille Pradel, Philippe Muller
arXiv:1907.08672 · cs.CL · submitted Jul 19, 2019 · updated Apr 4, 2022
abstract · pdf · html · Accepted at LREC2022

add comment on HN