about
Neural Text Summarization: A Critical Evaluation (arxiv.org)
3 points by sel1 on Aug 27, 2019 | hide | past | pdf | discuss on HN

In plain words: A close look at the datasets, scoring measures, and models behind automatic text summarization shows why progress has stalled. Each falls short: the data is noisy and vague, the scores barely match human judgment and skip factual errors, and models copy dataset layout tricks.

Abstract

Text summarization aims at compressing long documents into a shorter form that conveys the most important parts of the original document. Despite increased interest in the community and notable research effort, progress on benchmark datasets has stagnated. We critically evaluate key ingredients of the current research setup: datasets, evaluation metrics, and models, and highlight three primary shortcomings: 1) automatically collected datasets leave the task underconstrained and may contain noise detrimental to training and evaluation, 2) current evaluation protocol is weakly correlated with human judgment and does not account for important characteristics such as factual correctness, 3) models overfit to layout biases of current datasets and offer limited diversity in their outputs.

Wojciech Kryściński, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher
arXiv:1908.08960 · cs.CL · submitted Aug 23, 2019
abstract · pdf · html · To appear in EMNLP 2019, 13 pages, 2 figures, 6 tables

add comment on HN