about
Headline Generation: Learning from Decomposable Document Titles (arxiv.org)
2 points by the_decider on May 17, 2019 | hide | past | pdf | discuss on HN

In plain words: A network builds titles by answering a chain of questions about a document, using only words from it, and learns from millions of article-headline pairs. In a blind test, people rated its headlines as good or better than the human originals in most cases.

Abstract

We propose a novel method for generating titles for unstructured text documents. We reframe the problem as a sequential question-answering task. A deep neural network is trained on document-title pairs with decomposable titles, meaning that the vocabulary of the title is a subset of the vocabulary of the document. To train the model we use a corpus of millions of publicly available document-title pairs: news articles and headlines. We present the results of a randomized double-blind trial in which subjects were unaware of which titles were human or machine-generated. When trained on approximately 1.5 million news articles, the model generates headlines that humans judge to be as good or better than the original human-written headlines in the majority of cases.

Oleg Vasilyev, Tom Grek, John Bohannon
arXiv:1904.08455 · cs.CL · submitted Apr 17, 2019 · updated May 10, 2019
abstract · pdf · html · 10 pages, 9 figures, 1 table. v3: Better figures, tables and descriptions - by reviewer Anna Venancio-Marques

add comment on HN