In plain words: Built a collection of 203,000 sentence pairs where each complicated sentence is broken into the smallest self-contained statements, one idea each. Unlike earlier collections with only a few such splits, every sentence here is fully broken down, giving tools many examples to learn from.
Abstract
We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split examples, we present a dataset where each input sentence is broken down into a set of minimal propositions, i.e. a sequence of sound, self-contained utterances with each of them presenting a minimal semantic unit that cannot be further decomposed into meaningful propositions. This corpus is useful for developing sentence splitting approaches that learn how to transform sentences with a complex linguistic structure into a fine-grained representation of short sentences that present a simple and more regular structure which is easier to process for downstream applications and thus facilitates and improves their performance.
Christina Niklaus, Andre Freitas, Siegfried Handschuh
arXiv:1909.12131 · cs.CL · submitted Sep 26, 2019
abstract · pdf · html