In plain words: A new German text collection of about 211,000 web sentences, built to train systems that judge how easy a text is to read or rewrite it more simply. Unlike plain-text collections, it also records headings, fonts, and pictures in an extended standard format.
Abstract
In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it contains information on text structure, typography, and images, which can be exploited as part of machine learning approaches to readability assessment and text simplification. The focus of this publication is on representing such information as an extension to an existing corpus standard.
Alessia Battisti, Sarah Ebling
arXiv:1909.09067 · cs.CL · submitted Sep 19, 2019
abstract · pdf · html