about
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022) (arxiv.org)
2 points by tosh on Apr 16, 2024 | hide | past | pdf | discuss on HN

In plain words: They hand-checked 205 web-scraped text collections for hundreds of languages from five big public releases. Low-resource ones are often broken: at least 15 hold no usable text and many are under half acceptable sentences, problems easy to spot even without speaking the language.

Abstract · Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets

With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.

Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, et al.
arXiv:2103.12028 · cs.CL, cs.AI · submitted Mar 22, 2021 · updated Feb 21, 2022
abstract · pdf · html · Accepted at TACL; pre-MIT Press publication version

add comment on HN
Also discussed: Apr 2024 (3 points, 0 comments) · Mar 2021 (1 point, 1 comment)