about
Anetac: Arabic Named Entity Transliteration and Classification Dataset (arxiv.org)
2 points by sel1 on Jul 10, 2019 | hide | past | pdf | discuss on HN

In plain words: A free collection pairs English names with their Arabic spellings and a label of person, place, or organization, drawn from open translation texts. It holds 79,924 entries for training systems that write Arabic versions of names or sort them by type.

Abstract · ANETAC: Arabic Named Entity Transliteration and Classification Dataset

In this paper, we make freely accessible ANETAC our English-Arabic named entity transliteration and classification dataset that we built from freely available parallel translation corpora. The dataset contains 79,924 instances, each instance is a triplet (e, a, c), where e is the English named entity, a is its Arabic transliteration and c is its class that can be either a Person, a Location, or an Organization. The ANETAC dataset is mainly aimed for the researchers that are working on Arabic named entity transliteration, but it can also be used for named entity classification purposes.

Mohamed Seghir Hadj Ameur, Farid Meziane, Ahmed Guessoum
arXiv:1907.03110 · cs.CL, cs.IR · submitted Jul 6, 2019
abstract · pdf · html

add comment on HN