about
A Universal Part-of-Speech Tagset (arxiv.org)
2 points by cskau on Jan 15, 2012 | hide | past | pdf | discuss on HN

In plain words: A shared set of twelve word categories—noun, verb, and so on—comes with mappings from many existing labeling schemes, giving common labels for 22 languages. In a test, a system that learns sentence structure without correct word labels still reached competitive accuracy.

Abstract

To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset, we develop a mapping from 25 different treebank tagsets to this universal set. As a result, when combined with the original treebank data, this universal tagset and mapping produce a dataset consisting of common parts-of-speech for 22 different languages. We highlight the use of this resource via two experiments, including one that reports competitive accuracies for unsupervised grammar induction without gold standard part-of-speech tags.

Slav Petrov, Dipanjan Das, Ryan McDonald
arXiv:1104.2086 · cs.CL · submitted Apr 11, 2011
abstract · pdf · html

add comment on HN