about
Pushing the Limits of Low-Resource Morphological Inflection (arxiv.org)
1 point by sel1 on Aug 19, 2019 | hide | past | pdf | discuss on HN

In plain words: A decoder reading a word in two attention steps, plus patterns borrowed from related languages and invented practice words, learns word forms with few examples. It beat the prior best by 15 percentage points; transfer worked best when languages were similar and shared a writing system.

Abstract

Recent years have seen exceptional strides in the task of automatic morphological inflection generation. However, for a long tail of languages the necessary resources are hard to come by, and state-of-the-art neural methods that work well under higher resource settings perform poorly in the face of a paucity of data. In response, we propose a battery of improvements that greatly improve performance under such low-resource conditions. First, we present a novel two-step attention architecture for the inflection decoder. In addition, we investigate the effects of cross-lingual transfer from single and multiple languages, as well as monolingual data hallucination. The macro-averaged accuracy of our models outperforms the state-of-the-art by 15 percentage points. Also, we identify the crucial factors for success with cross-lingual transfer for morphological inflection: typological similarity and a common representation across languages.

Antonios Anastasopoulos, Graham Neubig
arXiv:1908.05838 · cs.CL · submitted Aug 16, 2019 · updated Aug 20, 2019
abstract · pdf · html · to appear at EMNLP 2019

add comment on HN