about
Cross-Lingual Dependency Parsing Using Code-Mixed TreeBank (arxiv.org)
2 points by sel1 on Sep 9, 2019 | hide | past | pdf | discuss on HN

In plain words: Instead of translating whole sentences and guessing which words line up, this approach translates only the words it is sure about, leaving a mixed-language treebank whose grammar carries over through shared word embeddings. It parsed better than full translation and rivaled other cross-lingual methods.

Abstract

Treebank translation is a promising method for cross-lingual transfer of syntactic dependency knowledge. The basic idea is to map dependency arcs from a source treebank to its target translation according to word alignments. This method, however, can suffer from imperfect alignment between source and target words. To address this problem, we investigate syntactic transfer by code mixing, translating only confident words in a source treebank. Cross-lingual word embeddings are leveraged for transferring syntactic knowledge to the target from the resulting code-mixed treebank. Experiments on University Dependency Treebanks show that code-mixed treebanks are more effective than translated treebanks, giving highly competitive performances among cross-lingual parsing methods.

Zhang Meishan, Zhang Yue, Fu Guohong
arXiv:1909.02235 · cs.CL · submitted Sep 5, 2019
abstract · pdf · html · 10 pages

add comment on HN