about
Cross-Lingual Contextual Word Embeddings Mapping with Multi-Sense Words in Mind (arxiv.org)
2 points by sel1 on Sep 21, 2019 | hide | past | pdf | discuss on HN

In plain words: Matching word meanings across two languages gets muddled by words with several senses, so this approach drops them or replaces them with averages of similar uses. For building bilingual dictionaries with no paired examples, it raised accuracy more than 10 points without hurting overall results.

Abstract · Cross-Lingual Contextual Word Embeddings Mapping With Multi-Sense Words In Mind

Recent work in cross-lingual contextual word embedding learning cannot handle multi-sense words well. In this work, we explore the characteristics of contextual word embeddings and show the link between contextual word embeddings and word senses. We propose two improving solutions by considering contextual multi-sense word embeddings as noise (removal) and by generating cluster level average anchor embeddings for contextual multi-sense word embeddings (replacement). Experiments show that our solutions can improve the supervised contextual word embeddings alignment for multi-sense words in a microscopic perspective without hurting the macroscopic performance on the bilingual lexicon induction task. For unsupervised alignment, our methods significantly improve the performance on the bilingual lexicon induction task for more than 10 points.

Zheng Zhang, Ruiqing Yin, Jun Zhu, Pierre Zweigenbaum
arXiv:1909.08681 · cs.CL · submitted Sep 18, 2019
abstract · pdf · html · 12 pages

add comment on HN