In plain words: They measure how much earlier systems rely on resources like a source-language Wikipedia and bilingual maps, then improve how candidates are found and picked using only limited data. On four very low-resource languages, linking accuracy rose 6 to 23 percent over the usual approach.
Abstract · Towards Zero-resource Cross-lingual Entity Linking
Cross-lingual entity linking (XEL) grounds named entities in a source language to an English Knowledge Base (KB), such as Wikipedia. XEL is challenging for most languages because of limited availability of requisite resources. However, much previous work on XEL has been on simulated settings that actually use significant resources (e.g. source language Wikipedia, bilingual entity maps, multilingual embeddings) that are unavailable in truly low-resource languages. In this work, we first examine the effect of these resource assumptions and quantify how much the availability of these resource affects overall quality of existing XEL systems. Next, we propose three improvements to both entity candidate generation and disambiguation that make better use of the limited data we do have in resource-scarce scenarios. With experiments on four extremely low-resource languages, we show that our model results in gains of 6-23% in end-to-end linking accuracy.
Shuyan Zhou, Shruti Rijhwani, Graham Neubig
arXiv:1909.13180 · cs.CL · submitted Sep 29, 2019 · updated Oct 1, 2019
abstract · pdf · html · Accepted by EMNLP DeepLo workshop 2019