In plain words: Word vectors are usually trained just from how words appear together in text. Here they are also nudged using labeled word-part and grammar information, so words sharing those features sit close together; tests on German show they capture word structure better than text-only training.
Abstract
Linguistic similarity is multi-faceted. For instance, two words may be similar with respect to semantics, syntax, or morphology inter alia. Continuous word-embeddings have been shown to capture most of these shades of similarity to some degree. This work considers guiding word-embeddings with morphologically annotated data, a form of semi-supervised learning, encouraging the vectors to encode a word's morphology, i.e., words close in the embedded space share morphological features. We extend the log-bilinear model to this end and show that indeed our learned embeddings achieve this, using German as a case study.
Ryan Cotterell, Hinrich Schütze
arXiv:1907.02423 · cs.CL · submitted Jul 4, 2019
abstract · pdf · html · Published at NAACL 2015