In plain words: Word co-occurrence counts can be drawn as a network where words link when they appear together; grouping words into communities shrinks the huge space of possible word meanings. The resulting word embeddings match the quality of the best current approaches.
Abstract
Most of the time, the first step to learn word embeddings is to build a word co-occurrence matrix. As such matrices are equivalent to graphs, complex networks theory can naturally be used to deal with such data. In this paper, we consider applying community detection, a main tool of this field, to the co-occurrence matrix corresponding to a huge corpus. Community structure is used as a way to reduce the dimensionality of the initial space. Using this community structure, we propose a method to extract word embeddings that are comparable to the state-of-the-art approaches.
Nicolas Dugué, Victor Connes
arXiv:1910.01489 · cs.CL, cs.LG · submitted Oct 3, 2019
abstract · pdf · html · in French