In plain words: A word's meaning is the contexts it appears in, written as a vector, and the same idea covers phrases and sentences and says when one meaning follows from another. It covers earlier ways of combining word vectors, like adding, multiplying, and tensor products.
Abstract · A Context-theoretic Framework for Compositionality in Distributional Semantics
Techniques in which words are represented as vectors have proved useful in many applications in computational linguistics, however there is currently no general semantic formalism for representing meaning in terms of vectors. We present a framework for natural language semantics in which words, phrases and sentences are all represented as vectors, based on a theoretical analysis which assumes that meaning is determined by context. In the theoretical analysis, we define a corpus model as a mathematical abstraction of a text corpus. The meaning of a string of words is assumed to be a vector representing the contexts in which it occurs in the corpus model. Based on this assumption, we can show that the vector representations of words can be considered as elements of an algebra over a field. We note that in applications of vector spaces to representing meanings of words there is an underlying lattice structure; we interpret the partial ordering of the lattice as describing entailment between meanings. We also define the context-theoretic probability of a string, and, based on this and the lattice structure, a degree of entailment between strings. We relate the framework to existing methods of composing vector-based representations of meaning, and show that our approach generalises many of these, including vector addition, component-wise multiplication, and the tensor product.
Daoud Clarke
arXiv:1101.4479 · cs.CL, cs.AI · submitted Jan 24, 2011
abstract · pdf · html · Submitted to Computational Linguistics on 20th January 2010 for review
Fixed-length representations are more useful, because we can use standard learning machinery (ML) to predict over them. Learning techniques over variable-length sequences are more primitive. e.g. a CRF for token sequences cannot incorporate long-distance dependencies, whereas a two-layer neural network can model ANY mathematical function.
I used to believe that a variable-length sequence would, ipso facto, require a variable-length representation. However, Leon Bottou argued there must be an upper bound (1000?) on the number of bits required to represent English sentences that a human could recognize and parse in the course of normal conversation. I'm not talking about a pathological grammatical case or some Old Testament-like inventory of someone's possessions. I mean simply a sentence that you could parse and repeat back to me in your own words.
My problem with the cited work is that it is purely theoretical, and does no empirical work to explore potential limitations of the framework. It is difficult, without throwing the approach at real data, to see if it is actually an effective model for practical use. I haven't evaluated the approach deeply enough to poke any specific holes in it.
The author writes "there is currently no general semantic formalism for representing meaning in terms of vectors". However, I believe this is untrue. The author is seemingly unaware of the entire connectionist literature on fixed-length representations, which are based upon recursive neural-networks. For example, the recursive auto-associative memory work (RAAM) by Pollack in 1988, the Labeled RAAM architecture, the holographic reduced representation (Plate, 1991), and the recursive nets used by Sperduti and collaborators in the mid-90's, these works are all highly germane, but remain uncited. In principle, these architectures are powerful enough to represent all meaning in fixed length vectors, and operate over these vectors effectively. The problem with these approaches isn't theoretical, it's practical. We simply don't know how to train these architectures effectively. I find it annoying when a theoretician makes claims on the basis of existing theoretician models, and is ignorant of existing empirical models.
RAAM in particular is pretty cool. It's a fixed-length machine trained to eat the input left-to-right. It is designed so that it can uncompress itself right-to-left. So it has two basic operations: Consume, and uncompress. Each time it eats an input, it outputs a new machine of the same fixed-length. Each time it uncompresses a token, it outputs the token and a new machine of the same length. Very cool!
As you can tell, I am more excited by purely empirical and data-driven vector-based methods. For vector-based word meanings, see the language model of Collobert + Weston, which I summarized in this paper: http://www.aclweb.org/anthology/P/P10/P10-1040.pdf You can also download some word representations and code to play with here: http://metaoptimize.com/projects/wordreprs/