about
Recovering human knowledge of multiple object features from word embeddings (arxiv.org)
2 points by PaulHoule on May 4, 2022 | hide | past | pdf | 1 comment on HN

In plain words: By lining up words along a scale between two opposites—like small to big—this technique pulls out one trait at a time from word spaces that normally only show overall similarity. It matched people's judgments of size, intelligence, and danger across object categories.

Abstract · Semantic projection: recovering human knowledge of multiple, distinct object features from word embeddings

The words of a language reflect the structure of the human mind, allowing us to transmit thoughts between individuals. However, language can represent only a subset of our rich and detailed cognitive architecture. Here, we ask what kinds of common knowledge (semantic memory) are captured by word meanings (lexical semantics). We examine a prominent computational model that represents words as vectors in a multidimensional space, such that proximity between word-vectors approximates semantic relatedness. Because related words appear in similar contexts, such spaces - called "word embeddings" - can be learned from patterns of lexical co-occurrences in natural language. Despite their popularity, a fundamental concern about word embeddings is that they appear to be semantically "rigid": inter-word proximity captures only overall similarity, yet human judgments about object similarities are highly context-dependent and involve multiple, distinct semantic features. For example, dolphins and alligators appear similar in size, but differ in intelligence and aggressiveness. Could such context-dependent relationships be recovered from word embeddings? To address this issue, we introduce a powerful, domain-general solution: "semantic projection" of word-vectors onto lines that represent various object features, like size (the line extending from the word "small" to "big"), intelligence (from "dumb" to "smart"), or danger (from "safe" to "dangerous"). This method, which is intuitively analogous to placing objects "on a mental scale" between two extremes, recovers human judgments across a range of object categories and properties. We thus show that word embeddings inherit a wealth of common knowledge from word co-occurrence statistics and can be flexibly manipulated to express context-dependent meanings.

Gabriel Grand, Idan Asher Blank, Francisco Pereira, Evelina Fedorenko
arXiv:1802.01241 · cs.CL · submitted Feb 5, 2018 · updated Mar 6, 2018
abstract · pdf

add comment on HN

On one hand I find it a fascinating result, but I also think Figure 1 shows how low the bar is on research on "embeddings", that is, the model thinks a mouse is smaller than an ant, an alligator is bigger than an elephant, etc.

If this was the performance of a human or a database or any other kind of system people would say "go back to the drawing board and try again".

There was a time that I messed around with glove and word2vec and scikit-learn and made a large number of graphs similar to the one above, I never "published" anything because it was all pretty sketchy.

I note that paper was submitted to arXiv in 2018, it got finally published in 2021 which is likely a sign it struggled in peer review.