about
CogniVal: A Framework for Cognitive Word Embedding Evaluation (arxiv.org)
3 points by sel1 on Sep 22, 2019 | hide | past | pdf | discuss on HN

In plain words: Word lists are scored by how well they match brain and eye signals recorded while people read, using 15 datasets from eye-tracking, EEG, and fMRI instead of the usual single small dataset. Scores agreed strongly across datasets, recording types, and real language tasks.

Abstract

An interesting method of evaluating word representations is by how much they reflect the semantic representations in the human brain. However, most, if not all, previous works only focus on small datasets and a single modality. In this paper, we present the first multi-modal framework for evaluating English word representations based on cognitive lexical semantics. Six types of word embeddings are evaluated by fitting them to 15 datasets of eye-tracking, EEG and fMRI signals recorded during language processing. To achieve a global score over all evaluation hypotheses, we apply statistical significance testing accounting for the multiple comparisons problem. This framework is easily extensible and available to include other intrinsic and extrinsic evaluation methods. We find strong correlations in the results between cognitive datasets, across recording modalities and to their performance on extrinsic NLP tasks.

Nora Hollenstein, Antonio de la Torre, Nicolas Langer, Ce Zhang
arXiv:1909.09001 · cs.CL · submitted Sep 19, 2019 · updated Oct 29, 2019
abstract · pdf · html · accepted at CoNLL 2019

add comment on HN