about
Language Modeling Using Tensor Trains (arxiv.org)
2 points by PaulHoule on May 11, 2024 | hide | past | pdf | discuss on HN

In plain words: It stores each sentence in a huge space built by combining its words, then computes the sentence's probability through a compact chain of small matrices so the math stays cheap. On real text it beat plain small recurrent networks at predicting the next word.

Abstract

We propose a novel tensor network language model based on the simplest tensor network (i.e., tensor trains), called `Tensor Train Language Model' (TTLM). TTLM represents sentences in an exponential space constructed by the tensor product of words, but computing the probabilities of sentences in a low-dimensional fashion. We demonstrate that the architectures of Second-order RNNs, Recurrent Arithmetic Circuits (RACs), and Multiplicative Integration RNNs are, essentially, special cases of TTLM. Experimental evaluations on real language modeling tasks show that the proposed variants of TTLM (i.e., TTLM-Large and TTLM-Tiny) outperform the vanilla Recurrent Neural Networks (RNNs) with low-scale of hidden units. (The code is available at https://github.com/shuishen112/tensortrainlm.)

Zhan Su, Yuqin Zhou, Fengran Mo, Jakob Grue Simonsen
arXiv:2405.04590 · cs.CL, cs.IR · submitted May 7, 2024
abstract · pdf · html

add comment on HN