about
Learning Mathematical Properties of Integers (arxiv.org)
54 points by Anon84 on Sep 18, 2021 | hide | past | pdf | 3 comments on HN

In plain words: Numbers are turned into vectors by training on mathematical sequences, like counting patterns, instead of English sentences, so the vectors can carry math facts. On numerical reasoning tasks, these math-trained number vectors did much better than ones learned from English text.

Abstract

Embedding words in high-dimensional vector spaces has proven valuable in many natural language applications. In this work, we investigate whether similarly-trained embeddings of integers can capture concepts that are useful for mathematical applications. We probe the integer embeddings for mathematical knowledge, apply them to a set of numerical reasoning tasks, and show that by learning the representations from mathematical sequence data, we can substantially improve over number embeddings learned from English text corpora.

Maria Ryskina, Kevin Knight
arXiv:2109.07230 · cs.CL, cs.LG · submitted Sep 15, 2021
abstract · pdf · html · BlackboxNLP 2021

add comment on HN

From the paper:

> While human aptitude questions focus on a specific range of patterns that are relatively easy for people to identify, we also want to test the predictive abilities on the OEIS test set sequences which showcase more sophisticated phenomena. The LSTM embeddings yield higher accuracy in that case.

Pretty cool to think that we might be able to use computers to spot elusive patterns in sequences in the (not too distant?) future. Kinda feels like asking an alien for help with a tricky problem on a math problem set. Many fascinating areas of math research originated from spotting patterns in sequences (best example I can think of is how the monstrous moonshine was discovered [1]); it would be nice to have a 'sixth sense' of sorts to help with that.

I don't know much about ML but I'd be interested to see if this ever gets applied to slightly more abstract mathematical objects like group presentations, knots/links, (co)homology groups—but I guess none of those have training sets even comparable to the OEIS...

[1] https://en.wikipedia.org/wiki/Monstrous_moonshine#History

Crazy, I just recommended this exact approach to a student in my data science class.
It's like, so Pythagorean