about
Parametric Matrix Models (arxiv.org)
66 points by evanb on Jul 16, 2024 | hide | past | pdf | 7 comments on HN

In plain words: Instead of copying how brain cells work, these models learn matrix equations like physics uses to describe a system, fitting them to data. They can approximate any function and gave accurate, interpretable results across varied tasks, even on inputs beyond the training range.

Abstract

We present a general class of machine learning algorithms called parametric matrix models. In contrast with most existing machine learning models that imitate the biology of neurons, parametric matrix models use matrix equations that emulate physical systems. Similar to how physics problems are usually solved, parametric matrix models learn the governing equations that lead to the desired outputs. Parametric matrix models can be efficiently trained from empirical data, and the equations may use algebraic, differential, or integral relations. While originally designed for scientific computing, we prove that parametric matrix models are universal function approximators that can be applied to general machine learning problems. After introducing the underlying theory, we apply parametric matrix models to a series of different challenges that show their performance for a wide range of problems. For all the challenges tested here, parametric matrix models produce accurate results within an efficient and interpretable computational framework that allows for input feature extrapolation.

Patrick Cook, Danny Jammooa, Morten Hjorth-Jensen, Daniel D. Lee, Dean Lee
arXiv:2401.11694 · cs.LG, cond-mat.dis-nn, nucl-th, physics.comp-ph, quant-ph · submitted Jan 22, 2024 · updated Jan 6, 2025
abstract · pdf · html

add comment on HN

Looks like a cool idea and it could benefit from a more complete and detailed illustration and explanation of the architecture.

Reads like the authors skip to implications before clarifying the design.

Also, a stylistic sidenote, narrower columns of text are much easier to read, newspapers and journals do this for good reason

First author here. I agree with your points, we were constrained by the format and (especially) writing style expected by the journal we're submitting to. The Methods sections contain more explicit explanations as well as other analyses.
Huh, interesting.

According to the authors, these "parameteric matrix models" or PMMs outperform:

* commonly used (zero- or low-parameter) regression models like XGBoost, random forests, kNN, and support vector machines on a variety of regression tasks, and

* DNNs with 10x to 100x more parameters on small-scale image classification tasks like MNIST variants, CIFAR-10, and CIFAR-100 -- albeit with a lot of feature engineering.

It looks promising, but I cannot find a link to the authors' code for replicating their experiments.

All of the code and data will be released with the peer-reviewed published version. If I remember, I'll come back to this thread and post the link.
Thank you for taking the time to update everyone on HN.

I've added your work to my reading list.

Of course - I'm glad to see people are interested. I look forward to any feedback on our work either public or private.
Morten was my PhD adviser. I'll ask him what's up.