about
Spare: A Single-Pass Neural Model for Relational Databases (arxiv.org)
3 points by PaulHoule on Nov 1, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Instead of treating a multi-table database as a graph and passing messages over it many times, this model reads the data in one pass, using the tables' regular structure to build row representations. It trains and runs faster while matching the graph approach's accuracy.

Abstract · SPARE: A Single-Pass Neural Model for Relational Databases

While there has been extensive work on deep neural networks for images and text, deep learning for relational databases (RDBs) is still a rather unexplored field. One direction that recently gained traction is to apply Graph Neural Networks (GNNs) to RBDs. However, training GNNs on large relational databases (i.e., data stored in multiple database tables) is rather inefficient due to multiple rounds of training and potentially large and inefficient representations. Hence, in this paper we propose SPARE (Single-Pass Relational models), a new class of neural models that can be trained efficiently on RDBs while providing similar accuracies as GNNs. For enabling efficient training, different from GNNs, SPARE makes use of the fact that data in RDBs has a regular structure, which allows one to train these models in a single pass while exploiting symmetries at the same time. Our extensive empirical evaluation demonstrates that SPARE can significantly speedup both training and inference while offering competitive predictive performance over numerous baselines.

Benjamin Hilprecht, Kristian Kersting, Carsten Binnig
arXiv:2310.13581 · cs.DB, cs.AI · submitted Oct 20, 2023
abstract · pdf · html

add comment on HN

> Therefore, one often follows an alternative approach in practice, namely, to simply join the data of different tables, materialize the output of the multi-way join in one large table, and train a predictive model on the resulting table (Kanter and Veeramachaneni 2015). This approach, however, not only comes with potentially high upfront costs of joining potentially many large tables but also the structure in a relational database is ignored, since all data is represented as a single flat table in the model.

This actually feels like the only valid path. It also gives you an abstraction layer such that you can modify your schema without having to retrain your models.

Spending 20 minutes creating a view upfront could save you a lifetime of getting a NN to sort this out for you. I understand the arguments where preprocessing is pointless (i.e. cleaning up images for machine vision), but I don't think those hold in this domain where we are dealing with an arbitrary number of dimensions.