about
Interpretable graph neural networks for tabular data (arxiv.org)
77 points by PaulHoule on Aug 26, 2023 | hide | past | pdf | 5 comments on HN

In plain words: A graph neural network for spreadsheet-style data is trained so every prediction can be traced step by step back to the original input features. It matched top tabular tools like XGBoost and random forests, and its feature-importance explanations matched the true ones at no extra cost.

Abstract · Interpretable Graph Neural Networks for Tabular Data

Data in tabular format is frequently occurring in real-world applications. Graph Neural Networks (GNNs) have recently been extended to effectively handle such data, allowing feature interactions to be captured through representation learning. However, these approaches essentially produce black-box models, in the form of deep neural networks, precluding users from following the logic behind the model predictions. We propose an approach, called IGNNet (Interpretable Graph Neural Network for tabular data), which constrains the learning algorithm to produce an interpretable model, where the model shows how the predictions are exactly computed from the original input features. A large-scale empirical investigation is presented, showing that IGNNet is performing on par with state-of-the-art machine-learning algorithms that target tabular data, including XGBoost, Random Forests, and TabNet. At the same time, the results show that the explanations obtained from IGNNet are aligned with the true Shapley values of the features without incurring any additional computational overhead.

Amr Alkhatib, Sofiane Ennadir, Henrik Boström, Michalis Vazirgiannis
arXiv:2308.08945 · cs.LG, cs.AI · submitted Aug 17, 2023 · updated Aug 13, 2024
abstract · pdf · html · Accepted at ECAI 2024

add comment on HN

> At the same time, the results show that the explanations obtained from IGNNet are aligned with the true Shapley values of the features without incurring any additional computational overhead

TabPFN: https://github.com/automl/TabPFN https://twitter.com/FrankRHutter/status/1583410845307977733

"TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second" (2022) https://arxiv.org/abs/2308.08945

FWIU TabPFN is Bayesian-calibrated/trained with better performance than xgboost for non-categorical data

Right, the significance of the original article and the related field of research is that ChatGPT-like models don't handle tabular data well and there's a lot of need for things that do.
There are multiple metrics to optimize for when optimizing.

FWIU, from the diagram in the photo in the linked tweet, which is similar to a diagram on page 16 of the TabPFN paper [1], on the OpenML-CC18, TabPFN has a better ROC Receiver Operating Characteristic after 1 second than XGboost, Catboost, LightGBM, KNN, SAINT, Reg. Cocktail, and Autogluon after any amount of time, but Auto-sklearn 2.0 required 5 minutes to reach ~ROC parity with TabPFN.

1. "TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second" (May 2023) https://arxiv.org/abs/2207.01848

2. "Interpretable Graph Neural Networks for Tabular Data" (Aug 2023) https://arxiv.org/abs/2308.08945

Microsoft LIDA seems something that can incorporate this for additional insight from tabular data sources
The paper sounds interesting, but without the code being available, it is difficult to say whether the idea is really viable.