about
Language Models Are Realistic Tabular Data Generators (arxiv.org)
3 points by PaulHoule on Apr 25, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Each table row is turned into text so a language model can write new rows, filling in any missing columns from the ones you give it. The synthetic rows matched or beat the best tabular generators across many real and fake datasets.

Abstract · Language Models are Realistic Tabular Data Generators

Tabular data is among the oldest and most ubiquitous forms of data. However, the generation of synthetic samples with the original data's characteristics remains a significant challenge for tabular data. While many generative models from the computer vision domain, such as variational autoencoders or generative adversarial networks, have been adapted for tabular data generation, less research has been directed towards recent transformer-based large language models (LLMs), which are also generative in nature. To this end, we propose GReaT (Generation of Realistic Tabular data), which exploits an auto-regressive generative LLM to sample synthetic and yet highly realistic tabular data. Furthermore, GReaT can model tabular data distributions by conditioning on any subset of features; the remaining features are sampled without additional overhead. We demonstrate the effectiveness of the proposed approach in a series of experiments that quantify the validity and quality of the produced data samples from multiple angles. We find that GReaT maintains state-of-the-art performance across numerous real-world and synthetic data sets with heterogeneous feature types coming in various sizes.

Vadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk, Gjergji Kasneci
arXiv:2210.06280 · cs.LG · submitted Oct 12, 2022 · updated Apr 22, 2023
abstract · pdf · html

add comment on HN

Nice work, well-written and easy to understand, and with lots of real-world applications.

The key insight is that if you reformat each row in your tabular data as a string of text (e.g., "income is $50,000, college education is True, age is 30 to 40, ..."), you can tokenize the text, feed it to a pretrained LLM, and finetune the LLM to generate synthetic records that are statistically similar to the real ones.