about
Are LLMs Naturally Good at Synthetic Tabular Data Generation? (arxiv.org)
2 points by belter on Jun 23, 2024 | hide | past | pdf | discuss on HN

In plain words: Large language models write table rows one column at a time, and shuffling column order during training breaks the fixed relationships between columns that real tables rely on. Making the model aware of column order fixes much of this and improves synthetic tables.

Abstract · Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)

Synthetic data generation is integral to ML pipelines, e.g., to augment training data, replace sensitive information, and even to power advanced platforms like DeepSeek. While LLMs fine-tuned for synthetic data generation are gaining traction, synthetic table generation -- a critical data type in business and science -- remains under-explored compared to text and image synthesis. This paper shows that LLMs, whether used as-is or after traditional fine-tuning, are inadequate for generating synthetic tables. Their autoregressive nature, combined with random order permutation during fine-tuning, hampers the modeling of functional dependencies and prevents capturing conditional mixtures of distributions essential for real-world constraints. We demonstrate that making LLMs permutation-aware can mitigate these issues.

Shengzhe Xu, Cho-Ting Lee, Mandar Sharma, Raquib Bin Yousuf, Nikhil Muralidhar, Naren Ramakrishnan
arXiv:2406.14541 · cs.LG · submitted Jun 20, 2024 · updated Mar 13, 2025
abstract · pdf · html

add comment on HN