about
LLM-Assisted Labeling Function Generation for Semantic Type Detection (arxiv.org)
1 point by PaulHoule on Sep 9, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of people hand-writing small rules that guess a column's type, a language model writes those rules to label training data automatically. Tests on real web tables show the rules help, and the analysis reveals which prompting choices work best.

Abstract · LLM-assisted Labeling Function Generation for Semantic Type Detection

Detecting semantic types of columns in data lake tables is an important application. A key bottleneck in semantic type detection is the availability of human annotation due to the inherent complexity of data lakes. In this paper, we propose using programmatic weak supervision to assist in annotating the training data for semantic type detection by leveraging labeling functions. One challenge in this process is the difficulty of manually writing labeling functions due to the large volume and low quality of the data lake table datasets. To address this issue, we explore employing Large Language Models (LLMs) for labeling function generation and introduce several prompt engineering strategies for this purpose. We conduct experiments on real-world web table datasets. Based on the initial results, we perform extensive analysis and provide empirical insights and future directions for researchers in this field.

Chenjie Li, Dan Zhang, Jin Wang
arXiv:2408.16173 · cs.DB, cs.AI · submitted Aug 28, 2024
abstract · pdf · html · VLDB'24-DATAI

add comment on HN