In plain words: To find database table and column names in English questions, it auto-labels extra examples from each question's matching database query code, then trains a text model to tag them. It beat two name taggers on precision and recall, with those examples adding over 10%.
Abstract
This paper addresses the challenge of Database Entity Recognition (DB-ER) in Natural Language Queries (NLQ). We present several key contributions to advance this field: (1) a human-annotated benchmark for DB-ER task, derived from popular text-to-sql benchmarks, (2) a novel data augmentation procedure that leverages automatic annotation of NLQs based on the corresponding SQL queries which are available in popular text-to-SQL benchmarks, (3) a specialized language model based entity recognition model using T5 as a backbone and two down-stream DB-ER tasks: sequence tagging and token classification for fine-tuning of backend and performing DB-ER respectively. We compared our DB-ER tagger with two state-of-the-art NER taggers, and observed better performance in both precision and recall for our model. The ablation evaluation shows that data augmentation boosts precision and recall by over 10%, while fine-tuning of the T5 backbone boosts these metrics by 5-10%.
Zikun Fu, Chen Yang, Kourosh Davoudi, Ken Q. Pu
arXiv:2508.19372 · cs.CL, cs.AI, cs.DB, cs.LG · submitted Aug 26, 2025
abstract · pdf · html · 6 pages, 5 figures. Accepted at IEEE 26th International Conference on Information Reuse and Integration for Data Science (IRI 2025), San Jose, California, August 6-8, 2025