about
AnnaParser: Semantic Parsing for Tabular Data Analysis (arxiv.org)
2 points by sel1 on Oct 24, 2019 | hide | past | pdf | discuss on HN

In plain words: It turns a question about any table into SQL by swapping columns for generic placeholders, then using hand-written rules and a trained scorer to pick the best. It beat the previous best system on the large WikiSQL set and handled harder English and Chinese questions.

Abstract · A Hybrid Semantic Parsing Approach for Tabular Data Analysis

This paper presents a novel approach to translating natural language questions to SQL queries for given tables, which meets three requirements as a real-world data analysis application: cross-domain, multilingualism and enabling quick-start. Our proposed approach consists of: (1) a novel data abstraction step before the parser to make parsing table-agnosticism; (2) a set of semantic rules for parsing abstracted data-analysis questions to intermediate logic forms as tree derivations to reduce the search space; (3) a neural-based model as a local scoring function on a span-based semantic parser for structured optimization and efficient inference. Experiments show that our approach outperforms state-of-the-art algorithms on a large open benchmark dataset WikiSQL. We also achieve promising results on a small dataset for more complex queries in both English and Chinese, which demonstrates our language expansion and quick-start ability.

Yan Gao, Jian-Guang Lou, Dongmei Zhang
arXiv:1910.10363 · cs.AI, cs.CL, cs.DB · submitted Oct 23, 2019 · updated Oct 24, 2019
abstract · pdf · html · 10 pages, 5 figures

add comment on HN