about
Modern Baselines for Sparql Semantic Parsing (arxiv.org)
2 points by PaulHoule on Sep 22, 2023 | hide | past | pdf | discuss on HN

In plain words: A system turns questions into database queries by arranging the given entities and relations with query language keywords, using pretrained text models. With careful word-splitting, it beats earlier purpose-built systems on two question sets and can copy words from the question into the query.

Abstract · Modern Baselines for SPARQL Semantic Parsing

In this work, we focus on the task of generating SPARQL queries from natural language questions, which can then be executed on Knowledge Graphs (KGs). We assume that gold entity and relations have been provided, and the remaining task is to arrange them in the right order along with SPARQL vocabulary, and input tokens to produce the correct SPARQL query. Pre-trained Language Models (PLMs) have not been explored in depth on this task so far, so we experiment with BART, T5 and PGNs (Pointer Generator Networks) with BERT embeddings, looking for new baselines in the PLM era for this task, on DBpedia and Wikidata KGs. We show that T5 requires special input tokenisation, but produces state of the art performance on LC-QuAD 1.0 and LC-QuAD 2.0 datasets, and outperforms task-specific models from previous works. Moreover, the methods enable semantic parsing for questions where a part of the input needs to be copied to the output query, thus enabling a new paradigm in KG semantic parsing.

Debayan Banerjee, Pranav Ajit Nair, Jivat Neet Kaur, Ricardo Usbeck, Chris Biemann
arXiv:2204.12793 · cs.IR, cs.CL · submitted Apr 27, 2022 · updated Sep 14, 2023
abstract · pdf · html · 5 pages, short paper, SIGIR 2022

add comment on HN