In plain words: It reads text, pulls a list of likely event frames from a stored set, then picks the best ones — no need to mark the trigger word first. It beat prior systems on both FrameNet test sets and helped a question-to-query translator handle reworded questions.
Abstract
Recent advancements in Natural Language Processing have significantly improved the extraction of structured semantic representations from unstructured text, especially through Frame Semantic Role Labeling (FSRL). Despite this progress, the potential of Retrieval-Augmented Generation (RAG) models for frame detection remains under-explored. In this paper, we present the first RAG-based approach for frame detection called RCIF (Retrieve Candidates and Identify Frames). RCIF is also the first approach to operate without the need for explicit target span and comprises three main stages: (1) generation of frame embeddings from various representations ; (2) retrieval of candidate frames given an input text; and (3) identification of the most suitable frames. We conducted extensive experiments across multiple configurations, including zero-shot, few-shot, and fine-tuning settings. Our results show that our retrieval component significantly reduces the complexity of the task by narrowing the search space thus allowing the frame identifier to refine and complete the set of candidates. Our approach achieves state-of-the-art performance on FrameNet 1.5 and 1.7, demonstrating its robustness in scenarios where only raw text is provided. Furthermore, we leverage the structured representation obtained through this method as a proxy to enhance generalization across lexical variations in the task of translating natural language questions into SPARQL queries.
Papa Abdou Karim Karou Diallo, Amal Zouaq
arXiv:2502.12210 · cs.CL, cs.AI, cs.LG · submitted Feb 17, 2025
abstract · pdf · html
The idea of detecting frames and using them to tease out the implicit meaning from text is quite nice. It seems there is a lot more to discover about using LLMs prior to RAG. Text is like code, you can't know what it does untill you run it, and in this case, until you annotate it. For example "10+10" won't embed close to "20". And "The fifth letter in this string" won't retrieve "f" by emmbedding similarity.