about
Enhancing Frame Detection with Retrieval Augmented Generation (arxiv.org)
37 points by PaulHoule on Feb 28, 2025 | hide | past | pdf | 12 comments on HN

In plain words: It reads text, pulls a list of likely event frames from a stored set, then picks the best ones — no need to mark the trigger word first. It beat prior systems on both FrameNet test sets and helped a question-to-query translator handle reworded questions.

Abstract

Recent advancements in Natural Language Processing have significantly improved the extraction of structured semantic representations from unstructured text, especially through Frame Semantic Role Labeling (FSRL). Despite this progress, the potential of Retrieval-Augmented Generation (RAG) models for frame detection remains under-explored. In this paper, we present the first RAG-based approach for frame detection called RCIF (Retrieve Candidates and Identify Frames). RCIF is also the first approach to operate without the need for explicit target span and comprises three main stages: (1) generation of frame embeddings from various representations ; (2) retrieval of candidate frames given an input text; and (3) identification of the most suitable frames. We conducted extensive experiments across multiple configurations, including zero-shot, few-shot, and fine-tuning settings. Our results show that our retrieval component significantly reduces the complexity of the task by narrowing the search space thus allowing the frame identifier to refine and complete the set of candidates. Our approach achieves state-of-the-art performance on FrameNet 1.5 and 1.7, demonstrating its robustness in scenarios where only raw text is provided. Furthermore, we leverage the structured representation obtained through this method as a proxy to enhance generalization across lexical variations in the task of translating natural language questions into SPARQL queries.

Papa Abdou Karim Karou Diallo, Amal Zouaq
arXiv:2502.12210 · cs.CL, cs.AI, cs.LG · submitted Feb 17, 2025
abstract · pdf · html

add comment on HN

I usually annotate my chunks with title, summary, keywords and multiple levels of hierarchical topics. More recently I thought about annotating intent, user values and tactics, especially for debate related text and LLM chats. So I would annotate state (the summary of the content itself), values (intent and values) and policy (tactics), taking inspiration from RL.

The idea of detecting frames and using them to tease out the implicit meaning from text is quite nice. It seems there is a lot more to discover about using LLMs prior to RAG. Text is like code, you can't know what it does untill you run it, and in this case, until you annotate it. For example "10+10" won't embed close to "20". And "The fifth letter in this string" won't retrieve "f" by emmbedding similarity.

I have this very stupid question and I haven't seen it answered anywhere.

Let's say LLM while training is ingesting a very long book.

The name of the author of the book would appear at the very beginning.

While inference, how does the llm determine that the last chapter of the book is written by so and so author and hence that chunk should be near that author's style.

The information is probably lost, training on very long inputs is expensive, we usually train on short inputs and only send longer ones at the end. So the author name would not appear together with all the chunks unless it appears on every page in the header, assuming the headers are not stripped out first.
I am wondering the same techniques we use to augment chunks for RAG also the same technique used to augment training data with rich metadata. A preprocessing step which ensures metadata rich text for creating foundational models.
It might not directly learn it, but it could still identify it if it's quoted by others or if there's enough text of the same style with the author's name as context.
I had to learn what a "frame" is to understand this. https://framenet.icsi.berkeley.edu/ is useful (FrameNet is the collection of 1,000 frames used in the paper).

An example of a frame is an "Event" - https://framenet.icsi.berkeley.edu/fnReports/data/frameIndex... - where:

> An Event takes place at a Place and Time.

So if you're extracting frames from a piece of text, that's one of the concepts you might be trying to identify - along with what the place and time are.

https://en.wikipedia.org/wiki/Frame_(artificial_intelligence...

> Frames are an artificial intelligence data structure used to divide knowledge into substructures by representing "stereotyped situations".

They are highly related to ontology systems and knowledge engineering.

Seems like a viable name for this data structure could have been “Thing”.
Cool paper but unsurprising results since anything benefits from RAG
Also, grep is really exciting with lots of data.
Love FrameNet. Awesome to see it's alive and in-use.
From the paper:

> Frames ... are conceptual structures that capture the semantic and syntactic relationships underlying language. They are helpful in providing a structured semantic context for understanding relationships between entities, enabling tasks like Machine Reading Comprehension and Information Extraction to be more accurate and contextually aware.