about
System 2 Attention (is something you might need too) (arxiv.org)
1 point by carlossouza on Nov 21, 2023 | hide | past | pdf | 2 comments on HN

In plain words: Before answering, the model rewrites its input to keep only the relevant parts, then responds from that cleaned version instead of the original. On tasks filled with opinions or irrelevant details, it gave more factual, objective answers and less flattery than the usual approach.

Abstract

Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these issues, we introduce System 2 Attention (S2A), which leverages the ability of LLMs to reason in natural language and follow instructions in order to decide what to attend to. S2A regenerates the input context to only include the relevant portions, before attending to the regenerated context to elicit the final response. In experiments, S2A outperforms standard attention-based LLMs on three tasks containing opinion or irrelevant information, QA, math word problems and longform generation, where S2A increases factuality and objectivity, and decreases sycophancy.

Jason Weston, Sainbayar Sukhbaatar
arXiv:2311.11829 · cs.CL, cs.AI, cs.LG · submitted Nov 20, 2023
abstract · pdf · html

add comment on HN
Also discussed: Nov 2023 (2 points, 0 comments) · Nov 2023 (3 points, 0 comments)

The study introduces System 2 Attention (S2A) as a solution to the problem of incorporating irrelevant information in Transformer-based Large Language Models (LLMs). S2A improves performance on tasks involving opinion or irrelevant information by regenerating the input context to only include relevant portions before attending to it. Result is increased factuality and objectivity and decreased sycophancy.
In Langroid (the agent-oriented LLM framework from ex-CMU/UW-Madison researchers), we call it Relevance Extraction — given a passage and a query, use the LLM to extract only the portions relevant to the query. In a RAG pipeline where you optimistically retrieve top k chunks (to improve recall), the chunks could be large and hence contain irrelevant/distracting text. We concurrently do relevance extraction from these k chunks: https://github.com/langroid/langroid/blob/main/langroid/agen...

One thing often missed in this is the un-necessary cost (latency and token-cost) of parroting out verbatim text from context. In Langroid we use a numbering trick to mitigate this: pre-annotate the passage sentences with numbers, and ask the LLM to simply specify the relevant sentence-numbers. We have an elegant implementation of this in our RelevanceExtractorAgent using tools/function-calling.

Here's a post I wrote about comparing Langroid's method with LangChain's naive equivalent of relevance extraction called `LLMChainExtractor.compress` , and no surprise Langroid's methos is far faster and cheaper: https://www.reddit.com/r/LocalLLaMA/comments/17k39es/relevan...