about
A Flexible Retrieval-Augmented Framework for Long-Text Query Processing (arxiv.org)
2 points by PaulHoule on Mar 17, 2025 | hide | past | pdf | discuss on HN

In plain words: A system that checks the task's state and decides as it goes how much of a long document to pull in or shrink, rather than feeding everything or using a fixed retrieval plan. It answered long-document questions more accurately and cheaply than usual approaches.

Abstract · OkraLong: A Flexible Retrieval-Augmented Framework for Long-Text Query Processing

Large Language Models (LLMs) encounter challenges in efficiently processing long-text queries, as seen in applications like enterprise document analysis and financial report comprehension. While conventional solutions employ long-context processing or Retrieval-Augmented Generation (RAG), they suffer from prohibitive input expenses or incomplete information. Recent advancements adopt context compression and dynamic retrieval loops, but still sacrifice critical details or incur iterative costs. To address these limitations, we propose OkraLong, a novel framework that flexibly optimizes the entire processing workflow. Unlike prior static or coarse-grained adaptive strategies, OkraLong adopts fine-grained orchestration through three synergistic components: analyzer, organizer and executor. The analyzer characterizes the task states, which guide the organizer in dynamically scheduling the workflow. The executor carries out the execution and generates the final answer. Experimental results demonstrate that OkraLong not only enhances answer accuracy but also achieves cost-effectiveness across a variety of datasets.

Yulong Hui, Yihao Liu, Yao Lu, Huanchen Zhang
arXiv:2503.02603 · cs.CL, cs.IR · submitted Mar 4, 2025 · updated Mar 5, 2025
abstract · pdf · html

add comment on HN