about
LMAR: Language Model Augmented Retriever for Domain-Specific Knowledge Indexing (arxiv.org)
2 points by PaulHoule on Aug 28, 2025 | hide | past | pdf | discuss on HN

In plain words: A language model writes and double-checks example question–document matches, which are then used to retrain a standard text retriever so it can find specialized knowledge without heavy computing. On several domain-specific benchmarks it beat baseline retrievers while staying fast and needing only modest hardware.

Abstract · LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing

Retrieval Augmented Generation (RAG) systems often struggle with domain-specific knowledge due to performance deterioration of pre-trained embeddings and prohibitive computational costs of large language model (LLM)-based retrievers. While fine-tuning data augmentation embedding models offers a promising direction, its effectiveness is limited by the need for high-quality training data and reliable chunking strategies that preserve contextual integrity. We propose LMAR (Language Model Augmented Retriever), a model-agnostic framework that addresses these challenges by combining LLM-guided data synthesis with contrastive embedding adaptation and efficient text clustering. LMAR consists of a two-stage pipeline: (1) Triplet sampling and synthetic data augmentation, where LLMs act as both labeler and validator to ensure high-fidelity supervision throughout the pipeline. Experimental results across multiple domain-specific benchmark datasets demonstrate that LMAR outperforms multiple baseline models, while maintaining moderate hardware requirements and low latency. Its model-agnostic nature further enables seamless integration with emerging RAG architectures and text embedding models, ensuring continual improvements without redesigning the pipeline. These results highlight LMAR as a practical and cost-effective solution for scalable domain-specific adaptation.

Yao Zhao, Yantian Ding, Zhiyue Zhang, Dapeng Yao, Yanxun Xu
arXiv:2508.05672 · cs.IR, cs.AI · submitted Aug 4, 2025 · updated Sep 12, 2025
abstract · pdf · html

add comment on HN