about
Training Plug-N-Play Knowledge Modules with Deep Context Distillation (arxiv.org)
2 points by PaulHoule on Mar 27, 2025 | hide | past | pdf | discuss on HN

In plain words: Small plug-in modules store a document's knowledge for adding to a model later, trained by copying the hidden signals and word guesses of a model that reads the document, not by predicting the next word. This beat plain next-word training on both test sets.

Abstract · Training Plug-n-Play Knowledge Modules with Deep Context Distillation

Dynamically integrating new or rapidly evolving information after (Large) Language Model pre-training remains challenging, particularly in low-data scenarios or when dealing with private and specialized documents. In-context learning and retrieval-augmented generation (RAG) face limitations, including their high inference costs and their inability to capture global document information. In this paper, we propose a way of modularizing knowledge by training document-level Knowledge Modules (KMs). KMs are lightweight components implemented as parameter-efficient LoRA modules, which are trained to store information about new documents and can be easily plugged into models on demand. We show that next-token prediction performs poorly as the training objective for KMs. We instead propose Deep Context Distillation: we learn KMs parameters such as to simulate hidden states and logits of a teacher that takes the document in context. Our method outperforms standard next-token prediction and pre-instruction training techniques, across two datasets. Finally, we highlight synergies between KMs and RAG.

Lucas Caccia, Alan Ansell, Edoardo Ponti, Ivan Vulić, Alessandro Sordoni
arXiv:2503.08727 · cs.LG, cs.AI · submitted Mar 11, 2025 · updated Aug 8, 2025
abstract · pdf · html · Accepted at the CONFERENCE ON LANGUAGE MODELING (COLM) 2025

add comment on HN