about
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models (arxiv.org)
20 points by mfiguiere on Sep 20, 2024 | hide | past | pdf | 3 comments on HN

In plain words: Instead of splitting a document into pieces and embedding each alone, this runs the document through a long-context model, then cuts it just before averaging word signals into vectors. The pieces keep surrounding context and beat separately embedded chunks on retrieval tests, without training.

Abstract

Many use cases require retrieving smaller portions of text, and dense vector-based retrieval systems often perform better with shorter text segments, as the semantics are less likely to be over-compressed in the embeddings. Consequently, practitioners often split text documents into smaller chunks and encode them separately. However, chunk embeddings created in this way can lose contextual information from surrounding chunks, resulting in sub-optimal representations. In this paper, we introduce a novel method called late chunking, which leverages long context embedding models to first embed all tokens of the long text, with chunking applied after the transformer model and just before mean pooling - hence the term late in its naming. The resulting chunk embeddings capture the full contextual information, leading to superior results across various retrieval tasks. The method is generic enough to be applied to a wide range of long-context embedding models and works without additional training. To further increase the effectiveness of late chunking, we propose a dedicated fine-tuning approach for embedding models.

Michael Günther, Isabelle Mohr, Daniel James Williams, Bo Wang, Han Xiao
arXiv:2409.04701 · cs.CL, cs.IR · submitted Sep 7, 2024 · updated Jul 7, 2025
abstract · pdf · html · 11 pages, 3rd draft

add comment on HN

This is an interesting development. A couple blog posts exposing on late chunking

https://weaviate.io/blog/late-chunking

https://jina.ai/news/late-chunking-in-long-context-embedding...

The late chunking idea originates from ColBERT, an embedding technique from 2020 https://arxiv.org/pdf/2004.12832

Another (simpler?) approach is to also split your chunks into sentences. So you’ll end up with chunk embeddings and sentence embeddings. Now you can do sentence level search. And also distill chunks down to their most relevant sentences at query time before you dump ‘em into your LLM’s context window. If you use Sentence Transformers, you get your chunk embeddings for free, because they are just the np.mean of the embeddings of all the sentences in that chunk.
I'm not sure how it's going to solve the same issue. If your chunk is too short for the context, it's still going to miss the meaning if I understand the description correctly. As in "it's" 2 paragraphs in is still going to miss context, whether you embed chunks or sentences