about
CoVE: Compressed Vocabulary Expansion Makes Better LLM-Based Recommender Systems (arxiv.org)
2 points by PaulHoule on Jul 1, 2025 | hide | past | pdf | discuss on HN

In plain words: Each item gets a unique ID in an expanded vocabulary, so the model reads a user's history as a sequence; the ID lookup table is compressed to stay small. It beat the usual approach of aligning language models to the task on several datasets.

Abstract · CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems

Recommender systems play a pivotal role in providing relevant content to users. With the rapid development of large language models (LLMs), researchers have begun utilizing LLMs to build more powerful recommender systems. However, existing approaches that focus on aligning LLMs with recommendation tasks do not fully leverage their sequential information processing capabilities, leading to suboptimal performance. In this paper, we propose a novel system called compressed vocabulary expansion (CoVE). In CoVE, each item is assigned a unique ID within the expanded vocabulary. Our framework effectively capitalizes on sequence understanding abilities of LLMs, significantly enhancing their performance on recommendation tasks. Additionally, we compress the embedding layer, making CoVE practical for large-scale industrial applications. The effectiveness and performance of CoVE are demonstrated through comprehensive experiments on multiple recommendation datasets and comparisons with prior works. Our code can be found at https://github.com/HaochenZhang717/CoVE-official-Repo.

Haochen Zhang, Tianyi Zhang, Junze Yin, Oren Gal, Anshumali Shrivastava, Vladimir Braverman
arXiv:2506.19993 · cs.IR, cs.LG · submitted Jun 24, 2025
abstract · pdf · html · Accepted by ACL 2025 Findings

add comment on HN