about
LLM4ES: Learning User Embeddings from Event Sequences via Large Language Models (arxiv.org)
1 point by PaulHoule on Sep 1, 2025 | hide | past | pdf | discuss on HN

In plain words: It turns each user's history of events into text and trains a language model to guess the next item, using its internal state as a compact user profile. This beat the usual embedding methods at classifying users in finance and other fields.

Abstract

This paper presents LLM4ES, a novel framework that exploits large pre-trained language models (LLMs) to derive user embeddings from event sequences. Event sequences are transformed into a textual representation, which is subsequently used to fine-tune an LLM through next-token prediction to generate high-quality embeddings. We introduce a text enrichment technique that enhances LLM adaptation to event sequence data, improving representation quality for low-variability domains. Experimental results demonstrate that LLM4ES achieves state-of-the-art performance in user classification tasks in financial and other domains, outperforming existing embedding methods. The resulting user embeddings can be incorporated into a wide range of applications, from user segmentation in finance to patient outcome prediction in healthcare.

Aleksei Shestov, Omar Zoloev, Maksim Makarenko, Mikhail Orlov, Egor Fadeev, Ivan Kireev, Andrey Savchenko
arXiv:2508.05688 · cs.IR · submitted Aug 6, 2025 · updated Dec 17, 2025
abstract · pdf · html

add comment on HN