about
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation (arxiv.org)
1 point by PaulHoule on Jan 30, 2025 | hide | past | pdf | discuss on HN

In plain words: The system teaches a language model to think through a user's past preferences and writing style before drafting, then keeps training it on its own best-written outputs. On four personalized long-form writing tasks, it beat the strongest existing approaches by an average of 14.5%.

Abstract

Personalized text generation requires a unique ability of large language models (LLMs) to learn from context that they often do not encounter during their standard training. One way to encourage LLMs to better use personalized context for generating outputs that better align with the user's expectations is to instruct them to reason over the user's past preferences, background knowledge, or writing style. To achieve this, we propose Reasoning-Enhanced Self-Training for Personalized Text Generation (REST-PG), a framework that trains LLMs to reason over personal data during response generation. REST-PG first generates reasoning paths to train the LLM's reasoning abilities and then employs Expectation-Maximization Reinforced Self-Training to iteratively train the LLM based on its own high-reward outputs. We evaluate REST-PG on the LongLaMP benchmark, consisting of four diverse personalized long-form text generation tasks. Our experiments demonstrate that REST-PG achieves significant improvements over state-of-the-art baselines, with an average relative performance gain of 14.5% on the benchmark.

Alireza Salemi, Cheng Li, Mingyang Zhang, Qiaozhu Mei, Weize Kong, Tao Chen, Zhuowan Li, Michael Bendersky, Hamed Zamani
arXiv:2501.04167 · cs.CL, cs.AI, cs.IR · submitted Jan 7, 2025
abstract · pdf · html

add comment on HN