about
Do LLMs Understand User Preferences? Evaluating LLMs on User Rating Prediction (arxiv.org)
2 points by PaulHoule on May 12, 2023 | hide | past | pdf | 2 comments on HN

In plain words: They tested language models on predicting a user's rating for an item from their past ratings, with no examples, a few, or fine-tuning. Without fine-tuning they fell behind classic recommenders, but after tuning on a small slice of ratings they matched or beat them.

Abstract · Do LLMs Understand User Preferences? Evaluating LLMs On User Rating Prediction

Large Language Models (LLMs) have demonstrated exceptional capabilities in generalizing to new tasks in a zero-shot or few-shot manner. However, the extent to which LLMs can comprehend user preferences based on their previous behavior remains an emerging and still unclear research question. Traditionally, Collaborative Filtering (CF) has been the most effective method for these tasks, predominantly relying on the extensive volume of rating data. In contrast, LLMs typically demand considerably less data while maintaining an exhaustive world knowledge about each item, such as movies or products. In this paper, we conduct a thorough examination of both CF and LLMs within the classic task of user rating prediction, which involves predicting a user's rating for a candidate item based on their past ratings. We investigate various LLMs in different sizes, ranging from 250M to 540B parameters and evaluate their performance in zero-shot, few-shot, and fine-tuning scenarios. We conduct comprehensive analysis to compare between LLMs and strong CF methods, and find that zero-shot LLMs lag behind traditional recommender models that have the access to user interaction data, indicating the importance of user interaction data. However, through fine-tuning, LLMs achieve comparable or even better performance with only a small fraction of the training data, demonstrating their potential through data efficiency.

Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta, Maheswaran Sathiamoorthy, Lichan Hong, Ed Chi, Derek Zhiyuan Cheng
arXiv:2305.06474 · cs.IR, cs.LG · submitted May 10, 2023
abstract · pdf · html

add comment on HN
Also discussed: Jun 2023 (1 point, 0 comments)

I've often pondered the idea of semantic recommendation but it is a dynamic problem with an aversion for repetition. I am looking forward to digging into this paper. A big question for me is if you can successfully avoid repeating different coverage while maximizing affinity.
My content-based recommender shows me one piece of content at a time, so like Tik Tok or Tinder, I get a clean signal. (Unlike the Youtube problem which is more like predicting which distractor from the image on https://tvtropes.org/pmwiki/pmwiki.php/Funny/Idiocracy you will click on)

My main evaluation metric is area under curve for predicting “will I like the item?” but that metric doesn’t necessarily correlate to satisfaction. I’ve done a round of improving the model to add another percentage point to my AUC but right now that’s a distraction from targeting annoyances such getting 10 articles about the same news event.

The real frontier is “sequential recommendation” where all the literature seems to come from China and India and I fear there will be a “missile gap” for e-commerce in the next few years. (e.g. the state of the art in e-commerce in the US is “You’ve subscribed to Prime for 15 years and they figure you’ll subscribe to Prime for 15 years even if two day shipping is downgraded to five day shipping”)