In plain words: A system guesses which of two posts will get more retweets by first having a chatbot write likely reactions, then feeding those to a model that picks the winner. The best version, using Claude's reactions, beat the same model run on post text alone.
Abstract
In the realm of social media, understanding and predicting post reach is a significant challenge. This paper presents a Crowd Reaction AssessMent (CReAM) task designed to estimate if a given social media post will receive more reaction than another, a particularly essential task for digital marketers and content writers. We introduce the Crowd Reaction Estimation Dataset (CRED), consisting of pairs of tweets from The White House with comparative measures of retweet count. The proposed Generator-Guided Estimation Approach (GGEA) leverages generative Large Language Models (LLMs), such as ChatGPT, FLAN-UL2, and Claude, to guide classification models for making better predictions. Our results reveal that a fine-tuned FLANG-RoBERTa model, utilizing a cross-encoder architecture with tweet content and responses generated by Claude, performs optimally. We further use a T5-based paraphraser to generate paraphrases of a given post and demonstrate GGEA's ability to predict which post will elicit the most reactions. We believe this novel application of LLMs provides a significant advancement in predicting social media post reach.
Sohom Ghosh, Chung-Chi Chen, Sudip Kumar Naskar
arXiv:2403.09702 · cs.CL · submitted Mar 8, 2024
abstract · pdf · html · Accepted for publication in The ACM Web Conference WWW'24 Companion Short Papers Track, May 13 to 17 2024, Singapore, DOI 10.1145/3589335.3651512
It is easy to collect the data for a model like this (compared to say, a sentiment model for social media) so it is a good project but it does have enough "danger" that I haven't shared my models yet.