In plain words: Several pre-trained transformer models were trained to label text with emotions, testing how much of the model is fine-tuned and how the text is cleaned beforehand. Cleaning backfired: stripping punctuation and common little words lowered accuracy, because those cues carry tone and emphasis the models rely on.
Abstract
In this study, we explore the application of transformer-based models for emotion classification on text data. We train and evaluate several pre-trained transformer models, on the Emotion dataset using different variants of transformers. The paper also analyzes some factors that in-fluence the performance of the model, such as the fine-tuning of the transformer layer, the trainability of the layer, and the preprocessing of the text data. Our analysis reveals that commonly applied techniques like removing punctuation and stop words can hinder model performance. This might be because transformers strength lies in understanding contextual relationships within text. Elements like punctuation and stop words can still convey sentiment or emphasis and removing them might disrupt this context.
Mahdi Rezapour
arXiv:2403.15454 · cs.CL, stat.AP · submitted Mar 18, 2024 · updated Jul 27, 2024
abstract · pdf