In plain words: ChatGPT labeled 2,382 tweets for relevance, stance, topics, and frames with no examples, and its answers were compared with those from paid crowd-workers and trained annotators. It beat crowd-workers on four of five tasks and cost about twenty times less per label.
Abstract · ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd-workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we demonstrate that ChatGPT outperforms crowd-workers for several annotation tasks, including relevance, stance, topics, and frames detection. Specifically, the zero-shot accuracy of ChatGPT exceeds that of crowd-workers for four out of five tasks, while ChatGPT's intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk. These results show the potential of large language models to drastically increase the efficiency of text classification.
Fabrizio Gilardi, Meysam Alizadeh, Maël Kubli
arXiv:2303.15056 · cs.CL, cs.CY · submitted Mar 27, 2023 · updated Jul 19, 2023
abstract · pdf · html · Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. 2023. "ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks". Proceedings of the National Academy of Sciences 120(30): e2305016120
One problem is that you don't have the ground truth for D. So you start by annotating D with the labels assinged by a third classifier, C:
Having thus established a modicum of "ground truth", ish, you proceed to annotate D with the two classifiers you are comparing, A and B: Then you compare the classifications D₂ and D₃ to D₁, and find that D₃ better approximates D₁ than D₂.What is the result of the experiment? I summarise it as follows:
Now we can name the three classifiers as they were used in the experiments in the linked article:A: Human annotators.
B: ChatGPT
C: Human annotators.
So the result of the paper is that, plugging in the names:
And that, is the finding of the paper.Which is clearly absurd and a cause to re-think methodology.