about
Everyone prefers human writers, even AI (arxiv.org)
5 points by freejoe76 350 days ago | hide | past | pdf | discuss on HN

In plain words: They gave people and AI judges literary passages, swapping the author labels to see if the name alone changes style ratings. Both preferred work marked human-written—AI judges by 34.3 percentage points versus 13.7 for people, even flipping judgments of identical features.

Abstract · The human-authorship halo: attribution bias in literary style evaluation by humans and AI

As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standards are elusive and judgments are inherently subjective. We conducted controlled experiments using Raymond Queneau's Exercises in Style (1947) to measure attribution bias across evaluators. Study 1 compared human participants (N=556) and AI models (N=13) evaluating literary passages from Queneau versus GPT-4-generated versions under three conditions: blind, accurately labeled, and counterfactually labeled. Study 2 tested bias generalization across a 14x14 matrix of AI evaluators and creators. Both studies revealed systematic pro-human attribution bias. Humans showed +13.7 percentage point (pp) bias (Cohen's h = 0.28, 95% CI: 0.19-0.37), while AI models showed +34.3 percentage point bias (h = 0.70, 95% CI: 0.62-0.78), a 2.5-fold stronger effect (P<0.001). Study 2 confirmed this bias operates across AI architectures (+25.8pp, 95% CI: 24.1-27.6%), demonstrating that AI systems systematically devalue creative content when labeled as "AI-generated" regardless of which AI created it. We also find that attribution labels lead evaluators to invert assessment criteria, with identical features receiving opposing evaluations based solely on perceived authorship. This suggests AI models have absorbed human cultural biases against artificial creativity during training, including through the preference signals on which they are aligned. Our study represents the first controlled comparison of attribution bias between human and artificial evaluators in aesthetic judgment, revealing that AI systems not only replicate but amplify this human tendency.

Wouter Haverals, Meredith Martin
arXiv:2510.08831 · cs.AI, cs.CL, cs.HC · submitted Oct 9, 2025 · updated Sep 4, 2026
abstract · pdf · html · 70 pages (main text + SI Appendix), 7 main-text figures. v2: revised and retitled after peer review

add comment on HN