In plain words: A controlled resume experiment tested whether AI hiring screeners favor resumes written by AI, especially their own kind. They did: applicants using the same AI as the screener were 23% to 60% more likely to be shortlisted than equally qualified people with human-written resumes.
Abstract
As artificial intelligence (AI) tools become widely adopted, large language models (LLMs) are increasingly involved on both sides of decision-making processes, ranging from hiring to content moderation. This dual adoption raises a critical question: do LLMs systematically favor content that resembles their own outputs? Prior research in computer science has identified self-preference bias -- the tendency of LLMs to favor their own generated content -- but its real-world implications have not been empirically evaluated. We focus on the hiring context, where job applicants often rely on LLMs to refine resumes, while employers deploy them to screen those same resumes. Using a large-scale controlled resume correspondence experiment, we find that LLMs consistently prefer resumes generated by themselves over those written by humans or produced by alternative models, even when content quality is controlled. The bias against human-written resumes is particularly substantial, with self-preference bias ranging from 67% to 82% across major commercial and open-source models. To assess labor market impact, we simulate realistic hiring pipelines across 24 occupations. These simulations show that candidates using the same LLM as the evaluator are 23% to 60% more likely to be shortlisted than equally qualified applicants submitting human-written resumes, with the largest disadvantages observed in business-related fields such as sales and accounting. We further demonstrate that this bias can be reduced by more than 50% through simple interventions targeting LLMs' self-recognition capabilities. These findings highlight an emerging but previously overlooked risk in AI-assisted decision making and call for expanded frameworks of AI fairness that address not only demographic-based disparities, but also biases in AI-AI interactions.
Jiannan Xu, Gujie Li, Jane Yi Jiang
arXiv:2509.00462 · cs.CY · submitted Aug 30, 2025 · updated Jun 6, 2026
abstract · pdf · html · This paper has been accepted as a non-archival submission at EAAMO 2025 and AIES 2025
"If I read the paper correctly, they don’t actually show that LLMs prefer resumes they generate.
Their actual method seems to be taking a human written resume, deleting the executive summary, having an LLM rewrite the executive summary based on the rest of the resume and then having another LLM rate the executive summary without the rest of the resume.
That’s likely to massively overstate any real impact, if you can even rely on it capturing a real effect.
I really wonder if I read that correctly, because I can’t come up with a justification for that study design."
[0] I couldn't help but mildly copy-edit before pasting here.
Edit: yes, the authors present a reason for their design, and an ideal version of my comment would've said that. I do not consider it much of a justification. See below: https://news.ycombinator.com/item?id=47987256#47987727.