about
People who frequently use ChatGPT for writing tasks can detect AI-generated text (arxiv.org)
12 points by gaws on Jul 17, 2025 | hide | past | pdf | 7 comments on HN

In plain words: People who often use AI tools for writing judged 300 articles as human- or machine-written and explained why. When five voted together, they got just 1 of 300 wrong, beating most software detectors even after the text was paraphrased to hide its origins.

Abstract · People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text

In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments show that annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback. In fact, the majority vote among five such "expert" annotators misclassifies only 1 of 300 articles, significantly outperforming most commercial and open-source detectors we evaluated even in the presence of evasion tactics like paraphrasing and humanization. Qualitative analysis of the experts' free-form explanations shows that while they rely heavily on specific lexical clues ('AI vocabulary'), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors. We release our annotated dataset and code to spur future research into both human and automated detection of AI-generated text.

Jenna Russell, Marzena Karpinska, Mohit Iyyer
arXiv:2501.15654 · cs.CL, cs.AI · submitted Jan 26, 2025 · updated May 19, 2025
abstract · pdf · html · ACL 2025 33 pages

add comment on HN
Also discussed: May 2026 (3 points, 0 comments) · Apr 2026 (11 points, 2 comments) · Jul 2025 (1 point, 2 comments) · May 2025 (1 point, 0 comments) · Jan 2025 (2 points, 0 comments)

LLMs by default use the same style every time and don't know have any drive to differentiate. I just wrote an article about this in the "web site design" space where the designed sites ended up looking all alike (https://www.jasonthorsness.com/29). It makes complete sense that people will start to recognize the "default style" of each LLM.

I wonder whether this style is specific to the LLM (Grok vs. ChatGPT) or if it will somehow arise from the raining data itself that they all share and be sort of a permanent "accent" the LLMs have.

> I wonder whether this style is specific to the LLM (Grok vs. ChatGPT) or if it will somehow arise from the raining data itself that they all share and be sort of a permanent "accent" the LLMs have.

It's very different. I used LLMs to brainstorm a business plan write-up hypothetical startup and Claude/Gemini/qwen3:30b-a3b, one-shot with a long background text. They all generated similar-but-slightly-non-overlapping ideas in very different language.

If I remember right it went something like this: Gemini used stilted business-speak and bold everywhere, but had the best structure for the document. Claude gave surprisingly good one-paragraph intros to every section. Claude and Gwen3 were roughly tied to how nicely written their bullet point content was. I made a new document with Gemini's base structure, Claude's intros, and bullet points mashed from Claude and Gwen, making sure I covered all the good ideas from Gemini that weren't mentioned. Then I edited everything for homogeneity and style.

I recommend experimenting with combining their work.

(Also, qwen3:30b-a3b is amazing for local LLM work, the MoE architecture makes it have the speed of a 3B parameter model!)

I think you don't even need to use chatGPT, who was using make-me-sound smart verbs like 'delve' frequently before? Now, in some publications; I see it regularly.
It's like the magic sunglasses in They Live. You're just seeing the intricate tapestry of nuances.
You're absolutely right!
Why would I investigate the labyrinth floor if the walls are shorter than my eyesight?

I feel like this is a fools errand. Let's use the marketing terms to frame this problem. Companies are saying LLMs are like "the new printing press". Should I worry about detecting if something was printed or written by hand? No, I should worry about the volume and contents. That's what matters.

This looks like AI slop, I can tell from the em dash and from seeing quite a lot of AI slop in my time

I mean, not a very surprising result? At work I can almost always immediately see who wrote which lines of code, since we don't enforce code formatting. From what I can gather most authors do have a distinctive writing style as well, which can be detected. Why would LLMs be different?