about
LLMs Do Not Grade Essays Like Humans (arxiv.org)
6 points by PretzelFisch 188 days ago | hide | past | pdf | 4 comments on HN

In plain words: Off-the-shelf AI models graded essays like human raters, with no task training. Their scores only weakly matched human grades—flattering short, thin essays and penalizing long ones with small typos—though their written feedback agreed with their own scores.

Abstract

Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-generated scores compare with human grades and analyze the grading behavior of several models from the GPT and Llama families in an out-of-the-box setting, without task-specific training. Our results show that agreement between LLM and human scores remains relatively weak and varies with essay characteristics. In particular, compared to human raters, LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors. We also find that the scores generated by LLMs are generally consistent with the feedback they generate: essays receiving more praise tend to receive higher scores, while essays receiving more criticism tend to receive lower scores. These results suggest that LLM-generated scores and feedback follow coherent patterns but rely on signals that differ from those used by human raters, resulting in limited alignment with human grading practices. Nevertheless, our work shows that LLMs produce feedback that is consistent with their grading and that they can be reliably used in supporting essay scoring.

Jerin George Mathew, Sumayya Taher, Anindita Kundu, Denilson Barbosa
arXiv:2603.23714 · cs.AI, cs.CL · submitted Mar 24, 2026
abstract · pdf · html

add comment on HN

I feel like this is actually human-like but like the average human in the pretraining data. Let's look:

1. They reward short or under-developed essays. I'd say most online content, especially with high upvotes next to the post, fits that. Social media surely does.

2. If it's longer posts, the system starts nitpicking it on minor details, like grammar. We see this even on Hacker News, a community valuing quality, with some longer submissions. It's also a debate tactic to derail opponents' better arguments in many discussions which are in their pretraining data.

3. Essays with more praise get higher scores and with more criticism get lower scores. "Get on the Bandwagon" Effect. Echo chambers. One person writes a thing followed by 5-20 people confirming it. That's probably in the pretraining data. It might survive some filtering/cleaning strategies, too.

So, no, I think these AI's are acting way too human. They need to fine-tune them to act like more, reasonable humans. That will initially take RLHF data for many types of situations. Given pretraining bias, they might also have to train them to drop the bad habits the article mentions.

School-type long essays only seem to exist in academia. I took a "business communication" class in college and we didn't write essays. My life experience since then has supported the "no essays" conclusion.

A long comment online now means either two things: it's written by a crank who has strong opinions, usually only tangentially related; or someone who has deep knowledge about the subject and has a lot of detail to provide. It's usually the former.

I agree with you on how their quality is spread out. But, this...

"School-type long essays only seem to exist in academia."

Does an AI know what an essay is? Would it consider any long, descriptive post an essay? Especially if pretraining data has many people describing long posts as essays or "essay-like?" Or only actual essays? And what is an actual essay again?

I think AI's might have different interpretations due to the above questions. They might also conflate essays with longer, detailed, or argumentative posts. We'd have to put a bunch of posts into a bunch of AI's to ask how they classify them.

Interesting point. Do you think this is more about training data limitations or evaluation methods?