In plain words: A new test scores a story's creativity on 14 yes-or-no checks of fluency, flexibility, originality, and detail, judged by professional writers. AI-written stories passed 3 to 10 times fewer checks than human stories, and AI judges didn't match the experts.
Abstract
Researchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT), which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose the Torrance Test of Creative Writing (TTCW) to evaluate creativity as a product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments.
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu
arXiv:2309.14556 · cs.CL, cs.AI, cs.HC · submitted Sep 25, 2023 · updated Mar 8, 2024
abstract · pdf · html · ACM CHI 2024
They didn't even bother using multiple passes to prompt better quality creative writing like "Write an outline for the story..." "Expand this outline focusing on improving style and tone..." "Here's a story written by an amateur writer. As a professional writer, give it a rewrite improving vocabulary, metaphor, and overall quality."
As with a lot of the LLM stuff I see these days, I have to wonder at how much of what's being measured is the capacity of the tool and how much the measurement of the capacity of its users.
I'd never imagine using a LLM zero shot in a single pass for any creative writing tasks, and measuring the inability to perform in suboptimal conditions isn't all that revealing or novel (pun intended).