Human readers saw little overall difference
Human evaluators largely saw the two sets of short stories as alike overall. AI-authored stories scored higher on Authenticity, averaging 3.50 versus 3.27 for human-authored stories on the five-point scale, and on Elaboration, averaging 3.54 versus 3.17. The other human-rated dimensions did not differ significantly.
The study compared 200 short stories paired with WritingPrompts prompts: 100 randomly selected human-generated stories and 100 machine-generated stories. Five large language models produced the machine-written set, with 20 stories from each model.
Human evaluation comprised 441 ratings from 115 participants, averaging 2.21 independent ratings per story. The human and LLM assessments were blind and covered 11 dimensions using a five-point scale.
The automated judge told a different story
The automated judge produced a sharply different result. Its scores for AI-generated and human-generated stories differed significantly on all 11 dimensions, always favoring the AI stories. On six dimensions—Effectiveness, Elaboration, Fluency, Originality, Value and overall Creativity—the AI stories received perfect scores with no variation.
When the researchers compared the order in which the judge and human evaluators rated the stories, agreement was weak at best. Elaboration showed the strongest match at 0.31, overall Creativity was 0.23, and Surprise was essentially unaligned at 0.01.
The disagreement also appeared in the way the ratings moved together. Human overall Creativity was associated with all 10 other dimensions, whereas the LLM's overall Creativity showed no association with Surprise or Usefulness and was most strongly linked to Elaboration.
Simple metrics missed the human signal
The study tested seven computed metrics against human ratings, using rank-based correlations to see whether higher machine scores tracked higher human scores. Most of the correlations fell between -0.2 and 0.2, indicating very weak or negligible relationships.
The Creativity Index was almost unrelated to human-rated Creativity, with a correlation of 0.07, and its correlation with the other 10 dimensions did not exceed 0.11. The strongest positive link from any automatic metric was between CR-POS and human-rated Elaboration, at 0.29, which the paper described as weak.
Perplexity, a token-level uncertainty score, was negatively associated with all 11 dimensions in the LLM judge's ratings. The paper described the links with Surprise and Creativity as weak and those with Originality and Novelty as moderate, pointing to a separation between token-level uncertainty and the narrative properties assigned by the judge.
A narrow result with a clear warning
The conclusion is limited to this evaluation setting. The study used one task and a 200-text corpus, while a single open-source LLM served as the judge, so it did not test whether other judge models would show the same pattern.
Within those boundaries, the results argue against treating automatic creativity metrics and LLM judges as interchangeable substitutes for human evaluation of short stories. They also do not establish that human writers are more creative overall, because human ratings were largely similar across the two story groups.
Paper data and sources
Original title: The Limits of Automatic Evaluation of Creativity in Large Language Models
Authors: Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text