On a new controlled dataset of corrupted short stories, eight LLMs produce mostly correct, specific writing feedback but often fail to identify the biggest writing issue and are poor at deciding when to say a story is perfect.
I don’t see any big problems in the story
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Help Me Write a Story: Evaluating LLMs' Ability to Generate Writing Feedback
On a new controlled dataset of corrupted short stories, eight LLMs produce mostly correct, specific writing feedback but often fail to identify the biggest writing issue and are poor at deciding when to say a story is perfect.