Pith. sign in

REVIEW 1 cited by

Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13769 v1 pith:Y6UOZHU3 submitted 2024-05-22 cs.CL

classification cs.CL
keywords automatichumanevaluationlanguagemodelslargellmsmeasures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Storytelling is an integral part of human experience and plays a crucial role in social interactions. Thus, Automatic Story Evaluation (ASE) and Generation (ASG) could benefit society in multiple ways, but they are challenging tasks which require high-level human abilities such as creativity, reasoning and deep understanding. Meanwhile, Large Language Models (LLM) now achieve state-of-the-art performance on many NLP tasks. In this paper, we study whether LLMs can be used as substitutes for human annotators for ASE. We perform an extensive analysis of the correlations between LLM ratings, other automatic measures, and human annotations, and we explore the influence of prompting on the results and the explainability of LLM behaviour. Most notably, we find that LLMs outperform current automatic measures for system-level evaluation but still struggle at providing satisfactory explanations for their answers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches

    cs.CL 2025-06 conditional novelty 6.0 of 10

    GPT-4o increases emotional lexicon and uses more questions and exclamations when asked to strengthen speeches, but follows a surface style rather than human-like persuasive argumentation.

Pith tools