A reference-based Likert scoring method with analyze-rate prompting improves LLM judges' agreement with human creativity rankings on the TTCW benchmark, though the headline result is partly fitted to the test set.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
A reference-based Likert scoring method with analyze-rate prompting improves LLM judges' agreement with human creativity rankings on the TTCW benchmark, though the headline result is partly fitted to the test set.