An LLM evaluator using Chain-of-Thought and few-shot prompting correlates with human quality ratings at 0.65 to 0.78, beating BLEU, ROUGE, METEOR and semantic similarity metrics.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Generated Commit Messages with Large Language Models
An LLM evaluator using Chain-of-Thought and few-shot prompting correlates with human quality ratings at 0.65 to 0.78, beating BLEU, ROUGE, METEOR and semantic similarity metrics.