REVIEW 5 cited by
A Closer Look into Automatic Evaluation Using Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Using large language models (LLMs) to evaluate text quality has recently gained popularity. Some prior works explore the idea of using LLMs for evaluation, while they differ in some details of the evaluation process. In this paper, we analyze LLM evaluation (Chiang and Lee, 2023) and G-Eval (Liu et al., 2023), and we discuss how those details in the evaluation process change how well the ratings given by LLMs correlate with human ratings. We find that the auto Chain-of-Thought (CoT) used in G-Eval does not always make G-Eval more aligned with human ratings. We also show that forcing the LLM to output only a numeric rating, as in G-Eval, is suboptimal. Last, we reveal that asking the LLM to explain its own ratings consistently improves the correlation between the ChatGPT and human ratings and pushes state-of-the-art (SoTA) correlations on two meta-evaluation datasets.
Forward citations
Cited by 5 Pith papers
-
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
A new benchmark and 3D decomposition paradigm for factuality evaluation of interpretive claims about contact center conversations, with best LLM-judge F1 of 0.86.
-
Can Large Language Models Serve as Evaluators for Code Summarization?
An LLM prompt that makes the model role-play a code reviewer scores code summaries with 81.59% Spearman correlation with human judgment, outperforming BLEU and BERTScore on a 300-sample benchmark.
-
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources
Refining an LLM auto-evaluator with expert-teacher themes and few-shot examples improved agreement with human scores on quiz quality, but only on the same questions used for refinement.
-
Do LLMs Agree on the Creativity Evaluation of Alternative Uses?
Four LLMs show high agreement when scoring and ranking alternative uses and do not favor their own outputs, but the accuracy benchmark is derived from the generation prompts rather than human judgment.
-
SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions
SRSA, a router that picks between direct, parallel, and planning searches for each query, improves informativeness and completeness on a new contextual-query benchmark while using fewer LLM inference steps than a ReAct agent.
Discussion (0). Continue with ORCID to comment.