REVIEW 3 cited by
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust.
Forward citations
Cited by 3 Pith papers
-
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam
Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...
-
CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
CAMB provides a seven-task, eight-dataset benchmark for assessing LLM and embedding model performance in civil aviation maintenance, with initial results showing large models top out near 69% on domain multiple-choice...
-
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.
Discussion (0). Sign in to comment.