Pith. sign in

REVIEW 3 cited by

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04069 v2 pith:MIGDVJSS submitted 2024-07-04 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords evaluationmodelsreviewcausingchallengescriticalensureevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 6.5 of 10

    Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...

  2. CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance

    cs.CL 2025-08 conditional novelty 5.0 of 10

    CAMB provides a seven-task, eight-dataset benchmark for assessing LLM and embedding model performance in civil aviation maintenance, with initial results showing large models top out near 69% on domain multiple-choice...

  3. Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey that classifies AI agent evaluation benchmarks along environment and capability axes, and proposes five traits that distinguish agents from chatbots.

Pith tools