Pith. sign in

REVIEW 2 cited by

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08008 v2 pith:WPAPRZMG submitted 2024-04-10 cs.LG cs.CLcs.HC

classification cs.LGcs.CLcs.HC
keywords humanevaluationdiscrepancylanguagellmsmethodsample-efficientcode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reliable evaluation of large language models (LLMs) is impeded by two key challenges: objective metrics often fail to reflect human perception of natural language, and exhaustive human labeling is prohibitively expensive. Here, we propose a sample-efficient human evaluation method for LLMs based on the principle of MAximum Discrepancy (MAD) Competition. Our method automatically and adaptively selects a compact set of input instructions that maximize semantic discrepancy between pairs of LLM responses. Human evaluators then perform three-alternative forced choices on these paired responses, which are aggregated into a global ranking using Elo rating. We apply our approach to compare eight widely used LLMs across four tasks: scientific knowledge understanding, mathematical reasoning, creative and functional writing, and code generation and explanation. Experimental results show that our sample-efficient evaluation method recovers "gold-standard" model rankings with a handful of MAD-selected instructions, reveals respective strengths and weaknesses of each LLM, and offers nuanced insights to guide future LLM development. Code is available at https://github.com/weiji-Feng/MAD-Eval .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

    cs.CL 2026-08 reject novelty 6.0 of 10

    SurveyReview is a dataset of 675 surveys with 1,630 reviews annotated into four quality dimensions, plus a fine-tuned evaluator (SurveyAlign) that reports lower error than GPT-5.2 but only reaches majority-class basel...

  2. How to Select Datapoints for Efficient Human Evaluation of NLG Models?

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Selecting human-evaluation items by metric variance, metric consistency, output diversity, or IRT-based informativeness matches random-sampling ranking accuracy with roughly 70% of the annotation budget in WMT23 and SummEval.

Pith tools