Pith. sign in

REVIEW 14 cited by

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18403 v3 pith:5UHH7F2S submitted 2024-06-26 cs.CL

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

classification cs.CL
keywords humanmodelsllmsacrossannotationscoveringdatasetsevaluated
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide JUDGE-BENCH, an extensible collection of 20 NLP datasets with human annotations covering a broad range of evaluated properties and types of data, and comprehensively evaluate 11 current LLMs, covering both open-weight and proprietary models, for their ability to replicate the annotations. Our evaluations show substantial variance across models and datasets. Models are reliable evaluators on some tasks, but overall display substantial variability depending on the property being evaluated, the expertise level of the human judges, and whether the language is human or model-generated. We conclude that LLMs should be carefully validated against human judgments before being used as evaluators.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Natural Language-Focused Software Engineering via Code-Documentation Equivalence

    cs.SE 2026-06 unverdicted novelty 7.0

    Defines documentation-to-code equivalence and introduces Documentary to generate matching docs for 53.4% of function snippets, raising LLM output prediction accuracy by 12.8-24.5% over human-written docs.

  2. Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication

    cs.CY 2026-05 unverdicted novelty 7.0

    An audit of one million Korean synthetic personas shows marginal demographic alignment does not preserve joint distributions, with three specific mismatches identified via a new Independence-Assumption Footprint method.

  3. Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench

    cs.AI 2026-04 conditional novelty 7.0

    AgentProp-Bench shows substring judging agrees with humans at kappa=0.049, LLM ensemble at 0.432, bad-parameter injection propagates with ~0.62 probability, rejection and recovery are independent, and a runtime fix cu...

  4. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 7.0

    The paper delivers a taxonomy of seven LLM study types in software engineering along with eight guidelines that separate mandatory requirements from recommended practices to address reproducibility challenges.

  5. Instructions Shape Production of Language, not Processing

    cs.CL 2026-05 unverdicted novelty 6.0

    Instructions trigger a production-centered mechanism in language models, with task-specific information stable in input tokens but varying strongly in output tokens and correlating with behavior.

  6. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 6.0

    A group of 22 researchers proposes seven study types and eight guidelines for empirical software engineering studies involving LLMs to enhance reproducibility and replicability.

  7. Instructions Shape Production of Language, not Processing

    cs.CL 2026-05 unverdicted novelty 5.0

    Instructions primarily shape the production stage of language models rather than the processing stage, with task-specific information and causal effects stronger in output tokens than input tokens.

  8. Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

    cs.LG 2026-05 unverdicted novelty 5.0

    Generative AI evaluation should prioritize real-world utility through stakeholder goals and longitudinal human outcome measurements instead of static benchmark performance.

  9. Mixed response geometry and critical crossover in the Ising model

    cond-mat.stat-mech 2026-04 unverdicted novelty 5.0

    In the 2D Ising model, the mixed response Ω_βh = −N cov(m,e) forms a localized ridge from criticality into finite-field crossover and collapses onto a susceptibility-constrained curve in normalized coordinates.

  10. Mixed response geometry and critical crossover in the Ising model

    cond-mat.stat-mech 2026-04 unverdicted novelty 5.0

    Thermodynamic curvature on the (β, h) manifold in the Ising model produces a ridge that geometrically identifies the Widom line as the locus of maximal response extending from the critical point.

  11. Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

    cs.LG 2026-05 unverdicted novelty 4.0

    Generative AI evaluation must shift from static benchmark scores to measuring sustained improvements in human capabilities within specific deployment contexts.

  12. Personalized AI Practice Replicates Learning Rate Regularity at Scale

    cs.CY 2026-03 unverdicted novelty 4.0

    Large-scale data from an AI platform confirms students have consistent learning rates (IQR 7.01-8.25 opportunities to 80% mastery) despite variable starting knowledge, replicating prior findings with automated knowled...

  13. A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models

    cs.CL 2026-07 conditional novelty 3.0

    A five-model panel ranked each other's anonymized responses, and the averaged ranks show which chatbots other chatbots prefer.

  14. LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

    cs.CL 2024-12 accept novelty 3.0

    A survey that organizes LLMs-as-judges research into functionality, methodology, applications, meta-evaluation, and limitations.