REVIEW 14 cited by
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
read the original abstract
There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide JUDGE-BENCH, an extensible collection of 20 NLP datasets with human annotations covering a broad range of evaluated properties and types of data, and comprehensively evaluate 11 current LLMs, covering both open-weight and proprietary models, for their ability to replicate the annotations. Our evaluations show substantial variance across models and datasets. Models are reliable evaluators on some tasks, but overall display substantial variability depending on the property being evaluated, the expertise level of the human judges, and whether the language is human or model-generated. We conclude that LLMs should be carefully validated against human judgments before being used as evaluators.
Forward citations
Cited by 14 Pith papers
-
Natural Language-Focused Software Engineering via Code-Documentation Equivalence
Defines documentation-to-code equivalence and introduces Documentary to generate matching docs for 53.4% of function snippets, raising LLM output prediction accuracy by 12.8-24.5% over human-written docs.
-
Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication
An audit of one million Korean synthetic personas shows marginal demographic alignment does not preserve joint distributions, with three specific mismatches identified via a new Independence-Assumption Footprint method.
-
Evaluating Tool-Using Language Agents: Judge Reliability, Propagation Cascades, and Runtime Mitigation in AgentProp-Bench
AgentProp-Bench shows substring judging agrees with humans at kappa=0.049, LLM ensemble at 0.432, bad-parameter injection propagates with ~0.62 probability, rejection and recovery are independent, and a runtime fix cu...
-
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
The paper delivers a taxonomy of seven LLM study types in software engineering along with eight guidelines that separate mandatory requirements from recommended practices to address reproducibility challenges.
-
Instructions Shape Production of Language, not Processing
Instructions trigger a production-centered mechanism in language models, with task-specific information stable in input tokens but varying strongly in output tokens and correlating with behavior.
-
Guidelines for Empirical Studies in Software Engineering involving Large Language Models
A group of 22 researchers proposes seven study types and eight guidelines for empirical software engineering studies involving LLMs to enhance reproducibility and replicability.
-
Instructions Shape Production of Language, not Processing
Instructions primarily shape the production stage of language models rather than the processing stage, with task-specific information and causal effects stronger in output tokens than input tokens.
-
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
Generative AI evaluation should prioritize real-world utility through stakeholder goals and longitudinal human outcome measurements instead of static benchmark performance.
-
Mixed response geometry and critical crossover in the Ising model
In the 2D Ising model, the mixed response Ω_βh = −N cov(m,e) forms a localized ridge from criticality into finite-field crossover and collapses onto a susceptibility-constrained curve in normalized coordinates.
-
Mixed response geometry and critical crossover in the Ising model
Thermodynamic curvature on the (β, h) manifold in the Ising model produces a ridge that geometrically identifies the Widom line as the locus of maximal response extending from the critical point.
-
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
Generative AI evaluation must shift from static benchmark scores to measuring sustained improvements in human capabilities within specific deployment contexts.
-
Personalized AI Practice Replicates Learning Rate Regularity at Scale
Large-scale data from an AI platform confirms students have consistent learning rates (IQR 7.01-8.25 opportunities to 80% mastery) despite variable starting knowledge, replicating prior findings with automated knowled...
-
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
A five-model panel ranked each other's anonymized responses, and the averaged ranks show which chatbots other chatbots prefer.
-
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
A survey that organizes LLMs-as-judges research into functionality, methodology, applications, meta-evaluation, and limitations.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.