Pith. sign in

hub

Large language models are inconsistent and biased evaluators

21 Pith papers cite this work, alongside 8 external citations. Polarity classification is still indexing.

21 Pith papers citing it
8 external citations · Pith
abstract

The zero-shot capability of Large Language Models (LLMs) has enabled highly flexible, reference-free metrics for various tasks, making LLM evaluators common tools in NLP. However, the robustness of these LLM evaluators remains relatively understudied; existing work mainly pursued optimal performance in terms of correlating LLM scores with human expert scores. In this paper, we conduct a series of analyses using the SummEval dataset and confirm that LLMs are biased evaluators as they: (1) exhibit familiarity bias-a preference for text with lower perplexity, (2) show skewed and biased distributions of ratings, and (3) experience anchoring effects for multi-attribute judgments. We also found that LLMs are inconsistent evaluators, showing low "inter-sample" agreement and sensitivity to prompt differences that are insignificant to human understanding of text quality. Furthermore, we share recipes for configuring LLM evaluators to mitigate these limitations. Experimental results on the RoSE dataset demonstrate improvements over the state-of-the-art LLM evaluators.

hub tools

citation-role summary

background 2

citation-polarity summary

roles

background 2

polarities

background 2

representative citing papers

GRASP: Deterministic argument ranking in interaction graphs

cs.LG · 2026-05-18 · unverdicted · novelty 7.0

GRASP aggregates stable local LLM interaction judgments into global argument rankings via a convergent attack-defense propagation operator on interaction graphs, yielding higher reproducibility than holistic judging and no correlation with human convincingness.

ProactBench: Beyond What The User Asked For

cs.LG · 2026-05-09 · unverdicted · novelty 7.0

ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.

Auditing Stance Asymmetry in Generative Explanations

cs.CL · 2026-05-27 · unverdicted · novelty 6.0

Introduces Symmetry Decomposition Evaluation (SDE) to audit stable stance asymmetries in generative explanations using paired situations, role rewrites, and evidence controls on a 32-family prototype suite.

Self-Preference Bias in LLM-as-a-Judge

cs.CL · 2024-10-29 · unverdicted · novelty 6.0

LLMs judge their own outputs higher because they assign better scores to lower-perplexity text, even when the text is not self-generated.

On Cost-Effective LLM-as-a-Judge Improvement Techniques

cs.CL · 2026-04-15 · conditional · novelty 5.0

Ensemble scoring plus task-specific criteria injection lifts LLM judges to 85.8% on RewardBench 2 (+13.5pp), dominating calibration and escalation on the cost–accuracy frontier.

citing papers explorer

Showing 21 of 21 citing papers.