Pith. sign in

REVIEW 1 cited by

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09269 v1 pith:3D2XJQS2 submitted 2024-12-12 cs.CL cs.AI

Towards Understanding the Robustness of LLM-based Evaluations under Perturbations

classification cs.CL cs.AI
keywords evaluatorsllmsmetricswhenexplorehumanperturbationsrobustness
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Traditional evaluation metrics like BLEU and ROUGE fall short when capturing the nuanced qualities of generated text, particularly when there is no single ground truth. In this paper, we explore the potential of Large Language Models (LLMs), specifically Google Gemini 1, to serve as automatic evaluators for non-standardized metrics in summarization and dialog-based tasks. We conduct experiments across multiple prompting strategies to examine how LLMs fare as quality evaluators when compared with human judgments on the SummEval and USR datasets, asking the model to generate both a score as well as a justification for the score. Furthermore, we explore the robustness of the LLM evaluator by using perturbed inputs. Our findings suggest that while LLMs show promise, their alignment with human evaluators is limited, they are not robust against perturbations and significant improvements are required for their standalone use as reliable evaluators for subjective metrics.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Learning When to Trust in Contextual Social Bandits

    cs.AI 2026-03 conditional novelty 7.0

    Sparse audits suffice to learn per-evaluator contextual trust boundaries that break sycophantic majorities, yielding sublinear latent regret matching the information-theoretic necessity of audits.