Pith. sign in

REVIEW 3 cited by

JuStRank: Benchmarking LLM Judges for System Ranking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09569 v2 pith:DPBKEQ7C submitted 2024-12-12 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords judgesystemjudgesrankingassessmentbiasfirstquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based judges a compelling solution for this challenge. Crucially, this approach requires first to validate the quality of the LLM judge itself. Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems. We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems. To address this gap, we conduct the first large-scale study of LLM judges as system rankers. System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking. Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    cs.CL 2026-05 conditional novelty 6.0 of 10

    On binary verdicts, Pearson, Spearman, Kendall's tau-b, phi, and the Matthews correlation are a single statistic, so most multi-metric agreement reports repeat one number under different names.

  2. CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

    cs.CL 2025-07 conditional novelty 6.0 of 10

    CLEAR converts per-instance LLM judge critiques into system-level error issues with prevalence counts and an interactive dashboard for exploration.

  3. Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Across five LLMs and ten no-consensus datasets, neutrality drops sharply when models act as pairwise judges, pointwise judges, or debaters compared to when they generate answers with an explicit neutral option.

Pith tools