Pith. sign in

REVIEW 14 cited by

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.06709 v1 pith:YLAUA24J submitted 2025-08-08 cs.CL cs.AI

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

classification cs.CL cs.AI
keywords modelsself-biasoutputsevaluationsjudgesmethodothersystematically
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings to their own outputs, a phenomenon known as self-bias, which can distort evaluations of true model performance. Previous studies often conflate genuine differences in model quality with bias or incorrectly assume that evaluations from LLMs and humans follow the same rating distributions. In this work, we present a statistical framework that explicitly formalizes assumptions under which self-bias can be identified and estimated. Our method models the difference in the scoring distribution that LLM-as-a-judge assigns to its own completions compared to other models, while accounting for the underlying quality of the completions provided by an independent, third-party judge (e.g., humans). Our method reliably isolates and quantifies self-bias, even when models vary in ability, ensuring that genuine performance differences are not mistaken for self-bias. We conduct an empirical analysis of self-bias on a large dataset (>5000 prompt-completion pairs) consisting of expert human annotations and judgments from nine different LLM judges. We find that some models, such as GPT-4o and Claude 3.5 Sonnet, systematically assign higher scores to their own outputs. These models also display family-bias; systematically assigning higher ratings to outputs produced by other models of the same family. Our findings highlight potential pitfalls of using LLM judges and offer practical guidance to mitigate biases when interpreting automated evaluations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

    cs.CL 2026-07 conditional novelty 7.0

    A 209-task Chinese benchmark across six dialogue-failure modes shows that even the top frontier model fully satisfies only 41.1% of long multi-turn requests.

  2. Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill

    cs.SE 2026-06 conditional novelty 7.0

    Controlled ablation finds Popperian code-generation skill adds no separable correctness benefit over labels-only scaffold; gains track structure not content.

  3. CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

    cs.CL 2026-06 conditional novelty 7.0

    CoEval generates task-specific benchmarks by rotating models through teacher, student, and judge roles, then weights questions by discriminative power and judges by panel consensus to recover accurate model rankings w...

  4. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

    cs.CL 2026-04 conditional novelty 7.0

    LLM judges over-approve their own outputs in rubric-based evaluation, even when rubrics are programmatically verifiable, and ensembling only partially corrects the bias.

  5. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

    cs.CL 2026-04 unverdicted novelty 7.0

    Rubric-based LLM judges show self-preference bias, incorrectly marking their own failed outputs as satisfied up to 50% more often on verifiable benchmarks and skewing scores by 10 points on subjective ones.

  6. When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

    cs.AI 2025-10 unverdicted novelty 7.0

    Anonymization in multi-agent debate reduces identity bias by equalizing self and peer weights in a Bayesian update model, quantified by the Identity Bias Coefficient.

  7. Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG

    cs.CL 2026-07 conditional novelty 6.5

    Answer-paired analysis of a 3×3 GPT/Grok/Gemini judge matrix finds near-zero same-model recall bias for induced RAG grounding errors; remaining flag gaps reflect label-task mismatch.

  8. Why Do Safety Guardrails Degrade Across Languages?

    cs.CL 2026-05 conditional novelty 6.0

    A latent variable IRT framework decouples four safety-driving factors across 61 model configurations and 10 languages using 1.9 million evaluations, revealing that safety is largely unidimensional and that high cross-...

  9. Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

    cs.CL 2026-04 conditional novelty 6.0

    Self-preference bias persists in rubric-based LLM evaluation even with fully objective, programmatically verifiable rubrics, and can shift subjective medical-chat scores by up to ~10 points.

  10. Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge

    cs.AI 2026-04 unverdicted novelty 6.0

    Both humans and LLMs trust content more when labeled human-authored than AI-generated, with LLMs showing denser attention to labels and higher uncertainty under AI labels, mirroring human heuristic patterns.

  11. Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

    cs.CL 2026-01 conditional novelty 6.0

    Using an output-matched proxy baseline across 16 models on 9 datasets, 89.6% of measured self-preference bias disappears and roughly half of prior findings lose significance.

  12. Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety

    cs.CL 2025-12 unverdicted novelty 6.0

    Distilling safe refusal behavior from OpenAI o1-mini into Llama-3, Gemma-2, and Qwen3 models via response-based LoRA on multilingual jailbreak data increases jailbreak success rates on MultiJail by up to 16.6 points.

  13. Extreme Self-Preference in Language Models

    cs.AI 2025-09 unverdicted novelty 6.0

    Eight LLMs exhibited massive self-preference that followed assigned identities rather than true ones, appearing in both simple word tasks and consequential evaluations of job candidates and AI technologies.

  14. RedNote-Vibe: A Dataset for Capturing Temporal Dynamics of AI-Generated Text in Lifestyle Social Media

    cs.CL 2025-09 unverdicted novelty 6.0

    RedNote-Vibe supplies a longitudinal dataset of AI versus human lifestyle posts from 2020 to mid-2025 plus the PLAD detection framework that applies cognitive psychology signatures for improved AI-text identification.