Pith. sign in

REVIEW 26 cited by

Verbosity Bias in Preference Labeling by Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10076 v1 pith:UZBW7OHE submitted 2023-10-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsbiasfeedbacklanguagelearninganswersgpt-4human
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, Large Language Models (LLMs) have witnessed a remarkable surge in prevalence, altering the landscape of natural language processing and machine learning. One key factor in improving the performance of LLMs is alignment with humans achieved with Reinforcement Learning from Human Feedback (RLHF), as for many LLMs such as GPT-4, Bard, etc. In addition, recent studies are investigating the replacement of human feedback with feedback from other LLMs named Reinforcement Learning from AI Feedback (RLAIF). We examine the biases that come along with evaluating LLMs with other LLMs and take a closer look into verbosity bias -- a bias where LLMs sometimes prefer more verbose answers even if they have similar qualities. We see that in our problem setting, GPT-4 prefers longer answers more than humans. We also propose a metric to measure this bias.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Does Multi-Agent Debate Improve AI Feedback on Research Papers?

    econ.GN 2026-07 accept novelty 7.0 of 10

    Authors of economics meta-analyses found a single-pass AI report more useful than two multi-agent debate tools that cost up to thirty times more to run.

  2. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

    cs.CL 2025-07 conditional novelty 7.0 of 10

    With prompt engineering (audio concatenation plus in-context examples), large audio models rank speech synthesis systems in line with human preferences, reaching up to 0.91 Spearman correlation.

  3. JuStRank: Benchmarking LLM Judges for System Ranking

    cs.CL 2024-12 conditional novelty 7.0 of 10

    JuStRank ranks AI judges by how well their aggregated scores reproduce the Chatbot Arena human system ranking, revealing that judge realization and bias, not just model size, determine ranking quality.

  4. Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

    cs.CL 2026-08 conditional novelty 6.0 of 10

    CUE-Bench provides 51,823 Chinese discourse instances annotated with a nine-way Affective Stance defined by explicit-implicit polarity, plus pragmatic intent and fine-grained emotion labels.

  5. Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Evidence-grounded persona panels with bounded-confidence deliberation raise GenUI judge–human correlation from 0.716 to 0.922, mostly from persona grounding rather than multi-prompt averaging.

  6. Test-Time Scaling via Error Localization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    TTEL uses feedback-induced token probability drops to localize the first error in a failed reasoning trace and branch a new generation from that prefix, improving pass@k per token on coding and math benchmarks.

  7. When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking

    cs.LG 2026-07 unverdicted novelty 6.0 of 10

    A bias-aware Bradley-Terry judge cannot identify the quality/bias split from comparisons alone; only prior assumptions, trusted anchors, or paired renderings can supply it.

  8. Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A cluster-conditioned Bayesian prompt ensemble improves the calibration and accuracy of multimodal LLM judges for text-to-image preference evaluation.

  9. Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A regression controlling for human-rated quality detects positive self- and family-bias in several LLM judges, including GPT-4o and Claude 3.5 Sonnet.

  10. Evaluating the Use of LLMs for Documentation to Code Traceability

    cs.SE 2025-06 conditional novelty 6.0 of 10

    LLMs identify documentation-to-code trace links with F1 up to 80.4%, outperforming TF-IDF, BM25, and CodeBERT, but their explanations and chain reconstructions need human oversight.

  11. Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

    cs.CL 2025-04 conditional novelty 6.0 of 10

    MPO uses a meta reward model to continuously rewrite the reward model's evaluation prompt during PPO training, and the resulting models beat static-prompt RLAIF baselines on four tasks.

  12. Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Pairwise LLM judgments flip in about 35% of cases when a stylistic distractor is added, versus only 9% for absolute scores, so pointwise scoring is more robust to manipulation.

  13. Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

    cs.SE 2025-02 conditional novelty 6.0 of 10

    Output-based LLM-as-a-judge methods using large LLMs achieve near-human correlation with human scores for code translation and generation, but not for code summarization or pairwise comparisons.

  14. A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A rubric-plus-question-answering LLM framework for literary translation evaluation beats traditional MT metrics but still trails human agreement, especially on Korean honorifics.

  15. Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals

    cs.HC 2024-11 accept novelty 6.0 of 10

    On-demand, user-activated AI audio descriptions give blind and low vision viewers control over timing and detail of video descriptions, but they increase cognitive load and are preferred more for instructional than en...

  16. Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A conceptual framework classifies human feedback to RL agents along nine dimensions and seven quality criteria, unifying human-centered, interface-centered, and model-centered design perspectives.

  17. When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Judge upgrades are not interchangeable: only Qwen3 1.7B→4B yields robust adjacent gains, MiniMax adjacent releases do not, and stronger judges reduce but do not remove bias or correlated jury errors.

  18. SCOPE: Selective Conformal Optimized Pairwise LLM Judging

    cs.CL 2026-02 conditional novelty 5.0 of 10

    A conformal calibration method (SCOPE) plus a bidirectional entropy score (BPE) lets LLM pairwise judges abstain selectively while keeping accepted-set error below a user-specified bound.

  19. Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

    cs.LG 2026-01 conditional novelty 5.0 of 10

    The optimal reward for KL-regularized LLM alignment is a threshold function—reward B above a prompt-dependent cutoff, 0 below—which can be estimated from base-model samples and integrated into decoding-time alignment.

  20. How Reliable is Multilingual LLM-as-a-Judge?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLM-as-a-Judge is inconsistent across languages: five models grading the same parallel ground-truth content in 25 languages agreed poorly (Fleiss' Kappa around 0.2 on average), and the proposed majority-vote ensemble ...

  21. From First Draft to Final Insight: A Multi-Agent Approach for Feedback Generation

    cs.HC 2025-05 reject novelty 5.0 of 10

    A generate-evaluate-regenerate loop with GPT-4o raised rubric scores for feedback on 208 quiz responses, but the second-round evaluation was done by the same model that rewrote the feedback, and the abstract misreport...

  22. Engineering AI Judge Systems

    cs.SE 2024-11 conditional novelty 5.0 of 10

    A constitution-based, search-driven development framework for AI judge systems improves judged accuracy by up to 6.2% on commit message generation, with about 58% of general principles reused across five languages.

  23. Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition

    cs.AI 2026-01 conditional novelty 4.0 of 10

    One engineer paired with ChatGPT and reached second place in ESA's ELOPE competition in about one week of work; the paper draws best-practice lessons from that experience.

  24. Exploring Modularity of Agentic Systems for Drug Discovery

    cs.LG 2025-06 conditional novelty 4.0 of 10

    On 26 chemistry questions, swapping the LLM, agent type, or prompt in an LLM agent changes its scores so much that the system cannot be treated as modular.

  25. Potential and Perils of Large Language Models as Judges of Unstructured Textual Data

    cs.CL 2025-01 conditional novelty 4.0 of 10

    LLM judges show only fair-to-moderate agreement with human raters on thematic summary alignment and consistently over-rate alignment compared to humans.

  26. A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A review of 58 papers finds that large language models pass some Theory of Mind tests but remain brittle and fall short of human performance.

Pith tools