REVIEW 18 cited by
Benchmarking Cognitive Biases in Large Language Models as Evaluators
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Language Models are cognitively biased judges. Large Language Models (LLMs) have recently been shown to be effective as automatic evaluators with simple prompting and in-context learning. In this work, we assemble 15 LLMs of four different size ranges and evaluate their output responses by preference ranking from the other LLMs as evaluators, such as System Star is better than System Square. We then evaluate the quality of ranking outputs introducing the Cognitive Bias Benchmark for LLMs as Evaluators (CoBBLEr), a benchmark to measure six different cognitive biases in LLM evaluation outputs, such as the Egocentric bias where a model prefers to rank its own outputs highly in evaluation. We find that LLMs are biased text quality evaluators, exhibiting strong indications on our bias benchmark (average of 40% of comparisons across all models) within each of their evaluations that question their robustness as evaluators. Furthermore, we examine the correlation between human and machine preferences and calculate the average Rank-Biased Overlap (RBO) score to be 49.6%, indicating that machine preferences are misaligned with humans. According to our findings, LLMs may still be unable to be utilized for automatic annotation aligned with human preferences. Our project page is at: https://minnesotanlp.github.io/cobbler.
Forward citations
Cited by 18 Pith papers
-
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
Self-preference bias persists in rubric-based LLM evaluation even with fully objective, programmatically verifiable rubrics, and can shift subjective medical-chat scores by up to ~10 points.
-
The Leaderboard Illusion
Chatbot Arena's rankings are systematically distorted by undisclosed private testing, selective score reporting, and data access asymmetries that favor large proprietary providers.
-
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
LLM-as-judge scoring biases concentrate in low-dimensional, type-specific activation subspaces that support bidirectional causal steering and cross-domain failure prediction.
-
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
A phase-structured multi-turn red-team framework reports 97.6–100% lenient ASR but only 66.7–78.6% full actionable ASR on six frontier LLMs, with success strongly depth-dependent.
-
What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces
LLM quality judgments are far more sensitive to structure and voice than to vocabulary, according to replication and cross-model degradation tests.
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
A regression controlling for human-rated quality detects positive self- and family-bias in several LLM judges, including GPT-4o and Claude 3.5 Sonnet.
-
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
A dynamic red-teaming audit reports that 94% of MedQA-correct answers fail under adversarial mutation, with 86% privacy leak rates, 81% bias shift rates, and 66-74% hallucination rates across 15 medical LLMs.
-
CRISP: Complex Reasoning with Interpretable Step-based Plans
A small model fine-tuned on CRISP, a filtered dataset of high-level plans, generates plans that improve downstream math and code benchmarks more than few-shot prompting of larger models.
-
AbsenceBench: Language Models Can't Tell What's Missing
LLMs that ace Needle-in-a-Haystack struggle to identify deliberately omitted content, a new benchmark called AbsenceBench shows.
-
Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP
RuleSHAP, which combines global SHAP values with XGBoost and LASSO rule selection, recovers injected univariate, conjunctive, and non-convex LLM behavior triggers more faithfully than RuleFit and SHAP alone.
-
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
V2R-Bench shows that 21 large vision-language models are markedly less accurate on simple object and direction tasks when object position, scale, orientation, or context is varied, and attributes the failure to multim...
-
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
CHIRP, a 104-question pairwise benchmark, exposes scaling trends in vision-language models that existing benchmarks miss, and its evaluations correlate better with training loss.
-
LLM-as-a-Verifier: A General-Purpose Verification Framework
Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.
-
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from lengt...
-
Effectively obtaining acoustic, visual and textual data from videos
A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.
-
Beyond the Surface: Measuring Self-Preference in LLM Judgments
The DBG metric measures LLM self-preference bias as the gap between a judge model's own win rate and the win rate assigned by an ensemble of gold judges.
-
Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
BetaConform estimates LLM ensemble judgment accuracy from few labeled samples via a mixture of Beta-Binomial distributions, conformal-style adaptive stopping, and text-similarity prior transfer, but its theoretical gu...
Discussion (0). Continue with ORCID to comment.