2403.00811 , primaryclass =

Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, Zexue He · 2024 · arXiv 2403.00811

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

representative citing papers

Mitigating Cognitive Bias in RLHF by Altering Rationality

cs.AI · 2026-05-07 · unverdicted · novelty 6.0

Dynamically adjusting beta via LLM-as-judge downweights biased comparisons to learn more rational reward models from flawed human preferences.

Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs

cs.AI · 2026-05-11 · unverdicted · novelty 5.0

Numeric anchors embedded in images systematically bias VLM quality judgments more than severe visual degradation, with layer-wise probing showing that anchor-saturated layers are suboptimal for quality prediction.

Inertia in Moral and Value Judgments of Large Language Models

cs.CL · 2024-08-16 · unverdicted · novelty 4.0

LLMs exhibit persistent inertia in value orientations, with harm avoidance and fairness remaining skewed across persona prompts.

Bias in Large Language Models: Origin, Evaluation, and Mitigation

cs.CL · 2024-11-16 · unverdicted · novelty 2.0

A literature review that categorizes bias in LLMs, surveys evaluation and mitigation techniques, and discusses ethical implications.

AMEL: Accumulated Message Effects on LLM Judgments

cs.AI · 2026-05-21

citing papers explorer

Showing 5 of 5 citing papers.

Mitigating Cognitive Bias in RLHF by Altering Rationality cs.AI · 2026-05-07 · unverdicted · none · ref 6
Dynamically adjusting beta via LLM-as-judge downweights biased comparisons to learn more rational reward models from flawed human preferences.
Don't Look at the Numbers: Visual Anchoring Bias and Layer-wise Representation in VLMs cs.AI · 2026-05-11 · unverdicted · none · ref 3
Numeric anchors embedded in images systematically bias VLM quality judgments more than severe visual degradation, with layer-wise probing showing that anchor-saturated layers are suboptimal for quality prediction.
Inertia in Moral and Value Judgments of Large Language Models cs.CL · 2024-08-16 · unverdicted · none · ref 17
LLMs exhibit persistent inertia in value orientations, with harm avoidance and fairness remaining skewed across persona prompts.
Bias in Large Language Models: Origin, Evaluation, and Mitigation cs.CL · 2024-11-16 · unverdicted · none · ref 25
A literature review that categorizes bias in LLMs, surveys evaluation and mitigation techniques, and discusses ethical implications.
AMEL: Accumulated Message Effects on LLM Judgments cs.AI · 2026-05-21 · unreviewed · ref 10

2403.00811 , primaryclass =

fields

years

verdicts

representative citing papers

citing papers explorer