Pith. sign in

Shaver, and Daphna Pereg

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Rater State Bias in RLHF Preference Data: An Audit Framework

cs.AI · 2026-04-14 · conditional · novelty 6.0

The paper hypothesizes that stressed RLHF raters systematically prefer emotionally validating responses, and proposes an audit framework with five falsifiable predictions to detect this bias in public models.

citing papers explorer

Showing 1 of 1 citing paper.

  • Rater State Bias in RLHF Preference Data: An Audit Framework cs.AI · 2026-04-14 · conditional · none · ref 29

    The paper hypothesizes that stressed RLHF raters systematically prefer emotionally validating responses, and proposes an audit framework with five falsifiable predictions to detect this bias in public models.