LLMs self-correct toward uniform answers in multi-turn repetition, and the gap between single-turn and multi-turn answer rates (B-score) flags biased answers better than verbalized confidence.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
B-score: Detecting biases in large language models using response history
LLMs self-correct toward uniform answers in multi-turn repetition, and the gap between single-turn and multi-turn answer rates (B-score) flags biased answers better than verbalized confidence.