BayesBench evaluates LLMs on multi-turn Bayesian estimation, prediction, and persona-framed tasks, finding scaling aids latent inference but not reliably downstream prediction.
who, what, when, where
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
BayesBench evaluates LLMs on multi-turn Bayesian estimation, prediction, and persona-framed tasks, finding scaling aids latent inference but not reliably downstream prediction.