A Gaussian information-gain metric in embedding space quantifies semantic progress in dialogues via uncertainty reduction and shows competitive agreement with human judgments on MT-Bench and UltraFeedback.
A com- prehensive analysis of the effectiveness of large language models as automatic dialogue eval- uators
2 Pith papers cite this work, alongside 11 external citations. Polarity classification is still indexing.
2
Pith papers citing it
11
external citations · OpenAlex
years
2026 2representative citing papers
Bayesian deep learning method rankings are unreliable under data scarcity, reversing across datasets and sample sizes, and a hierarchical Bayesian framework with predictive detectability curves is needed to assess evaluation sufficiency.
citing papers explorer
-
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation
Bayesian deep learning method rankings are unreliable under data scarcity, reversing across datasets and sample sizes, and a hierarchical Bayesian framework with predictive detectability curves is needed to assess evaluation sufficiency.