Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
Simulations show standard CI methods underperform for classifier metrics in small and nested datasets, while Agresti-Coull, Wilson, Clopper-Pearson, and a new pseudo-count regularized bootstrap perform better, with specific adjustments needed for nested structures.