A risk score built from confidence and vote entropy is proposed for triaging LLM qualitative coding, but the central R-squared=0.979 result is inflated by the definitions of agreement and diversity.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
A risk score built from confidence and vote entropy is proposed for triaging LLM qualitative coding, but the central R-squared=0.979 result is inflated by the definitions of agreement and diversity.