A personalized calibration network that combines LLM answers to multiple rubric questions predicted human judges' overall satisfaction scores on dialogues about twice as accurately as the uncalibrated LLM.
In Proceedings of the 45th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval , SIGIR ’22, page 3360–3362, New York, NY , USA
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
A personalized calibration network that combines LLM answers to multiple rubric questions predicted human judges' overall satisfaction scores on dialogues about twice as accurately as the uncalibrated LLM.