Nugget-based fact-recall scores on Search Arena battles correlate with human preferences, with about 54% agreement, and offer finer-grained diagnostic signals than raw win/loss votes.
Pencils down! automatic rubric-based evaluation of re- trieve/generate systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses
Nugget-based fact-recall scores on Search Arena battles correlate with human preferences, with about 54% agreement, and offer finer-grained diagnostic signals than raw win/loss votes.