LLM-generated relevance judgments misrank top retrieval systems and produce many false significant differences relative to human judgments.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation
LLM-generated relevance judgments misrank top retrieval systems and produce many false significant differences relative to human judgments.