REVIEW 8 cited by
Can We Use Large Language Models to Fill Relevance Judgment Holes?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Incomplete relevance judgments limit the re-usability of test collections. When new systems are compared against previous systems used to build the pool of judged documents, they often do so at a disadvantage due to the ``holes'' in test collection (i.e., pockets of un-assessed documents returned by the new system). In this paper, we take initial steps towards extending existing test collections by employing Large Language Models (LLM) to fill the holes by leveraging and grounding the method using existing human judgments. We explore this problem in the context of Conversational Search using TREC iKAT, where information needs are highly dynamic and the responses (and, the results retrieved) are much more varied (leaving bigger holes). While previous work has shown that automatic judgments from LLMs result in highly correlated rankings, we find substantially lower correlates when human plus automatic judgments are used (regardless of LLM, one/two/few shot, or fine-tuned). We further find that, depending on the LLM employed, new runs will be highly favored (or penalized), and this effect is magnified proportionally to the size of the holes. Instead, one should generate the LLM annotations on the whole document pool to achieve more consistent rankings with human-generated labels. Future work is required to prompt engineering and fine-tuning LLMs to reflect and represent the human annotations, in order to ground and align the models, such that they are more fit for purpose.
Forward citations
Cited by 8 Pith papers
-
Criteria-Based LLM Relevance Judgments
A multi-criteria prompting framework, where an LLM grades exactness, coverage, topicality, and contextual fit separately and then aggregates them, produces relevance judgments whose system rankings closely match human...
-
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
LLM relevance judgments align with physician trainees on only about 45 to 66 percent of sentences, and pruning contexts to physician-labeled relevant sentences improves LLM and trainee accuracy.
-
Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search
A production VLM-based relevance-labeling pipeline at Pinterest search produces human-aligned sDCG@K metrics and about a 6× smaller minimum detectable effect in A/B tests.
-
Leveraging LLMs to Evaluate Usefulness of Document
A cascade of LLM judges, fed with search context and behavior, produces multilevel usefulness labels for clicked documents and improves search satisfaction prediction.
-
Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems
The paper adds Type II error metrics to the evaluation of relevance judgment sets and shows that balanced accuracy and Matthews correlation can summarize qrels' discriminative power in one number.
-
When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search
Two LLMs disagree on about 16% of SDG relevance labels, and the disagreement is lexically systematic and changes top-20 retrieval results.
-
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.
-
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
LLM-based three-class gender labeling agrees with human annotations better than the lexical NFaiRR score, and the proposed CWEx metric combines neutral exposure with male-female exposure disparity for ranking fairness...
Discussion (0). Sign in to comment.