Pith. sign in

REVIEW 7 cited by

Can We Use Large Language Models to Fill Relevance Judgment Holes?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05600 v1 pith:AO76HISH submitted 2024-05-09 cs.IR cs.CL

classification cs.IRcs.CL
keywords holesjudgmentshighlyhumanmodelstestannotationsautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Incomplete relevance judgments limit the re-usability of test collections. When new systems are compared against previous systems used to build the pool of judged documents, they often do so at a disadvantage due to the ``holes'' in test collection (i.e., pockets of un-assessed documents returned by the new system). In this paper, we take initial steps towards extending existing test collections by employing Large Language Models (LLM) to fill the holes by leveraging and grounding the method using existing human judgments. We explore this problem in the context of Conversational Search using TREC iKAT, where information needs are highly dynamic and the responses (and, the results retrieved) are much more varied (leaving bigger holes). While previous work has shown that automatic judgments from LLMs result in highly correlated rankings, we find substantially lower correlates when human plus automatic judgments are used (regardless of LLM, one/two/few shot, or fine-tuned). We further find that, depending on the LLM employed, new runs will be highly favored (or penalized), and this effect is magnified proportionally to the size of the holes. Instead, one should generate the LLM annotations on the whole document pool to achieve more consistent rankings with human-generated labels. Future work is required to prompt engineering and fine-tuning LLMs to reflect and represent the human annotations, in order to ground and align the models, such that they are more fit for purpose.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Criteria-Based LLM Relevance Judgments

    cs.IR 2025-07 conditional novelty 6.0 of 10

    A multi-criteria prompting framework, where an LLM grades exactness, coverage, topicality, and contextual fit separately and then aggregates them, produces relevance judgments whose system rankings closely match human...

  2. Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

    cs.IR 2026-08 conditional novelty 5.0 of 10

    A production VLM-based relevance-labeling pipeline at Pinterest search produces human-aligned sDCG@K metrics and about a 6× smaller minimum detectable effect in A/B tests.

  3. Leveraging LLMs to Evaluate Usefulness of Document

    cs.IR 2025-06 conditional novelty 5.0 of 10

    A cascade of LLM judges, fed with search context and behavior, produces multilevel usefulness labels for clicked documents and improves search satisfaction prediction.

  4. Measuring Hypothesis Testing Errors in the Evaluation of Retrieval Systems

    cs.IR 2025-07 conditional novelty 4.0 of 10

    The paper adds Type II error metrics to the evaluation of relevance judgment sets and shows that balanced accuracy and Matthews correlation can summarize qrels' discriminative power in one number.

  5. When LLMs Disagree: Diagnosing Relevance Filtering Bias and Retrieval Divergence in SDG Search

    cs.IR 2025-07 conditional novelty 4.0 of 10

    Two LLMs disagree on about 16% of SDG relevance labels, and the disagreement is lexically systematic and changes top-20 retrieval results.

  6. Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications

    cs.IR 2025-07 reject novelty 4.0 of 10

    LLMs alone annotate search clarifications unreliably; adding confidence-based selective human review cuts effort 24-45% in simulation, but the evaluation is partly built from the ground truth it predicts.

  7. Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement

    cs.IR 2025-06 conditional novelty 4.0 of 10

    LLM-based three-class gender labeling agrees with human annotations better than the lexical NFaiRR score, and the proposed CWEx metric combines neutral exposure with male-female exposure disparity for ranking fairness...

Pith tools