Pith. sign in

REVIEW 3 cited by

Variations in Relevance Judgments and the Shelf Life of Test Collections

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20937 v2 pith:SWXFQ3FD submitted 2025-02-28 cs.IR

classification cs.IR
keywords collectionsrelevancetestjudgmentssystemassessordisagreementhowever
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards neural retrieval models affected the characteristics of modern test collections, e.g., documents are short, judged with four grades of relevance, and information needs have no descriptions or narratives. Under these changes, it is unclear whether assessor disagreement remains negligible for system comparisons. We investigate this aspect under the additional condition that the few modern test collections are heavily re-used. Given more possible query interpretations due to less formalized information needs, an ``expiration date'' for test collections might be needed if top-effectiveness requires overfitting to a single interpretation of relevance. We run a reproducibility study and re-annotate the relevance judgments of the 2019~TREC Deep Learning track. We can reproduce prior work in the neural retrieval setting, showing that assessor disagreement does not affect system rankings. However, we observe that some models substantially degrade with our new relevance judgments, and some have already reached the effectiveness of humans as rankers, providing evidence that test collections can expire.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Robustness and Reliability of Benchmark-Based Evaluation of LLMs

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Paraphrasing benchmark questions keeps LLM rankings stable but reduces their accuracy, suggesting static benchmarks overestimate model robustness.

  2. Disentangling Locality and Entropy in Ranking Distillation

    cs.IR 2025-05 reject novelty 6.0 of 10

    Under ranking distillation, complex hard-negative sampling pipelines yield little or no benefit over BM25 sampling, while intermediate teacher score entropy improves in-domain effectiveness and the paper's generalizat...

  3. Rank-K: Test-Time Reasoning for Listwise Reranking

    cs.IR 2025-05 conditional novelty 6.0 of 10

    Rank-K, a reasoning-model-based listwise reranker distilled from DeepSeek R1 traces, beats RankZephyr on several benchmarks but only marginally on TREC DL 2019/2020.

Pith tools