Relevance Context Learning generates explicit relevance narratives from judged examples to guide LLM assessors, outperforming zero-shot and standard in-context learning for IR relevance judgments.
In: Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.IR 10years
2026 10roles
background 3polarities
background 3representative citing papers
Scoring functions are sub-optimal for all utility-fairness trade-offs in ranking under a generic fairness formulation, but semi-greedy post-processing can approach the performance of exhaustive post-processing.
MIRA is a new benchmark for multi-category integrated retrieval built from real queries on a social science platform, with LLM assistance for topic descriptions and relevance labeling across four item categories.
Synthetically formalizing information needs into topics with descriptions and narratives improves LLM relevance assessor agreement with humans and reduces over-labeling of relevant documents on TREC Deep Learning and Robust04.
LLMs consistently overrate relevance of inadequate passages in IR evaluations due to biases toward length and lexical features rather than true content match.
ChunkGroupSHAP clusters semantically related chunks into shared cross-document features for listwise Shapley explanations of embedding-based rankings.
Adaptive Re-Ranking trains a classifier to route queries to BM25, MiniLM-L6-v2, or BGE-v2-m3 based on a utility label, yielding 1.15-53x lower median latency and competitive nDCG@10 versus always using the heaviest model.
LLM-generated reference documents serve as relevance pivots for dynamic ranked-list truncation and adaptive/parallel listwise reranking, reportedly beating prior RLT methods and cutting LLM reranking cost by up to 66%.
LLMs judge document relevance at a level comparable to humans but frequently highlight different passages, indicating they are often not right for the right reasons and cannot fully replace human assessors.
The authors propose an evaluation framework for LLM-generated structured search summaries and describe plans for implementing and testing it.
citing papers explorer
-
Hybrid Pooling with LLMs via Relevance Context Learning
Relevance Context Learning generates explicit relevance narratives from judged examples to guide LLM assessors, outperforming zero-shot and standard in-context learning for IR relevance judgments.
-
Scoring Is Not Enough: Addressing Gaps in Utility-fairness Trade-offs for Ranking
Scoring functions are sub-optimal for all utility-fairness trade-offs in ranking under a generic fairness formulation, but semi-greedy post-processing can approach the performance of exhaustive post-processing.
-
MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval
MIRA is a new benchmark for multi-category integrated retrieval built from real queries on a social science platform, with LLM assistance for topic descriptions and relevance labeling across four item categories.
-
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
Synthetically formalizing information needs into topics with descriptions and narratives improves LLM relevance assessor agreement with humans and reduces over-labeling of relevant documents on TREC Deep Learning and Robust04.
-
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
LLMs consistently overrate relevance of inadequate passages in IR evaluations due to biases toward length and lexical features rather than true content match.
-
Listwise Explanation of Embedding-Based Rankings via Semantic Chunk Grouping
ChunkGroupSHAP clusters semantically related chunks into shared cross-document features for listwise Shapley explanations of embedding-based rankings.
-
Adaptive Re-Ranking
Adaptive Re-Ranking trains a classifier to route queries to BM25, MiniLM-L6-v2, or BGE-v2-m3 based on a utility label, yielding 1.15-53x lower median latency and competitive nDCG@10 versus always using the heaviest model.
-
Dynamic Ranked List Truncation for Reranking Pipelines via LLM-generated Reference-Documents
LLM-generated reference documents serve as relevance pivots for dynamic ranked-list truncation and adaptive/parallel listwise reranking, reportedly beating prior RLT methods and cutting LLM reranking cost by up to 66%.
-
LLMs as Assessors: Right for the Right Reason?
LLMs judge document relevance at a level comparable to humans but frequently highlight different passages, indicating they are often not right for the right reasons and cannot fully replace human assessors.
-
Plans for Evaluating Structured Generative Search Summaries
The authors propose an evaluation framework for LLM-generated structured search summaries and describe plans for implementing and testing it.