DoGMaTiQ automates creation of QA-based nuggets via document-grounded generation, paraphrase clustering, and quality subselection, producing report evaluations that correlate strongly with human judgments on cross-lingual TREC tasks.
CoRRabs/2411.08275(2024)
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
method 1polarities
use method 1representative citing papers
Relevance Context Learning generates explicit relevance narratives from judged examples to guide LLM assessors, outperforming zero-shot and standard in-context learning for IR relevance judgments.
Synthetically formalizing information needs into topics with descriptions and narratives improves LLM relevance assessor agreement with humans and reduces over-labeling of relevant documents on TREC Deep Learning and Robust04.
LLMs consistently overrate relevance of inadequate passages in IR evaluations due to biases toward length and lexical features rather than true content match.
A prompt perturbation approach builds comparison graphs from LLM judgments, filters inconsistent cycles or ties, and aggregates more reliable rankings.
The authors propose an evaluation framework for LLM-generated structured search summaries and describe plans for implementing and testing it.
citing papers explorer
-
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
DoGMaTiQ automates creation of QA-based nuggets via document-grounded generation, paraphrase clustering, and quality subselection, producing report evaluations that correlate strongly with human judgments on cross-lingual TREC tasks.
-
Hybrid Pooling with LLMs via Relevance Context Learning
Relevance Context Learning generates explicit relevance narratives from judged examples to guide LLM assessors, outperforming zero-shot and standard in-context learning for IR relevance judgments.
-
Formalized Information Needs Improve Large-Language-Model Relevance Judgments
Synthetically formalizing information needs into topics with descriptions and narratives improves LLM relevance assessor agreement with humans and reduces over-labeling of relevant documents on TREC Deep Learning and Robust04.
-
When LLM Judges Inflate Scores: Exploring Overrating in Relevance Assessment
LLMs consistently overrate relevance of inadequate passages in IR evaluations due to biases toward length and lexical features rather than true content match.
-
Prompt Perturbation for Reliable LLM Evaluation over Comparison Graphs
A prompt perturbation approach builds comparison graphs from LLM judgments, filters inconsistent cycles or ties, and aggregates more reliable rankings.
-
Plans for Evaluating Structured Generative Search Summaries
The authors propose an evaluation framework for LLM-generated structured search summaries and describe plans for implementing and testing it.