Pith. sign in

REVIEW 4 cited by

Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09607 v1 pith:GO6XJF5M submitted 2024-11-14 cs.IR cs.CL

classification cs.IRcs.CL
keywords evaluationnuggettrecautomaticfullyinitialnuggetsprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report provides an initial look at partial results from the TREC 2024 Retrieval-Augmented Generation (RAG) Track. We have identified RAG evaluation as a barrier to continued progress in information access (and more broadly, natural language processing and artificial intelligence), and it is our hope that we can contribute to tackling the many challenges in this space. The central hypothesis we explore in this work is that the nugget evaluation methodology, originally developed for the TREC Question Answering Track in 2003, provides a solid foundation for evaluating RAG systems. As such, our efforts have focused on "refactoring" this methodology, specifically applying large language models to both automatically create nuggets and to automatically assign nuggets to system answers. We call this the AutoNuggetizer framework. Within the TREC setup, we are able to calibrate our fully automatic process against a manual process whereby nuggets are created by human assessors semi-manually and then assigned manually to system answers. Based on initial results across 21 topics from 45 runs, we observe a strong correlation between scores derived from a fully automatic nugget evaluation and a (mostly) manual nugget evaluation by human assessors. This suggests that our fully automatic evaluation process can be used to guide future iterations of RAG systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAVine: Reality-Aligned Evaluation for Agentic Search

    cs.CL 2025-07 conditional novelty 6.0 of 10

    RAVine is an attributable nugget-based benchmark with process metrics that shows current agentic search models have low citation recall and rely heavily on internal knowledge.

  2. Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations

    cs.IR 2026-06 unverdicted novelty 5.0 of 10

    Presents a three-phase human-in-the-loop nugget annotation workflow and tool for accountable LLM-as-a-judge evaluations of AI outputs.

  3. UiS-IAI@LiveRAG: Retrieval-Augmented Information Nugget-Based Generation of Responses

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A nugget-based RAG pipeline with query rewriting and cluster-based summarization is applied to the LiveRAG challenge, where few rewrites plus the original query improve recall and larger document cutoffs hit diminishi...

  4. A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A research agenda calling for geo-temporal reasoning in deep research systems, with no experiments or system implementation.

Pith tools