Pith. sign in

REVIEW 4 cited by

RAGGED: Towards Informed Design of Scalable and Stable RAG Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.09040 v3 pith:I62S7CFC submitted 2024-03-14 cs.CL

classification cs.CL
keywords retrievalgenerationraggedsystemsdatasetsdegradedepthframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Retrieval-augmented generation (RAG) enhances language models by integrating external knowledge, but its effectiveness is highly dependent on system configuration. Improper retrieval settings can degrade performance, making RAG less reliable than closed-book generation. In this work, we introduce RAGGED, a framework for systematically evaluating RAG systems across diverse retriever-reader configurations, retrieval depths, and datasets. Our analysis reveals that reader robustness to noise is the key determinant of RAG stability and scalability. Some readers benefit from increased retrieval depth, while others degrade due to their sensitivity to distracting content. Through large-scale experiments on open-domain, multi-hop, and specialized-domain datasets, we show that retrievers, rerankers, and prompts influence performance but do not fundamentally alter these reader-driven trends. By providing a principled framework and new metrics to assess RAG stability and scalability, RAGGED enables systematic evaluation of retrieval-augmented generation systems, guiding future research on optimizing retrieval depth and model robustness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation

    cs.LG 2026-07 accept novelty 6.0 of 10

    A render-matched control explains away almost all (+0.159 of +0.184) of a fine-grained revision-ledger's apparent advantage over a flat baseline, leaving a near-zero mechanism residual and making coarse invalidation t...

  2. Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A configurable benchmark with four retrieval-noise types shows RAG accuracy drops sharply beyond 50% noise and that noise type, not just quantity, determines failure patterns.

  3. Investigating the Robustness of Retrieval-Augmented Generation at the Query Level

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Retrieval-augmented generation performance drops noticeably under minor query perturbations, with end-to-end results often tracking retriever behavior.

  4. SemRAG: Semantic Knowledge-Augmented RAG for Improved Question-Answering

    cs.CL 2025-07 reject novelty 4.0 of 10

    SemRAG combines semantic chunking with knowledge graph community retrieval and reports mixed, partly contradictory gains over naive RAG on MultiHopRAG and Wikipedia QA datasets.

Pith tools