REVIEW 4 cited by
Evaluating the Retrieval Robustness of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Evaluating the Retrieval Robustness of Large Language Models
read the original abstract
Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG may also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness). We focus on three research questions: (1) whether RAG is always better than non-RAG; (2) whether more retrieved documents always lead to better performance; (3) and whether document orders impact results. To facilitate this study, we establish a benchmark of 1500 open-domain questions, each with retrieved documents from Wikipedia. We introduce three robustness metrics, each corresponds to one research question. Our comprehensive experiments, involving 11 LLMs and 3 prompting strategies, reveal that all of these LLMs exhibit surprisingly high retrieval robustness; nonetheless, different degrees of imperfect robustness hinders them from fully utilizing the benefits of RAG.
Forward citations
Cited by 4 Pith papers
-
The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise
The Power-of-Noise effect in RAG is an artifact of restrictive prompting and decoding, not a general benefit of random documents.
-
The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise
The Power-of-Noise effect in RAG is reproducible only under the original constrained setup and disappears or weakens once instruction templates, longer outputs, and modern LLMs are used.
-
Lost in the Evidence? Reproducing Document Position and Context Size Effects in RAG
Reproducibility study shows position and context size effects in RAG depend on topic sampling and retrieval quality, proposes calibration for stable trends, and releases code after finding discrepancies with prior ind...
-
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering
A RAG pipeline using BM25 retrieval and GPT-o3 generation achieves 91% Precision@1 and 85% truthfulness for Nepali legal question answering on a curated 100-query benchmark.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.