Pith. sign in

REVIEW 5 major objections 6 minor 3 cited by

Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers

T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A survey of retrieval-augmented generation organizes the field into four architectural families and shows which family delivers the largest gains on question answering.

desk verdict A solid, incremental RAG survey with mostly reproducible tables; one real transcription error and a weak comparability story, but worth refereeing. read the letter →

arxiv 2506.00054 v1 pith:Q32VU7BR submitted 2025-05-28 cs.IR cs.CL

classification cs.IRcs.CL
keywords Retrieval-AugmentedGenerationarchitecturetaxonomyqueryreformulationcontextfilteringrerankingmulti-hopreasoningrobustnessevaluationbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a survey that tries to establish a dependable map of retrieval-augmented generation (RAG) design space. It proposes a taxonomy of four architectural families, retriever-centric, generator-centric, hybrid, and robustness-oriented, and then uses relative performance gains across short-form and multi-hop question answering to show which design choices tend to pay off. The value, if the survey is right, is that a practitioner can look at the architecture family and get a realistic expectation of how much improvement to expect over a raw LLM or a standard retrieval baseline, before running their own evaluation. The analysis also claims that retrieval optimization is the most consistent driver of large multi-hop gains, while closed-loop verification is what produces robustness.

What carries the argument

The organizing device is a taxonomy that sorts RAG systems by where the core innovation sits: retriever-centric, generator-centric, hybrid, or robustness-oriented. Carrying the empirical argument is a relative-gain normalization that converts raw reported scores (F1, EM, Accuracy, FactScore) into improvements over two baselines, the raw backbone LLM (B) and the same backbone with standard retrieval (B+R), enabling cross-paper comparison in Table 2 and Table 3. The formal foundation is the conditional decomposition $P(y|x) = \sum_{d} P(y|x,d) P(d|x)$, which the taxonomy operationalizes by asking which component the system improves.

What would settle it

Recompute each relative improvement in Table 2 from the appendix's raw scores and check against the original papers; if a cornerstone value such as RQ-RAG's reported ~850% HotpotQA F1 gain fails to reproduce, the comparative ranking of design families collapses.

Watch

Extended reading notes

Core claim

The central claim is that the design space of RAG systems is well captured by four architectural orientations, and that reported performance differences across these families form a coherent pattern: retriever-side innovations such as query decomposition, reranking, and granularity control produce the largest and most consistent relative gains on multi-hop QA; generator-side mechanisms like self-reflection help short-form factual recall; and hybrid systems that add correctness-checking loops deliver the strongest robustness improvements in factual consistency (FactScore), with Self-CRAG reporting +0.456 over its retrieval-augmented baseline on Biography. The survey further claims that retrieval alone is insufficient for robust generation, and that the strongest systems couple retrieval, generation, and verification in iterative loops.

Load-bearing premise

The load-bearing premise is that relative improvements computed from numbers reported in different original papers remain comparable across backbones, prompts, and datasets, even though the evaluation protocols were never standardized.

Editorial extensions

If this is right

  • Retrieval-side investment is the highest-leverage move for multi-hop accuracy: query decomposition and reranking systems (RQ-RAG, LQR, RankRAG) report the largest relative gains over raw backbones.
  • Generator-side mechanisms such as self-reflection and context compression produce solid short-form gains but often lose to standard retrieval baselines when compression is extreme (xRAG's negative gains).
  • Robustness gains require closed-loop verification: critique-based hybrids (Self-CRAG, CRAG, SELF-RAG) lead FactScore improvements, while retrieval without verification (Stochastic RAG) shows near-zero gains.
  • The relative-gain comparison framework lets practitioners transfer expected performance improvements across different backbones and datasets, as long as they accept the normalization assumptions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: a practitioner-facing heuristic follows — for knowledge-intensive multi-hop tasks, improve retrieval first (decomposition, reranking, granularity), and add a self-critique loop when factual consistency matters more than latency.
  • Editorial extension: the relative-gain normalization could be stress-tested by re-evaluating a handful of highlighted systems on one shared benchmark; the rank order of the four design families should persist if the survey's comparative conclusions are genuinely about architecture rather than evaluation luck.
  • Editorial extension: the negative gains of extreme context compression (xRAG) suggest a testable hypothesis that compression trades away evidence the generator needs for multi-hop composition; a controlled sweep of compression ratios across HotpotQA and MuSiQue would quantify that trade-off.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This survey proposes a taxonomy of retrieval-augmented generation (RAG) architectures—retriever-centric, generator-centric, hybrid, and robustness-oriented—and reviews enhancements in retrieval, filtering, efficiency, robustness, and reranking. The paper presents comparative performance analyses of representative RAG frameworks on short-form and multi-hop QA, a robustness comparison, and a review of evaluation benchmarks. The central evidence for the comparative claims is Table 2 (relative gains over raw LLM and LLM+retrieval baselines) and Table 3 (robustness gains), with Appendix Tables 5–7 providing the original reported scores.

Significance. If the comparative analysis were reliable, the survey would provide a useful map of the RAG design space and a starting point for model selection. The proposed taxonomy is reasonable and the appendix's inclusion of raw reported scores is a welcome transparency measure. However, the central tables contain several transcription, labeling, and convention errors that currently prevent the paper's main comparative claims from being verified from its own appendix. The paper's contribution is therefore weakened until these issues are resolved. The survey also covers a broad and current set of systems, which is valuable for the community, but the presence of self-inconsistencies in the empirical core makes the 'comprehensive' and 'comparative' claims premature.

major comments (5)
  1. [Table 2 / Appendix Table 6] The central comparative table (Table 2) is not reproducible from the appendix for at least two rows. In the R2AG/LLaMA2-7B row, the 2Wiki entry '34.52/4.445' mixes an absolute F1 score (34.52, from Table 6) with a relative-gain value (4.445 = (34.52-6.34)/6.34). Since Table 2 promises relative improvements in both columns, this entry would be read as a 3452% gain rather than the correct 444.5% gain. In the same table, the SimRAG/ARC entries (0.145 and 0.034) are placed in the 'B' (raw-LLM) column even though Appendix Table 5 supplies no raw-LLM score for SimRAG and the values correspond to (81.4-71.08)/71.08 and (88.65-85.75)/85.75, i.e., gains over LLM+retrieval. These errors directly affect the Section 5.1 and 5.2 narratives that attribute large gains to retriever-based designs.
  2. [Section 5.3 / Table 3 / Table 7] Section 5.3 and Table 3 contain several label/value mismatches that undermine the robustness comparison. (i) For FILCO/FEVER, Table 3 reports 3.25 in a column of relative improvements, where Flare-Direct's 0.216 is interpreted by the text as 21.6%; read consistently, 3.25 would be a 325% relative gain, but the text says '+3.25%'. (ii) Table 3 lists two 'SELF-RAG LLaMA2-7B ASQA' rows with different values; the second row's values match the LLaMA2-13B scores in Table 7. (iii) The sentence 'Self-RAG and CRAG, reporting +0.372 and +0.252 gains on the same dataset' is unsupported: no Self-RAG/Biography row exists in either Table 3 or Table 7. These issues cloud the claimed ordering of robustness gains.
  3. [Section 5.1 / Table 2 / Table 5] The claimed empirical basis of Section 5.1 is incomplete: RAAT is discussed as showing a 116% raw-baseline and 27% retrieval-baseline improvement on RAG-Bench, yet RAAT is absent from Table 2, which is the table the section opens by referencing. The values are present in Appendix Table 5, so this is fixable, but as written the headline claim is not traceable to the table that the section says it summarizes.
  4. [Section 5 / Table 2 / Table 6] The relative-gain comparisons are computed from highly heterogeneous raw scores, which can make the rankings fragile; for example, SELF-RAG's PopQA gains are computed from a raw-LLM accuracy of 14.7, while RQ-RAG's HotpotQA gain is computed from a raw F1 of 6.6. The paper should provide explicit protocol-compatibility caveats, confidence intervals, or error bars, or restrict claims to qualitative ordering. Without this, a single outlier such as the R2AG 2Wiki entry can materially change the conclusion that retriever-based systems are the most consistent winners.
  5. [Tables 2 and 3 / Appendix Tables 6 and 7] The metric labels in Tables 2 and 3 need to be reconciled with the appendix. Stochastic RAG/HotpotQA is labeled F1 in Table 2 but EM in Table 6; Re2G KGI1/TriviaQA Precision in Table 3 has no corresponding appendix row (Table 7 lists KGI0 rows and a TriviaQA recall, not precision). These mismatches prevent a reader from verifying the reported gains and should be corrected systematically.
minor comments (6)
  1. [Section 4.4 / Table 5] The text states that RAAT improves F1/EM by 20–30%, but Appendix Table 5 shows a 116% improvement over the raw baseline and 27% over the retrieval baseline; please clarify which baseline the 20–30% figure refers to.
  2. [Section 5.1 / Table 2] The sentence 'xRAG achieves 10–29% improvements over raw LLM baselines' seems to exclude the Mixtral-8x7B TriviaQA entry, which is 0.043 (4.3%); the range should be 4–29%.
  3. [Section 4.1 / Table 1] The claim that TA-ARE reduces redundant retrievals by 14.9% is not supported by an appendix row or a specific citation in the text; please provide the source or add an appendix entry.
  4. [Section 4.3 / Table 1] Table 1 reports a 20–50% TTFT reduction for Speculative Pipelining, while the text says 20–30%; please reconcile these numbers and report the actual range from the cited paper.
  5. [Section 4.3 / Table 1] The Table 1 entry for RAGCache says it 'eliminates recomputation,' while the text correctly limits this to high-throughput workloads; please soften the table description to match the text.
  6. [Section 2.2 / Equation (2)] Equation (2) presents the truncated sum as an approximation of Equation (1), but it is not a normalized approximation since P(d|x) is not renormalized over the top-k set; please clarify the nature of the approximation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: comparative analysis is a restatement of externally published benchmark scores, not a fit or self-citation chain.

full rationale

This is a survey with no new derivation. The only formal content is the standard RAG decomposition P(y|x)=sum_d P(y|x,d)P(d|x), which is background definition, not used to derive any claimed result. The load-bearing comparative claims in Section 5 are relative improvements computed from raw scores listed in the appendix. Spot-checking confirms the values are reproducible percentage changes rather than fits: RQ-RAG HotpotQA (62.6-6.6)/6.6=8.485 and (62.6-16.7)/16.7=2.749; SELF-RAG LLaMA2-13B PopQA (55.8-14.7)/14.7=2.796 and (55.8-45.7)/45.7=0.221; xRAG Mixtral-8x7B NQ (47.28-41.99)/41.99=0.126. Every source score is attributed to an external cited paper (RQ-RAG [6], SELF-RAG [1], xRAG [10], etc.), and the author's own prior work is not invoked to justify any central claim. The proposed taxonomy is an organizing scheme, not a theorem forced by definition. Any column-labeling ambiguities in Table 2 (e.g., SimRAG's missing raw baseline) are correctness or presentation issues, not circularity, because the underlying values trace to published external results. No fitted input is relabeled as a prediction, no uniqueness theorem is imported from the authors, and no ansatz is smuggled in via self-citation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no invented entities. Its central claim rests on the accuracy of cited numbers, the suitability of the standard RAG probability decomposition, and the coverage of the proposed taxonomy.

assumptions (3)
  • domain assumption The reported scores in the cited papers are accurately transcribed and comparable.
    Section 5 and the Appendix rely on numbers extracted from original publications. The survey does not independently re-run any experiment, so any transcription error or protocol difference propagates into the comparative analysis.
  • standard math The mathematical formulation of RAG in Eq (1) and Eq (2) is an adequate model for all surveyed systems.
    Section 2.2 states this conditional decomposition as the foundation of RAG. The survey never questions whether some systems (e.g., graph-based or tool-based RAG) violate this formulation.
  • domain assumption The proposed four-category taxonomy covers the design space of RAG.
    Section 3 assigns every surveyed system to one of four categories. The survey acknowledges overlaps (e.g., CRAG is placed in 'hybrid' despite corrective filtering) but does not validate that the taxonomy is exhaustive or non-overlapping.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers." pith.science (2026). https://pith.science/paper/Q32VU7BR

@misc{pith2026250600054,
  author       = {Pith},
  title        = {Pith review of: Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q32VU7BR}},
  note         = {Machine review of arXiv:2506.00054}
}
read the original abstract

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm to enhance large language models (LLMs) by conditioning generation on external evidence retrieved at inference time. While RAG addresses critical limitations of parametric knowledge storage-such as factual inconsistency and domain inflexibility-it introduces new challenges in retrieval quality, grounding fidelity, pipeline efficiency, and robustness against noisy or adversarial inputs. This survey provides a comprehensive synthesis of recent advances in RAG systems, offering a taxonomy that categorizes architectures into retriever-centric, generator-centric, hybrid, and robustness-oriented designs. We systematically analyze enhancements across retrieval optimization, context filtering, decoding control, and efficiency improvements, supported by comparative performance analyses on short-form and multi-hop question answering tasks. Furthermore, we review state-of-the-art evaluation frameworks and benchmarks, highlighting trends in retrieval-aware evaluation, robustness testing, and federated retrieval settings. Our analysis reveals recurring trade-offs between retrieval precision and generation flexibility, efficiency and faithfulness, and modularity and coordination. We conclude by identifying open challenges and future research directions, including adaptive retrieval architectures, real-time retrieval integration, structured reasoning over multi-hop evidence, and privacy-preserving retrieval mechanisms. This survey aims to consolidate current knowledge in RAG research and serve as a foundation for the next generation of retrieval-augmented language modeling systems.

Figures

Figures reproduced from arXiv: 2506.00054 by the authors.

Figure 1
Figure 1. Retrieval-Augmented Generation (RAG) workflow. A user query is processed by the retriever, which may perform query expansion before retrieving documents from external knowledge sources (e.g., databases, APIs, or document stores). Retrieved documents are re-ranked by relevance, and the Top-K are passed to the generator as factual context. The generator synthesizes a response conditioned on both the query and retrieve… view at source ↗
Figure 2
Figure 2. Taxonomy of Retrieval-Augmented Generation (RAG) Systems. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing

    cs.AI 2026-08 conditional novelty 5.0 of 10

    The paper proposes a MARL-centered three-layer reference architecture for LLM augmentation in smart manufacturing, with a conditional allocation: MARL for frequent coordination, LLMs for semantic, reward, and planning roles.

  2. TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization

    cs.LG 2026-07 reject novelty 5.0 of 10

    A case study of TurboQuant for enterprise RAG reports a large recall advantage over product quantization, but the advantage depends on an unequal memory comparison.

  3. Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Hybrid RAG over UK public health guidance sharply raises MCQA accuracy and free-form faithfulness, letting smaller open models match larger closed models without retrieval.

Reference graph

Works this paper leans on

88 extracted references · 25 canonical work pages · cited by 3 Pith papers

  1. [1]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In The Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=hSyW5go0v8

  2. [2]

    Orlando Ayala and Patrice Bechard. 2024. Reducing hallucination in structured outputs via Retrieval-Augmented Generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), Yi Yang, Aida Davani, Avi Sil, and Anoop Kumar (Eds.). Associ...

  3. [3]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https://arxiv.org/abs/1611.09268 Manuscript subm...

  4. [4]

    Scott Barnett, Stefanus Kurniawan, Srikanth Thudumu, Zach Brannelly, and Mohamed Abdelrazek. 2024. Seven Failure Points When Engineering a Retrieval Augmented Generation System. In Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering - Software Engineering for AI (Lisbon, Portugal) (CAIN ’24). Association for Computing Machinery, New...

  5. [5]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  6. [6]

    Chi-Min Chan, Chunpu Xu, Ruibin Yuan, Hongyin Luo, Wei Xue, Yike Guo, and Jie Fu. 2024. RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation. In First Conference on Language Modeling . https://openreview.net/forum?id=tzE7VqsaJ4

  7. [7]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelli...

  8. [8]

    Pengzhou Cheng, Yidong Ding, Tianjie Ju, Zongru Wu, Wei Du, Haodong Zhao, Ping Yi, Zhuosheng Zhang, and Gongshen Liu. 2024. TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models. https://openreview.net/forum?id=RfYD6v829Y

Show all 88 references
  1. [9]

    Xin Cheng, Di Luo, Xiuying Chen, Lemao Liu, Dongyan Zhao, and Rui Yan. 2023. Lift yourself up: retrieval-augmented text generation with self-memory. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). ...

  2. [10]

    Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, and Dongyan Zhao. 2024. xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems . ht...

  3. [11]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...

  4. [12]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457 [cs.AI] https://arxiv.org/abs/1803.05457

  5. [13]

    Cormack, Charles L A Clarke, and Stefan Buettcher

    Gordon V. Cormack, Charles L A Clarke, and Stefan Buettcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (Boston, MA,...

  6. [14]

    Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri

  7. [15]

    Nguyen Nam Doan, Aki Härmä, Remzi Celebi, and Valeria Gottardo. 2024. A Hybrid Retrieval Approach for Advancing Retrieval-Augmented Generation Systems. In Proceedings of the 7th International Conference on Natural Language and Speech Processing (ICNLSP 2024) , Mourad Abbas and...

  8. [16]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2025. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130 [cs.CL] https://arxiv....

  9. [17]

    Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations , Niko...

  10. [18]

    Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. 2024. Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (...

  11. [19]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval- augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 2 (2023), 1

  12. [20]

    Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Naik, Pengshan Cai, and Alfio Gliozzo. 2022. Re2G: Retrieve, Rerank, Generate. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langu...

  13. [21]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Man...

  14. [22]

    Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh V Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-Retriever: Retrieval- Augmented Generation for Textual Graph Understanding and Question Answering. In The Thirty-eighth Annual Conference on Neural Informati...

  15. [23]

    Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps. In Proceedings of the 28th International Conference on Computational Linguistics , Donia Scott, Nuria Bel, and Chengqing ...

  16. [24]

    Sebastian Hofstätter, Jiecao Chen, Karthik Raman, and Hamed Zamani. 2023. FiD-Light: Efficient and Effective Retrieval-Augmented Text Generation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan)...

  17. [25]

    Jie Huang, Mo Wang, Yunpeng Cui, Juan Liu, Li Chen, Ting Wang, Huan Li, and Jinming Wu. 2024. Layered Query Retrieval: An Adaptive Framework for Retrieval-Augmented Generation in Complex Question Answering for Large Language Models. Applied Sciences 14, 23 (2024). doi:10.3390/...

  18. [26]

    Wenyu Huang, Mirella Lapata, Pavlos Vougiouklis, Nikos Papasarantopoulos, and Jeff Pan. 2023. Retrieval Augmented Generation with Rich Answer Encoding. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-P...

  19. [27]

    Gautier Izacard and Edouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , Paola Merlo, Jorg Tied...

  20. [28]

    Jisoo Jang and Wen-Syan Li. 2024. AU-RAG: Agent-based Universal Retrieval Augmented Generation. InProceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (Tokyo, Japan) (SIGIR-AP 2024). A...

  21. [29]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  22. [30]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  23. [31]

    Ziyan Jiang, Xueguang Ma, and Wenhu Chen. 2024. LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs. arXiv:2406.15319 [cs.CL] https://arxiv.org/abs/2406.15319

  24. [32]

    Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Augmented Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Jua...

  25. [33]

    Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, and Xin Jin. 2024. RAGCache: Efficient Knowledge Caching for Retrieval- Augmented Generation. arXiv:2404.12457 [cs.DC] https://arxiv.org/abs/2404.12457

  26. [34]

    Hailey Joren, Jianyi Zhang, Chun-Sung Ferng, Da-Cheng Juan, Ankur Taly, and Cyrus Rashtchian. 2025. Sufficient Context: A New Lens on Retrieval Augmented Generation Systems. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=...

  27. [35]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Re...

  28. [36]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)...

  29. [37]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

  30. [38]

    Hyunji Lee, Sohee Yang, Hanseok Oh, and Minjoon Seo. 2022. Generative Multi-hop Retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguist...

  31. [39]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58t...

  32. [40]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing s...

  33. [41]

    Zhuowan Li, Cheng Li, Mingyang Zhang, Qiaozhu Mei, and Michael Bendersky. 2024. Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry T...

  34. [42]

    Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. InProceedings of the 61st Annual Meeting of the Association for Comput...

  35. [43]

    Nicholas Matsumoto, Jay Moran, Hyunjun Choi, Miguel E Hernandez, Mythreye Venkatesan, Paul Wang, and Jason H Moore. 2024. KRAGEN: a knowledge graph-enhanced RAG framework for biomedical problem solving using large language models. Bioinformatics 40, 6 (2024), btae353

  36. [45]

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the 2023 Conference on Empiric...

  37. [46]

    Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024. RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computat...

  38. [47]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  39. [48]

    Shintaro Ozaki, Yuta Kato, Siyuan Feng, Masayo Tomita, Kazuki Hayashi, Wataru Hashimoto, Ryoma Obara, Masafumi Oyamada, Katsuhiko Hayashi, Hidetaka Kamigaito, and Taro Watanabe. 2025. Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the...

  40. [49]

    Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Task...

  41. [50]

    Zackary Rackauckas. 2024. Rag-Fusion: A New Take on Retrieval Augmented Generation. International Journal on Natural Language Computing 13, 1 (Feb. 2024), 37–47. doi:10.5121/ijnlc.2024.13103

  42. [51]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67

  43. [53]

    Lillicrap, Jean-Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M

    Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy P. Lillicrap, Jean-Baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, Andrew M. Dai, Katie Millican, Ethan Dyer, Mia...

  44. [54]

    Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends® in Information Retrieval 3, 4 (2009), 333–389

  45. [55]

    Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: ...

  46. [57]

    Alireza Salemi and Hamed Zamani. 2024. Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington D...

  47. [58]

    Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-Augmented Black-Box Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa...

  48. [59]

    Zhengliang Shi, Shuo Zhang, Weiwei Sun, Shen Gao, Pengjie Ren, Zhumin Chen, and Zhaochun Ren. 2024. Generate-then-Ground in Retrieval- Augmented Generation for Multi-hop Question Answering. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguisti...

  49. [60]

    Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid Questions Meet Long-Form Answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for...

  50. [61]

    Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V...

  51. [62]

    Viju Sudhi, Sinchana Ramakanth Bhat, Max Rudat, and Roman Teucher. 2024. RAG-Ex: A Generic Framework for Explaining Retrieval Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC...

  52. [63]

    Yixuan Tang and Yi Yang. 2024. MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries. InFirst Conference on Language Modeling. https://openreview.net/forum?id=t4eB3zYWBK

  53. [64]

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex...

  54. [65]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human La...

  55. [66]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  56. [67]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics 10 (2022), 539–554. doi:10.1162/tacl_a_00475

  57. [68]

    Shuai Wang, Ekaterina Khramtsova, Shengyao Zhuang, and Guido Zuccon. 2024. FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval ...

  58. [69]

    Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023. Learning to filter context for retrieval-augmented generation. arXiv preprint arXiv:2311.08377 (2023)

  59. [70]

    Zheng Wang, Shu Teo, Jieer Ouyang, Yongjun Xu, and Wei Shi. 2024. M-RAG: Reinforcing Large Language Model Performance through Retrieval- Augmented Generation with Multiple Partitions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V...

  60. [71]

    Zilong Wang, Zifeng Wang, Long Le, Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, and Tomas Pfister. 2025. Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting. InThe Thirteenth Internati...

  61. [72]

    Junde Wu, Jiayuan Zhu, and Yunli Qi. 2024. Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation. CoRR abs/2408.04187 (2024). https://doi.org/10.48550/arXiv.2408.04187

  62. [73]

    Ho, Carl Yang, and Qi He

    Ran Xu, Hui Liu, Sreyashi Nag, Zhenwei Dai, Yaochen Xie, Xianfeng Tang, Chen Luo, Yang Li, Joyce C. Ho, Carl Yang, and Qi He. 2025. SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains. In Proceedings of the 2025 Confe...

  63. [74]

    Sheng Xu, Mike Chen, and Shuwen Chen. 2024. Enhancing Retrieval-Augmented Generation Models with Knowledge Graphs: Innovative Practices Through a Dual-Pathway Approach. In Advanced Intelligent Computing Technology and Applications: 20th International Conference, ICIC 2024, Tia...

  64. [75]

    Shicheng Xu, Liang Pang, Jun Xu, Huawei Shen, and Xueqi Cheng. 2024. List-aware Reranking-Truncation Joint Model for Search and Retrieval- augmented Generation. In Proceedings of the ACM Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Computing Machinery,...

  65. [76]

    Shicheng Xu, Liang Pang, Mo Yu, Fandong Meng, Huawei Shen, Xueqi Cheng, and Jie Zhou. 2024. Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational ...

  66. [77]

    Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Deshpande, Xiaofeng Wang, and Zheng Li. 2024. Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering. In Proceedings of the 47th International ACM SIGIR Conference on Research...

  67. [78]

    Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou. 2024. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models. https://openreview.net/forum?id=G2p8TLuJgy

  68. [79]

    Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective Retrieval Augmented Generation. https://openreview.net/forum?id= JnWJbrnaUE

  69. [80]

    Diji Yang, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo, Yawen Zhang, Jie Yang, and Yi Zhang. 2024. IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Info...

  70. [81]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...

  71. [82]

    Fuda Ye, Shuangyin Li, Yongqi Zhang, and Lei Chen. 2024. R2AG: Incorporating Retrieval Information into Retrieval Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2024 , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Associat...

  72. [83]

    Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Mohammad Shoeybi, and Bryan Catanzaro. 2024. RankRAG: Uni- fying Context Ranking with Retrieval-Augmented Generation in LLMs. In NeurIPS. http://papers.nips.cc/paper_files/paper/2024/hash/ db93ccb6cf392f352570dd...

  73. [84]

    Hamed Zamani and Michael Bendersky. 2024. Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (S...

  74. [85]

    Shenglai Zeng, Jiankun Zhang, Pengfei He, Yiding Liu, Yue Xing, Han Xu, Jie Ren, Yi Chang, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang. 2024. The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG). In Findings of the Association for Computation...

  75. [86]

    Zihan Zhang, Meng Fang, and Ling Chen. 2024. RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering. In Findings of the Association for Computational Linguistics: ACL 2024 , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Ed...

  76. [87]

    Xinping Zhao, Dongfang Li, Yan Zhong, Boren Hu, Yibin Chen, Baotian Hu, and Min Zhang. 2024. SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, ...

  77. [88]

    Yuanhang Zheng, Peng Li, Wei Liu, Yang Liu, Jian Luan, and Bin Wang. 2024. ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COL...

  78. [89]

    Kun Zhu, Xiaocheng Feng, Xiyuan Du, Yuxuan Gu, Weijiang Yu, Haotian Wang, Qianglong Chen, Zheng Chu, Jingchang Chen, and Bing Qin. 2024. An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generation. In Proceedings of the 62nd Annual Mee...

  79. [90]

    Yun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu, Liangchen Luo, Lei Meng, Bang Liu, and Jindong Chen. 2025. Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection. In The Thirteenth International Conference on Lea...

  80. [2019]

    Transactions of the Association for Computational Linguistics 7 (2019), 452–466

    Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics 7 (2019), 452–466. doi:10.1162/tacl_a_00276

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.