REVIEW 3 major objections 8 minor 140 references
Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
T0 review · 3 major / 8 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read RAG's core ideas — retrieval plus generation, knowledge augmentation, answer verification, and iterative query refinement — were already built and tested in early-2000s IR and QA research, with 2002's QUALIFIER as an early Agentic RAG.
desk verdict A genuinely useful historical reframing of RAG that overreaches slightly on the QUALIFIER-to-Agentic-RAG analogy; the core thesis that most RAG components predate LLMs holds and the paper deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the QUALIFIER system of 2002–2003, and in particular its successive constraint relaxation (SCR) loop. QUALIFIER modeled a question as a structured event with slots (time, location, agent, action), encoded the known slots as a deliberately strict Boolean query, and retrieved with the MG indexing system; when an over-constrained query returned no documents, it removed a portion of the expanded terms and retried, repeating for up to five iterations and returning NIL rather than a low-confidence guess if nothing survived. The paper's interpretive move is to see this relax-and-retry cycle as retrieval shaped by intermediate reasoning states — retrieve, assess whether the evidence suffices, adapt, retrieve again — which is the same loop that defines Agentic RAG in current systems. Supporting machinery includes the stage-by-stage mapping of the classical QA pipeline (question analysis, retrieval, answer synthesis with verification) onto modern RAG, QUALIFIER's Answer Event Score for ranking candidate passages, and the TREC 2002 benchmark result (290 of 500 questions correct, second to the logic-based PowerAnswer system) that anchors the claim that iterative refinement carried practical weight.
What would settle it
Re-run QUALIFIER's TREC 2002 task with the feedback loop disabled: submit the initial strict query once, and on an empty result issue a single fixed relaxation with no intermediate verification. If accuracy on the 500 questions stays at the reported 290 correct answers, the 'retrieval shaped by intermediate reasoning states' interpretation gains no support, and the proto-Agentic claim reduces to a mechanical artifact of Boolean retrieval.
Extended reading notes
Core claim
The central claim is that the core ideas underlying RAG, including Agentic RAG, are not new: integrating retrieval with language generation, augmenting knowledge with external sources, verifying answers against evidence, and iteratively refining queries or prompts were already studied and instantiated in IR and QA research dating back to the early 2000s, before large language models existed. The paper makes this case by mapping each stage of modern RAG onto a TREC-era antecedent: query rewriting and expansion onto classical relevance feedback and query expansion; structured querying onto QUALIFIER's event-based constraint encoding; evidence-conditioned generation onto retrieval-then-synthesis QA pipelines; answer verification onto mechanisms such as PIQUANT's sanity checker and web-redundancy scoring; and iterative, 'agentic' retrieval onto QUALIFIER's successive constraint relaxation, in which a maximally strict query is loosened only when it returns no answer. The authors' reading is that LLMs are best understood as a new interface layer on a decades-old architecture, and that this continuity has gone under-recognized because of community fragmentation, shifting terminology, and recency bias.
Load-bearing premise
The argument hinges on reading QUALIFIER's iterated 'retrieve, check, relax, retry' loop as a true forerunner of Agentic RAG; if that loop was merely a workaround forced by the fact that strict term-matching returns nothing when a query is too long, the strong historical thesis collapses into the far weaker claim that some RAG components existed earlier.
Editorial extensions
If this is right
- Classical IR work becomes a direct design source for next-generation RAG: query expansion, relevance feedback, and structured queries are ready-made solutions to problems RAG is currently re-solving from scratch.
- RAG evaluation should absorb IR's user-centered discipline — bias-aware feedback modeling, interaction studies, and long-term user-experience measures — rather than relying on static reference-based metrics alone.
- Each of the paper's four proposed directions — personalized RAG, proactive RAG, governance-aware RAG, and user-centered evaluation — inherits an established IR research line (user modeling, query clarification, rule-based governance and auditing, click-bias-aware evaluation) that current RAG has underused.
- Because QUALIFIER's iterative loop outperformed single-shot systems at TREC 2002 with no language model in the loop, the account implies that iterative 'agentic' retrieval has measurable value independent of LLM capability — a claim worth testing directly in modern stacks.
Reading between the lines
- One testable extension the authors do not pursue: transplant QUALIFIER's constrain-then-relax loop onto a modern retriever–LLM stack and compare against single-pass RAG; if the loop still buys accuracy, the historical continuity is functional rather than merely conceptual.
- Implicit in the 'LLMs as interface layer' framing is a prediction the authors do not state: if the field absorbs this history, pre-LLM QA and IR citations should start appearing in RAG papers as design justifications rather than background, and a few years of citation data could measure whether the field actually stops rediscovering.
- The historical thesis plausibly extends to a neighbouring problem: agentic tool-use and web-search agents are reconfigurations of interactive IR ideas (mixed-initiative interaction, relevance feedback, session modeling), so the same 'forgotten history' pattern may hold across the broader agentic web agenda.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that the core ideas underlying modern RAG and Agentic RAG—integrating retrieval and generation, knowledge augmentation, answer verification, and iterative query/refinement—were already studied and instantiated in IR and QA research from the early 2000s, particularly in TREC QA systems such as QUALIFIER. It traces the conceptual lineage from classical retrieval-based QA through RAG to Agentic RAG, proposes viewing LLMs as a new interface layer atop a decades-old QA architecture, and identifies four IR-informed future directions (personalization, proactive interaction, governance, and user-centered evaluation). The historical narrative is grounded in TREC proceedings and Voorhees's overviews, and the paper includes a conceptual mapping diagram and a comparison of TREC 2002 results.
Significance. If the historical thesis holds, the paper provides a valuable corrective to the prevailing LLM-centric narrative of RAG: it identifies concrete antecedents for RAG's modular pipeline, surfaces underutilized IR work that could inform next-generation RAG design, and offers a falsifiable account of continuity. The paper is well referenced on the historical facts, and the forward-looking Section 5 is constructive and clearly tied to established IR concerns. However, the most distinctive claim—that QUALIFIER is a genuine precursor to Agentic RAG, so that even the agentic variant is 'not new'—rests on an interpretation of QUALIFIER's successive constraint relaxation that the manuscript itself renders ambiguous. The comparative performance evidence in Section 3.2 is also used to draw a causal conclusion that the presented results do not support. These issues are central to the strength of the historical claim, though fixable within the manuscript's scope.
major comments (3)
- [§3.1, §3.2] The claim that QUALIFIER is an early Agentic RAG system depends on whether its successive constraint relaxation (SCR) was triggered by failed answer verification or by an empty Boolean hit set. The manuscript states in §3.1 that 'Because Boolean retrieval only returns documents matching all query terms, an overly expanded query could return no results at all. To handle this, they introduced SCR,' which suggests a mechanical recall-repair fallback rather than retrieval 'iteratively shaped by intermediate reasoning states' (§3.1). Please report directly from [108] and [110] what condition actually triggered a relaxation: zero retrieved documents, failure to extract any candidate answer, or explicit constraint/verification failure. If SCR fired only on empty hit sets or empty extraction output, then the proto-Agentic framing is overstated and should be weakened to 'iterative query refinement for recall recovery.'
- [§3.2, Table 1] The sentence 'QUALIFIER's competitive performance suggests that agentic RAG with iterative refinement can offer practical advantages over simple, single-pass RAG' is not supported by the evidence presented. In Table 1, QUantifier ranks second with 290/500, but the top system, LCC PowerAnswer (415/500), is not characterized as a simple single-pass system and does not fit the paper's iterative-refinement frame. Without an ablation or process analysis that isolates the effect of the iterative loop from other QUALIFIER design choices (e.g., the structured event model, external resources, answer justification), the performance comparison cannot carry this causal conclusion. Please qualify the claim to 'competitive at the time' or provide direct evidence tying the iterative mechanism to the measured outcome.
- [Abstract, §2.1] The claim that 'integrating retrieval and language generation' was 'instantiated' in early 2000s QA systems needs a definitional guard. The systems described perform extractive answer selection, pattern-based candidate ranking, and template-style answer construction; none performs open-ended text generation in the LLM sense. If 'generation' is intended broadly to cover any evidence-conditioned answer synthesis, the claim is less distinctive; if it is intended to match modern generative language modeling, the historical examples do not instantiate it. The paper should explicitly state which sense of 'generation' it uses and adjust the continuity claim in the abstract and Section 2.1 accordingly.
minor comments (8)
- [§3.1] The system name is misspelled as 'QUALIFER' in the sentence 'QUALIFER's underlying design philosophy' in Section 3.1.
- [Table 1] The column header 'Answer Correctness' is unclear; it should read 'Number of correctly answered questions (out of 500)' or equivalent.
- [§3.2] The phrase 'dominated the field for the next decade, before community QA approaches emerged' is vague and does not align with the TREC QA track timeline (1999–2007); please clarify the intended period and what 'community QA approaches' refers to.
- [References [108]–[111]] Several of the key QUALIFIER references share a co-author with the current paper, yet the manuscript does not disclose this relationship; a brief acknowledgment or neutral note would help readers calibrate the interpretive weight placed on this specific system.
- [§2.5, Figure 1] Figure 1 is not explicitly cited where the 'two levels of conceptual continuity' are discussed in Section 2.5; please add a reference to the figure at that point.
- [§2.2, Ref [71]] Characterizing LLMs as 'a huge knowledge base' by citing [71] is a contested framing; please qualify it (e.g., 'can be viewed as approximate knowledge stores') to avoid overstating the analogy.
- [§4] The sentence 'The main difference lies not in the underlying problem structure, but in the implementation layer' appears in both Section 2.1 and Section 4; consider keeping only one instance or adding a cross-reference.
- [§5.3.2] The heading 'Governance-A ware RAG' contains a typo; it should read 'Governance-Aware RAG.'
Circularity Check
No circularity: QUALIFIER's behavior and TREC results are externally documented, so the historical thesis does not reduce to its inputs.
full rationale
This is a historical and interpretive position paper, not a derivation with fitted parameters or predicted quantities. No equation is reused as an output, and no parameter is fitted to a subset of data and then reported as a prediction. The central claim — that core RAG ideas, including iterative query refinement, existed in early IR/QA work — rests on public artifacts: TREC QA track proceedings, the official Voorhees overview of TREC 2002 results (ref [96]), and the QUALIFIER system papers (refs [108–111]). Although some of those system papers are authored by co-author Tat-Seng Chua, the cited behavior (Boolean retrieval over the MG index, successive constraint relaxation triggered when over-expanded queries return no hits, up to five relax-and-retry iterations, and NIL on failure) is publicly documented in NIST proceedings and externally evaluated, so the self-citation is verifiable evidence rather than a self-referential premise. The skeptic's objection that SCR may be an artifact of strict Boolean matching is a challenge to the strength of the historical analogy, not evidence of circularity; the paper's claim does not become true by definition or by the authors' own assertion. Under the hard rules, externally documented and falsifiable cited results count as independent support, so no circular step is identified and the appropriate score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The components of early TREC QA systems correspond to the stages of modern RAG pipelines.
- domain assumption QUALIFIER's iterative constraint relaxation is functionally equivalent to the agentic RAG loop.
- domain assumption The intellectual lineage can be traced through the specific systems selected (JAVELIN, MultiText, DIOGENE, QUALIFIER, etc.).
Cite this review
Pith. "Pith review of Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective." pith.science (2026). https://pith.science/paper/V4YARIZM
@misc{pith2026260808445,
author = {Pith},
title = {Pith review of: Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/V4YARIZM}},
note = {Machine review of arXiv:2608.08445}
}
read the original abstract
Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core ideas underlying RAG are not new: foundational concepts such as integrating retrieval and language generation, knowledge augmentation, answer verification, and iterative query (or prompt) refinement had already been studied and instantiated in information retrieval (IR) and question answering (QA) research dating back to the early 2000s, well before the emergence of LLMs. We make this case by systematically tracing the intellectual lineage of modern RAG and Agentic RAG back to their classical IR and QA antecedents, and examining why this continuity has gone under-recognized -- a consequence of community fragmentation, shifting terminology, and the recency bias endemic to fast-moving fields. Rather than treating LLMs as the origin point of retrieval-augmented intelligence, we propose viewing them as a new interface layer atop a decades-old QA architecture. This reframing is not merely historical: by situating RAG within the longer trajectory of IR research, we surface underutilized prior work -- on user modeling, answer validation, and query refinement -- that can directly inform next-generation RAG design, reducing unintentional rediscovery and fostering genuine cross-community integration.
Figures
Reference graph
Works this paper leans on
- [108]
-
[110]
Nikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanoulas, and Maarten De Ri- jke. 2020. Query resolution for conversational search with limited supervision. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 921–930
work page 2020
-
[1]
Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, and W Bruce Croft
-
[2]
Gianni Amati and Cornelis Joost Van Rijsbergen. 2002. Probabilistic models of information retrieval based on measuring the divergence from randomness. ACM Transactions on Information Systems (TOIS)20, 4 (2002), 357–389
2002
-
[3]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. (2024)
2024
-
[4]
Sumit Bhatia, Debapriyo Majumdar, and Prasenjit Mitra. 2011. Query sugges- tions in the absence of query logs. InProceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. 795–804
2011
-
[5]
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Bal- dini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, et al. 2022. Attributed question answering: Evaluation and modeling for attributed large language models.arXiv preprint arXiv:2212.08037 (2022)
arXiv 2022
-
[6]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Ruther- ford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bog- dan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. InInternational conference on machine learning. PMLR, 2206–2240
2022
Show all 140 references
-
[7]
Eric Brill, Susan Dumais, and Michele Banko. 2002. An analysis of the AskMSR question-answering system. InProceedings of the 2002 conference on empirical methods in natural language processing (EMNLP 2002). 257–264
2002
-
[8]
Eric Brill, Jimmy Lin, Michele Banko, Susan T Dumais, Andrew Y Ng, et al. 2001. Data-Intensive Question Answering.. InTREC, Vol. 56. 90
2001
-
[9]
Lin, Michele Banko, Susan T
Eric Brill, Jimmy J. Lin, Michele Banko, Susan T. Dumais, and Andrew Y. Ng
-
[10]
Fei Cai, Maarten De Rijke, et al. 2016. A survey of query auto completion in information retrieval.Foundations and Trends®in Information Retrieval10, 4 (2016), 273–363
2016
-
[11]
James P Callan, W Bruce Croft, and Stephen M Harding. 1992. The INQUERY retrieval system. InDatabase and Expert Systems Applications: Proceedings of the International Conference in Valencia, Spain, 1992. Springer, 78–83
1992
-
[12]
Jaime Carbonell and Jade Goldstein. 1998. The use of MMR, diversity-based reranking for reordering documents and producing summaries. InProceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. 335–336
1998
-
[13]
Claudio Carpineto and Giovanni Romano. 2012. A survey of automatic query expansion in information retrieval.Acm Computing Surveys (CSUR)44, 1 (2012), 1–50
2012
-
[14]
Olivier Chapelle and Ya Zhang. 2009. A dynamic bayesian network click model for web search ranking. InProceedings of the 18th international conference on World wide web. 1–10
2009
-
[15]
Prager, Chris A
Jennifer Chu-Carroll, John M. Prager, Chris A. Welty, David Ferrucci, and David A. Boguraev. 2002. A Multi-Strategy and Multi-Source Approach to Question Answering. InProceedings of the Eleventh Text Retrieval Conference, TREC 2002, Gaithersburg, Maryland, USA, November 19-22,...
2002
-
[16]
Charles LA Clarke, Gordon V Cormack, and Thomas R Lynam. 2001. Exploiting redundancy in question answering. InProceedings of the 24th annual inter- national ACM SIGIR conference on Research and development in information retrieval. 358–365
2001
-
[17]
Charles L. A. Clarke, Gordon V. Cormack, Graeme Kemkes, M. Laszlo, Thomas R. Lynam, Egidio L. Terra, and Philip L. Tilker. 2002. Statistical Selection of Exact Answers (MultiText Experiments for TREC 2002). InProceedings of the Eleventh Text REtrieval Conference, TREC 2002, Ga...
2002
-
[18]
Nick Craswell, Onno Zoeter, Michael Taylor, and Bill Ramsey. 2008. An exper- imental comparison of click position-bias models. InProceedings of the 2008 international conference on web search and data mining. 87–94
2008
-
[19]
W Bruce Croft. 1995. What do people want from information retrieval.D-Lib magazine1, 5 (1995)
1995
-
[20]
Steve Cronen-Townsend, Yun Zhou, and W Bruce Croft. 2004. A framework for selective query expansion. InProceedings of the thirteenth ACM international conference on Information and knowledge management. 236–237
2004
-
[21]
Michail Dadopoulos, Anestis Ladas, Stratos Moschidis, and Ioannis Negkakis
-
[22]
Fernando Diaz, Bhaskar Mitra, and Nick Craswell. 2016. Query Expansion with Locally-Trained Word Embeddings. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 367–377
2016
-
[23]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query- focused summarization.arXiv preprint arXiv:2404.16130(2024)
2024 arXiv
-
[24]
Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. 2019. Can you unpack that? learning to rewrite questions-in-context.Can You Unpack That? Learning to Rewrite Questions-in-Context(2019)
2019
-
[25]
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. ...
2024
-
[26]
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse lexical and expansion model for first stage ranking. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2288–2292
2021
-
[27]
Furnas, Thomas K
George W. Furnas, Thomas K. Landauer, Louis M. Gomez, and Susan T. Dumais
-
[28]
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Cha- ganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al
-
[29]
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023. Precise zero- shot dense retrieval without relevance labels. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1762–1777
2023
-
[30]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997 2, 1 (2023)
2023 arXiv
-
[31]
Gurbinder Gill, Ritvik Gupta, Denis Lusson, Anand Chandrashekar, and Donald Nguyen. 2025. From Search to Reasoning: a Five-Level Rag Capability Frame- work for Enterprise Data. In2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware). IEEE, 01–09
2025
-
[32]
Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Naik, Pengshan Cai, and Alfio Gliozzo. 2022. Re2G: Retrieve, rerank, generate. InPro- ceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...
2022
-
[33]
Pei Guo, Enjie Liu, Ruichao Zhong, Mochi Gao, Yunzhi Tan, Bo Hu, and Zang Li
-
[34]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang
-
[35]
Kailash A Hambarde and Hugo Proenca. 2023. Information retrieval: recent advances and beyond.IEEE Access11 (2023), 76581–76604
2023
-
[36]
Marti A Hearst. 2006. Clustering versus faceted categories for information exploration.Commun. ACM49, 4 (2006), 59–61
2006
-
[37]
Ulf Hermjakob, Abdessamad Echihabi, and Daniel Marcu. 2002. Natural Lan- guage Based Reformulation Resource and Wide Exploitation for Question An- swering. InProceedings of the Eleventh Text REtrieval Conference (TREC 2002). https://trec.nist.gov/pubs/trec11/papers/usc.hermjakob.pdf
2002
-
[38]
Jerry R Hobbs. 1978. Resolving pronoun references.Lingua44, 4 (1978), 311– 338
1978
-
[39]
InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)
DSRAG: A Double-Stream Retrieval-Augmented Generation Framework for Countless Intent Detection. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track). 318–328
2025
-
[40]
2005.The turn: Integration of information seeking and retrieval in context
Peter Ingwersen and Kalervo Järvelin. 2005.The turn: Integration of information seeking and retrieval in context. Springer
2005
-
[41]
Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. InProceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume. 874–880
2021
-
[42]
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models.Journal of Machine Learning Research24, 251 (2...
2023
-
[43]
Thorsten Joachims. 2002. Optimizing search engines using clickthrough data. InProceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining. 133–142
2002
-
[44]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering.. InEMNLP (1). 6769–6781
2020
-
[45]
Jaana Kekäläinen and Kalervo Järvelin. 2000. The co-effects of query struc- ture and expansion on retrieval performance in probabilistic text retrieval. Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective Information retrieval1, 4 (2000), 329–344
2000
-
[46]
1992.Information retrieval interaction
Peter Ingwersen. 1992.Information retrieval interaction. Vol. 246. Taylor Graham London
1992
-
[47]
To Eun Kim and Fernando Diaz. 2025. Towards fair rag: On the impact of fair ranking in retrieval-augmented generation. InProceedings of the 2025 Interna- tional ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR). 33–43
2025
-
[48]
Jon M Kleinberg. 1999. Authoritative sources in a hyperlinked environment. Journal of the ACM (JACM)46, 5 (1999), 604–632
1999
-
[49]
Cody CT Kwok, Oren Etzioni, and Daniel S Weld. 2001. Scaling question answering to the web. InProceedings of the 10th international conference on World Wide Web. 150–161
2001
-
[50]
Shalom Lappin and Herbert J Leass. 1994. An algorithm for pronominal anaphora resolution.Computational linguistics20, 4 (1994), 535–561
1994
-
[51]
Victor Lavrenko and W Bruce Croft. 2001. Relevance-Based Language Models. (2001)
2001
-
[52]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing ...
2020
-
[53]
Diane Kelly et al. 2009. Methods for evaluating interactive information retrieval systems with users.Foundations and Trends®in Information Retrieval3, 1–2 (2009), 1–224
2009
-
[54]
Jintao Liang, Huifeng Lin, You Wu, Rui Zhao, Ziyue Li, et al. 2025. Reasoning rag via system 1 or system 2: A survey on reasoning agentic retrieval-augmented generation for industry challenges. InProceedings of the 14th International Joint Conference on Natural Language Proces...
2025
-
[55]
Hui Lin and Jeff Bilmes. 2011. A class of submodular functions for document summarization. InProceedings of the 49th annual meeting of the association for computational linguistics: human language technologies. 510–520
2011
-
[56]
Jimmy Lin, Aaron Fernandes, Boris Katz, Gregory Marton, and Stefanie Tellex
-
[57]
Tie-Yan Liu et al. 2009. Learning to rank for information retrieval.Foundations and Trends®in Information Retrieval3, 3 (2009), 225–331
2009
-
[58]
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting in retrieval-augmented large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5303– 5315
2023
-
[59]
Bernardo Magnini, Matteo Negri, Roberto Prevete, and Hristo Tanev. 2002. Mining Knowledge from Repeated Co-Occurrences: DIOGENE at TREC 2002. In Proceedings of the Eleventh Text REtrieval Conference (TREC-2002). NIST Special Publication, 577–586
2002
-
[60]
Zhicong Li, Jiahao Wang, Zhishu Jiang, Hangyu Mao, Zhongxia Chen, Jiazhen Du, Yuanxing Zhang, Fuzheng Zhang, Di Zhang, and Yong Liu. 2024. Dmqr-rag: Diverse multi-query rewriting for rag.arXiv preprint arXiv:2411.13154(2024)
2024 arXiv
-
[61]
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2021. Generation-augmented retrieval for open-domain question answering. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inter...
2021
-
[62]
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al. 2022. Teaching language models to support answers with verified quotes.arXiv preprint arXiv:2203.11147(2022)
2022 arXiv
-
[63]
George A Miller. 1992. WordNet: A Lexical Database for English. InSpeech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992
1992
-
[64]
Mandar Mitra, Amit Singhal, and Chris Buckley. 1998. Improving automatic query expansion. InProceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. 206–214
1998
-
[65]
Moldovan, S
D. Moldovan, S. Harabagiu, R. Girju, P. Morarescu, F. Lacatusu, A. Novischi, A. Badulescu, and O. Bolohan. 2001. LCC Tools for Question Answering. In Proceedings of the Tenth Text REtrieval Conference (TREC 2001). National Institute of Standards and Technology (NIST), 111–120
2001
-
[66]
Dan Moldovan, Marius Pasca, Sanda Harabagiu, and Mihai Surdeanu. 2002. Performance Issues and Error Analysis in an Open-Domain Question Answer- ing System. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 33–40
2002
-
[67]
Marco Morik, Ashudeep Singh, Jessica Hong, and Thorsten Joachims. 2020. Controlling fairness and bias in dynamic learning-to-rank. InProceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval. 429–438
2020
-
[68]
2008.Introduction to information retrieval
Christopher D Manning. 2008.Introduction to information retrieval. Syngress Publishing. Chapter 9: Relevance feedback and query expansion
2008
-
[69]
Eric Nyberg, Teruko Mitamura, Jaime Carbonell, Jamie Callan, Kevyn Collins- Thompson, Karen Czuba, Michael Duggan, L Hiyakumoto, N Hu, Y Huang, et al
-
[70]
Alexandra Olteanu, Jean Garcia-Gathright, Maarten de Rijke, Michael D Ek- strand, Adam Roegiest, Aldo Lipani, Alex Beutel, Alexandra Olteanu, Ana Lucic, Ana-Andreea Stoica, et al. 2021. FACTS-IR: fairness, accountability, confiden- tiality, transparency, and safety in informat...
2021
-
[71]
Miller, and Sebastian Riedel
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2019. Language Models as Knowledge Bases? arXiv:1909.01066 [cs.CL] https://arxiv.org/abs/1909.01066
2019 arXiv
-
[72]
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. InFindings of the Association for Computational Linguistics: EMNLP
2023
-
[73]
Yonggang Qiu and Hans-Peter Frei. 1993. Concept based query expansion. In Proceedings of the 16th annual international ACM SIGIR conference on Research and development in information retrieval. 160–169
1993
-
[74]
Zackary Rackauckas. 2024. Rag-fusion: a new take on retrieval-augmented generation.arXiv preprint arXiv:2402.03367(2024)
2024 arXiv
-
[75]
Dragomir Radev, Weiguo Fan, Hong Qi, Harris Wu, and Amardeep Grewal
-
[76]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332(2021)
2021 arXiv
-
[77]
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented lan- guage models.Transactions of the Association for Computational Linguistics11 (2023), 1316–1331
2023
-
[78]
InProceedings of the Eleventh Text Retrieval Conference (TREC 2002)
The JAVELIN Question-Answering System at TREC 2002. InProceedings of the Eleventh Text Retrieval Conference (TREC 2002). National Institute of Standards and Technology (NIST), 112–121
2002
-
[79]
Joseph John Rocchio Jr. 1971. Relevance feedback in information retrieval.The SMART retrieval system: experiments in automatic document processing(1971)
1971
-
[80]
Alireza Salemi, Surya Kallumadi, and Hamed Zamani. 2024. Optimization meth- ods for personalizing large language models through retrieval augmentation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 752–762
2024
-
[81]
Alireza Salemi and Hamed Zamani. 2024. Towards a search engine for machines: Unified ranking for multiple retrieval-augmented large language models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 741–751
2024
-
[82]
Gerard Salton and Chris Buckley. 1990. Improving retrieval performance by relevance feedback.Journal of the American society for information science41, 4 (1990), 288–297
1990
-
[83]
Gerard Salton, Edward A Fox, and Harry Wu. 1983. Extended boolean informa- tion retrieval.Commun. ACM26, 11 (1983), 1022–1036
1983
-
[84]
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024. Raptor: Recursive abstractive processing for tree-organized retrieval. InThe Twelfth International Conference on Learning Representations
2024
-
[85]
InProceedings of the 11th international conference on World Wide Web
Probabilistic question answering on the web. InProceedings of the 11th international conference on World Wide Web. 408–419
-
[86]
Dragomir R Radev, Hongyan Jing, Małgorzata Styś, and Daniel Tam. 2004. Centroid-based summarization of multiple documents.Information Processing & Management40, 6 (2004), 919–938
2004
-
[87]
Xuehua Shen, Bin Tan, and ChengXiang Zhai. 2005. Implicit user modeling for personalized search. InProceedings of the 14th ACM international conference on Information and knowledge management. 824–831
2005
-
[88]
Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond.Foundations and trends®in information retrieval 3, 4 (2009), 333–389
2009
-
[89]
Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2025. Agen- tic retrieval-augmented generation: A survey on agentic rag.arXiv preprint arXiv:2501.09136(2025)
2025 arXiv
-
[90]
Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William Cohen. 2018. Open domain question answer- ing using early fusion of knowledge bases and text. InProceedings of the 2018 conference on empirical methods in natural language processin...
2018
-
[91]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[92]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[93]
Howard Turtle and W Bruce Croft. 1989. Inference networks for document retrieval. InProceedings of the 13th annual international ACM SIGIR conference on Research and development in information retrieval. 1–24
1989
-
[94]
Ellen M Voorhees. 1994. Query expansion using lexical-semantic relations. InSIGIR’94: Proceedings of the Seventeenth Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, organised by Dublin City University. Springer, 61–69
1994
-
[95]
Mahmoud F Sayed and Douglas W Oard. 2019. Jointly modeling relevance and sensitivity for search among sensitive content. InProceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval. 615–624
2019
-
[96]
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy. InFindings of the Association for Com- putational Linguistics: EMNLP 2023. 9248–9274
2023
-
[97]
Voorhees and Dawn M
Ellen M. Voorhees and Dawn M. Tice. 2000. The TREC-8 Question Answer- ing Track. InProceedings of the Second International Conference on Language Resources and Evaluation (LREC 2000)
2000
-
[98]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. InFindings of the Association for Computational Linguistics: EMNLP 2021. 3784–3803
2021
-
[99]
Jonas Wallat, Maria Heuss, Maarten de Rijke, and Avishek Anand. 2025. Cor- rectness is not Faithfulness in Retrieval Augmented Generation Attributions. InProceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (IC...
2025
-
[100]
Haoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian, Zhengyang Wang, Chen Luo, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, and Jing Gao. 2024. Blend- filter: Advancing retrieval-augmented large language models via query genera- tion blending and knowledge filtering. InProceeding...
2024
-
[101]
Lidan Wang, Jimmy Lin, and Donald Metzler. 2011. A cascade ranking model for efficient ranked retrieval. InProceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. 105–114
2011
-
[102]
Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query Expansion with Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 9414–9423
2023
-
[103]
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor nega- tive contrastive learning for dense text retrieval.arXiv preprint arXiv:2007.00808 (2020)
2020 arXiv
-
[104]
InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers)
Interleaving retrieval with chain-of-thought reasoning for knowledge- intensive multi-step questions. InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers). 10014–10037
-
[105]
Jinxi Xu and W Bruce Croft. 1996. Query expansion using local and global document analysis. InProceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval. 4–11
1996
-
[106]
Xiwei Xu, Hans Weytjens, Dawen Zhang, Qinghua Lu, Ingo Weber, and Liming Zhu. 2025. RAGOps: Operating and Managing Retrieval-Augmented Generation Pipelines.arXiv preprint arXiv:2506.03401(2025)
2025 arXiv
-
[107]
Voorhees
Ellen M. Voorhees. 2001. The TREC Question Answering Track.Natural Language Engineering7, 4 (2001), 361–378
2001
-
[109]
Hui Yang and Tat-Seng Chua. 2003. Qualifier: question answering by lexical fabric and external resources. In10th Conference of the European Chapter of the Association for Computational Linguistics
2003
-
[111]
Hui Yang, Hang Cui, Mstislav Maslennikov, Long Qiu, Min-Yen Kan, and Tat- Seng Chua. 2003. QUALIFIER in TREC-12 QA Main Task. InProceedings of the Twelfth Text REtrieval Conference (TREC 2003). NIST
2003
-
[112]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning represen- tations
2022
-
[113]
Raquib Bin Yousuf, Shengzhe Xu, Mandar Sharma, Andrew Neeser, Chris La- timer, and Naren Ramakrishnan. 2026. Utilizing Metadata for Better Retrieval- Augmented Generation.arXiv preprint arXiv:2601.11863(2026)
2026
-
[114]
Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Moham- mad Shoeybi, and Bryan Catanzaro. 2024. Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Advances in Neural Information Processing Systems37 (2024), 121156–121184
2024
-
[115]
Hamed Zamani and W Bruce Croft. 2017. Relevance-based word embedding. InProceedings of the 40th international acm sigir conference on research and development in information retrieval. 505–514
2017
-
[116]
W Xiong, P Lewis, S Riedel, XL Li, W Wang, S Iyer, Y Mehdad, D Kiela, J Du, WT Yih, et al. 2020. Answering Complex Open-Domain Questions with Multi- Hop Dense Retrieval. InICLR 2021-9th International Conference on Learning Representations, Vol. 2021. ICLR
2020
-
[117]
Hamed Zamani, Susan Dumais, Nick Craswell, Paul Bennett, and Gord Lueck
-
[118]
Hamed Zamani, Bhaskar Mitra, Everest Chen, Gord Lueck, Fernando Diaz, Paul N Bennett, Nick Craswell, and Susan T Dumais. 2020. Analyzing and learning from user interactions for search clarification. InProceedings of the 43rd international acm sigir conference on research and d...
2020
-
[119]
Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. (2024)
2024
-
[120]
Hui Yang and Tat-Seng Chua. 2002. The Integration of Lexical Knowledge and External Resources for Question Answering. InProceedings of the Eleventh Text REtrieval Conference (TREC 2002). Gaithersburg, Maryland, USA, 59–69
2002
-
[121]
Kepu Zhang, Teng Shi, Weijie Yu, and Jun Xu. 2025. Prlm: Learning explicit reasoning for personalized rag via contrastive reward optimization. InProceed- ings of the 34th ACM International Conference on Information and Knowledge Management. 5484–5488
2025
-
[122]
Hui Yang, Tat-Seng Chua, Shuguang Wang, and Chun-Keat Koh. 2003. Struc- tured use of external knowledge for event-based open domain question answer- ing. InProceedings of the 26th annual international ACM SIGIR conference on Research and development in information retrieval. 33–40
2003
-
[123]
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi, Quoc V Le, and Denny Zhou. 2023. Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models.CoRR(2023)
2023
-
[124]
Zhiping Zheng. 2002. AnswerBus question answering system. InHuman Lan- guage Technology Conference (HLT 2002), Vol. 27
2002
-
[125]
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al . 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625(2022)
2022 arXiv
-
[126]
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chen- long Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey.arXiv preprint arXiv:2308.07107(2023)
2023
-
[127]
2025.{PoisonedRAG}: Knowledge corruption attacks to{Retrieval-Augmented} generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025.{PoisonedRAG}: Knowledge corruption attacks to{Retrieval-Augmented} generation of large language models. In34th USENIX Security Symposium (USENIX Security 25). 3827–3844
2025
-
[128]
Hamed Zamani, Fernando Diaz, Mostafa Dehghani, Donald Metzler, and Michael Bendersky. 2022. Retrieval-enhanced machine learning. InProceedings of the 45th international ACM SIGIR conference on research and development in infor- mation retrieval. 2875–2886
2022
-
[130]
InProceedings of the web conference 2020
Generating clarifying questions for information retrieval. InProceedings of the web conference 2020. 418–428
2020
-
[132]
George Zerveas, Navid Rekabsaz, Daniel Cohen, and Carsten Eickhoff. 2022. Mitigating bias in search results through contextual document reranking and neutrality regularization. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Informat...
2022
-
[133]
Chengxiang Zhai and John Lafferty. 2001. Model-based feedback in the lan- guage modeling approach to information retrieval. InProceedings of the tenth international conference on Information and knowledge management. 403–410
2001
-
[135]
Chen Zhao, Chenyan Xiong, Jordan Boyd-Graber, and Hal Daumé III. 2021. Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...
2021
-
[1987]
ACM30, 11 (1987), 964–971
The vocabulary problem in human-system communication.Commun. ACM30, 11 (1987), 964–971
1987
-
[2001]
InText Retrieval Conference (TREC)
Data-Intensive Question Answering. InText Retrieval Conference (TREC). https://microsoft.com
-
[2002]
InProceedings of the Eleventh Text Retrieval Conference (TREC)
Extracting Answers from the Web Using Knowledge Annotation and Knowledge Mining Techniques. InProceedings of the Eleventh Text Retrieval Conference (TREC). 571–580
-
[2018]
In Proceedings of the 2018 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)
FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 809–819
2018
-
[2019]
InProceedings of the 42nd international acm sigir conference on research and development in information retrieval
Asking clarifying questions in open-domain information-seeking conver- sations. InProceedings of the 42nd international acm sigir conference on research and development in information retrieval. 475–484
-
[2020]
InInternational con- ference on machine learning
Retrieval augmented language model pre-training. InInternational con- ference on machine learning. PMLR, 3929–3938
-
[2023]
InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Rarr: Researching and revising what language models say, using lan- guage models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 16477–16508
-
[2025]
Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering.arXiv preprint arXiv:2510.24402(2025)
2025
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.