Pith. sign in

REVIEW 3 major objections 8 minor 140 references

Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective

T0 review · 3 major / 8 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read RAG's core ideas — retrieval plus generation, knowledge augmentation, answer verification, and iterative query refinement — were already built and tested in early-2000s IR and QA research, with 2002's QUALIFIER as an early Agentic RAG.

desk verdict A genuinely useful historical reframing of RAG that overreaches slightly on the QUALIFIER-to-Agentic-RAG analogy; the core thesis that most RAG components predate LLMs holds and the paper deserves a serious referee. read the letter →

arxiv 2608.08445 v1 pith:V4YARIZM submitted 2026-08-09 cs.AI

classification cs.AI
keywords retrieval-augmentedgenerationagenticRAGinformationretrievalquestionansweringTRECQAtrackQUALIFIERqueryrefinementuser-centeredevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that retrieval-augmented generation (RAG) is not a novel approach invented to fix the shortcomings of large language models, but a modern reconfiguration of a question-answering architecture that information retrieval (IR) researchers designed and tested two decades earlier. The authors walk through the TREC QA track from 1999 to 2007 and show that its systems already combined query reformulation, evidence retrieval, answer extraction, and answer validation — the same stages that structure today's RAG pipelines. Their central historical exhibit is the QUALIFIER system, whose successive constraint relaxation loop (retrieve, check for an answer, relax the query, retry) they read as an early instance of Agentic RAG. A sympathetic reader should care because, if the continuity holds, classical IR work stops being a curiosity and becomes a direct design source for next-generation RAG — especially for personalization, proactive interaction, governance, and user-centered evaluation — and the field can stop rediscovering old ideas under new names. The reframing is that LLMs are a new interface layer executing a decades-old QA architecture, not the origin of retrieval-augmented intelligence.

What carries the argument

The load-bearing object is the QUALIFIER system of 2002–2003, and in particular its successive constraint relaxation (SCR) loop. QUALIFIER modeled a question as a structured event with slots (time, location, agent, action), encoded the known slots as a deliberately strict Boolean query, and retrieved with the MG indexing system; when an over-constrained query returned no documents, it removed a portion of the expanded terms and retried, repeating for up to five iterations and returning NIL rather than a low-confidence guess if nothing survived. The paper's interpretive move is to see this relax-and-retry cycle as retrieval shaped by intermediate reasoning states — retrieve, assess whether the evidence suffices, adapt, retrieve again — which is the same loop that defines Agentic RAG in current systems. Supporting machinery includes the stage-by-stage mapping of the classical QA pipeline (question analysis, retrieval, answer synthesis with verification) onto modern RAG, QUALIFIER's Answer Event Score for ranking candidate passages, and the TREC 2002 benchmark result (290 of 500 questions correct, second to the logic-based PowerAnswer system) that anchors the claim that iterative refinement carried practical weight.

What would settle it

Re-run QUALIFIER's TREC 2002 task with the feedback loop disabled: submit the initial strict query once, and on an empty result issue a single fixed relaxation with no intermediate verification. If accuracy on the 500 questions stays at the reported 290 correct answers, the 'retrieval shaped by intermediate reasoning states' interpretation gains no support, and the proto-Agentic claim reduces to a mechanical artifact of Boolean retrieval.

Watch

Extended reading notes

Core claim

The central claim is that the core ideas underlying RAG, including Agentic RAG, are not new: integrating retrieval with language generation, augmenting knowledge with external sources, verifying answers against evidence, and iteratively refining queries or prompts were already studied and instantiated in IR and QA research dating back to the early 2000s, before large language models existed. The paper makes this case by mapping each stage of modern RAG onto a TREC-era antecedent: query rewriting and expansion onto classical relevance feedback and query expansion; structured querying onto QUALIFIER's event-based constraint encoding; evidence-conditioned generation onto retrieval-then-synthesis QA pipelines; answer verification onto mechanisms such as PIQUANT's sanity checker and web-redundancy scoring; and iterative, 'agentic' retrieval onto QUALIFIER's successive constraint relaxation, in which a maximally strict query is loosened only when it returns no answer. The authors' reading is that LLMs are best understood as a new interface layer on a decades-old architecture, and that this continuity has gone under-recognized because of community fragmentation, shifting terminology, and recency bias.

Load-bearing premise

The argument hinges on reading QUALIFIER's iterated 'retrieve, check, relax, retry' loop as a true forerunner of Agentic RAG; if that loop was merely a workaround forced by the fact that strict term-matching returns nothing when a query is too long, the strong historical thesis collapses into the far weaker claim that some RAG components existed earlier.

Editorial extensions

If this is right

  • Classical IR work becomes a direct design source for next-generation RAG: query expansion, relevance feedback, and structured queries are ready-made solutions to problems RAG is currently re-solving from scratch.
  • RAG evaluation should absorb IR's user-centered discipline — bias-aware feedback modeling, interaction studies, and long-term user-experience measures — rather than relying on static reference-based metrics alone.
  • Each of the paper's four proposed directions — personalized RAG, proactive RAG, governance-aware RAG, and user-centered evaluation — inherits an established IR research line (user modeling, query clarification, rule-based governance and auditing, click-bias-aware evaluation) that current RAG has underused.
  • Because QUALIFIER's iterative loop outperformed single-shot systems at TREC 2002 with no language model in the loop, the account implies that iterative 'agentic' retrieval has measurable value independent of LLM capability — a claim worth testing directly in modern stacks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable extension the authors do not pursue: transplant QUALIFIER's constrain-then-relax loop onto a modern retriever–LLM stack and compare against single-pass RAG; if the loop still buys accuracy, the historical continuity is functional rather than merely conceptual.
  • Implicit in the 'LLMs as interface layer' framing is a prediction the authors do not state: if the field absorbs this history, pre-LLM QA and IR citations should start appearing in RAG papers as design justifications rather than background, and a few years of citation data could measure whether the field actually stops rediscovering.
  • The historical thesis plausibly extends to a neighbouring problem: agentic tool-use and web-search agents are reconfigurations of interactive IR ideas (mixed-initiative interaction, relevance feedback, session modeling), so the same 'forgotten history' pattern may hold across the broader agentic web agenda.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. This paper argues that the core ideas underlying modern RAG and Agentic RAG—integrating retrieval and generation, knowledge augmentation, answer verification, and iterative query/refinement—were already studied and instantiated in IR and QA research from the early 2000s, particularly in TREC QA systems such as QUALIFIER. It traces the conceptual lineage from classical retrieval-based QA through RAG to Agentic RAG, proposes viewing LLMs as a new interface layer atop a decades-old QA architecture, and identifies four IR-informed future directions (personalization, proactive interaction, governance, and user-centered evaluation). The historical narrative is grounded in TREC proceedings and Voorhees's overviews, and the paper includes a conceptual mapping diagram and a comparison of TREC 2002 results.

Significance. If the historical thesis holds, the paper provides a valuable corrective to the prevailing LLM-centric narrative of RAG: it identifies concrete antecedents for RAG's modular pipeline, surfaces underutilized IR work that could inform next-generation RAG design, and offers a falsifiable account of continuity. The paper is well referenced on the historical facts, and the forward-looking Section 5 is constructive and clearly tied to established IR concerns. However, the most distinctive claim—that QUALIFIER is a genuine precursor to Agentic RAG, so that even the agentic variant is 'not new'—rests on an interpretation of QUALIFIER's successive constraint relaxation that the manuscript itself renders ambiguous. The comparative performance evidence in Section 3.2 is also used to draw a causal conclusion that the presented results do not support. These issues are central to the strength of the historical claim, though fixable within the manuscript's scope.

major comments (3)
  1. [§3.1, §3.2] The claim that QUALIFIER is an early Agentic RAG system depends on whether its successive constraint relaxation (SCR) was triggered by failed answer verification or by an empty Boolean hit set. The manuscript states in §3.1 that 'Because Boolean retrieval only returns documents matching all query terms, an overly expanded query could return no results at all. To handle this, they introduced SCR,' which suggests a mechanical recall-repair fallback rather than retrieval 'iteratively shaped by intermediate reasoning states' (§3.1). Please report directly from [108] and [110] what condition actually triggered a relaxation: zero retrieved documents, failure to extract any candidate answer, or explicit constraint/verification failure. If SCR fired only on empty hit sets or empty extraction output, then the proto-Agentic framing is overstated and should be weakened to 'iterative query refinement for recall recovery.'
  2. [§3.2, Table 1] The sentence 'QUALIFIER's competitive performance suggests that agentic RAG with iterative refinement can offer practical advantages over simple, single-pass RAG' is not supported by the evidence presented. In Table 1, QUantifier ranks second with 290/500, but the top system, LCC PowerAnswer (415/500), is not characterized as a simple single-pass system and does not fit the paper's iterative-refinement frame. Without an ablation or process analysis that isolates the effect of the iterative loop from other QUALIFIER design choices (e.g., the structured event model, external resources, answer justification), the performance comparison cannot carry this causal conclusion. Please qualify the claim to 'competitive at the time' or provide direct evidence tying the iterative mechanism to the measured outcome.
  3. [Abstract, §2.1] The claim that 'integrating retrieval and language generation' was 'instantiated' in early 2000s QA systems needs a definitional guard. The systems described perform extractive answer selection, pattern-based candidate ranking, and template-style answer construction; none performs open-ended text generation in the LLM sense. If 'generation' is intended broadly to cover any evidence-conditioned answer synthesis, the claim is less distinctive; if it is intended to match modern generative language modeling, the historical examples do not instantiate it. The paper should explicitly state which sense of 'generation' it uses and adjust the continuity claim in the abstract and Section 2.1 accordingly.
minor comments (8)
  1. [§3.1] The system name is misspelled as 'QUALIFER' in the sentence 'QUALIFER's underlying design philosophy' in Section 3.1.
  2. [Table 1] The column header 'Answer Correctness' is unclear; it should read 'Number of correctly answered questions (out of 500)' or equivalent.
  3. [§3.2] The phrase 'dominated the field for the next decade, before community QA approaches emerged' is vague and does not align with the TREC QA track timeline (1999–2007); please clarify the intended period and what 'community QA approaches' refers to.
  4. [References [108]–[111]] Several of the key QUALIFIER references share a co-author with the current paper, yet the manuscript does not disclose this relationship; a brief acknowledgment or neutral note would help readers calibrate the interpretive weight placed on this specific system.
  5. [§2.5, Figure 1] Figure 1 is not explicitly cited where the 'two levels of conceptual continuity' are discussed in Section 2.5; please add a reference to the figure at that point.
  6. [§2.2, Ref [71]] Characterizing LLMs as 'a huge knowledge base' by citing [71] is a contested framing; please qualify it (e.g., 'can be viewed as approximate knowledge stores') to avoid overstating the analogy.
  7. [§4] The sentence 'The main difference lies not in the underlying problem structure, but in the implementation layer' appears in both Section 2.1 and Section 4; consider keeping only one instance or adding a cross-reference.
  8. [§5.3.2] The heading 'Governance-A ware RAG' contains a typo; it should read 'Governance-Aware RAG.'

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: QUALIFIER's behavior and TREC results are externally documented, so the historical thesis does not reduce to its inputs.

full rationale

This is a historical and interpretive position paper, not a derivation with fitted parameters or predicted quantities. No equation is reused as an output, and no parameter is fitted to a subset of data and then reported as a prediction. The central claim — that core RAG ideas, including iterative query refinement, existed in early IR/QA work — rests on public artifacts: TREC QA track proceedings, the official Voorhees overview of TREC 2002 results (ref [96]), and the QUALIFIER system papers (refs [108–111]). Although some of those system papers are authored by co-author Tat-Seng Chua, the cited behavior (Boolean retrieval over the MG index, successive constraint relaxation triggered when over-expanded queries return no hits, up to five relax-and-retry iterations, and NIL on failure) is publicly documented in NIST proceedings and externally evaluated, so the self-citation is verifiable evidence rather than a self-referential premise. The skeptic's objection that SCR may be an artifact of strict Boolean matching is a challenge to the strength of the historical analogy, not evidence of circularity; the paper's claim does not become true by definition or by the authors' own assertion. Under the hard rules, externally documented and falsifiable cited results count as independent support, so no circular step is identified and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's argument rests on interpretive domain assumptions about the equivalence of historical QA components and modern RAG components, and on the selection of QUALIFIER as the key example. No mathematical axioms or fitted parameters are used. The main assumption is that component-level analogies (query rewriting = query reformulation, answer validation = faithfulness control) are meaningful enough to support the continuity claim.

assumptions (3)
  • domain assumption The components of early TREC QA systems correspond to the stages of modern RAG pipelines.
    This mapping is the central interpretive claim of Section 2.1; it is not proven, only illustrated with selected examples.
  • domain assumption QUALIFIER's iterative constraint relaxation is functionally equivalent to the agentic RAG loop.
    Section 3.2 argues this; it is the load-bearing analogy. The paper does not provide a formal definition of 'agentic' against which to test it.
  • domain assumption The intellectual lineage can be traced through the specific systems selected (JAVELIN, MultiText, DIOGENE, QUALIFIER, etc.).
    Selection of examples is non-exhaustive and may be biased toward those that fit the narrative of continuity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective." pith.science (2026). https://pith.science/paper/V4YARIZM

@misc{pith2026260808445,
  author       = {Pith},
  title        = {Pith review of: Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V4YARIZM}},
  note         = {Machine review of arXiv:2608.08445}
}
read the original abstract

Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core ideas underlying RAG are not new: foundational concepts such as integrating retrieval and language generation, knowledge augmentation, answer verification, and iterative query (or prompt) refinement had already been studied and instantiated in information retrieval (IR) and question answering (QA) research dating back to the early 2000s, well before the emergence of LLMs. We make this case by systematically tracing the intellectual lineage of modern RAG and Agentic RAG back to their classical IR and QA antecedents, and examining why this continuity has gone under-recognized -- a consequence of community fragmentation, shifting terminology, and the recency bias endemic to fast-moving fields. Rather than treating LLMs as the origin point of retrieval-augmented intelligence, we propose viewing them as a new interface layer atop a decades-old QA architecture. This reframing is not merely historical: by situating RAG within the longer trajectory of IR research, we surface underutilized prior work -- on user modeling, answer validation, and query refinement -- that can directly inform next-generation RAG design, reducing unintentional rediscovery and fostering genuine cross-community integration.

Figures

Figures reproduced from arXiv: 2608.08445 by the authors.

Figure 1
Figure 1. Conceptual continuity from classical QA to modern RAG. Classical QA and RAG share a single-pass retrieval–answering [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Four IR-informed directions for future RAG: personalization, proactive interaction, governance-aware architectures, [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

140 extracted references · 45 canonical work pages

  1. [108]

    Voorhees

    Ellen M. Voorhees. 2002. Overview of the TREC 2002 Question Answering Track. InProceedings of the Eleventh Text REtrieval Conference (TREC 2002), Ellen M. Voorhees and Donna K. Harman (Eds.). National Institute of Standards and Technology (NIST), 42–51

  2. [110]

    Nikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanoulas, and Maarten De Ri- jke. 2020. Query resolution for conversational search with limited supervision. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 921–930

  3. [1]

    Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, and W Bruce Croft

  4. [2]

    Gianni Amati and Cornelis Joost Van Rijsbergen. 2002. Probabilistic models of information retrieval based on measuring the divergence from randomness. ACM Transactions on Information Systems (TOIS)20, 4 (2002), 357–389

  5. [3]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. (2024)

  6. [4]

    Sumit Bhatia, Debapriyo Majumdar, and Prasenjit Mitra. 2011. Query sugges- tions in the absence of query logs. InProceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. 795–804

  7. [5]

    Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Bal- dini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, et al. 2022. Attributed question answering: Evaluation and modeling for attributed large language models.arXiv preprint arXiv:2212.08037 (2022)

  8. [6]

    Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Ruther- ford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bog- dan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. InInternational conference on machine learning. PMLR, 2206–2240

Show all 140 references
  1. [7]

    Eric Brill, Susan Dumais, and Michele Banko. 2002. An analysis of the AskMSR question-answering system. InProceedings of the 2002 conference on empirical methods in natural language processing (EMNLP 2002). 257–264

  2. [8]

    Eric Brill, Jimmy Lin, Michele Banko, Susan T Dumais, Andrew Y Ng, et al. 2001. Data-Intensive Question Answering.. InTREC, Vol. 56. 90

  3. [9]

    Lin, Michele Banko, Susan T

    Eric Brill, Jimmy J. Lin, Michele Banko, Susan T. Dumais, and Andrew Y. Ng

  4. [10]

    Fei Cai, Maarten De Rijke, et al. 2016. A survey of query auto completion in information retrieval.Foundations and Trends®in Information Retrieval10, 4 (2016), 273–363

  5. [11]

    James P Callan, W Bruce Croft, and Stephen M Harding. 1992. The INQUERY retrieval system. InDatabase and Expert Systems Applications: Proceedings of the International Conference in Valencia, Spain, 1992. Springer, 78–83

  6. [12]

    Jaime Carbonell and Jade Goldstein. 1998. The use of MMR, diversity-based reranking for reordering documents and producing summaries. InProceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. 335–336

  7. [13]

    Claudio Carpineto and Giovanni Romano. 2012. A survey of automatic query expansion in information retrieval.Acm Computing Surveys (CSUR)44, 1 (2012), 1–50

  8. [14]

    Olivier Chapelle and Ya Zhang. 2009. A dynamic bayesian network click model for web search ranking. InProceedings of the 18th international conference on World wide web. 1–10

  9. [15]

    Prager, Chris A

    Jennifer Chu-Carroll, John M. Prager, Chris A. Welty, David Ferrucci, and David A. Boguraev. 2002. A Multi-Strategy and Multi-Source Approach to Question Answering. InProceedings of the Eleventh Text Retrieval Conference, TREC 2002, Gaithersburg, Maryland, USA, November 19-22,...

  10. [16]

    Charles LA Clarke, Gordon V Cormack, and Thomas R Lynam. 2001. Exploiting redundancy in question answering. InProceedings of the 24th annual inter- national ACM SIGIR conference on Research and development in information retrieval. 358–365

  11. [17]

    Charles L. A. Clarke, Gordon V. Cormack, Graeme Kemkes, M. Laszlo, Thomas R. Lynam, Egidio L. Terra, and Philip L. Tilker. 2002. Statistical Selection of Exact Answers (MultiText Experiments for TREC 2002). InProceedings of the Eleventh Text REtrieval Conference, TREC 2002, Ga...

  12. [18]

    Nick Craswell, Onno Zoeter, Michael Taylor, and Bill Ramsey. 2008. An exper- imental comparison of click position-bias models. InProceedings of the 2008 international conference on web search and data mining. 87–94

  13. [19]

    W Bruce Croft. 1995. What do people want from information retrieval.D-Lib magazine1, 5 (1995)

  14. [20]

    Steve Cronen-Townsend, Yun Zhou, and W Bruce Croft. 2004. A framework for selective query expansion. InProceedings of the thirteenth ACM international conference on Information and knowledge management. 236–237

  15. [21]

    Michail Dadopoulos, Anestis Ladas, Stratos Moschidis, and Ioannis Negkakis

  16. [22]

    Fernando Diaz, Bhaskar Mitra, and Nick Craswell. 2016. Query Expansion with Locally-Trained Word Embeddings. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 367–377

  17. [23]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query- focused summarization.arXiv preprint arXiv:2404.16130(2024)

  18. [24]

    Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. 2019. Can you unpack that? learning to rewrite questions-in-context.Can You Unpack That? Learning to Rewrite Questions-in-Context(2019)

  19. [25]

    Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining. ...

  20. [26]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse lexical and expansion model for first stage ranking. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2288–2292

  21. [27]

    Furnas, Thomas K

    George W. Furnas, Thomas K. Landauer, Louis M. Gomez, and Susan T. Dumais

  22. [28]

    Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Cha- ganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al

  23. [29]

    Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023. Precise zero- shot dense retrieval without relevance labels. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1762–1777

  24. [30]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997 2, 1 (2023)

  25. [31]

    Gurbinder Gill, Ritvik Gupta, Denis Lusson, Anand Chandrashekar, and Donald Nguyen. 2025. From Search to Reasoning: a Five-Level Rag Capability Frame- work for Enterprise Data. In2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware). IEEE, 01–09

  26. [32]

    Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Naik, Pengshan Cai, and Alfio Gliozzo. 2022. Re2G: Retrieve, rerank, generate. InPro- ceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...

  27. [33]

    Pei Guo, Enjie Liu, Ruichao Zhong, Mochi Gao, Yunzhi Tan, Bo Hu, and Zang Li

  28. [34]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang

  29. [35]

    Kailash A Hambarde and Hugo Proenca. 2023. Information retrieval: recent advances and beyond.IEEE Access11 (2023), 76581–76604

  30. [36]

    Marti A Hearst. 2006. Clustering versus faceted categories for information exploration.Commun. ACM49, 4 (2006), 59–61

  31. [37]

    Ulf Hermjakob, Abdessamad Echihabi, and Daniel Marcu. 2002. Natural Lan- guage Based Reformulation Resource and Wide Exploitation for Question An- swering. InProceedings of the Eleventh Text REtrieval Conference (TREC 2002). https://trec.nist.gov/pubs/trec11/papers/usc.hermjakob.pdf

  32. [38]

    Jerry R Hobbs. 1978. Resolving pronoun references.Lingua44, 4 (1978), 311– 338

  33. [39]

    InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)

    DSRAG: A Double-Stream Retrieval-Augmented Generation Framework for Countless Intent Detection. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track). 318–328

  34. [40]

    2005.The turn: Integration of information seeking and retrieval in context

    Peter Ingwersen and Kalervo Järvelin. 2005.The turn: Integration of information seeking and retrieval in context. Springer

  35. [41]

    Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. InProceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume. 874–880

  36. [42]

    Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models.Journal of Machine Learning Research24, 251 (2...

  37. [43]

    Thorsten Joachims. 2002. Optimizing search engines using clickthrough data. InProceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining. 133–142

  38. [44]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering.. InEMNLP (1). 6769–6781

  39. [45]

    Jaana Kekäläinen and Kalervo Järvelin. 2000. The co-effects of query struc- ture and expansion on retrieval performance in probabilistic text retrieval. Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective Information retrieval1, 4 (2000), 329–344

  40. [46]

    1992.Information retrieval interaction

    Peter Ingwersen. 1992.Information retrieval interaction. Vol. 246. Taylor Graham London

  41. [47]

    To Eun Kim and Fernando Diaz. 2025. Towards fair rag: On the impact of fair ranking in retrieval-augmented generation. InProceedings of the 2025 Interna- tional ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR). 33–43

  42. [48]

    Jon M Kleinberg. 1999. Authoritative sources in a hyperlinked environment. Journal of the ACM (JACM)46, 5 (1999), 604–632

  43. [49]

    Cody CT Kwok, Oren Etzioni, and Daniel S Weld. 2001. Scaling question answering to the web. InProceedings of the 10th international conference on World Wide Web. 150–161

  44. [50]

    Shalom Lappin and Herbert J Leass. 1994. An algorithm for pronominal anaphora resolution.Computational linguistics20, 4 (1994), 535–561

  45. [51]

    Victor Lavrenko and W Bruce Croft. 2001. Relevance-Based Language Models. (2001)

  46. [52]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing ...

  47. [53]

    Diane Kelly et al. 2009. Methods for evaluating interactive information retrieval systems with users.Foundations and Trends®in Information Retrieval3, 1–2 (2009), 1–224

  48. [54]

    Jintao Liang, Huifeng Lin, You Wu, Rui Zhao, Ziyue Li, et al. 2025. Reasoning rag via system 1 or system 2: A survey on reasoning agentic retrieval-augmented generation for industry challenges. InProceedings of the 14th International Joint Conference on Natural Language Proces...

  49. [55]

    Hui Lin and Jeff Bilmes. 2011. A class of submodular functions for document summarization. InProceedings of the 49th annual meeting of the association for computational linguistics: human language technologies. 510–520

  50. [56]

    Jimmy Lin, Aaron Fernandes, Boris Katz, Gregory Marton, and Stefanie Tellex

  51. [57]

    Tie-Yan Liu et al. 2009. Learning to rank for information retrieval.Foundations and Trends®in Information Retrieval3, 3 (2009), 225–331

  52. [58]

    Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting in retrieval-augmented large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5303– 5315

  53. [59]

    Bernardo Magnini, Matteo Negri, Roberto Prevete, and Hristo Tanev. 2002. Mining Knowledge from Repeated Co-Occurrences: DIOGENE at TREC 2002. In Proceedings of the Eleventh Text REtrieval Conference (TREC-2002). NIST Special Publication, 577–586

  54. [60]

    Zhicong Li, Jiahao Wang, Zhishu Jiang, Hangyu Mao, Zhongxia Chen, Jiazhen Du, Yuanxing Zhang, Fuzheng Zhang, Di Zhang, and Yong Liu. 2024. Dmqr-rag: Diverse multi-query rewriting for rag.arXiv preprint arXiv:2411.13154(2024)

  55. [61]

    Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, and Weizhu Chen. 2021. Generation-augmented retrieval for open-domain question answering. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inter...

  56. [62]

    Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al. 2022. Teaching language models to support answers with verified quotes.arXiv preprint arXiv:2203.11147(2022)

  57. [63]

    George A Miller. 1992. WordNet: A Lexical Database for English. InSpeech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992

  58. [64]

    Mandar Mitra, Amit Singhal, and Chris Buckley. 1998. Improving automatic query expansion. InProceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. 206–214

  59. [65]

    Moldovan, S

    D. Moldovan, S. Harabagiu, R. Girju, P. Morarescu, F. Lacatusu, A. Novischi, A. Badulescu, and O. Bolohan. 2001. LCC Tools for Question Answering. In Proceedings of the Tenth Text REtrieval Conference (TREC 2001). National Institute of Standards and Technology (NIST), 111–120

  60. [66]

    Dan Moldovan, Marius Pasca, Sanda Harabagiu, and Mihai Surdeanu. 2002. Performance Issues and Error Analysis in an Open-Domain Question Answer- ing System. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 33–40

  61. [67]

    Marco Morik, Ashudeep Singh, Jessica Hong, and Thorsten Joachims. 2020. Controlling fairness and bias in dynamic learning-to-rank. InProceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval. 429–438

  62. [68]

    2008.Introduction to information retrieval

    Christopher D Manning. 2008.Introduction to information retrieval. Syngress Publishing. Chapter 9: Relevance feedback and query expansion

  63. [69]

    Eric Nyberg, Teruko Mitamura, Jaime Carbonell, Jamie Callan, Kevyn Collins- Thompson, Karen Czuba, Michael Duggan, L Hiyakumoto, N Hu, Y Huang, et al

  64. [70]

    Alexandra Olteanu, Jean Garcia-Gathright, Maarten de Rijke, Michael D Ek- strand, Adam Roegiest, Aldo Lipani, Alex Beutel, Alexandra Olteanu, Ana Lucic, Ana-Andreea Stoica, et al. 2021. FACTS-IR: fairness, accountability, confiden- tiality, transparency, and safety in informat...

  65. [71]

    Miller, and Sebastian Riedel

    Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2019. Language Models as Knowledge Bases? arXiv:1909.01066 [cs.CL] https://arxiv.org/abs/1909.01066

  66. [72]

    Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. InFindings of the Association for Computational Linguistics: EMNLP

  67. [73]

    Yonggang Qiu and Hans-Peter Frei. 1993. Concept based query expansion. In Proceedings of the 16th annual international ACM SIGIR conference on Research and development in information retrieval. 160–169

  68. [74]

    Zackary Rackauckas. 2024. Rag-fusion: a new take on retrieval-augmented generation.arXiv preprint arXiv:2402.03367(2024)

  69. [75]

    Dragomir Radev, Weiguo Fan, Hong Qi, Harris Wu, and Amardeep Grewal

  70. [76]

    Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332(2021)

  71. [77]

    Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-context retrieval-augmented lan- guage models.Transactions of the Association for Computational Linguistics11 (2023), 1316–1331

  72. [78]

    InProceedings of the Eleventh Text Retrieval Conference (TREC 2002)

    The JAVELIN Question-Answering System at TREC 2002. InProceedings of the Eleventh Text Retrieval Conference (TREC 2002). National Institute of Standards and Technology (NIST), 112–121

  73. [79]

    Joseph John Rocchio Jr. 1971. Relevance feedback in information retrieval.The SMART retrieval system: experiments in automatic document processing(1971)

  74. [80]

    Alireza Salemi, Surya Kallumadi, and Hamed Zamani. 2024. Optimization meth- ods for personalizing large language models through retrieval augmentation. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 752–762

  75. [81]

    Alireza Salemi and Hamed Zamani. 2024. Towards a search engine for machines: Unified ranking for multiple retrieval-augmented large language models. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 741–751

  76. [82]

    Gerard Salton and Chris Buckley. 1990. Improving retrieval performance by relevance feedback.Journal of the American society for information science41, 4 (1990), 288–297

  77. [83]

    Gerard Salton, Edward A Fox, and Harry Wu. 1983. Extended boolean informa- tion retrieval.Commun. ACM26, 11 (1983), 1022–1036

  78. [84]

    Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher D Manning. 2024. Raptor: Recursive abstractive processing for tree-organized retrieval. InThe Twelfth International Conference on Learning Representations

  79. [85]

    InProceedings of the 11th international conference on World Wide Web

    Probabilistic question answering on the web. InProceedings of the 11th international conference on World Wide Web. 408–419

  80. [86]

    Dragomir R Radev, Hongyan Jing, Małgorzata Styś, and Daniel Tam. 2004. Centroid-based summarization of multiple documents.Information Processing & Management40, 6 (2004), 919–938

  81. [87]

    Xuehua Shen, Bin Tan, and ChengXiang Zhai. 2005. Implicit user modeling for personalized search. InProceedings of the 14th ACM international conference on Information and knowledge management. 824–831

  82. [88]

    Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond.Foundations and trends®in information retrieval 3, 4 (2009), 333–389

  83. [89]

    Aditi Singh, Abul Ehtesham, Saket Kumar, and Tala Talaei Khoei. 2025. Agen- tic retrieval-augmented generation: A survey on agentic rag.arXiv preprint arXiv:2501.09136(2025)

  84. [90]

    Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William Cohen. 2018. Open domain question answer- ing using early fusion of knowledge bases and text. InProceedings of the 2018 conference on empirical methods in natural language processin...

  85. [91]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal

  86. [92]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal

  87. [93]

    Howard Turtle and W Bruce Croft. 1989. Inference networks for document retrieval. InProceedings of the 13th annual international ACM SIGIR conference on Research and development in information retrieval. 1–24

  88. [94]

    Ellen M Voorhees. 1994. Query expansion using lexical-semantic relations. InSIGIR’94: Proceedings of the Seventeenth Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, organised by Dublin City University. Springer, 61–69

  89. [95]

    Mahmoud F Sayed and Douglas W Oard. 2019. Jointly modeling relevance and sensitivity for search among sensitive content. InProceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval. 615–624

  90. [96]

    Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy. InFindings of the Association for Com- putational Linguistics: EMNLP 2023. 9248–9274

  91. [97]

    Voorhees and Dawn M

    Ellen M. Voorhees and Dawn M. Tice. 2000. The TREC-8 Question Answer- ing Track. InProceedings of the Second International Conference on Language Resources and Evaluation (LREC 2000)

  92. [98]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. InFindings of the Association for Computational Linguistics: EMNLP 2021. 3784–3803

  93. [99]

    Jonas Wallat, Maria Heuss, Maarten de Rijke, and Avishek Anand. 2025. Cor- rectness is not Faithfulness in Retrieval Augmented Generation Attributions. InProceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (IC...

  94. [100]

    Haoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian, Zhengyang Wang, Chen Luo, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, and Jing Gao. 2024. Blend- filter: Advancing retrieval-augmented large language models via query genera- tion blending and knowledge filtering. InProceeding...

  95. [101]

    Lidan Wang, Jimmy Lin, and Donald Metzler. 2011. A cascade ranking model for efficient ranked retrieval. InProceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. 105–114

  96. [102]

    Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query Expansion with Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 9414–9423

  97. [103]

    Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor nega- tive contrastive learning for dense text retrieval.arXiv preprint arXiv:2007.00808 (2020)

  98. [104]

    InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers)

    Interleaving retrieval with chain-of-thought reasoning for knowledge- intensive multi-step questions. InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers). 10014–10037

  99. [105]

    Jinxi Xu and W Bruce Croft. 1996. Query expansion using local and global document analysis. InProceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrieval. 4–11

  100. [106]

    Xiwei Xu, Hans Weytjens, Dawen Zhang, Qinghua Lu, Ingo Weber, and Liming Zhu. 2025. RAGOps: Operating and Managing Retrieval-Augmented Generation Pipelines.arXiv preprint arXiv:2506.03401(2025)

  101. [107]

    Voorhees

    Ellen M. Voorhees. 2001. The TREC Question Answering Track.Natural Language Engineering7, 4 (2001), 361–378

  102. [109]

    Hui Yang and Tat-Seng Chua. 2003. Qualifier: question answering by lexical fabric and external resources. In10th Conference of the European Chapter of the Association for Computational Linguistics

  103. [111]

    Hui Yang, Hang Cui, Mstislav Maslennikov, Long Qiu, Min-Yen Kan, and Tat- Seng Chua. 2003. QUALIFIER in TREC-12 QA Main Task. InProceedings of the Twelfth Text REtrieval Conference (TREC 2003). NIST

  104. [112]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. InThe eleventh international conference on learning represen- tations

  105. [113]

    Raquib Bin Yousuf, Shengzhe Xu, Mandar Sharma, Andrew Neeser, Chris La- timer, and Naren Ramakrishnan. 2026. Utilizing Metadata for Better Retrieval- Augmented Generation.arXiv preprint arXiv:2601.11863(2026)

  106. [114]

    Yue Yu, Wei Ping, Zihan Liu, Boxin Wang, Jiaxuan You, Chao Zhang, Moham- mad Shoeybi, and Bryan Catanzaro. 2024. Rankrag: Unifying context ranking with retrieval-augmented generation in llms.Advances in Neural Information Processing Systems37 (2024), 121156–121184

  107. [115]

    Hamed Zamani and W Bruce Croft. 2017. Relevance-based word embedding. InProceedings of the 40th international acm sigir conference on research and development in information retrieval. 505–514

  108. [116]

    W Xiong, P Lewis, S Riedel, XL Li, W Wang, S Iyer, Y Mehdad, D Kiela, J Du, WT Yih, et al. 2020. Answering Complex Open-Domain Questions with Multi- Hop Dense Retrieval. InICLR 2021-9th International Conference on Learning Representations, Vol. 2021. ICLR

  109. [117]

    Hamed Zamani, Susan Dumais, Nick Craswell, Paul Bennett, and Gord Lueck

  110. [118]

    Hamed Zamani, Bhaskar Mitra, Everest Chen, Gord Lueck, Fernando Diaz, Paul N Bennett, Nick Craswell, and Susan T Dumais. 2020. Analyzing and learning from user interactions for search clarification. InProceedings of the 43rd international acm sigir conference on research and d...

  111. [119]

    Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. (2024)

  112. [120]

    Hui Yang and Tat-Seng Chua. 2002. The Integration of Lexical Knowledge and External Resources for Question Answering. InProceedings of the Eleventh Text REtrieval Conference (TREC 2002). Gaithersburg, Maryland, USA, 59–69

  113. [121]

    Kepu Zhang, Teng Shi, Weijie Yu, and Jun Xu. 2025. Prlm: Learning explicit reasoning for personalized rag via contrastive reward optimization. InProceed- ings of the 34th ACM International Conference on Information and Knowledge Management. 5484–5488

  114. [122]

    Hui Yang, Tat-Seng Chua, Shuguang Wang, and Chun-Keat Koh. 2003. Struc- tured use of external knowledge for event-based open domain question answer- ing. InProceedings of the 26th annual international ACM SIGIR conference on Research and development in information retrieval. 33–40

  115. [123]

    Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi, Quoc V Le, and Denny Zhou. 2023. Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models.CoRR(2023)

  116. [124]

    Zhiping Zheng. 2002. AnswerBus question answering system. InHuman Lan- guage Technology Conference (HLT 2002), Vol. 27

  117. [125]

    Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al . 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625(2022)

  118. [126]

    Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chen- long Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, and Ji-Rong Wen. 2023. Large language models for information retrieval: A survey.arXiv preprint arXiv:2308.07107(2023)

  119. [127]

    2025.{PoisonedRAG}: Knowledge corruption attacks to{Retrieval-Augmented} generation of large language models

    Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025.{PoisonedRAG}: Knowledge corruption attacks to{Retrieval-Augmented} generation of large language models. In34th USENIX Security Symposium (USENIX Security 25). 3827–3844

  120. [128]

    Hamed Zamani, Fernando Diaz, Mostafa Dehghani, Donald Metzler, and Michael Bendersky. 2022. Retrieval-enhanced machine learning. InProceedings of the 45th international ACM SIGIR conference on research and development in infor- mation retrieval. 2875–2886

  121. [130]

    InProceedings of the web conference 2020

    Generating clarifying questions for information retrieval. InProceedings of the web conference 2020. 418–428

  122. [132]

    George Zerveas, Navid Rekabsaz, Daniel Cohen, and Carsten Eickhoff. 2022. Mitigating bias in search results through contextual document reranking and neutrality regularization. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Informat...

  123. [133]

    Chengxiang Zhai and John Lafferty. 2001. Model-based feedback in the lan- guage modeling approach to information retrieval. InProceedings of the tenth international conference on Information and knowledge management. 403–410

  124. [135]

    Chen Zhao, Chenyan Xiong, Jordan Boyd-Graber, and Hal Daumé III. 2021. Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...

  125. [1987]

    ACM30, 11 (1987), 964–971

    The vocabulary problem in human-system communication.Commun. ACM30, 11 (1987), 964–971

  126. [2001]

    InText Retrieval Conference (TREC)

    Data-Intensive Question Answering. InText Retrieval Conference (TREC). https://microsoft.com

  127. [2002]

    InProceedings of the Eleventh Text Retrieval Conference (TREC)

    Extracting Answers from the Web Using Knowledge Annotation and Knowledge Mining Techniques. InProceedings of the Eleventh Text Retrieval Conference (TREC). 571–580

  128. [2018]

    In Proceedings of the 2018 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)

    FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 809–819

  129. [2019]

    InProceedings of the 42nd international acm sigir conference on research and development in information retrieval

    Asking clarifying questions in open-domain information-seeking conver- sations. InProceedings of the 42nd international acm sigir conference on research and development in information retrieval. 475–484

  130. [2020]

    InInternational con- ference on machine learning

    Retrieval augmented language model pre-training. InInternational con- ference on machine learning. PMLR, 3929–3938

  131. [2023]

    InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

    Rarr: Researching and revising what language models say, using lan- guage models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 16477–16508

  132. [2025]

    Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering.arXiv preprint arXiv:2510.24402(2025)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.