REVIEW 4 major objections 6 minor 107 references
Adaptive Orchestration of Modular Generative Information Access Systems
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that generative information access systems should reconfigure their processing pipeline per query, and demonstrates the idea with a contextual bandit that beats a static pipeline.
desk verdict The framework and taxonomy are genuinely useful, but the time-aware AQA result contradicts the paper's own reward equation, so the experimental claim should not be trusted as reported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a graph-based representation of pipelines combined with a contextual bandit. Each pipeline is a directed acyclic graph; nodes are task modules (NoR, OneR, IRCoT, Aggregate), executors (Flan-T5-XL as LLM agent, BM25 as retriever, a majority-vote aggregator), and resources (Wikipedia and a multihop passage corpus), and edges encode both the flow of information and the assignment of executors and resources to tasks. The set of seven valid graphs forms the discrete action set of a LinUCB contextual bandit, with the query's complexity label as the context vector and a reward $r_t = \beta P_t - (1-\beta) T_t$ that makes the effectiveness-efficiency trade-off explicit. The static baseline, GPTSwarm, is optimized with REINFORCE on the same module set and converges to one fixed graph, which cannot specialize by query.
What would settle it
Run LinUCB on the same seven graphs but with complexity labels predicted by a lightweight classifier (or with raw query text features only), and compare on the same test set; if per-query selection no longer beats the static GPTSwarm graph on F1 and latency, the paper's evidence for adaptive orchestration over static optimization collapses. Alternatively, collect user-judged complexity labels and show they disagree with the Flan-T5-XL labels, breaking the mapping the bandit learns.
Extended reading notes
Core claim
The central claim is that the architecture of future generative information access systems will be dynamic and self-organizing: for each user input, the system should decide which tasks to run, which executors (agents or tools) should perform them, and which resources to consult, instead of applying a one-size-fits-all pipeline. The paper formalizes this as a directed acyclic graph whose nodes are module types and whose edges encode ordering and assignment, and it treats pipeline construction as an optimization problem with a composite objective over effectiveness and cost. In the proof-of-concept instantiation, LinUCB learns to map question complexity to one of seven valid graphs: simple questions are routed to direct generation (NoR), harder ones to single-step retrieval (OneR) or interleaved retrieval-chain-of-thought (IRCoT), and an aggregate node merges answers. With F1 and time as rewards, the adaptive system reaches 0.697 F1 at 9.99 ms (time-agnostic) or 0.687 F1 at 8.89 ms (time-based), while the static GPTSwarm-optimized graph reaches 0.502 F1 at 12.78 ms. The authors conclude that adaptive orchestration can construct pipelines adapted to incoming questions and optimized for a stated efficiency-effectiveness trade-off.
Load-bearing premise
The demonstrated gains rely on the bandit receiving the dataset's gold complexity labels (A/B/C) as context, a signal real systems do not have at query time and that was produced by an LLM annotator rather than validated against user needs.
Editorial extensions
If this is right
- The same bandit-based selection generalizes to any modular system whose feasible pipeline graphs can be enumerated, making the framework applicable beyond question answering to retrieval, conversational search, and tool-use systems.
- Adaptive orchestration can incorporate cost beyond latency, such as compute, API fees, and environmental impact, because any such cost can enter the composite reward.
- Systems that adapt per query can integrate newly added modules over time: a new module only adds nodes and graphs to the action space, while the bandit re-learns their utility.
- The demonstrated gap between adaptive and static orchestration (F1 0.697 vs 0.502, lower latency) implies that static pipeline optimization leaves substantial accuracy and efficiency on the table when query complexity varies.
- Scaling to large module sets will require planning or general RL methods rather than exhaustive graph enumeration, as the paper notes.
Reading between the lines
- The paper's use of gold complexity labels as context is a favorable condition: in deployment, a complexity predictor trained without such labels would be needed, and the reported advantage over static orchestration may shrink if that predictor is noisy.
- Because the action space is fixed to seven whole graphs, 'adaptivity' here is strategy selection rather than true on-the-fly construction; a more granular formulation with node-by-node additions could support pipelines that branch based on intermediate retrieval results, which the paper lists as an extension.
- A direct test of the vision would replace the gold labels with a learned query-complexity classifier, or with bandit features derived purely from the query text, and check whether per-query selection still beats a static graph.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that future generative information access (GenIA) systems should be modular and adaptively orchestrated, constructing a new pipeline graph for each user query. It proposes a graph-based framework whose nodes are tasks, executors, and resources, and suggests planning and reinforcement-learning methods for construction. The framework is instantiated with a LinUCB contextual bandit on a QA task, using seven pre-computed pipeline graphs as arms, complexity labels as context, and a reward combining F1 with a latency penalty. Experiments compare this adaptive QA system (AQA) against GPTSwarm and report that AQA adapts to query complexity and achieves a better effectiveness-efficiency trade-off. The paper's central claim is that the instantiation 'successfully constructs system pipelines that are adapted to the characteristics of incoming questions and optimized for a given efficiency-effectiveness trade-off' (Section 5.4.1).
Significance. The perspective and framework are timely and useful: the module taxonomy and the formulation of pipeline construction as a contextual bandit problem are clear contributions, and the authors provide a public code repository, which strengthens reproducibility. However, the experimental demonstration currently does not support the central claim. The time-included reward specified in Eq. (3) is inconsistent with the reported behavior of AQA(T) for Context C, the context features are gold complexity labels unavailable at query time, the evaluation rests on a single run with no statistical significance testing, and the GPTSwarm baseline optimizes a different objective. These issues are substantial but fixable; the framework contribution is independent of the experimental validation.
major comments (4)
- [§5.3 and Table 4 (Eq. 3)] The time-included results are internally inconsistent with the stated reward. For Context C, Table 3 gives OneR F1=0.146, S=6.41 s and IRCoT F1=0.458, S=184.85 s. Substituting into Eq. (3) with β=0.5 gives r(OneR) ≈ 0.5·0.146 − 0.5·(6.41/10000) ≈ 0.0727 and r(IRCoT) ≈ 0.5·0.458 − 0.5·(184.85/50) ≈ −1.62. The optimal arm under Eq. (3) for Context C is therefore OneR, yet Table 4 reports AQA(T) with F1=0.523 and time 11.75 (natural-log ms, ≈126 s), which matches an IRCoT-containing pipeline and is identical to AQA(NT). After roughly 1,167 Context-C samples, LinUCB should easily separate a reward gap of about 1.7. The reported result implies that the reward was implemented differently from Eq. (3), the policy did not converge, or the table is misreported; in any case, the conclusion in §5.4.1 that the time-based model 'successfully' optimizes the efficiency-effectiveness trade-off is not supported by the published data.
- [§5.3 (context features)] The contextual features are 'the complexity labels provided in the dataset,' i.e., the gold A/B/C labels described in §5.1, which were generated with Flan-T5-XL. Real systems do not have access to gold complexity labels at query time, so the demonstration of adaptivity presupposes an oracle signal. The paper should either use features computable from the query itself (e.g., text embeddings or a trained complexity classifier) or explicitly frame the results as an oracle-upper-bound proof-of-concept. Section 4.4 acknowledges that 'contextual features may not be sufficiently informative,' but it does not mention that the features used here are gold labels, which is a stronger and more specific concern. This directly affects the transferability of the claim that the system adapts to 'characteristics of incoming questions.'
- [§5.4 (experimental protocol)] The experimental evaluation consists of a single run of the LinUCB process for 3,500 timesteps, evaluated on a single test set of 51 questions, with no variance, confidence intervals, or significance testing. The differences in Table 4 (e.g., AQA(NT) vs AQA(T): F1 0.697 vs 0.687, time 9.99 vs 8.89) and the differences versus GPTSwarm could be within run-to-run noise. Because the paper's central conclusion is that AQA 'successfully' adapts and outperforms static orchestration, the authors should report multiple seeds or bootstrapped intervals and, ideally, significance tests for the test-set comparisons.
- [§5.4.2 (baseline comparison)] The GPTSwarm baseline is optimized for F1 only, without any time cost (the paper notes that GPTSwarm does not accommodate context or time cost), whereas AQA(T) is optimized for the combined reward of Eq. (3). The higher latency of GPTSwarm is therefore expected and does not by itself demonstrate the value of adaptivity. To support the claim that adaptive orchestration achieves a superior efficiency-effectiveness trade-off over static optimization, the static baseline should be optimized for the same objective (e.g., a static graph optimizer with a time-sensitive reward) or the comparison should be explicitly limited to an F1-only static baseline.
minor comments (6)
- [§4.3.2] Items (2) and (3) in the context-features list are both titled 'User profile features' and contain nearly identical text; they should be merged or renumbered.
- [Figure 2] The legend is difficult to read in grayscale, and the color coding of the seven arms is not discernible in print; additionally, the legend label 'NOR' should be 'NoR' to match the text.
- [Table 4] The caption should state that the 'Time' column reports natural-log-transformed milliseconds; this information appears only in the text.
- [Eq. (3)] The definition of T_t is written with an awkward bracket structure; please rewrite it with explicit cases, e.g., T_t = S_t/10000 if 1 < S_t ≤ 10, and T_t = S_t/50 if S_t > 10.
- [References] Several references use '[n. d.]' instead of a year or venue (e.g., [9], [10], [55], [56]); these should be completed for the camera-ready version.
- [§5.4.1] The text refers to 'dotted red lines' in Figure 2, but the figure appears to be in color only in the electronic version; please ensure the dashed/dotted lines are also distinguishable in grayscale.
Circularity Check
No circularity found: the paper's AQA instantiation is an empirical bandit experiment, and its central claim rests on held-out evaluation rather than on a derivation that reduces to its inputs.
full rationale
The paper's main contribution is a perspective and framework, with a proof-of-concept instantiation (AQA) that uses LinUCB to select among seven pre-enumerated pipeline graphs. The optimization objective (Eq. 3) combines F1 with a time penalty, and the evaluation reports F1 and time on a held-out test set. This is an empirical comparison, not a derivation from first principles, so there is no self-definitional reduction: the reward is not constructed from the test-set outcomes, and the reported conclusions are about learned behavior on held-out data. The contextual features are the dataset's gold complexity labels ('Our contextual features are the complexity labels provided in the dataset'), which is a data-leakage / external-validity concern, not circularity, because the labels are external annotations and are not derived from the model's own outputs. The paper's self-citations (e.g., refs. 25, 34, 54, 86) appear in background, feature-taxonomy, and feedback-bias discussions; none is load-bearing for the central claim. The skeptical observation that the time-included results in Table 4 appear inconsistent with the reward defined by Eq. 3 is a correctness/consistency issue, not circularity: it does not show that the conclusion was forced by the inputs. Overall, the paper does not fit any of the enumerated circularity patterns, so the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- beta (reward weight) =
0.5
- latency cost thresholds =
T_t = S/10000 for 1 < S <= 10; T_t = 1/50 for S > 10; T_t = 0 for S <= 1
- training/test split =
210 training, 51 test questions
assumptions (4)
- domain assumption The Flan-T5-XL derived complexity labels (A/B/C) are a valid and query-time-available description of complexity.
- domain assumption The seven precomputed graphs of {NoR, OneR, IRCoT, Aggregate} cover the relevant pipeline space for the QA task.
- domain assumption F1 and the latency-derived cost T_t together represent the user's utility.
- domain assumption LinUCB's linear model adequately predicts the reward of each graph given complexity labels.
Cite this review
Pith. "Pith review of Adaptive Orchestration of Modular Generative Information Access Systems." pith.science (2026). https://pith.science/paper/PSDL45PH
@misc{pith2026250417454,
author = {Pith},
title = {Pith review of: Adaptive Orchestration of Modular Generative Information Access Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/PSDL45PH}},
note = {Machine review of arXiv:2504.17454}
}
read the original abstract
Advancements in large language models (LLMs) have driven the emergence of complex new systems to provide access to information, that we will collectively refer to as modular generative information access (GenIA) systems. They integrate a broad and evolving range of specialized components, including LLMs, retrieval models, and a heterogeneous set of sources and tools. While modularity offers flexibility, it also raises critical challenges: How can we systematically characterize the space of possible modules and their interactions? How can we automate and optimize interactions among these heterogeneous components? And, how do we enable this modular system to dynamically adapt to varying user query requirements and evolving module capabilities? In this perspective paper, we argue that the architecture of future modular generative information access systems will not just assemble powerful components, but enable a self-organizing system through real-time adaptive orchestration -- where components' interactions are dynamically configured for each user input, maximizing information relevance while minimizing computational overhead. We give provisional answers to the questions raised above with a roadmap that depicts the key principles and methods for designing such an adaptive modular system. We identify pressing challenges, and propose avenues for addressing them in the years ahead. This perspective urges the IR community to rethink modular system designs for developing adaptive, self-optimizing, and future-ready architectures that evolve alongside their rapidly advancing underlying technologies.
Figures
Reference graph
Works this paper leans on
-
[1]
Leonard Adolphs, Benjamin Börschinger, Christian Buck, Michelle Chen Hueb- scher, Massimiliano Ciaramita, Lasse Espeholt, Thomas Hofmann, Yannic Kilcher, Sascha Rothe, Pier Giuseppe Sessa, and Lierni Sestorain. 2022. Boosting Search Engines with Interactive Agents. Trans. Mach. Learn. Res. (2022)
2022
-
[2]
Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton, and Mikhail Burtsev. 2021. Building and Evaluating Open-Domain Dialogue Corpora with Clarifying Questions. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
-
[3]
James Allan, Eunsol Choi, Daniel P Lopresti, and Hamed Zamani. 2024. Future of Information Retrieval Research in the Age of Generative AI. arXiv preprint arXiv:2412.02043 (2024)
work page Pith review arXiv 2024
-
[4]
Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...
2023
-
[5]
Gaurav Arora, Shreya Jain, and Srujana Merugu. 2024. Intent Detection in the Age of LLMs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
2024
-
[6]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In The Twelfth International Conference on Learning Representations, ICLR
2024
-
[7]
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary G. Ives. 2007. DBpedia: A Nucleus for a Web of Open Data. In The Semantic Web, 6th International Semantic Web Conference
2007
-
[8]
Barman, Sascha Caron, Emily Sullivan, Henk W
Kristian G. Barman, Sascha Caron, Emily Sullivan, Henk W. de Regt, Roberto Ruiz de Austri, Mieke Boon, Michael Färber, Stefan Fröse, Faegheh Hasibi, Andreas Ipp, Rukshak Kapoor, Gregor Kasieczka, Daniel Kostić, Michael Krämer, Tobias Golling, Luis G. Lopez, Jesus Marco, Sydney Otten, Pawel Pawlowski, Pietro Vischia, Erik Weber, and Christoph Weniger. 2025...
arXiv 2025
Show all 107 references
-
[9]
Jakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand, Nathalie Bradley-Schmieg, Michael Chang, Natalie Clay, Adrian Collister, Vibhavari Dasagi, Lucy Gonzalez, Karol Gregor, Edward Hughes, Sheleem Kashem, Maria Loks-Thompson, Hannah Openshaw, Jack Parker-Holde...
2023
-
[10]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. [n. d.]. Language Models are Few-shot Learners. In Advances in neural information processing systems
-
[11]
Lucas, Peter I
Cameron Browne, Edward Jack Powley, Daniel Whitehouse, Simon M. Lucas, Peter I. Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez Liebana, Spyridon Samothrakis, and Simon Colton. 2012. A Survey of Monte Carlo Tree Search Methods. IEEE Trans. Comput. Intell. AI Games (2012)
2012
-
[12]
H. Chase. 2022. LangChain. https://github.com/hwchase17/langchain. Accessed: 2025-02-07
2022
-
[13]
Zheng Chu, Jingchang Chen, Qianglong Chen, Haotian Wang, Kun Zhu, Xiyuan Du, Weijiang Yu, Ming Liu, and Bing Qin. 2024. BeamAggR: Beam Aggregation Reasoning over Multi-source Knowledge for Multi-hop Question Answering. In Proceedings of the 62nd Annual Meeting of the Associati...
2024
-
[14]
Zhao, Yanping Huang, Andrew M
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Web- son, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robin...
2024
-
[15]
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The Power of Noise: Redefining Retrieval for RAG Systems. In Proceedings of the 47th International ACM SIGIR Conference o...
2024
-
[16]
Shane Culpepper, Fernando Diaz, and Mark D
J. Shane Culpepper, Fernando Diaz, and Mark D. Smucker. 2018. Research Fron- tiers in Information Retrieval: Report from the Third Strategic Workshop on Information Retrieval in Lorne (SWIRL 2018). SIGIR Forum (2018)
2018
-
[17]
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. 2024. Vision Transformers Need Registers. In The Twelfth International Conference on Learning Representations, ICLR 2024
2024
-
[18]
Leonardo Mendonça de Moura and Nikolaj S. Bjørner. 2008. Z3: An Efficient SMT Solver. In Tools and Algorithms for the Construction and Analysis of Systems, 14th International Conference, TACAS
2008
-
[19]
Yang Deng, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, and Tat-Seng Chua
-
[20]
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024. Detecting hallucinations in large language models using semantic entropy. Nat. (2024)
2024
-
[21]
Shahla Farzana, Qunzhi Zhou, and Petar Ristoski. 2023. Knowledge Graph- Enhanced Neural Query Rewriting. In Companion Proceedings of the ACM Web Conference 2023
2023
-
[22]
Wenlong Fei, Xiaohua Wang, Min Hu, Qingyu Zhang, and Hongbo Li. 2024. MTLS: Making Texts into Linguistic Symbols. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
2024
-
[23]
Debargha Ganguly, Srinivasan Iyengar, Vipin Chaudhary, and Shivkumar Kalya- naraman. 2024. PROOF OF THOUGHT : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning. In The First Workshop on System-2 Reasoning at Scale, NeurIPS’24
2024
-
[24]
Yunfan Gao, Yun Xiong, Meng Wang, and Haofen Wang. 2024. Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks. CoRR abs/2407.21059 (2024)
2024 arXiv
-
[25]
Shashank Gupta, Philipp Hager, Jin Huang, Ali Vardasbi, and Harrie Oosterhuis
-
[26]
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborat...
2024
-
[27]
Eric Horvitz, Andy Jacobs, and David Hovel. 1999. Attention-Sensitive Alerting. In UAI ’99: Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence, Stockholm, Sweden, July 30 - August 1, 1999
1999
-
[28]
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park
-
[29]
Salim, Falk Scholer, and Damiano Spina
Kaixin Ji, Danula Hettiachchi, Flora D. Salim, Falk Scholer, and Damiano Spina
-
[30]
Kaixin Ji, Damiano Spina, Danula Hettiachchi, Flora Dilys Salim, and Falk Scholer
-
[31]
InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL
2024
-
[32]
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active Retrieval Aug- mented Generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Ju...
2023
-
[33]
In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2024)
Characterizing Information Seeking Processes with Multiple Physiological Signals. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2024)
2024
-
[34]
Hideaki Joko, Shubham Chatterjee, Andrew Ramsay, Arjen P De Vries, Jeff Dal- ton, and Faegheh Hasibi. 2024. Doing personal laps: Llm-augmented dialogue construction for personalized multi-session conversational search. In Proceedings of the 47th International ACM SIGIR Confere...
2024
-
[35]
In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
Examining the Impact of Uncontrolled Variables on Physiological Signals in User Studies for Information Processing Activities. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
-
[36]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...
2024
-
[37]
Konda and John N
Vijay R. Konda and John N. Tsitsiklis. 1999. Actor-Critic Algorithms. InAdvances in Neural Information Processing Systems 12, [NIPS Conference]
1999
-
[38]
Jiajie Jin, Yutao Zhu, Xinyu Yang, Chenghao Zhang, and Zhicheng Dou. 2024. FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research. CoRR (2024). To appear in the Web Conference 2025
2024
-
[39]
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Gener- ation for Knowledge-Intensive NLP Tasks. In Ad...
2020
-
[40]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ...
2020
-
[41]
Pimentel
Pooya Khandel, Andrew Yates, Ana Lucia Varbanescu, Maarten de Rijke, and Andy D. Pimentel. 2024. Distillation vs. Sampling for Efficient Training of Learn- ing to Rank Models. InProceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR
2024
-
[42]
Schapire
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. 2010. A Contextual- bandit Approach to Personalized News Article Recommendation. In Proceedings of the 19th International Conference on World Wide Web . 661–670
2010
-
[43]
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In 9th International Conference on Learning Representati...
2021
-
[44]
Xingxuan Li, Ruochen Zhao, Yew Ken Chia, Bosheng Ding, Shafiq Joty, Soujanya Poria, and Lidong Bing. 2024. Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources. In The Twelfth International Conference on Learning Represe...
2024
-
[45]
Li, Sewon Min, Srinivasan Iyer, Yashar Mehdad, and Wen-tau Yih
Belinda Z. Li, Sewon Min, Srinivasan Iyer, Yashar Mehdad, and Wen-tau Yih. 2020. Efficient One-Pass End-to-End Entity Linking for Questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
-
[46]
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing S...
2023
-
[47]
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-Play Composi- tional Reasoning with Large Language Models. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural...
2023
-
[48]
Wenzhe Li, Hao Luo, Zichuan Lin, Chongjie Zhang, Zongqing Lu, and Deheng Ye. 2023. A Survey on Transformers in Reinforcement Learning. Trans. Mach. Learn. Res. 2023 (2023)
2023
-
[49]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories. In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023
-
[50]
Zihao Li, Yuyi Ao, and Jingrui He. 2024. SpherE: Expressive and Interpretable Knowledge Graph Embedding for Set Retrieval. In Proceedings of the 47th In- ternational ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR. 2629–2634
2024
-
[51]
Defu Lian, Xu Huang, Xiaolong Chen, Jin Chen, Xingmei Wang, Yankai Wang, Haoran Jin, Rui Fan, Zheng Liu, Le Wu, and Enhong Chen. 2023. RecStudio: Towards a Highly-Modularized Recommender System. In Proceedings of the 46th International ACM SIGIR Conference on Research and Deve...
2023
-
[52]
Marvin Minsky. 1988. Society of mind
1988
-
[53]
Chaitanya Malaviya, Peter Shaw, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2023. QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2023
-
[54]
Harrie Oosterhuis and Maarten de Rijke. 2020. Policy-Aware Unbiased Learning to Rank for Top-k Rankings. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR
2020
-
[55]
W. McCune. [n. d.]. Prover9 and Mace4. ([n. d.])
-
[56]
Niall McGuire and Yashar Moshfeghi. 2024. DEEPER: Dense Electroencephalog- raphy Passage Retrieval. CoRR abs/2412.06695 (2024)
2024 arXiv
-
[57]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu
-
[58]
Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy. 2023. LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers. In Proceedings of the 2023 Conference on Empirical...
2023
-
[59]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...
2021
-
[60]
OpenAI. 2023. OpenAI O1 System Card. https://openai.com/index/openai-o1- system-card/
2023
-
[61]
OpenAI. 2023. OpenAI O3 Mini System Card. https://openai.com/index/o3-mini- system-card/
2023
-
[62]
Robertson and Hugo Zaragoza
Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. (2009)
2009
-
[63]
IEEE Trans
Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Trans. Knowl. Data Eng. (2024)
2024
-
[64]
In-Context Learn- ing
Andrew Parry, Debasis Ganguly, and Manish Chandra. 2024. "In-Context Learn- ing" or: How I learned to stop worrying and love "Applied Information Retrieval". In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
2024
-
[65]
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. InAdvances in Neural Information Processing Systems 36: Annual ...
2023
-
[66]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event
2021
-
[67]
Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford
Stephen E. Robertson, Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and Mike Gatford. 1994. Okapi at TREC-3. In Proceedings of The Third Text REtrieval Conference, TREC 1994, Gaithersburg, Maryland, USA, November 2-4, 1994
1994
-
[68]
Xiaqiang Tang, Qiang Gao, Jian Li, Nan Du, Qi Li, and Sihong Xie. 2025. MBA- RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity. In Proceedings of the 31st International Conference on Com- putational Linguistics, COLING
2025
-
[69]
Lee, and Eunho Yang
Hyun Ryu, Gyeongman Kim, Hyemin S. Lee, and Eunho Yang. 2025. Divide and Translate: Compositional First-Order Logic Translation and Verification for Complex Logical Reasoning. In The Thirteenth International Conference on Learning Representations
2025
-
[70]
Harrisen Scells, Shengyao Zhuang, and Guido Zuccon. 2022. Reduce, Reuse, Recycle: Green Information Retrieval Research. InSIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
2022
-
[71]
van Hulst, Faegheh Hasibi, Koen Dercksen, Krisztian Balog, and Arjen P
Johannes M. van Hulst, Faegheh Hasibi, Koen Dercksen, Krisztian Balog, and Arjen P. de Vries. 2020. REL: An Entity Linker Standing on the Shoulders of Giants. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SI...
2020
-
[72]
Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024
-
[73]
Szymanski and Michael D
Peter T. Szymanski and Michael D. Lemmon. 1993. Adaptive mixtures of local experts are source coding solutions. In Proceedings of International Conference on Neural Networks (ICNN’88)
1993
-
[74]
Christopher J. C. H. Watkins and Peter Dayan. 1992. Technical Note Q-Learning. Mach. Learn. (1992)
1992
-
[75]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation ...
2023
-
[76]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[77]
In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge- Intensive Multi-Step Questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL
-
[78]
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Framework. CoRR abs/2308.08155 (2023)
2023 arXiv
-
[79]
Sambasivan, Wenhao Lu, Yajun Wang, Mehul Parsana, Purushottam Kar, and Manik Varma
Hemanth Vemuri, Sheshansh Agrawal, Shivam Mittal, Deepak Saini, Akshay Soni, Abhinav V. Sambasivan, Wenhao Lu, Yajun Wang, Mehul Parsana, Purushottam Kar, and Manik Varma. 2023. Personalized Retrieval over Millions of Items. In Proceedings of the 46th International ACM SIGIR C...
2023
-
[80]
Xuezhi Wang and Denny Zhou. 2024. Chain-of-Thought Reasoning Without Prompting. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS
2024
-
[81]
Ziyi Ye, Qingyao Ai, Yiqun Liu, Maarten de Rijke, Min Zhang, Christina Lioma, and Tuukka Ruotsalo. 2025. Generative language reconstruction from brain recordings. Communications Biology 8, 1 (March 2025), 346
2025
-
[82]
Lawrie, and Benjamin Van Durme
Orion Weller, Dawn J. Lawrie, and Benjamin Van Durme. 2024. NevIR: Negation in Neural Information Retrieval. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL
2024
-
[83]
Williams
Ronald J. Williams. 1992. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Mach. Learn. (1992)
1992
-
[84]
Ian Wu, Sravan Jayanthi, Vijay Viswanathan, Simon Rosenberg, Sina Pakazad, Tongshuang Wu, and Graham Neubig. 2024. Synthetic Multimodal Question Generation. In Findings of the Association for Computational Linguistics: EMNLP
2024
-
[85]
Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi Fung, Hao Peng, and Heng Ji. 2024. CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets. In The Twelfth International Conference on Learning Representations
2024
-
[86]
Zilin Xiao, MING GONG, Jie Wu, Xingyao Zhang, Linjun Shou, and Daxin Jiang
-
[87]
In The 2023 Conference on Empirical Methods in Natural Language Processing
Instructed Language Models with Retrievers Are Powerful Entity Linkers. In The 2023 Conference on Empirical Methods in Natural Language Processing . 11 SIGIR ’25, July 13–18,2025, Padua, Italy Mohanna Hoveyda, Harrie Oosterhuis, Arjen P. de Vries, Maarten de Rijke, and Faegheh Hasibi
2023
-
[88]
Cohen, Ruslan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Langu...
2018
-
[89]
An Zhang, Yang Deng, Yankai Lin, Xu Chen, Ji-Rong Wen, and Tat-Seng Chua
-
[90]
Ziyi Ye, Jingtao Zhan, Qingyao Ai, Yiqun Liu, Maarten de Rijke, Christina Lioma, and Tuukka Ruotsalo. 2024. Query Augmentation with Brain Signals. In Proceed- ings of the 32nd ACM International Conference on Multimedia . Association for Computing Machinery
2024
-
[91]
Wenpeng Yin, Muhao Chen, Rui Zhang, Ben Zhou, Fei Wang, and Dan Roth
-
[92]
In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts
Enhancing LLM Capabilities Beyond Scaling Up. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts
2024
-
[93]
Ori Yoran, Tomer Wolfson, Ori Ram, and Jonathan Berant. 2024. Making Retrieval- Augmented Language Models Robust to Irrelevant Context. In The Twelfth Inter- national Conference on Learning Representations, ICLR
2024
-
[94]
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, An- drew M Dai, Quoc V Le, James Laudon, et al. 2022. Mixture-of-experts with Expert Choice Routing. Advances in Neural Information Processing Systems (2022)
2022
-
[95]
Yifei Yuan, Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke, and Wai Lam. 2024. Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search. In Proceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, May 13-17, 2024
2024
-
[96]
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024. The Shift from Models to Compound AI Systems. https: //bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/
2024
-
[97]
ChengXiang Zhai. 2024. Large Language Models and Future of Information Retrieval: Opportunities and Challenges. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
2024
-
[99]
In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24)
Large Language Model Powered Agents for Information Retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24) . Association for Computing Machinery
-
[100]
Tianhua Zhang, Jiaxin Ge, Hongyin Luo, Yung-Sung Chuang, Mingye Gao, Yuan Gong, Yoon Kim, Xixin Wu, Helen Meng, and Jim Glass. 2024. Natural Language Embedded Programs for Hybrid Language Symbolic Reasoning. In Findings of the Association for Computational Linguistics: NAACL
2024
-
[101]
Jujia Zhao, Wenjie Wang, Yiyan Xu, Teng Sun, Fuli Feng, and Tat-Seng Chua. 2024. Denoising Diffusion Recommender Model. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
2024
-
[102]
Wenxuan Zhou, Qiang Ning, Heba Elfardy, Kevin Small, and Muhao Chen. 2022. Answer Consolidation: Formulation and Benchmarking. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2022
-
[103]
Yingxue Zhou, Jie Hao, Mukund Rungta, Yang Liu, Eunah Cho, Xing Fan, Yan- bin Lu, Vishal Vasudevan, Kellen Gillespie, and Zeynab Raeesy. 2023. Unified Contextual Query Rewriting. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume...
2023
-
[105]
Yunhua Zhou, Guofeng Quan, and Xipeng Qiu. 2023. A Probabilistic Framework for Discovering New Intents. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2023
-
[106]
Mingchen Zhuge, Haozhe Liu, Francesco Faccio, Dylan R. Ashley, Róbert Csordás, Anand Gopalakrishnan, Abdullah Hamdi, Hasan Abed Al Kader Hammoud, Vincent Herrmann, Kazuki Irie, Louis Kirsch, Bing Li, Guohao Li, Shuming Liu, Jinjie Mai, Piotr Piekos, Aditya Ramesh, Imanol Schla...
2023
-
[107]
Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jürgen Schmidhuber. 2024. GPTSwarm: Language Agents as Optimizable Graphs. In Forty-first International Conference on Machine Learning, ICML. 12
2024
-
[2023]
InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
Recent Advances in the Foundations and Applications of Unbiased Learning to Rank. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR
-
[2024]
In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024
Towards Human-centered Proactive Conversational Agents. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.