REVIEW 5 major objections 5 minor 73 references
MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A per-query mixture of sparse, dense, and human retrievers can outperform every single retriever, including 7B LLM-based ones, without any training.
desk verdict Useful ensemble weighting for RAG; the 'zero-shot' claim is softer than the abstract implies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the weight allocation function $f(q, R_i, D)$, built from two families of zero-shot signals. The pre-retrieval signal $V_{\mathrm{pre}}$ clusters the corpus in each retriever's embedding space and measures the size-weighted distance from the query vector to the cluster centroids, so a query far from all document regions is down-weighted. The post-retrieval signals are the Moran coefficient, a scalar measure of how clustered the top retrieved documents are, and $V_{\mathrm{post}}$, the average of $V_{\mathrm{pre}}$ values over the top 20 documents, which says how familiar those documents are to the retriever's own corpus geometry. These signals are combined with fixed coefficients and then multiplied into the per-retriever relevance scores before a final weighted sum and re-ranking; the multi-granularity expansion comes from decomposing queries and documents into sub-questions and propositions, effectively quadrupling the number of retriever variants without retraining any model.
What would settle it
On a held-out sample of any of the four datasets, compute the per-query rank correlation between each retriever's $V_{\mathrm{pre}}$ or $V_{\mathrm{post}}$ score on a query and that retriever's actual NDCG@20 on the same query; if most correlations are near zero or negative, the weights are not tracking retrieval quality and the reported gains would not generalize. A second check is to sweep the coefficients $(a,b,c)$ widely: if performance collapses sharply outside the hand-picked point, the method is effectively tuned rather than zero-shot.
Extended reading notes
Core claim
MoR's central claim is that retriever quality is not a fixed property of a retriever but a per-query property, and that it can be estimated without ground-truth labels from the geometry of the query and the retrieved documents. For each query $q$, every retriever $R_i$ returns scores; MoR reweights them with $f(q, R_i, D) = a \cdot V_{\mathrm{pre}} + b \cdot I_{\mathrm{Moran}} + c \cdot V_{\mathrm{post}}$, using coefficients $(a,b,c)=(0.1,0.3,0.6)$, then sums the reweighted scores across retrievers and re-ranks. Here $V_{\mathrm{pre}}$ measures how close the query embedding lands to size-weighted centroids of corpus clusters in the retriever's embedding space, $V_{\mathrm{post}}$ averages those pre-retrieval familiarity values over the top 20 retrieved documents, and $I_{\mathrm{Moran}}$ is the Moran coefficient quantifying spatial autocorrelation among retrieved documents. In addition, each retriever is expanded across four textual granularities, such as original queries and documents versus sub-questions and propositions, before fusion. With this per-query weighting, the roughly 0.8B-parameter pool beats its best individual member by 10.8% on average on NDCG@20, beats the best supervised component by 12.2%, and beats the 7B GritLM baseline by 3.9% on average.
Load-bearing premise
The load-bearing premise is that the geometric signals, query-to-cluster distance before retrieval and average document familiarity after retrieval, track a retriever's true per-query quality closely enough that one fixed weighted combination transfers across queries and datasets without ground-truth labels.
Editorial extensions
If this is right
- Practitioners would no longer need to pick a single retriever by heuristic: an ensemble of cheap, BERT-sized retrievers can replace a much larger LLM-based retriever component in a RAG pipeline.
- The same per-query trust weights can serve as a calibration signal for human-in-the-loop systems, weighting human-provided documents by corpus familiarity rather than by declared expertise.
- Because a threshold on pre-retrieval weights lets MoR keep most of its performance while using only about 20% of retrievers per query, the approach can be made cheap enough for latency-sensitive serving.
- Retriever subset selection should prioritize complementarity over individual accuracy, since mixing just two complementary retrievers can match the performance of the full eight-retriever pool.
- Better retrieval from mixing transfers to generation: MoR improves exact-match accuracy on SciFact and SciQ RAG tasks over both individual retrievers and the 7B baselines without retraining the reader model.
Reading between the lines
- If the geometric signals transfer to new corpora, the same recipe could rank retrievers before any labeled data exists in a new domain, but the paper's evidence is limited to four scientific-domain datasets, so cross-domain transfer is an open test.
- The fixed coefficients $(0.1, 0.3, 0.6)$ treat every query the same, yet the paper itself notes that optimal coefficients could vary per query; learning query-specific coefficients is a natural extension that could push performance closer to the reported Route Oracle upper bound of +13.5% over GritLM.
- A concrete testable extension is to feed MoR's per-query weights as features into a lightweight learned router, which might combine the transparency of the geometric signals with the headroom of supervised routing.
- The human-retriever experiments simulate out-of-domain experts as random rankers; real human errors may be correlated and domain-specific, so genuine human-LLM collaboration gains would need testing with actual annotators rather than this oracle simulation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Mixture of Retrievers (MoR), a method for combining multiple heterogeneous retrievers (sparse, dense, and simulated human) with per-query weights computed without ground-truth labels. Weights are derived from a pre-retrieval corpus-familiarity signal (Vpre), post-retrieval Moran's I and Vpost signals, and combined linearly with fixed coefficients (a,b,c). The authors report on four scientific retrieval datasets that MoR-post outperforms all component BERT-size retrievers and two 7B LLM retrievers on average NDCG@20, and that adding simulated human experts yields further gains. They also include ablations on signal design, deep fusion, and efficiency via retriever subset selection.
Significance. If the reported results hold, MoR provides evidence that a parameter-light mixture of small retrievers can match or exceed a 7B retriever on some scientific-domain tasks, which is practically valuable. The proposed signals connect to established ideas in query performance prediction and aggregated search, and the paper is transparent about model checkpoints and hyperparameters. However, the zero-shot claim is weakened by the empirical selection of the combination coefficients on the evaluation suite, and the human-retriever experiments use oracle simulations that do not match the 'non-oracle' claim. These issues, together with an ill-defined Vpre formula and a text contradiction with Table 3, mean the main conclusions need revision before the paper can be accepted.
major comments (5)
- [Section 4.2, Vpre definition] The displayed definition of Vpre in Section 4.2 computes a sum of vectors (unit vectors v_hat_k scaled by cluster weights and inverse squared norms), not a scalar. Since Vpost is then defined as the mean of Vpre over the top-20 documents, and fpost is a linear combination a*Vpre + b*IMoran + c*Vpost, the weighted sum in Section 4 (unnumbered equation) would combine a vector-valued weight with the scalar relevance scores s_i(q,d_j), which is mathematically ill-posed. The authors must state the intended scalar-valued definition—for example, the norm of that sum or the sum of cluster-size-weighted reciprocal squared distances—and verify that Vpost and fpost are scalars. This is a central derivation issue because all subsequent weight allocation uses Vpre/Vpost.
- [Section 4.3, coefficient selection] In Section 4.3, the coefficients (a,b,c) = (0.1,0.3,0.6) are selected 'empirically' without a held-out development split for MoR itself. The only development set described in Appendix A.4 (100 queries) is used for the Performance Normalization baseline. If the coefficients were chosen by inspecting the test-set results in Table 3, then the 'zero-shot' claim in the abstract is not supported, and the comparison against GritLM/RepLLaMA is not apples-to-apples. The authors should either tune the coefficients on a separate development split and report the selected values, provide a sensitivity analysis demonstrating robust performance over a range of (a,b,c), or explicitly revise the claim to acknowledge the tuned hyperparameters.
- [Section 5.1, Table 3 vs text] Section 5.1 states that 'MoR-post achieves better performance than GritLM on NFCorpus, SciFact, and SciQ,' but Table 3 shows MoR-post at 73.2 NDCG@20 on SciFact versus GritLM at 79.8, a 6.6-point deficit. The later statement that MoR-post is 'comparable performance to GritLM' on SciFact is also inaccurate at this gap. Moreover, the reported +3.9% average improvement over GritLM is dominated by SciQ (92.8 vs 79.7), with much smaller gains on NFCorpus and SciDocs. The text should be corrected, and the per-dataset variability should be discussed honestly, ideally with confidence intervals or significance tests.
- [Section 5.2 and abstract, human retrievers] The abstract claims MoR can incorporate 'specialized non-oracle human information sources,' but the simulation in Section 5.2 assigns each expert oracle retrieval (gold documents ranked at top) on its own domain and random ranks on all other domains. There is no non-oracle or noisy human retriever in the experiments; hence the conclusions about 'human trustworthiness estimation' and the +58.9% relative gain over humans alone apply only to oracle-simulated experts. The claims should be reworded, or the simulation should be replaced with a realistic noisy human model (e.g., imperfect ranks on the expert domain).
- [Abstract and Section 3.1, parameter accounting] The abstract's 'Despite totaling just 0.8B parameters' counts only the eight base retriever models (0.836B). The full MoR pipeline also includes the propositioner (Flan-T5-Large, ~780M) for deep fusion and, in the human experiments, an additional MPNet encoder. The 4x granularity expansion also increases the index and search cost. The authors should report the end-to-end parameter count and describe the additional inference overhead, or scope the claim to the base retrievers explicitly.
minor comments (5)
- [Section 4.1] The list of granularity variants names 'Rsq-p' twice (sub-questions and passage; sub-questions and propositions). The second is likely a typo (e.g., Rsq-prop); please fix the notation.
- [Table 3] The column group 'Average across datasets' with merged ND@5/ND@20 headers is difficult to parse; consider restructuring the table so each dataset's NDCG@5 and NDCG@20 columns are clearly separated.
- [Throughout] No error bars, standard deviations, or significance tests are reported anywhere in the paper; given that several of the headline gaps are small (e.g., +0.4 NDCG@20 on SciDocs over GritLM), the authors should provide multiple-run variability or at least discuss seed sensitivity.
- [Section 2] The phrase 'for differen purposes' should read 'for different purposes.'
- [Figure 1] The text and weight matrix in Figure 1 are very small, and the example query 'What Oxidant and Reductants can accept electrons?' is awkward; please enlarge and proofread the figure.
Circularity Check
No significant circularity: MoR's per-query weights are label-free corpus-geometry and QPP signals; the hand-set combination coefficients and self-citations do not force the reported result.
full rationale
MoR's central claim is an empirical result rather than a derivation. The per-query weights are generated from label-free signals: Vpre measures query-to-cluster distance in each retriever's embedding space, Vpost averages Vpre over the top-20 retrieved documents, and I_Moran is the standard Moran coefficient from the query performance prediction literature. The final score is a weighted sum of normalized retriever scores (Section 4), and the weights are not computed from ground-truth labels or from a model trained to maximize the reported NDCG. Route Oracle (Table 2) is explicitly an oracle upper bound, not a zero-shot prediction. The main near-input-dependent choice is the parametric combination (a,b,c)=(0.1,0.3,0.6) in Section 4.3, selected 'empirically' with no documented held-out split for MoR itself; the 100-query development set in Appendix A.4 is used only for the Performance Normalization baseline. If those coefficients were chosen on the test suite, the 'zero-shot' label would be weakened and the comparison to GritLM/RepLLaMA would not be apples-to-apples, but this is a reproducibility and correctness risk rather than constructional circularity: three scalar hyperparameters do not encode the per-query routing outcomes, and the individual weight signals still have independent content. Self-citations to Chen et al. (2023b), Cai et al. (2024), and Thrust (Zhao et al., 2023) supply the propositioner, the multi-granularity design idea, and baseline signals, respectively; none is invoked as an external proof of MoR's effectiveness, and no uniqueness theorem or renamed known result appears. The Limitations section appropriately notes that the parametric combination uses one universal set of coefficients, but that admission does not make the central comparison circular. Hence no significant circularity is identified; score 1 reflects only the minor coefficient-selection caveat.
Assumptions & free parameters
free parameters (4)
- (a, b, c) in MoR-post =
0.1, 0.3, 0.6
- |Dq| top documents for Vpost =
20
- KMeans cluster count K =
max(ceil(|D|^(1/4)), 3)
- Retriever pruning threshold =
95th percentile
assumptions (5)
- domain assumption Cluster hypothesis: documents that are close in the retriever's embedding space tend to be relevant to the same query.
- ad hoc to paper Query-to-corpus centroid distance predicts retriever effectiveness.
- ad hoc to paper The three signals are linearly combinable with fixed coefficients across queries and datasets.
- domain assumption Normalizing each retriever's scores to [0,1] makes weighted sums across heterogeneous retrievers valid.
- ad hoc to paper Simulated human experts with oracle in-domain ranks approximate real specialized human retrievers.
Cite this review
Pith. "Pith review of MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers." pith.science (2026). https://pith.science/paper/IUXK5WG5
@misc{pith2026250615862,
author = {Pith},
title = {Pith review of: MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers},
year = {2026},
howpublished = {\url{https://pith.science/paper/IUXK5WG5}},
note = {Machine review of arXiv:2506.15862}
}
read the original abstract
Retrieval-augmented Generation (RAG) is powerful, but its effectiveness hinges on which retrievers we use and how. Different retrievers offer distinct, often complementary signals: BM25 captures lexical matches; dense retrievers, semantic similarity. Yet in practice, we typically fix a single retriever based on heuristics, which fails to generalize across diverse information needs. Can we dynamically select and integrate multiple retrievers for each individual query, without the need for manual selection? In our work, we validate this intuition with quantitative analysis and introduce mixture of retrievers: a zero-shot, weighted combination of heterogeneous retrievers. Extensive experiments show that such mixtures are effective and efficient: Despite totaling just 0.8B parameters, this mixture outperforms every individual retriever and even larger 7B models by +10.8% and +3.9% on average, respectively. Further analysis also shows that this mixture framework can help incorporate specialized non-oracle human information sources as retrievers to achieve good collaboration, with a 58.9% relative performance improvement over simulated humans alone.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Negar Arabzadeh, Chuan Meng, Mohammad Aliannejadi, and Ebrahim Bagheri. 2024. https://doi.org/10.1145/3673791.3698438 Query performance prediction: Techniques and applications in modern information retrieval . In Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Regio...
arXiv 2024
-
[3]
Jaime Arguello and Fernando Diaz. 2013. Relevance Ranking of Vertical Search Engines, chapter Vertical Selection and Aggregation. Elsevier
work page 2013
-
[4]
Jaime Arguello and 1 others. 2017. Aggregated search. Foundations and Trends in Information Retrieval , 10(5):365--502
work page 2017
-
[5]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, and 1 others. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268
arXiv 2016
-
[6]
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38, pages 716--722. Springer
2016
-
[7]
Fengyu Cai, Xinran Zhao, Tong Chen, Sihao Chen, Hongming Zhang, Iryna Gurevych, and Heinz Koeppl. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.579 M ix GR : Enhancing retriever generalization for scientific domain through complementary granularity . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10369...
-
[8]
James P Callan, Zhihong Lu, and W Bruce Croft. 1995. Searching distributed collections with inference networks. In Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval, pages 21--28
work page 1995
Show all 73 references
-
[9]
Jamie Callan. 2002. Distributed information retrieval. In Advances in information retrieval: recent research from the center for intelligent information retrieval, pages 127--150. Springer
2002
-
[10]
Hsinchun Chen, Haiyan Fan, Michael Chau, and Daniel Zeng. 2001. Metaspider: Meta-searching and categorization on the web. Journal of the American Society for Information Science and Technology, 52(13):1134--1147
2001
-
[11]
Sihao Chen, Hongming Zhang, Tong Chen, Ben Zhou, Wenhao Yu, Dian Yu, Baolin Peng, Hongwei Wang, Dan Roth, and Dong Yu. 2023 a . https://arxiv.org/pdf/2311.04335.pdf Sub-sentence encoder: Contrastive learning of propositional semantic representations . arXiv preprint arXiv:2311.04335
2023 arXiv
-
[12]
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Dong Yu, and Hongming Zhang. 2023 b . Dense x retrieval: What retrieval granularity should we use? arXiv preprint arXiv:2312.06648
2023 arXiv
-
[13]
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021. Decontextualization: Making sentences stand-alone. Transactions of the Association for Computational Linguistics, 9:447--461
2021
-
[14]
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, and 1 others. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1--53
2024
-
[15]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020. https://doi.org/10.18653/v1/2020.acl-main.207 SPECTER : Document-level representation learning using citation-informed transformers . In Proceedings of the 58th Annual Meeting of the Association for C...
2020 doi
-
[16]
Gordon V Cormack, Charles LA Clarke, and Stefan Buettcher. 2009. Reciprocal rank fusion outperforms condorcet and individual rank learning methods. In Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, pages 758--759
2009
-
[17]
Zhuyun Dai, Yubin Kim, and Jamie Callan. 2017. Learning to rank resources. In Proceedings of the 40th International ACM SIGIR conference on research and development in information retrieval, pages 837--840
2017
-
[18]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[19]
Fernando Diaz. 2007. https://doi.org/10.1145/1277741.1277841 Performance prediction using spatial autocorrelation . In Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '07, page 583–590, New York, NY,...
2007
-
[20]
Fernando Diaz, Mounia Lalmas, and Milad Shokouhi. 2010. From federated to aggregated search. In Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval, pages 910--910
2010
-
[21]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.552 S im CSE : Simple contrastive learning of sentence embeddings . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894--6910, Online ...
2021 doi
-
[22]
Eric J Glover, Steve Lawrence, William P Birmingham, and C Lee Giles. 1999. Architecture of a metasearch engine that supports user information needs. In Proceedings of the eighth international conference on Information and knowledge management, pages 210--216
1999
-
[23]
Rachid Guerraoui, Anne-Marie Kermarrec, Diana Petrescu, Rafael Pires, Mathis Randl, and Martijn de Vos. 2025. https://doi.org/10.1145/3721146.3721942 Efficient federated search for retrieval-augmented generation . In Proceedings of the 5th Workshop on Machine Learning and Syst...
2025
-
[24]
Harris, K
Charles R. Harris, K. Jarrod Millman, St \' e fan van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime ...
2020 doi
-
[25]
Sebastian Hofst \"a tter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently teaching an effective dense retriever with balanced topic aware sampling. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in In...
2021
-
[26]
John D Hunter. 2007. Matplotlib: A 2d graphics environment. Computing in science & engineering, 9(03):90--95
2007
-
[27]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2022. https://openreview.net/forum?id=jKN1pXi7b0 Unsupervised dense information retrieval with contrastive learning . Transactions on Machine Learning Research
2022
-
[28]
Jacobs, Michael I
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. 1991. https://doi.org/10.1162/neco.1991.3.1.79 Adaptive mixtures of local experts . Neural Computation, 3(1):79--87
1991 doi
-
[29]
Jardine and C.J
N. Jardine and C.J. van Rijsbergen . 1971. https://doi.org/10.1016/0020-0271(71)90051-9 The use of hierarchic clustering in information retrieval . Information Storage and Retrieval, 7(5):217--240
1971 doi
-
[30]
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park. 2024. https://doi.org/10.18653/v1/2024.naacl-long.389 Adaptive- RAG : Learning to adapt retrieval-augmented large language models through question complexity . In Proceedings of the 2024 Conference of the N...
2024 doi
-
[31]
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.495 Active retrieval augmented generation . In Proceedings of the 2023 Conference on Empirical Methods in...
2023 doi
-
[33]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen - tau Yih. 2020 b . https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference...
2020 doi
-
[34]
Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktashmotlagh, Xi Wang, and Guido Zuccon. 2023. https://doi.org/10.1145/3624918.3625330 Selecting which dense retriever to use for zero-shot search . In Proceedings of the Annual International ACM SIGIR Conference on Research and D...
2023
-
[35]
To Eun Kim and Fernando Diaz. 2025. https://arxiv.org/abs/2506.13743 Ltrr: Learning to rank retrievers for llms . Preprint, arXiv:2506.13743
2025 arXiv
-
[36]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019 doi
-
[37]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...
2023
-
[38]
Langchain. 2025. https://python.langchain.com/docs/how_to/ensemble_retriever/ Langchain ensembleretriever . Accessed: 04-May-2025
2025
-
[39]
Hyunji Lee, Luca Soldaini, Arman Cohan, Minjoon Seo, and Kyle Lo. 2024. Routerretriever: Exploring the benefits of routing over multiple expert embedding models. arXiv preprint arXiv:2409.02685
2024 arXiv
-
[40]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, and 1 others. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Informati...
2020
- [41]
-
[42]
Bo Long and Yi Chang. 2014. Relevance Ranking for Vertical Search Engines, 1st edition. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA
2014
-
[43]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. https://arxiv.org/abs/2310.08319 Fine-tuning llama for multi-stage text retrieval
2023 arXiv
-
[44]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.546 When not to trust language models: Investigating effectiveness of parametric and non-parametric memories . In Proceedings of the 6...
2023 doi
-
[45]
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.741 FA ct S core: Fine-grained atomic evaluation of factual precision in long form text generatio...
2023 doi
-
[46]
Feiteng Mu, Yong Jiang, Liwen Zhang, Liuchu Liuchu, Wenjie Li, Pengjun Xie, and Fei Huang. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.598 Query routing for homogeneous tools: An instantiation in the RAG scenario . In Findings of the Association for Computational Lin...
2024 doi
-
[47]
Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. 2024. Generative representational instruction tuning
2024
-
[48]
Niklas Muennighoff, Nouamane Tazi, Lo \" c Magne, and Nils Reimers. 2022. Mteb: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316
2022 arXiv
-
[49]
Giang Ngo, Rodney Beard, and Rohitash Chandra. 2022. https://doi.org/10.1016/j.neucom.2022.08.055 Evolutionary bagging for ensemble learning . Neurocomputing, 510:1–14
2022 doi
-
[50]
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.669 Large dual encoders are generalizable retrievers . In Proceedings of the 2022 Con...
2022 doi
-
[51]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K \" o pf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner...
2019
-
[52]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. https://doi.org/10.18653/v1/D16-1264 SQ u AD : 100,000+ questions for machine comprehension of text . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383-...
2016 doi
-
[53]
Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference...
2019 doi
-
[54]
Robertson and Hugo Zaragoza
Stephen E. Robertson and Hugo Zaragoza. 2009. https://doi.org/10.1561/1500000019 The probabilistic relevance framework: BM25 and beyond . Found. Trends Inf. Retr., 3(4):333--389
2009 doi
-
[55]
Haggai Roitman. 2017. https://doi.org/10.1145/3077136.3080665 An enhanced approach to query performance prediction using reference lists . In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '17, page 869–87...
2017
-
[56]
Kunal Sawarkar, Abhilasha Mangal, and Shivam Raj Solanki. 2024. https://doi.org/10.1109/MIPR62202.2024.00031 Blended rag: Improving rag (retriever-augmented generation) accuracy with semantic search and hybrid query-based retrievers . In 2024 IEEE 7th International Conference ...
2024
-
[57]
Tran, Yi Tay, and Donald Metzler
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. 2022. https://openreview.net/forum?id=uLYc4L3C81A Confident adaptive language modeling . In Advances in Neural Information Processing Systems
2022
-
[58]
Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. https://openreview.net/forum?id=B1ckMDqlg Outrageously large neural networks: The sparsely-gated mixture-of-experts layer . In International Conference on Learning ...
2017
-
[59]
Ashutosh Singh, Debasis Ganguly, Suchana Datta, and Craig McDonald. 2023. https://doi.org/10.1145/3539618.3592082 Unsupervised query performance prediction for neural models with pairwise rank preferences . In Proceedings of the 46th International ACM SIGIR Conference on Resea...
2023
-
[60]
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. arXiv preprint arXiv:2004.09297
2020 arXiv
-
[61]
Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. 2021. https://openreview.net/forum?id=wCu6T5xFjeJ BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models . In Thirty-fifth Conference on Neural Info...
2021
-
[62]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, and 1 others. 2023 a . Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[63]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton - Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu,...
2023 arXiv
-
[64]
Andrew Trotman, Antti Puurula, and Blake Burgess. 2014. https://doi.org/10.1145/2682862.2682863 Improvements to bm25 and language models examined . In Proceedings of the 2014 Australasian Document Computing Symposium, ADCS '14, page 58–65, New York, NY, USA. Association for Co...
2014
-
[65]
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.609 Fact or fiction: Verifying scientific claims . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[66]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. https://arxiv.org/abs/2203.11171 Self-consistency improves chain of thought reasoning in language models
2023 arXiv
-
[67]
Liu, and Matt Gardner
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. https://doi.org/10.18653/v1/W17-4413 Crowdsourcing multiple choice science questions . In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 94--106, Copenhagen, Denmark. Association for Computational Linguistics
2017 doi
-
[68]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R \'e mi Louf, Morgan Funtowicz, and 1 others. 2019. https://arxiv.org/abs/1910.03771 Huggingface's transformers: State-of-the-art natural language processing ....
2019 arXiv
-
[69]
Bennett, Junaid Ahmed, and Arnold Overwijk
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. https://openreview.net/forum?id=zeFrfgyZln Approximate nearest neighbor negative contrastive learning for dense text retrieval . In International Conference o...
2021
-
[70]
Woongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek, and Sung Ju Hwang. 2025. Universalrag: Retrieval-augmented generation over multiple corpora with diverse modalities and granularities. arXiv preprint arXiv:2504.20734
2025 arXiv
-
[71]
Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. https://arxiv.org/abs/2102.09690 Calibrate before use: Improving few-shot performance of language models
2021 arXiv
-
[72]
Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, and Jianshu Chen. 2023. Thrust: Adaptively propels large language models with external knowledge. arXiv preprint arXiv:2307.10442
2023 arXiv
-
[73]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[74]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.