REVIEW 48 references
Optimizing Compound Retrieval Systems
T0 review · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Optimized compound retrieval systems, which learn both which LLM predictions to gather and how to combine them, beat single-stage cascade reranking baselines on TREC-DL at most efficiency levels.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The authors combine BM25 with Gemini 1.5 Flash used in two modes: pointwise (is this document relevant?) and pairwise (is document A better than B?). An optimization procedure decides where to spend LLM calls and how to weight the answers, under a penalty for each call. In one variant it trains with human relevance labels; in another it tries to imitate an expensive pairwise ranking without labels.
On TREC-DL queries, the learned systems reach about the same or better nDCG with roughly ten times fewer LLM calls than standard pairwise prompting, for most budget levels. The learned call patterns are not simple top-K reranking; they include partial pairwise comparisons and mixtures of pointwise and pairwise judgments. The biggest caveats are that the comparison uses only single-stage cascades as baselines, efficiency is counted as number of calls rather than latency or tokens, and the test sets are small.
Extended reading notes
Core claim
The load-bearing claim is that optimized compound retrieval systems provide better effectiveness-efficiency trade-offs than cascading systems. Section 7.1 states: 'our optimized compound retrieval systems provide significantly better overall effectiveness-efficiency trade-off curves.' The evidence includes Table 1, where supervised compound at N=10^4 LLM calls reaches nDCG@100 0.519, while PRP at N=10^5 reaches 0.497, a roughly 10x efficiency gain at equal or better effectiveness. If true, retrieval system design should be treated as an optimizable policy over prediction gathering, not a fixed cascade sequence.
Load-bearing premise
The conclusion that compound systems beat cascading approaches assumes the tested baselines represent the cascade paradigm. Section 6.6 uses only single re-ranking-step cascades, pointwise and PRP top-K rerankers over BM25; Section 9 concedes the system 'was also only compared with single re-ranking cascading systems.' If tuned multi-stage cascades with learned dynamic cutoffs, such as the cited Culpepper et al. 2016 or Gallagher et al. 2019, close the gap, the broad claim over cascading systems is unsupported. This premise is load-bearing and is not tested.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (3)
- Score aggregation variables A_r, B_r, C_r for pointwise and pairwise components =
Learned, values not reported
- Selection policy probabilities pi(s_point[r]=1) and pi(s_pair[r,r']=1) =
Learned, values not reported
- Weights of the two 3x64 MLPs generating A, B, C, and pi from rank indices =
Learned, values not reported
assumptions (5)
- standard math Relevance labels are bounded in [0, V] so the distillation lower bound can be minimized componentwise.
- domain assumption LLM pointwise and pairwise prediction probabilities are useful relevance signals for ranking.
- domain assumption Efficiency is adequately measured by the expected number of LLM calls, with each call treated as equal cost.
- domain assumption BM25 top-1000 contains the relevant documents for TREC-DL queries; only re-ranking within this set is considered.
- domain assumption The straight-through estimator provides gradient signals sufficient for policy learning.
Cite this review
Pith. "Pith review of Optimizing Compound Retrieval Systems." pith.science (2026). https://pith.science/paper/SXAFUDSK
@misc{pith2026250412063,
author = {Pith},
title = {Pith review of: Optimizing Compound Retrieval Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/SXAFUDSK}},
note = {Machine review of arXiv:2504.12063}
}
read the original abstract
Modern retrieval systems do not rely on a single ranking model to construct their rankings. Instead, they generally take a cascading approach where a sequence of ranking models are applied in multiple re-ranking stages. Thereby, they balance the quality of the top-K ranking with computational costs by limiting the number of documents each model re-ranks. However, the cascading approach is not the only way models can interact to form a retrieval system. We propose the concept of compound retrieval systems as a broader class of retrieval systems that apply multiple prediction models. This encapsulates cascading models but also allows other types of interactions than top-K re-ranking. In particular, we enable interactions with large language models (LLMs) which can provide relative relevance comparisons. We focus on the optimization of compound retrieval system design which uniquely involves learning where to apply the component models and how to aggregate their predictions into a final ranking. This work shows how our compound approach can combine the classic BM25 retrieval model with state-of-the-art (pairwise) LLM relevance predictions, while optimizing a given ranking metric and efficiency target. Our experimental results show optimized compound retrieval systems provide better trade-offs between effectiveness and efficiency than cascading approaches, even when applied in a self-supervised manner. With the introduction of compound retrieval systems, we hope to inspire the information retrieval field to more out-of-the-box thinking on how prediction models can interact to form rankings.
Figures
Reference graph
Works this paper leans on
-
[1]
Nima Asadi and Jimmy Lin. 2013. Document vector representations for feature extraction in multi-stage document ranking. Information retrieval 16 (2013), 747–768
work page 2013
-
[2]
Nima Asadi and Jimmy Lin. 2013. Effectiveness/efficiency tradeoffs for candidate generation in multi-stage retrieval architectures. In Proceedings of the 36th in- ternational ACM SIGIR conference on Research and development in information retrieval. 997–1000
work page 2013
-
[3]
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432 (2013)
arXiv 2013
-
[4]
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45
2024
-
[5]
Ruey-Cheng Chen, Luke Gallagher, Roi Blanco, and J Shane Culpepper. 2017. Efficient cost-aware cascade ranking in multi-stage retrieval. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval. 445–454
work page 2017
-
[6]
Charles LA Clarke, J Shane Culpepper, and Alistair Moffat. 2016. Assessing efficiency–effectiveness tradeoffs in multi-stage retrieval systems without using relevance judgments. Information Retrieval Journal 19 (2016), 351–377
work page 2016
-
[7]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Ellen M Voorhees, and Ian Soboroff. 2021. TREC deep learning track: Reusable test collections in the large data regime. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval . 2369–2375
work page 2021
-
[8]
W Bruce Croft, Donald Metzler, and Trevor Strohman. [n. d.]. Search engines: Information retrieval in practice . Vol. 520
Show all 48 references
-
[9]
J Shane Culpepper, Charles LA Clarke, and Jimmy Lin. 2016. Dynamic cutoff prediction in multi-stage retrieval systems. In Proceedings of the 21st Australasian Document Computing Symposium. 17–24
2016
-
[10]
Luke Gallagher, Ruey-Cheng Chen, Roi Blanco, and J Shane Culpepper. 2019. Joint optimization of cascade ranking models. In Proceedings of the twelfth ACM international conference on web search and data mining . 15–23
2019
-
[11]
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Rethink training of BERT rerankers in multi-stage retrieval pipeline. In Advances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43 . Spring...
2021
-
[12]
Jiafeng Guo, Yinqiong Cai, Yixing Fan, Fei Sun, Ruqing Zhang, and Xueqi Cheng
-
[13]
Shashank Gupta, Harrie Oosterhuis, and Maarten de Rijke. 2024. Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank. In Pro- ceedings of the 33rd ACM International Conference on Information and Knowledge Management. 737–747
2024
-
[14]
Rolf Jagerman, Xuanhui Wang, Honglei Zhuang, Zhen Qin, Michael Bendersky, and Marc Najork. 2022. Rax: composable learning-to-rank using Jax. In Pro- ceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3051–3060
2022
-
[15]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems (TOIS) 20, 4 (2002), 422–446
2002
-
[16]
Diederik P Kingma. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[17]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Ben- jamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cos- grove, Christopher D Manning, Christopher Re, Dia...
2023
-
[18]
Shichen Liu, Fei Xiao, Wenwu Ou, and Luo Si. 2017. Cascade ranking for opera- tional e-commerce search. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1557–1565
2017
-
[19]
Weiwen Liu, Yunjia Xi, Jiarui Qin, Fei Sun, Bo Chen, Weinan Zhang, Rui Zhang, and Ruiming Tang. 2022. Neural Re-ranking in Multi-stage Recom- mender Systems: A Review. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud ...
2022 doi
- [20]
-
[21]
Adian Liusie, Vatsal Raina, Yassir Fathullah, and Mark Gales. 2024. Efficient LLM Comparative Assessment: A Product of Experts Framework for Pairwise Comparisons. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohi...
2024 doi
-
[22]
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024. Ranked List Truncation for Large Language Model-based Re-Ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 141–151
2024
-
[23]
Alistair Moffat and Justin Zobel. 1996. Self-indexing inverted files for fast text retrieval. ACM Transactions on Information Systems (TOIS) 14, 4 (1996), 349–379
1996
-
[24]
Harrie Oosterhuis. 2021. Computationally efficient optimization of plackett-luce ranking models for relevance and fairness. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1023–1032
2021
-
[25]
Andrew Parry, Sean MacAvaney, and Debasis Ganguly. 2024. Top-Down Parti- tioning for Efficient List-Wise Ranking. arXiv preprint arXiv:2405.14589 (2024)
2024 arXiv
-
[26]
Matthias Petri, J Shane Culpepper, and Alistair Moffat. 2013. Exploring the magic of WAND. In Proceedings of the 18th Australasian Document Computing Symposium. 58–65
2013
-
[27]
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021. The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv preprint arXiv:2101.05667 (2021)
2021 arXiv
-
[28]
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! arXiv preprint arXiv:2312.02724 (2023)
2023 arXiv
-
[29]
Jiarui Qin, Jiachen Zhu, Bo Chen, Zhirong Liu, Weiwen Liu, Ruiming Tang, Rui Zhang, Yong Yu, and Weinan Zhang. 2022. Rankflow: Joint optimization of multi- stage cascade ranking systems as flows. In Proceedings of the 45th International ACM SIGIR Conference on Research and Dev...
2022
-
[30]
Tao Qin, Tie-Yan Liu, and Hang Li. 2010. A general approximation framework for direct optimization of information retrieval measures. Information retrieval 13 (2010), 375–397
2010
-
[31]
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al. 2024. Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting. In Findings of the Association for Computational Linguistics:...
2024
-
[32]
Zhen Qin, Rolf Jagerman, Rama Kumar Pasumarthi, Honglei Zhuang, He Zhang, Aijun Bai, Kai Hui, Le Yan, and Xuanhui Wang. 2023. RD-Suite: a benchmark for ranking distillation. Advances in Neural Information Processing Systems 36 (2023)
2023
-
[33]
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109
1995
-
[34]
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investi- gating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...
2023
-
[35]
Jiaxi Tang and Ke Wang. 2018. Ranking distillation: Learning compact ranking models with high performance for recommender system. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 2289–2298
2018
-
[36]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)
2023 arXiv
-
[37]
Nicola Tonellotto, Craig Macdonald, Iadh Ounis, et al . 2018. Efficient query processing for scalable web search. Foundations and Trends ® in Information Retrieval 12, 4-5 (2018), 319–500
2018
-
[38]
Lidan Wang, Jimmy Lin, and Donald Metzler. 2011. A cascade ranking model for efficient ranked retrieval. In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval . 105–114
2011
-
[39]
Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent Abilities of Large Language Mod...
2022
-
[40]
Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (1992), 229–256
1992
-
[41]
Jun Xu and Hang Li. 2007. Adarank: a boosting algorithm for information retrieval. In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval . 391–398. 10 Optimizing Compound Retrieval Systems SIGIR ’25, July 13–1...
2007
-
[42]
Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, and Harrie Oosterhuis. 2024. Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing. In Proceedings of the 2024 Conference on Empirical Methods in Natura...
2024 doi
-
[43]
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024. The Shift from Models to Compound AI Systems. https: //bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/
2024
-
[44]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021. Optimizing dense retrieval model training with hard negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1503–1512
2021
-
[45]
Yue Zhang, ChengCheng Hu, Yuqi Liu, Hui Fang, and Jimmy Lin. 2021. Learning to rank in the age of muppets: Effectiveness–efficiency tradeoffs in multi-stage ranking. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing. 64–73
2021
-
[46]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)
2023 arXiv
-
[47]
Kai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang, Na Mou, Yanan Niu, Yang Song, Hongning Wang, and Kun Gai. 2024. Full Stage Learning to Rank: A Unified Framework for Multi-Stage Systems. InProceedings of the ACM on Web Conference
2024
-
[2022]
ACM Transactions on Information Systems (TOIS) 40, 4 (2022), 1–42
Semantic models for the first-stage retrieval: A comprehensive review. ACM Transactions on Information Systems (TOIS) 40, 4 (2022), 1–42
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.