Pith. sign in

REVIEW 48 references

Optimizing Compound Retrieval Systems

T0 review · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Optimized compound retrieval systems, which learn both which LLM predictions to gather and how to combine them, beat single-stage cascade reranking baselines on TREC-DL at most efficiency levels.

arxiv 2504.12063 v1 pith:SXAFUDSK submitted 2025-04-16 cs.IR cs.AIcs.LG

classification cs.IRcs.AIcs.LG
keywords retrievalmodelscompoundsystemsrankingcascadingapproachmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Search systems today usually work like a funnel. A cheap model, such as BM25, finds a thousand candidate documents, and then a more expensive model re-ranks a small slice of them. This paper argues that the funnel is only one way to combine models. It defines a compound retrieval system in which a system learns two things at once: which documents and pairs of documents should get expensive LLM judgments, and how to add those judgments to the cheap scores to form the final ranking.

The authors combine BM25 with Gemini 1.5 Flash used in two modes: pointwise (is this document relevant?) and pairwise (is document A better than B?). An optimization procedure decides where to spend LLM calls and how to weight the answers, under a penalty for each call. In one variant it trains with human relevance labels; in another it tries to imitate an expensive pairwise ranking without labels.

On TREC-DL queries, the learned systems reach about the same or better nDCG with roughly ten times fewer LLM calls than standard pairwise prompting, for most budget levels. The learned call patterns are not simple top-K reranking; they include partial pairwise comparisons and mixtures of pointwise and pairwise judgments. The biggest caveats are that the comparison uses only single-stage cascades as baselines, efficiency is counted as number of calls rather than latency or tokens, and the test sets are small.

Extended reading notes

Core claim

The load-bearing claim is that optimized compound retrieval systems provide better effectiveness-efficiency trade-offs than cascading systems. Section 7.1 states: 'our optimized compound retrieval systems provide significantly better overall effectiveness-efficiency trade-off curves.' The evidence includes Table 1, where supervised compound at N=10^4 LLM calls reaches nDCG@100 0.519, while PRP at N=10^5 reaches 0.497, a roughly 10x efficiency gain at equal or better effectiveness. If true, retrieval system design should be treated as an optimizable policy over prediction gathering, not a fixed cascade sequence.

Load-bearing premise

The conclusion that compound systems beat cascading approaches assumes the tested baselines represent the cascade paradigm. Section 6.6 uses only single re-ranking-step cascades, pointwise and PRP top-K rerankers over BM25; Section 9 concedes the system 'was also only compared with single re-ranking cascading systems.' If tuned multi-stage cascades with learned dynamic cutoffs, such as the cited Culpepper et al. 2016 or Gallagher et al. 2019, close the gap, the broad claim over cascading systems is unsupported. This premise is load-bearing and is not tested.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces a conceptual framework rather than new physical entities. Its learnable components are the score aggregation variables, the selection policy probabilities, and the MLP weights that generate them. The main domain assumptions are the bounded-relevance bound, the usefulness of LLM probe predictions, the call-count cost model, the BM25 top-1000 recall assumption, and the straight-through estimator's adequacy.

free parameters (3)
  • Score aggregation variables A_r, B_r, C_r for pointwise and pairwise components = Learned, values not reported
    Eqs. 15-17 and Eq. 31 define final scores as linear functions of these variables; they are optimized by SGD under the combined ranking and cost loss.
  • Selection policy probabilities pi(s_point[r]=1) and pi(s_pair[r,r']=1) = Learned, values not reported
    Section 4.4; these determine where LLM predictions are gathered, optimized with straight-through gradient estimates and made deterministic by best-of-250 validation selection.
  • Weights of the two 3x64 MLPs generating A, B, C, and pi from rank indices = Learned, values not reported
    Section 6.4; introduced for regularity and generalization, not intrinsic to the framework, but they determine the learned aggregation and selection behavior.
assumptions (5)
  • standard math Relevance labels are bounded in [0, V] so the distillation lower bound can be minimized componentwise.
    Assumed in Section 5.1 to derive Eq. 23; without bounded labels the min-v step is unbounded.
  • domain assumption LLM pointwise and pairwise prediction probabilities are useful relevance signals for ranking.
    The framework treats M1 through M7 as components and relies on prior PRP results (Qin et al. 2024, Liang et al. 2023) rather than validating the LLM signals independently.
  • domain assumption Efficiency is adequately measured by the expected number of LLM calls, with each call treated as equal cost.
    Eq. 11 defines Lcost as the count of selected pointwise and pairwise calls; pairwise prompts contain two passages and likely consume more tokens and latency, so this assumption can misorder systems on true efficiency.
  • domain assumption BM25 top-1000 contains the relevant documents for TREC-DL queries; only re-ranking within this set is considered.
    Section 6.1 fixes a top-1000 re-ranking task; if relevant documents fall outside the BM25 top-1000, both cascades and compound systems miss them equally.
  • domain assumption The straight-through estimator provides gradient signals sufficient for policy learning.
    Section 4.4 replaces non-differentiable sampling with a biased gradient estimate; the paper acknowledges the bias and this may explain poor performance at very low call budgets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimizing Compound Retrieval Systems." pith.science (2026). https://pith.science/paper/SXAFUDSK

@misc{pith2026250412063,
  author       = {Pith},
  title        = {Pith review of: Optimizing Compound Retrieval Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SXAFUDSK}},
  note         = {Machine review of arXiv:2504.12063}
}
read the original abstract

Modern retrieval systems do not rely on a single ranking model to construct their rankings. Instead, they generally take a cascading approach where a sequence of ranking models are applied in multiple re-ranking stages. Thereby, they balance the quality of the top-K ranking with computational costs by limiting the number of documents each model re-ranks. However, the cascading approach is not the only way models can interact to form a retrieval system. We propose the concept of compound retrieval systems as a broader class of retrieval systems that apply multiple prediction models. This encapsulates cascading models but also allows other types of interactions than top-K re-ranking. In particular, we enable interactions with large language models (LLMs) which can provide relative relevance comparisons. We focus on the optimization of compound retrieval system design which uniquely involves learning where to apply the component models and how to aggregate their predictions into a final ranking. This work shows how our compound approach can combine the classic BM25 retrieval model with state-of-the-art (pairwise) LLM relevance predictions, while optimizing a given ranking metric and efficiency target. Our experimental results show optimized compound retrieval systems provide better trade-offs between effectiveness and efficiency than cascading approaches, even when applied in a self-supervised manner. With the introduction of compound retrieval systems, we hope to inspire the information retrieval field to more out-of-the-box thinking on how prediction models can interact to form rankings.

Figures

Figures reproduced from arXiv: 2504.12063 by the authors.

Figure 1
Figure 1. Overview of the compound retrieval system de [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Effectiveness-efficiency trade-off curves (averages over 50 runs), created by optimizing compound systems with various [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Examples of (deterministic) selection policies of our optimized compound retrieval systems. Top bars display the [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 19 canonical work pages

  1. [1]

    Nima Asadi and Jimmy Lin. 2013. Document vector representations for feature extraction in multi-stage document ranking. Information retrieval 16 (2013), 747–768

  2. [2]

    Nima Asadi and Jimmy Lin. 2013. Effectiveness/efficiency tradeoffs for candidate generation in multi-stage retrieval architectures. In Proceedings of the 36th in- ternational ACM SIGIR conference on Research and development in information retrieval. 997–1000

  3. [3]

    Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432 (2013)

  4. [4]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al . 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  5. [5]

    Ruey-Cheng Chen, Luke Gallagher, Roi Blanco, and J Shane Culpepper. 2017. Efficient cost-aware cascade ranking in multi-stage retrieval. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval. 445–454

  6. [6]

    Charles LA Clarke, J Shane Culpepper, and Alistair Moffat. 2016. Assessing efficiency–effectiveness tradeoffs in multi-stage retrieval systems without using relevance judgments. Information Retrieval Journal 19 (2016), 351–377

  7. [7]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Ellen M Voorhees, and Ian Soboroff. 2021. TREC deep learning track: Reusable test collections in the large data regime. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval . 2369–2375

  8. [8]

    W Bruce Croft, Donald Metzler, and Trevor Strohman. [n. d.]. Search engines: Information retrieval in practice . Vol. 520

Show all 48 references
  1. [9]

    J Shane Culpepper, Charles LA Clarke, and Jimmy Lin. 2016. Dynamic cutoff prediction in multi-stage retrieval systems. In Proceedings of the 21st Australasian Document Computing Symposium. 17–24

  2. [10]

    Luke Gallagher, Ruey-Cheng Chen, Roi Blanco, and J Shane Culpepper. 2019. Joint optimization of cascade ranking models. In Proceedings of the twelfth ACM international conference on web search and data mining . 15–23

  3. [11]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Rethink training of BERT rerankers in multi-stage retrieval pipeline. In Advances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part II 43 . Spring...

  4. [12]

    Jiafeng Guo, Yinqiong Cai, Yixing Fan, Fei Sun, Ruqing Zhang, and Xueqi Cheng

  5. [13]

    Shashank Gupta, Harrie Oosterhuis, and Maarten de Rijke. 2024. Practical and Robust Safety Guarantees for Advanced Counterfactual Learning to Rank. In Pro- ceedings of the 33rd ACM International Conference on Information and Knowledge Management. 737–747

  6. [14]

    Rolf Jagerman, Xuanhui Wang, Honglei Zhuang, Zhen Qin, Michael Bendersky, and Marc Najork. 2022. Rax: composable learning-to-rank using Jax. In Pro- ceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3051–3060

  7. [15]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems (TOIS) 20, 4 (2002), 422–446

  8. [16]

    Diederik P Kingma. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)

  9. [17]

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Ben- jamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cos- grove, Christopher D Manning, Christopher Re, Dia...

  10. [18]

    Shichen Liu, Fei Xiao, Wenwu Ou, and Luo Si. 2017. Cascade ranking for opera- tional e-commerce search. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1557–1565

  11. [19]

    Weiwen Liu, Yunjia Xi, Jiarui Qin, Fei Sun, Bo Chen, Weinan Zhang, Rui Zhang, and Ruiming Tang. 2022. Neural Re-ranking in Multi-stage Recom- mender Systems: A Review. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud ...

  12. [20]

    Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulic, Anna Ko- rhonen, and Nigel Collier. 2024. Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators. CoRR abs/2403.16950 (2024). https://doi.org/10.48550/arXiv.2403.16950

  13. [21]

    Adian Liusie, Vatsal Raina, Yassir Fathullah, and Mark Gales. 2024. Efficient LLM Comparative Assessment: A Product of Experts Framework for Pairwise Comparisons. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohi...

  14. [22]

    Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2024. Ranked List Truncation for Large Language Model-based Re-Ranking. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 141–151

  15. [23]

    Alistair Moffat and Justin Zobel. 1996. Self-indexing inverted files for fast text retrieval. ACM Transactions on Information Systems (TOIS) 14, 4 (1996), 349–379

  16. [24]

    Harrie Oosterhuis. 2021. Computationally efficient optimization of plackett-luce ranking models for relevance and fairness. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1023–1032

  17. [25]

    Andrew Parry, Sean MacAvaney, and Debasis Ganguly. 2024. Top-Down Parti- tioning for Efficient List-Wise Ranking. arXiv preprint arXiv:2405.14589 (2024)

  18. [26]

    Matthias Petri, J Shane Culpepper, and Alistair Moffat. 2013. Exploring the magic of WAND. In Proceedings of the 18th Australasian Document Computing Symposium. 58–65

  19. [27]

    Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021. The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv preprint arXiv:2101.05667 (2021)

  20. [28]

    Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! arXiv preprint arXiv:2312.02724 (2023)

  21. [29]

    Jiarui Qin, Jiachen Zhu, Bo Chen, Zhirong Liu, Weiwen Liu, Ruiming Tang, Rui Zhang, Yong Yu, and Weinan Zhang. 2022. Rankflow: Joint optimization of multi- stage cascade ranking systems as flows. In Proceedings of the 45th International ACM SIGIR Conference on Research and Dev...

  22. [30]

    Tao Qin, Tie-Yan Liu, and Hang Li. 2010. A general approximation framework for direct optimization of information retrieval measures. Information retrieval 13 (2010), 375–397

  23. [31]

    Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al. 2024. Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting. In Findings of the Association for Computational Linguistics:...

  24. [32]

    Zhen Qin, Rolf Jagerman, Rama Kumar Pasumarthi, Honglei Zhuang, He Zhang, Aijun Bai, Kai Hui, Le Yan, and Xuanhui Wang. 2023. RD-Suite: a benchmark for ranking distillation. Advances in Neural Information Processing Systems 36 (2023)

  25. [33]

    Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109

  26. [34]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investi- gating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...

  27. [35]

    Jiaxi Tang and Ke Wang. 2018. Ranking distillation: Learning compact ranking models with high performance for recommender system. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 2289–2298

  28. [36]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)

  29. [37]

    Nicola Tonellotto, Craig Macdonald, Iadh Ounis, et al . 2018. Efficient query processing for scalable web search. Foundations and Trends ® in Information Retrieval 12, 4-5 (2018), 319–500

  30. [38]

    Lidan Wang, Jimmy Lin, and Donald Metzler. 2011. A cascade ranking model for efficient ranked retrieval. In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval . 105–114

  31. [39]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent Abilities of Large Language Mod...

  32. [40]

    Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (1992), 229–256

  33. [41]

    Jun Xu and Hang Li. 2007. Adarank: a boosting algorithm for information retrieval. In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval . 391–398. 10 Optimizing Compound Retrieval Systems SIGIR ’25, July 13–1...

  34. [42]

    Le Yan, Zhen Qin, Honglei Zhuang, Rolf Jagerman, Xuanhui Wang, Michael Bendersky, and Harrie Oosterhuis. 2024. Consolidating Ranking and Relevance Predictions of Large Language Models through Post-Processing. In Proceedings of the 2024 Conference on Empirical Methods in Natura...

  35. [43]

    Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024. The Shift from Models to Compound AI Systems. https: //bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/

  36. [44]

    Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma. 2021. Optimizing dense retrieval model training with hard negatives. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1503–1512

  37. [45]

    Yue Zhang, ChengCheng Hu, Yuqi Liu, Hui Fang, and Jimmy Lin. 2021. Learning to rank in the age of muppets: Effectiveness–efficiency tradeoffs in multi-stage ranking. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing. 64–73

  38. [46]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)

  39. [47]

    Kai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang, Na Mou, Yanan Niu, Yang Song, Hongning Wang, and Kun Gai. 2024. Full Stage Learning to Rank: A Unified Framework for Multi-Stage Systems. InProceedings of the ACM on Web Conference

  40. [2022]

    ACM Transactions on Information Systems (TOIS) 40, 4 (2022), 1–42

    Semantic models for the first-stage retrieval: A comprehensive review. ACM Transactions on Information Systems (TOIS) 40, 4 (2022), 1–42

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.