Pith. sign in

REVIEW 3 major objections 4 minor 5 cited by

This paper builds an eight-task benchmark of complex, multi-constraint retrieval queries and finds that even the strongest current models produce weak results, with the best average nDCG@10 at 0.346.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

CRUMB is a new benchmark for complex, multi-aspect retrieval tasks on which state-of-the-art retrieval models score poorly, and query rewriting does not rescue the best models.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection A useful and transparent complex-retrieval benchmark; the exact headline number is shaky but the qualitative 'models struggle' conclusion holds across tasks. the 3 major comments →

arxiv 2509.07253 v1 pith:Q5ADV6G5 submitted 2025-09-08 cs.IR cs.AIcs.CL

Benchmarking Information Retrieval Models on Complex Retrieval Tasks

classification cs.IR cs.AIcs.CL
keywords complex retrievalmulti-aspect queriesbenchmarkCRUMBLLM query rewritinginstruction-following retrievaldense and sparse retrievalevaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Retrieval models are still built for single-aspect queries, but people increasingly expect search systems to handle requests with several parts, constraints, or conditions. This paper tries to establish that current state-of-the-art retrieval models cannot yet satisfy such complex information needs. It introduces CRUMB, a benchmark of eight diverse complex retrieval tasks, and reports that the best model averages only 0.346 nDCG@10 and 0.587 R@100 across them. It also finds that LLM-based query rewriting helps weaker models but consistently hurts the strongest one. If the results hold, progress in retrieval needs benchmarks and training methods aimed at multi-constraint queries, not just more reranking.

Core claim

The paper's central claim is that complex retrieval—queries with multiple aspects, constraints, or set-based logic—is still largely unsolved by current retrieval models. To show this, it constructs CRUMB from eight existing realistic tasks: tip-of-the-tongue movie search, multi-aspect paper search, set-operation entity queries, state-specific legal statute retrieval, multi-constraint theorem retrieval, reasoning-heavy Stack Exchange questions, clinical trial matching from patient histories, and code retrieval from problem statements. On a standardized passage collection, with documents in unified markdown and contextualized chunking, the strongest model reaches just 0.346 average nDCG@10 and

What carries the argument

CRUMB is the central object. It is a set of eight retrieval tasks chosen so that queries contain multiple parts expressed in different ways—long multi-detail descriptions, set operations such as and/not, numerical constraints, geographic constraints—with documents standardized into unified markdown and, where appropriate, split into contextualized chunks that carry their heading path. The benchmark is what carries the argument: by varying query/document vocabulary overlap, corpus scale, and label granularity, it lets the paper attribute model failures to specific task features rather than to a single dataset artifact.

Load-bearing premise

The benchmark's difficulty and rankings rest on the transformed relevance labels being correct: queries without a document matching all aspects were removed, StackExchange queries whose relevant chunk was split across two chunks were dropped, and in three tasks every chunk of a relevant document was counted as relevant.

What would settle it

Re-judge a sample of Tip-of-the-Tongue, Clinical Trial, and SetOps queries with human annotators at chunk level instead of treating every chunk of a relevant document as relevant, then recompute nDCG@10; if scores rise substantially, the reported ceiling is partly a labeling artifact. Alternatively, run the StackExchange queries that were excluded because their relevant chunks split across chunk boundaries; if models do well on them, the benchmark overstates task difficulty.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Best current models cannot be trusted to put relevant results on the first page for multi-constraint queries: the top average nDCG@10 is 0.346.
  • First-stage recall is a bottleneck: with R@1000 well below 1.0 on several tasks, even a perfect reranker cannot recover the missing relevant documents.
  • LLM query rewriting is not a universal fix; for the strongest model, every tested rewriting technique reduced both precision and recall.
  • Instruction-following capability, model size, and diverse training data are the attributes that separate the best models, suggesting where future training investments should go.
  • CRUMB's per-task validation splits enable few-shot and tuning methods that earlier general benchmarks did not support.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the differential effect of rewriting—answer-style rewriting improves nDCG@10 for many models while document-style rewriting improves recall—suggests an adaptive or hybrid rewriting strategy could outperform any single technique; the paper does not test this.
  • Editorial extension: the label transformations (all chunks of a relevant document treated as relevant; only queries with fully matching aspects kept) may make CRUMB's absolute scores pessimistic; a chunk-level human re-judgment study would show whether the ceiling is real.
  • Editorial extension: the strong performance of an instruction-tuned model trained only on MSMARCO plus instruction-based hard negatives implies that training data diversity may be replaceable by instruction-style negatives, a hypothesis worth testing on other backbones.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper introduces CRUMB, a benchmark of eight complex retrieval tasks compiled from existing datasets (Tip-of-the-Tongue, StackExchange, Paper Retrieval, SetOps, Clinical Trial, Legal QA, Theorem Retrieval, Code Retrieval), with standardized markdown formatting, contextualized chunking, and validation splits. The authors evaluate nine retrieval models (BM25, Snowflake, GTE Qwen 1.5B/7B, Lion SB/DS 1B/8B, Promptriever) on nDCG@10, R@100, and R@1000, and test three LLM-based query rewriting techniques with Gemma-3 27B. They report that the best model (GTE Qwen 7B) achieves only 0.346 average nDCG@10 and 0.587 R@100, and that rewriting consistently degrades the strongest model while sometimes helping weaker ones. They also analyze per-task failure patterns and the role of task-specific instructions.

Significance. If the measurements hold, CRUMB is a valuable community resource: it broadens complex-retrieval evaluation beyond reasoning-heavy QA to include set operations, numeric constraints, and diverse vocabularies, and it makes a falsifiable claim about the ceiling of current retrieval models. The unified markdown format and contextualized chunks are practical contributions, as are the validation splits and detailed per-task analysis. The strongest models' robustness to instruction changes and the finding that LLM rewriting hurts them are interesting, testable results. The main results are presented transparently with limitations acknowledged, and the data/code are to be released.

major comments (3)
  1. [§3.1, §6.3, Table 6] The Paper Retrieval task uses DORIS-MAE labels generated by GPT-3.5 over a candidate pool built by lexical/semantic/citation retrieval, and the original authors explicitly excluded reference papers from the pool. The CRUMB filtering additionally removes queries without a fully-satisfying document. Since Table 6's macro-average weights Paper Retrieval equally with the other seven tasks, incomplete labels could bias the 0.346 nDCG@10 headline and could change the conclusion that rewriting hurts GTE Qwen 7B (Tables 8–10). The paper's own limitation statement acknowledges this risk. Please report the averages with Paper Retrieval removed, and/or audit the excluded reference papers and pool with additional systems.
  2. [§3.1 vs Tables 1/2] The text says the query filter 'results in 79 queries', but both Table 1 and Table 2 report Q=72 for Paper Retrieval. This unresolved discrepancy makes the final dataset definition ambiguous and hinders reproducibility. Please clarify the exact query count and the filtering steps.
  3. [§3.1, StackExchange] Keeping only StackExchange queries whose original relevant chunks are fully encapsulated in one CRUMB chunk may systematically select for queries whose answers are conveniently located, potentially altering task difficulty and the measured ranking. Please report the number of excluded queries and compare basic statistics (query length, topic distribution, number of relevant documents) between included and excluded queries, or state explicitly that the task is a convenience subset.
minor comments (4)
  1. [§5.3] The statement 'the chunked versions of these collections always contain the highest nDCG and recall' is contradicted by Clinical Trial: GTE Qwen 7B nDCG@10 is 0.370 in Table 6 (chunked) vs 0.403 in Table 7 (full-document).
  2. [Table 2] Table 2 omits Paper Retrieval, Theorem Retrieval, and Code Retrieval even though Section 5.3 states they are included in the full-document results; please align the table with the reported experiments.
  3. [§4.3] The text says 'retrieving 2000 documents for each query' for the chunked version. Since the retrieval units are passages for chunked corpora, please clarify whether 2000 refers to chunks or documents and how this interacts with the MaxP aggregation.
  4. [§5.6] Typo: 'Promptriver' should be 'Promptriever' in the instruction-ablation paragraph.

Circularity Check

0 steps flagged

No circularity: CRUMB is an empirical benchmark built from external datasets; the central claims are measurements, not derived from fitted parameters or self-citations.

full rationale

The paper's central claims are empirical measurements on a newly assembled benchmark. No parameter is fitted and then reported as a prediction; the transformations of existing datasets (DORIS-MAE, QUEST, BRIGHT, TREC, APPS, etc.) are explicit and do not encode the target results. The only self-citations (Lion, Hypencoder, Search-R1) appear as baseline model choices or related work and are not load-bearing: the Lion models are independently released systems included as baselines, and the explanation citing the Lion paper is post hoc analysis, not a premise of the benchmark's difficulty claim. The acknowledged limitations (e.g., GPT-3.5 labels and candidate pooling in Paper Retrieval, Section 6.3) are correctness risks, not circularity: they concern label quality, not the derivation of results from the labels. Since the benchmark is self-contained and evaluated against external models and datasets, the appropriate circularity score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The central claim is an empirical measurement, so the ledger mostly contains design choices and inherited labeling assumptions rather than fitted parameters or invented theoretical constructs. The filtering choices are notable because they affect the composition and difficulty of the benchmark.

free parameters (4)
  • Chunk size (512 BERT tokens)
    The chunking strategy in Section 3.2 stops adding content when a chunk would exceed 512 BERT tokens, a hand-picked threshold that affects all chunked datasets.
  • Legal QA heading depth (3) = 3
    In Section 3.1, Legal QA documents are split so chunks share the first three levels of headings, a manual granularity choice.
  • SetOps relevance threshold = 1.5
    In Section 3.1, an entity is considered relevant if the average rater score is above 1.5, a threshold derived from the QUEST dataset that determines the label set.
  • Paper Retrieval query filter
    In Section 3.1, queries without at least one document satisfying all aspects are excluded, which changes task difficulty and composition.
axioms (4)
  • domain assumption Source dataset relevance labels are correct.
    The benchmark inherits relevance judgments from TREC, BRIGHT, QUEST, DORIS-MAE, and others, treating them as ground truth without independent verification (Section 3.1).
  • domain assumption Chunk-level relevance propagates from document-level relevance.
    For Tip-of-the-tongue, Clinical Trial, and SetOps, every chunk from a relevant document is labeled relevant, assuming relevance is uniformly distributed across the document (Section 3.1).
  • domain assumption MaxP aggregation recovers document relevance from passage scores.
    The main passage-level results use MaxP to produce final document lists (Section 4.3), which assumes the highest-scoring passage is sufficient to judge document relevance.
  • domain assumption The nine selected models are representative of the state of the art.
    Models were chosen from MTEB and other sources (Section 4.2.1); this set may not capture the full range of currently available retrieval systems.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking Information Retrieval Models on Complex Retrieval Tasks." pith.science (2026). https://pith.science/paper/Q5ADV6G5

@misc{pith2026250907253,
  author       = {Pith},
  title        = {Pith review of: Benchmarking Information Retrieval Models on Complex Retrieval Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q5ADV6G5}},
  note         = {Machine review of arXiv:2509.07253}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language models (LLMs) are incredible and versatile tools for text-based tasks that have enabled countless, previously unimaginable, applications. Retrieval models, in contrast, have not yet seen such capable general-purpose models emerge. To achieve this goal, retrieval models must be able to perform complex retrieval tasks, where queries contain multiple parts, constraints, or requirements in natural language. These tasks represent a natural progression from the simple, single-aspect queries that are used in the vast majority of existing, commonly used evaluation sets. Complex queries naturally arise as people expect search systems to handle more specific and often ambitious information requests, as is demonstrated by how people use LLM-based information systems. Despite the growing desire for retrieval models to expand their capabilities in complex retrieval tasks, there exist limited resources to assess the ability of retrieval models on a comprehensive set of diverse complex tasks. The few resources that do exist feature a limited scope and often lack realistic settings making it hard to know the true capabilities of retrieval models on complex real-world retrieval tasks. To address this shortcoming and spur innovation in next-generation retrieval models, we construct a diverse and realistic set of complex retrieval tasks and benchmark a representative set of state-of-the-art retrieval models. Additionally, we explore the impact of LLM-based query expansion and rewriting on retrieval quality. Our results show that even the best models struggle to produce high-quality retrieval results with the highest average nDCG@10 of only 0.346 and R@100 of only 0.587 across all tasks. Although LLM augmentation can help weaker models, the strongest model has decreased performance across all metrics with all rewriting techniques.

Figures

Figures reproduced from arXiv: 2509.07253 by Hamed Zamani, Julian Killingback.

Figure 1
Figure 1. Figure 1: An overview of our proposed benchmark: CRUMB. Each card is dedicated to one of the complex tasks [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Example of our contextualized chunking strategy. The original markdown document content is [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Plot of average nDCG@10 on the chunked-document version of CRUMB with various query rewriting [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Plot of average R@1000 on the chunked-document version of CRUMB with various query rewriting [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison of change in nDCG@10 when a generic instruction is used instead of a task-specific instruc [PITH_FULL_IMAGE:figures/full_fig_p030_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Comparison of change in R@1000 when a generic instruction is used instead of a task-specific [PITH_FULL_IMAGE:figures/full_fig_p031_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Prompt used for Query-as-Doc query rewriting. [PITH_FULL_IMAGE:figures/full_fig_p038_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Prompt used for Query-as-Answer query rewriting. [PITH_FULL_IMAGE:figures/full_fig_p038_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Prompt used for Query-as-Reasoning-Step query rewriting. [PITH_FULL_IMAGE:figures/full_fig_p039_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LeanSearch v2: Global Premise Retrieval for Lean 4 Theorem Proving

    cs.IR 2026-05 conditional novelty 7.0

    LeanSearch v2 recovers 46.1% of ground-truth premise groups on research-level Mathlib theorems and raises fixed-loop proof success from 4% to 20% via embedding-reranker plus iterative sketch-retrieve-reflect retrieval.

  2. LeanSearch v2: Global Premise Retrieval for Lean 4 Theorem Proving

    cs.IR 2026-05 conditional novelty 7.0

    LeanSearch v2 recovers 46.1% of ground-truth premise groups for research-level Lean 4 theorems within 10 candidates and raises fixed-loop proof success to 20%.

  3. Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation

    cs.IR 2026-04 unverdicted novelty 7.0

    An LLM simulation framework generates multilingual tip-of-the-tongue queries, validated by rank correlation with real queries, producing the first large-scale ToT benchmarks for four languages.

  4. Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines

    cs.IR 2026-04 unverdicted novelty 5.0

    QPP methods can select query variants that boost end-to-end RAG quality over the original query, though retrieval-optimized variants often fail to produce the best generated answers, revealing a utility gap.

  5. Reproducing Adaptive Reranking for Reasoning-Intensive IR

    cs.IR 2026-04 unverdicted novelty 2.0

    Reproducing GAR on BRIGHT shows it boosts reasoning-intensive retrieval effectiveness with low overhead when the reranker's signal quality is strong.

Reference graph

Works this paper leans on

95 extracted references · 33 canonical work pages · cited by 4 Pith papers · 2 internal anchors

  1. [1]

    https://openai.com/index/introducing-o3-and-o4-mini/

    2025. https://openai.com/index/introducing-o3-and-o4-mini/

  2. [2]

    Saling, Mark Sanderson, Falk Scholer, Damiano Spina, and Ryen W

    Marwah Alaofi, Luke Gallagher, Dana Mckay, Lauren L. Saling, Mark Sanderson, Falk Scholer, Damiano Spina, and Ryen W. White. 2022. Where Do Queries Come From?. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, Madrid Spain, 2850–2862. doi:10.1145/3477495.3531711

  3. [3]

    Voorhees

    James Allan, Donna Harman, Evangelos Kanoulas, Dan Li, Christophe Van Gysel, and Ellen M. Voorhees. 2017. TREC 2017 Common Core Track Overview. InProceedings of The Twenty-Sixth Text REtrieval Conference, TREC 2017, 38 Julian Killingback and Hamed Zamani Query-as-Doc Prompt Follow the below steps to produce a document that is relevant to the main informat...

  4. [4]

    Identify the type of document the user is trying to retrieve (and thus what should be output)

  5. [5]

    Analyze the request, what kinds of information is the user looking for? What kinds of documents would exist that would satisfy the request?

  6. [6]

    Do additional brainstorming about what a relevant document would be

  7. [7]

    The document should be in the style and format mentioned in the first line of the information request

    You must do the brainstorming and analysis steps! Once the full analysis and brainstorm- ing steps have been completed output ##final output followed by the final document (and no additional comments, information, or anything else) on a new line which would be relevant based on the information request. The document should be in the style and format mentio...

  8. [8]

    Analyze the request, what kinds of information is the user looking for? What kinds of information would a complete answer have?

  9. [9]

    Do additional brainstorming about what might be relevant and what might not be

  10. [10]

    Information Request:<instruction> <query> Fig

    You must do the brainstorming and analysis steps! Once the full analysis and brainstorm- ing steps have been completed output ##final output followed by the final answer (and no additional comments, information, or anything else) on a new line which should answer the information request. Information Request:<instruction> <query> Fig. 8. Prompt used for Qu...

  11. [11]

    Identify the essential problem in the information request

  12. [12]

    Think step by step to reason about what should be included in relevant documents

  13. [13]

    Information Request:<instruction> <query> Fig

    Draft an answer. Information Request:<instruction> <query> Fig. 9. Prompt used for Query-as-Reasoning-Step query rewriting

  14. [14]

    Jaime Arguello, Samarth Bhargav, Fernando Diaz, Evangelos Kanoulas, To Eun Kim, Yifan He, and Bhaskar Mitra. 2025. Overview of the TREC 2024 Tip-of-the-Tongue Track. InProceedings of the Thirty-Third Text REtrieval Conference

  15. [15]

    Jaime Arguello, Samarth Bhargav, Fernando Diaz, Evangelos Kanoulas, and Bhaskar Mitra. 2023. Overview of the TREC 2023 Tip-of-the-Tongue Track. InThe Thirty-Second Text REtrieval Conference Proceedings (TREC 2023), Gaithersburg, MD, USA, November 14-17, 2023 (NIST Special Publication, Vol. 500-xxx), Ian Soboroff and Angela Ellis (Eds.). National Institute...

  16. [16]

    Jaime Arguello, Adam Ferguson, Emery Fine, Bhaskar Mitra, Hamed Zamani, and Fernando Diaz. 2021. Tip of the Tongue Known-Item Retrieval: A Case Study in Movie Identification. InCHIIR ’21: ACM SIGIR Conference on Human Information Interaction and Retrieval, Canberra, ACT, Australia, March 14-19, 2021, Falk Scholer, Paul Thomas, David Elsweiler, Hideo Joho,...

  17. [17]

    Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2023. Task-aware Retrieval with Instructions. InFindings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Co...

  18. [18]

    Adrien Barbaresi. 2021. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. InProceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations. Association for Computation...

  19. [19]

    Shariq Bashir and Andreas Rauber. 2011. On the relationship between query characteristics and IR functions retrieval bias.Journal of the American Society for Information Science and Technology62, 8 (2011), 1515–1532. doi:10.1002/asi.21549 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/asi.21549

  20. [20]

    Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. 2023. TheoremQA: A Theorem-driven Question Answering Dataset. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 78...

  21. [21]

    Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020. HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data. InFindings of the Association for Computa- tional Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 1026–1036. doi...

  22. [22]

    Bennett, Fernando Diaz, Charlie Clarke, and Ellen M

    Kevyn Collins-Thompson, Paul N. Bennett, Fernando Diaz, Charlie Clarke, and Ellen M. Voorhees. 2013. TREC 2013 Web Track Overview. InProceedings of The Twenty-Second Text REtrieval Conference, TREC 2013, Gaithersburg, Maryland, USA, November 19-22, 2013 (NIST Special Publication, Vol. 500-302), Ellen M. Voorhees (Ed.). National Institute of Standards and ...

  23. [23]

    Bennett, Fernando Diaz, and Ellen M

    Kevyn Collins-Thompson, Craig Macdonald, Paul N. Bennett, Fernando Diaz, and Ellen M. Voorhees. 2014. TREC 2014 Web Track Overview. InProceedings of The Twenty-Third Text REtrieval Conference, TREC 2014, Gaithersburg, Maryland, USA, November 19-21, 2014 (NIST Special Publication, Vol. 500-308), Ellen M. Voorhees and Angela Ellis (Eds.). National Institute...

  24. [24]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020. Overview of the TREC 2020 Deep Learning Track. InProceedings of the Twenty-Ninth Text REtrieval Conference, TREC 2020, Virtual Event [Gaithersburg, Maryland, 40 Julian Killingback and Hamed Zamani Table 11. Example query and relevant passage for the Clinical Trial task. Query Patient is ...

  25. [25]

    Voorhees

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2020. Overview of the TREC 2019 deep learning track.CoRRabs/2003.07820 (2020). arXiv:2003.07820 https://arxiv.org/abs/2003.07820

  26. [26]

    Zhuyun Dai and Jamie Callan. 2019. Deeper Text Understanding for IR with Contextual Neural Language Modeling. InProceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2019, Paris, France, July 21-25, 2019, Benjamin Piwowarski, Max Chevalier, Éric Gaussier, Yoelle Maarek, Jian-Yun Nie, and Fal...

  27. [27]

    data.txt

    Zhuyun Dai, Vincent Y. Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B. Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot Dense Retrieval From 8 Examples. InThe Eleventh International Julian Killingback and Hamed Zamani 41 Table 12. Example query and relevant passage for the Code Retrieval task. Query On the way to school...

  28. [28]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai D...

  29. [29]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Vol...

  30. [30]

    Abhimanyu Dubey et al. 2024. The Llama 3 Herd of Models.CoRRabs/2407.21783 (2024). arXiv:2407.21783 doi:10. 48550/ARXIV.2407.21783

  31. [31]

    Kenneth C. Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzeminski, Genta Indra Winata, Saba Sturua, Saiteja Utpala, Mathieu Ciancone, Marion Schaeffer, Diganta Misra, Shreeya Dhakal, Jonathan Rystrøm, Roman Solomatin, Ömer Veysel Çagatan, Akash Kundu, and et al. Julian Killingback and...

  32. [32]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research.Tr...

  33. [33]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. InSIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, Fernando Diaz, Chirag Shah, Torsten Suel, Pablo Castells, Rosie Jones...

  34. [34]

    Simeng Han, Frank Palma Gomez, Tu Vu, Zefei Li, Daniel Cer, Hansi Zeng, Chris Tar, Arman Cohan, and Gus- tavo Hernandez Abrego. 2025. ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models. doi:10.48550/arXiv.2502.16766 arXiv:2502.16766 [cs]

  35. [35]

    Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2...

  36. [36]

    Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot Learning with Retrieval Augmented Language Models.J. Mach. Learn. Res.24 (2023), 251:1–251:43. https://jmlr.org/papers/v24/23-0037.html

  37. [37]

    Rolf Jagerman, Honglei Zhuang, Zhen Qin, Xuanhui Wang, and Michael Bendersky. 2023. Query Expansion by Prompting Large Language Models.CoRRabs/2305.03653 (2023). arXiv:2305.03653 doi:10.48550/ARXIV.2305.03653

  38. [38]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated gain-based evaluation of IR techniques.ACM Trans. Inf. Syst. 20, 4 (2002), 422–446. doi:10.1145/582415.582418

  39. [39]

    Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025. Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.CoRRabs/2503.09516 (2025). arXiv:2503.09516 doi:10.48550/ARXIV.2503.09516

  40. [40]

    Haocheng Ju and Bin Dong. 2025. MIRB: Mathematical Information Retrieval Benchmark.CoRRabs/2505.15585 (2025). arXiv:2505.15585 doi:10.48550/ARXIV.2505.15585

  41. [41]

    Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean-Bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Beyer, Xiaohai ...

  42. [42]

    What kind of car do you have?\

    Julian Killingback, Hansi Zeng, and Hamed Zamani. 2025. Hypencoder: Hypernetworks for Information Retrieval. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (Padua, Italy)(SIGIR ’25). Association for Computing Machinery, New York, NY, USA, 2372âĂŞ2383. doi:10.1145/ 46 Julian Killingback and...

  43. [43]

    Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. InProceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co-located with the 30th Annual Conference on Neural Information Processing Systems...

  44. [44]

    Lei Li, Xiao Zhou, and Zheng Liu. 2025. R2MED: A Benchmark for Reasoning-Driven Medical Retrieval.CoRR abs/2505.14558 (2025). arXiv:2505.14558 doi:10.48550/ARXIV.2505.14558

  45. [45]

    Xiangyang Li, Kuicai Dong, Yi Quan Lee, Wei Xia, Hao Zhang, Xinyi Dai, Yasheng Wang, and Ruiming Tang. 2025. CoIR: A Comprehensive Benchmark for Code Information Retrieval Models. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Wanxiang ...

  46. [46]

    Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards General Text Embeddings with Multi-stage Contrastive Learning.CoRRabs/2308.03281 (2023). arXiv:2308.03281 doi:10.48550/ ARXIV.2308.03281

  47. [47]

    Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. InProceedings of the 44th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 48 Julian Kill...

  48. [48]

    Kevin Lin, Kyle Lo, Joseph Gonzalez, and Dan Klein. 2023. Decomposing Complex Queries for Tip-of-the-tongue Retrieval. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5521–5533. doi:10.18653/v1/2023.findings- emnlp.367

  49. [49]

    Chaitanya Malaviya, Peter Shaw, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2023. QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, Anna Rogers, Jordan L. Bo...

  50. [50]

    Christopher Malon and Bing Bai. 2020. Generating Followup Questions for Interpretable Multi-hop Question Answering. CoRRabs/2002.12344 (2020). arXiv:2002.12344 https://arxiv.org/abs/2002.12344

  51. [51]

    Yazdan Mansourian and Nigel Ford. 2007. Web searchers’ attributions of success and failure: an empirical study. Journal of Documentation63, 5 (Sept. 2007), 659–679. doi:10.1108/00220410710827745

  52. [52]

    Carlo Merola and Jaspinder Singh. 2025. Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation.CoRRabs/2504.19754 (2025). arXiv:2504.19754 doi:10.48550/ARXIV.2504.19754

  53. [53]

    Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023. MTEB: Massive Text Embedding Benchmark. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023, Andreas Vlachos and Isabelle Augenstein (Eds.). Association for Computational Linguistics...

  54. [54]

    Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen

  55. [55]

    Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, H...

  56. [56]

    Loc Pham, Tung Luu, Thu Vo, Minh Nguyen, and Viet Hoang. 2025. VN-MTEB: Vietnamese Massive Text Embedding Benchmark. arXiv:2507.21500 [cs.CL] https://arxiv.org/abs/2507.21500

  57. [57]

    Rafal Poswiata, Slawomir Dadas, and Michal Perelkiewicz. 2024. PL-MTEB: Polish Massive Text Embedding Benchmark. CoRRabs/2405.10138 (2024). arXiv:2405.10138 doi:10.48550/ARXIV.2405.10138

  58. [58]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. https://arxiv.org/abs/1908.10084

  59. [59]

    Snowflake AI Research. 2024. Arctic embed: The best open source embedding models. Snowflake Blog. https://www. snowflake.com/blog/arctic-embed-open-source-embedding-models/ Accessed: June 11, 2025. See also arXiv:2412.04506 for technical details

  60. [60]

    Voorhees, Lucy Lu Wang, and William R

    Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, and William R. Hersh. 2021. Searching for scientific evidence in a pandemic: An overview of TREC-COVID.J. Biomed. Informatics121 (2021), 103865. doi:10.1016/J.JBI.2021.103865

  61. [61]

    Voorhees, Steven Bedrick, and William R

    Kirk Roberts, Dina Demner-Fushman, Ellen M. Voorhees, Steven Bedrick, and William R. Hersh. 2022. Overview of the TREC 2022 Clinical Trials Track. InProceedings of the Thirty-First Text REtrieval Conference, TREC 2022, online, November 15-19, 2022 (NIST Special Publication, Vol. 500-338), Ian Soboroff and Angela Ellis (Eds.). National Institute of Standar...

  62. [62]

    Robertson and Ian Soboroff

    Stephen E. Robertson and Ian Soboroff. 2001. The TREC 2001 Filtering Track Report. InProceedings of The Tenth Text REtrieval Conference, TREC 2001, Gaithersburg, Maryland, USA, November 13-16, 2001 (NIST Special Publication, Vol. 500-250), Ellen M. Voorhees and Donna K. Harman (Eds.). National Institute of Standards and Technology (NIST). http://trec.nist...

  63. [63]

    S. E. Robertson and S. Walker. 1994. Some simple effective approximations to the 2-Poisson model for probabilistic weighted retrieval. InProceedings of the 17th Annual International ACM SIGIR Conference on Research and Development Julian Killingback and Hamed Zamani 49 in Information Retrieval(Dublin, Ireland)(SIGIR ’94). Springer-Verlag, Berlin, Heidelbe...

  64. [64]

    Rulin Shao, Rui Qiao, Varsha Kishore, Niklas Muennighoff, Xi Victoria Lin, Daniela Rus, Bryan Kian Hsiang Low, Sewon Min, Wen-tau Yih, Pang Wei Koh, and Luke Zettlemoyer. 2025. ReasonIR: Training Retrievers for Reasoning Tasks.CoRRabs/2504.20595 (2025). arXiv:2504.20595 doi:10.48550/ARXIV.2504.20595

  65. [65]

    Jianyou (Andre) Wang, Kaicheng Wang, Xiaoyue Wang, Prudhviraj Naidu, Leon Bergen, and Ramamohan Paturi. 2023. Scientific Document Retrieval using Multi-level Aspect-based Queries. InAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 38404–38419. h...

  66. [66]

    arXiv:2503.05592 doi:10.48550/ARXIV.2503.05592

    R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.CoRRabs/2503.05592 (2025). arXiv:2503.05592 doi:10.48550/ARXIV.2503.05592

  67. [67]

    Smith, Luke Zettlemoyer, and Tao Yu

    Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. One Embedder, Any Task: Instruction-Finetuned Text Embeddings. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Comput...

  68. [68]

    Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan Ö

    Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Quan Shi, Zachary S. Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen, and Tao Yu. 2025. BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval. InThe Thirteenth International Conference on Learning Representation...

  69. [69]

    Weiwei Sun, Zhengliang Shi, Wu Long, Lingyong Yan, Xinyu Ma, Yiding Liu, Min Cao, Dawei Yin, and Zhaochun Ren. 2024. MAIR: A Massive Benchmark for Evaluating Instructed Retrieval. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Yaser Al-Onaizan, Mohit Bansal, and Y...

  70. [70]

    Alon Talmor and Jonathan Berant. 2018. The Web as a Knowledge-Base for Answering Complex Questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Marilyn Walker, Heng Ji, and Amanda Stent (Eds.). Association for Computational Lingui...

  71. [71]

    Katherine Thai, Yapei Chang, Kalpesh Krishna, and Mohit Iyyer. 2022. RELiC: Retrieving Evidence for Literary Claims. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, ...

  72. [72]

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2). https://openreview.net/forum?id=wCu6T5xFjeJ

  73. [73]

    Trippas, Sara Fahad Dawood Al Lawati, Joel Mackenzie, and Luke Gallagher

    Johanne R. Trippas, Sara Fahad Dawood Al Lawati, Joel Mackenzie, and Luke Gallagher. 2024. What do Users Really Ask Large Language Models? An Initial Log Analysis of Google Bard Interactions in the Wild. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 1...

  74. [74]

    Coen van den Elsen, Francien Barkhof, Thijmen Nijdam, Simon Lupart, and Mohammad Aliannejadi. 2025. Reproducing NevIR: Negation in Neural Information Retrieval. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025, Padua, Italy, July 13-18, 2025, Nicola Ferro, Maria Maistro, Gabriell...

  75. [75]

    Srihari Vemuru, Eric John, and Shrisha Rao. 2021. Handling Complex Queries Using Query Trees. doi:10.36227/techrxiv. 14845212.v1

  76. [76]

    Ellen Voorhees. 2005. Overview of the TREC 2004 Robust Retrieval Track. doi:10.6028/NIST.SP.500-261

  77. [77]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (...

  78. [78]

    Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query Expansion with Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6- 10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, 9414–9423. doi:10.18653/V1/2023....

  79. [79]

    Xiaoyue Wang, Jianyou Wang, Weili Cao, Kaicheng Wang, Ramamohan Paturi, and Leon Bergen. 2024. BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives.CoRRabs/2402.14151 (2024). arXiv:2402.14151 doi:10.48550/ARXIV.2402.14151 50 Julian Killingback and Hamed Zamani

  80. [80]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. (2023). https: //openreview.net/forum?id=1PL1NIMMrw

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.