REVIEW 3 major objections 4 minor 3 cited by
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation
T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Rankify claims to unify retrieval, re-ranking, and retrieval-augmented generation in one modular Python toolkit, backed by 40 pre-retrieved datasets and prebuilt Wikipedia and MS MARCO indexes, with retrieval outputs matching Pyserini and…
desk verdict Useful toolkit paper, but it overstates what it ships and the validation tables have several correctable internal inconsistencies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a standardized data schema — each query carries Question, Answers, and Contexts, where Contexts is a ranked list of retrieved documents — combined with a small set of unified Python interfaces (Dataset, Retriever, Reranker, Generator, Metrics). The prebuilt Wikipedia and MS MARCO indexes and the pre-retrieved files of 1,000 documents per query are what let users skip indexing and reproduce experiments immediately; the unified interfaces are what let them swap methods without rewriting pipelines.
What would settle it
Download the package and the linked pre-retrieved dataset collection, pick one query from each of the 40 named datasets, and check that 1,000 documents exist for each of BM25, DPR, ANCE, BGE, Contriever, and ColBERT, and that the advertised MS MARCO index can be loaded and queried; any missing combination would mean the 20,160-configuration claim overstates what the toolkit currently delivers.
Extended reading notes
Core claim
The paper's central claim is that Rankify is a modular, extensible toolkit that makes the standard retriever-then-reranker pipeline, and RAG on top of it, a single coherent workflow. Concretely, Rankify provides precomputed retrieval corpora for Wikipedia and MS MARCO, standardized access to 40 benchmark datasets across QA, multi-hop QA, fact verification, entity linking, dialog, and summarization tasks, six dense and one sparse retriever, and 24 re-ranking models spanning pointwise, pairwise, and listwise strategies, plus zero-shot, FiD, and in-context RAG generation. The paper states that because BM25 and DPR use the Pyserini backend, Rankify reproduces Pyserini's results exactly, and that independently implemented Contriever and ColBERT retrievers match their official GitHub results, with ANCE close. On that basis it presents the toolkit as a consistent environment for benchmarking retrieval and ranking methods.
Load-bearing premise
The load-bearing premise is that every advertised artifact actually exists in usable form: all 40 datasets ship with 1,000 pre-retrieved documents per query for each listed retriever, and the prebuilt Wikipedia and MS MARCO indexes load correctly; the paper's own footnote 8 says the MS MARCO pre-retrieval is still being processed.
Editorial extensions
If this is right
- A researcher can go from pip install rankify to comparing dense and sparse retrieval plus re-ranking on standard QA benchmarks without building or storing indexes.
- BM25 and DPR results reported by Rankify are directly comparable to Pyserini-based runs, since they share the same indexing backend.
- Because every reranker sees the same 1,000 pre-retrieved candidates, accuracy differences in re-ranking benchmarks reflect model differences rather than retrieval setup differences.
- The prebuilt Wikipedia and MS MARCO indexes remove one of the main cost barriers to reproducing large-scale retrieval experiments.
Reading between the lines
- If the full set of pre-retrieved files is completed, this collection could become a shared substrate for reranking research, because all methods would re-rank the same candidate lists rather than lists produced by different local setups.
- The validation of BM25 and DPR is inherited from Pyserini by design, so the independent correctness evidence rests mainly on the Contriever and ColBERT matches; a fuller validation would also reproduce ANCE exactly and add reranker output checks.
- The stated 20,160 configurations assume every dataset-retriever combination is shipped; until the MS MARCO processing in footnote 8 is done, the genuinely available combinations are a subset, and the effective count is lower.
- A testable extension the paper does not run: verify that the pre-retrieved BM25 lists are identical to live BM25 retrieval on the bundled index, since any drift would propagate into every reranking result built on the pre-retrieved files.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper describes Rankify, a modular Python toolkit that unifies sparse and dense retrieval, re-ranking, and retrieval-augmented generation (RAG). The toolkit is advertised as including 40 pre-retrieved datasets, 7 retrievers, 24 re-ranking models, and 3 RAG methods, together with prebuilt Wikipedia and MS MARCO indexes. The authors report BM25 retrieval accuracy across many datasets, compare dense retrievers on NQ, WebQ, and TriviaQA, validate their outputs against Pyserini and official model implementations, evaluate re-ranking models on BM25 top-100 outputs, and present RAG exact-match scores for several LLMs. The central claim is that Rankify provides a comprehensive, unified, and reproducible benchmarking environment for retrieval, re-ranking, and RAG.
Significance. If the announced artifacts are complete, Rankify is a useful contribution: it offers a single installable package with a consistent API over a wide range of retrievers, re-rankers, and RAG methods, and the pre-retrieved datasets would save substantial preprocessing effort. The paper ships open-source code, a PyPI package, documentation, and sanity checks against Pyserini and official model repositories, which is a valuable verification step for a resource paper. The main risks are the incompleteness of the advertised pre-retrieved resource and imprecision in some validation claims; both are addressable and do not, in my view, invalidate the underlying toolkit idea.
major comments (3)
- [Section 3.2, footnote 8] Footnote 8 states: 'We are currently processing different retrievers based on MS MARCO to generate and store 1,000 top-ranked documents per query for each dataset.' This directly contradicts the contribution bullet claiming '40 datasets, each with 1,000 pre-retrieved documents per query' and undermines Table 1's 20,160 configuration count, which assumes all dataset-retriever combinations are available. Since the pre-retrieved datasets are the core reproducibility resource, please either complete the resource, provide a per-dataset/per-retriever inventory of what is currently available, or revise the claims to reflect the current state.
- [Table 2 and Section 1] By my count, Table 2 lists 34 named datasets (14 QA, 4 multi-hop QA, 2 temporal QA, 2 long-form QA, 5 multiple-choice, 2 entity-linking, 2 slot filling, 1 dialog generation, 1 fact verification, 1 summarization), while the abstract, the contribution bullet, and Table 1 all state 40 datasets. This discrepancy is material to the advertised scope; please correct the count or complete the table, and ensure consistency across the abstract, Section 1, and Table 1.
- [Table 5 and Section 4.3] The caption of Table 5 says that 'Rankify implementation achieves results identical to Pyserini across all retrievers,' and the text claims identical results for DPR and BM25, but the table itself shows non-identical values: Contriever_Rankify TriviaQA Top-20 is 80.3 versus 80.4 for the official result, and ANCE_Rankify NQ Top-20/Top-100 are 82.5/88.5 versus 82.1/87.9 for the official ANCE row, with WebQ values reported only for the Rankify row. Please report the exact per-dataset differences and explain the ANCE discrepancy (e.g., index version, query processing, or evaluation split) or re-run the comparison so that the validation claim is precise.
minor comments (4)
- [Section 5] The final section contains a duplicated sentence: 'Future work will focus on improving retrieval efficiency, optimizing re-ranking strategies, and advancing RAG capabilities' appears twice, and the second appearance should be removed.
- [References and Sections 1, 3.3] DPR is cited as [109] in Sections 1 and 3.3, but reference [109] is the WSDM 2021 paper by Yates, Nogueira, and Lin; the correct citation for DPR is the Karpukhin et al. paper listed as [42]. Please fix these citations.
- [Figure 3] In Figure 3, the retriever list contains 'BGB', which appears to be a typo for 'BGE'; please correct the label.
- [Section 4.2] The text in Section 4.2 mentions 'MSS-DPR' and 'MSS' models, but Table 4 does not include these models; please align the text with the reported experiments.
Circularity Check
No significant circularity: the toolkit's benchmarks are measured against external baselines or official implementations, and the one self-referential Pyserini-backend check is explicitly disclosed and non-load-bearing.
full rationale
Rankify is a software/toolkit paper, not a derivation: there are no fitted parameters, no equations whose outputs are defined by their inputs, and no theoretical claim that is forced by a self-citation. The retrieval and re-ranking numbers in Tables 3-6 are either measured on standard datasets with the paper's own pipeline or benchmarked against Pyserini and official model repositories (e.g., Table 5 compares DPR, BM25, Contriever, ColBERT, and ANCE rows against Pyserini and official GitHub results). The only self-referential element is Table 5's statement that Rankify's DPR/BM25 numbers are identical to Pyserini 'as we use Pyserini as the indexing backend'; because the equality is by construction for those two rows, it is not independent validation, but the paper discloses this exactly and the toolkit's central usability claim does not rest on that equality being an external discovery. Footnote 8's admission that MS MARCO-based 1,000-document pre-retrievals are still being generated is a completeness/reproducibility limitation, not circularity. The few self-citations (ASRank, DynRank, ChroniclingAmericaQA, ArchivalQA) are context references and do not carry the argument. Verdict: no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Pyserini and the official model repositories provide correct reference results.
- domain assumption Top-k retrieval accuracy can be measured by whether a retrieved passage contains the answer string.
- ad hoc to paper The pre-retrieved dataset collection is complete and correctly generated.
Cite this review
Pith. "Pith review of Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation." pith.science (2026). https://pith.science/paper/HCIT3CN3
@misc{pith2026250202464,
author = {Pith},
title = {Pith review of: Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/HCIT3CN3}},
note = {Machine review of arXiv:2502.02464}
}
read the original abstract
Retrieval, re-ranking, and retrieval-augmented generation (RAG) are critical components of modern applications in information retrieval, question answering, or knowledge-based text generation. However, existing solutions are often fragmented, lacking a unified framework that easily integrates these essential processes. The absence of a standardized implementation, coupled with the complexity of retrieval and re-ranking workflows, makes it challenging for researchers to compare and evaluate different approaches in a consistent environment. While existing toolkits such as Rerankers and RankLLM provide general-purpose reranking pipelines, they often lack the flexibility required for fine-grained experimentation and benchmarking. In response to these challenges, we introduce Rankify, a powerful and modular open-source toolkit designed to unify retrieval, re-ranking, and RAG within a cohesive framework. Rankify supports a wide range of retrieval techniques, including dense and sparse retrievers, while incorporating state-of-the-art re-ranking models to enhance retrieval quality. Additionally, Rankify includes a collection of pre-retrieved datasets to facilitate benchmarking, available at Huggingface (https://huggingface.co/datasets/abdoelsayed/reranking-datasets-light). To encourage adoption and ease of integration, we provide comprehensive documentation (http://rankify.readthedocs.io/), an open-source implementation on GitHub (https://github.com/DataScienceUIBK/rankify), and a PyPI package for easy installation (https://pypi.org/project/rankify/). As a unified and lightweight framework, Rankify allows researchers and practitioners to advance retrieval and re-ranking methodologies while ensuring consistency, scalability, and ease of use.
Figures
Forward citations
Cited by 3 Pith papers
-
How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
On a new benchmark of post-April 2025 queries, LLM rerankers show a 5-15% performance drop compared with familiar benchmarks, and lightweight models match them on efficiency and sometimes accuracy.
-
Shifting from Ranking to Set Selection for Retrieval Augmented Generation
SETR identifies a query's information requirements with chain-of-thought reasoning and selects a compact passage set, improving multi-hop RAG accuracy over fixed-top-k reranking baselines.
-
RankLLM: A Python Package for Reranking with LLMs
RankLLM is an open-source Python package that modularly supports pointwise, pairwise, and listwise LLM rerankers, with integrated retrieval, evaluation, training, and response analysis, and reproduces results from Ran...
Reference graph
Works this paper leans on
-
[1]
Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Man- junath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: a system f...
2016
-
[2]
Amar Abane, Anis Bekri, and Abdella Battou. 2024. FastRAG: Retrieval Aug- mented Generation for Semi-structured Data. arXiv preprint arXiv:2411.13773 (2024)
work page Pith review arXiv 2024
-
[3]
Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani, and Adam Jatowt
-
[4]
Abdelrahman Elsayed Mahmoud Abdallah, Jamshid Mozafari, Bhawna Piryani, Mohammed M.Abdelgwad, and Adam Jatowt. 2025. DynRank: Improve Passage Retrieval with Dynamic Zero-Shot Prompting Based on Question Classification. In Proceedings of the 31st International Conference on Computational Linguis- tics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-...
2025
-
[5]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[6]
Davide Baldelli, Junfeng Jiang, Akiko Aizawa, and Paolo Torroni. 2024. TWOLAR: a TWO-step LLM-Augmented distillation method for passage Rerank- ing. In European Conference on Information Retrieval . Springer, 470–485
2024
-
[7]
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bah- danau, Nicolas Chapados, and Siva Reddy. 2024. Llm2vec: Large language models are secretly powerful text encoders. arXiv preprint arXiv:2404.05961 (2024)
arXiv 2024
-
[8]
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic Parsing on Freebase from Question-Answer Pairs. InProceedings of the 2013 Con- ference on Empirical Methods in Natural Language Processing , David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu, and Steven Bethard (Eds.). As- sociation for Computational Linguistics, Seattl...
2013
Show all 125 references
-
[9]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. PIQA: Reasoning about Physical Commonsense in Natural Language. In AAAI Conference on Artificial Intelligence. https://api.semanticscholar.org/CorpusID: 208290939
2019
-
[10]
Christopher JC Burges. 2010. From ranknet to lambdarank to lambdamart: An overview. Learning 11, 23-581 (2010), 81
2010
-
[11]
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning . 129–136
2007
-
[12]
Harrison Chase. 2022. LangChain: Building applications with LLMs through composability. https://www.langchain.com Available at https://www.langchain. com
2022
-
[14]
Shijie Chen, Bernal Jiménez Gutiérrez, and Yu Su. 2024. Attention in Large Lan- guage Models Yields Efficient Zero-Shot Re-Rankers. arXiv:2410.02642 [cs.CL] https://arxiv.org/abs/2410.02642
2024 arXiv
-
[15]
Tao Chen, Mingyang Zhang, Jing Lu, Michael Bendersky, and Marc Najork
-
[16]
Gobinda G Chowdhury. 2010. Introduction to modern information retrieval. Facet publishing
2010
-
[17]
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In NAACL
2019
-
[18]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question An- swering? Try ARC, the AI2 Reasoning Challenge. CoRR abs/1803.05457 (2018). arXiv:1803.05457 http://arxiv.org/abs/1803.05457
2018 arXiv
-
[19]
Benjamin Clavié. 2024. rerankers: A Lightweight Python Library to Unify Ranking Methods. arXiv:2408.17344 [cs.IR] https://arxiv.org/abs/2408.17344
2024 arXiv
-
[20]
W Bruce Croft, Donald Metzler, and Trevor Strohman. 2010. Search engines: Information retrieval in practice . Vol. 520. Addison-Wesley Reading
2010
-
[21]
Damodaran
P. Damodaran. 2024. FlashRank, Lightest and Fastest 2nd Stage Reranker for search pipelines. https://doi.org/10.5281/zenodo.11093524
2024 doi
-
[22]
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of Wikipedia: Knowledge-Powered Conversational Agents. In International Conference on Learning Representations . https://openreview. net/forum?id=r1l73iRqKm
2019
-
[23]
Hare, Frédérique Laforest, and Elena Simperl
Hady ElSahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon S. Hare, Frédérique Laforest, and Elena Simperl. 2018. T-REx: A Large Scale Alignment of Natural Language with Knowledge Base Triples. In Proceedings of the Eleventh International Conference on Langu...
2018
-
[24]
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long Form Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Ko- rhonen, David Traum, and Lluís Màrquez (Eds....
2019 doi
-
[25]
Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu
-
[26]
Luyu Gao, Zhuyun Dai, Tongfei Chen, Zhen Fan, Benjamin Van Durme, and Jamie Callan. 2021. Complement Lexical Retrieval Model with Semantic Residual Embeddings. In Advances in Information Retrieval - 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28 - ...
2021
-
[27]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997 (2023)
2023 arXiv
-
[28]
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Be- rant. 2021. Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies. Transactions of the Association for Computational Linguistics 9 (2021), 346–361. https://do...
2021 doi
-
[29]
Hiroaki Hayashi, Prashant Budania, Peng Wang, Chris Ackerson, Raj Neer- vannan, and Graham Neubig. 2020. WikiAsp: A Dataset for Multi-domain Aspect-based Summarization. Transactions of the Association for Computational Linguistics (TACL) (2020). https://arxiv.org/abs/2011.07832
2020 arXiv
-
[30]
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. Aligning AI With Shared Human Values. Proceedings of the International Conference on Learning Representations (ICLR) (2021)
2021
-
[31]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Under- standing. Proceedings of the International Conference on Learning Representations (ICLR) (2021)
2021
-
[32]
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reason- ing Steps. In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computatio...
2020
-
[33]
Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Man- fred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum
-
[34]
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al
-
[35]
Chao-Wei Huang and Yun-Nung Chen. 2024. PairDistill: Pairwise Relevance Distillation for Dense Retrieval. arXiv preprint arXiv:2410.01383 (2024)
2024 arXiv
- [36]
-
[37]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised Dense Infor- mation Retrieval with Contrastive Learning. arXiv:2112.09118 (2021)
2021 arXiv
-
[38]
Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282 (2020)
2020 arXiv
-
[39]
Jiajie Jin, Yutao Zhu, Xinyu Yang, Chenghao Zhang, and Zhicheng Dou. 2024. FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research. arXiv preprint arXiv:2405.13576 (2024)
2024 arXiv
-
[40]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. Trivi- aQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Com- prehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ...
2017 doi
-
[41]
Ashwin Kalyan, Abhinav Kumar, Arjun Chandrasekaran, Ashish Sabharwal, and Peter Clark. 2021. How Much Coffee Was Consumed During EMNLP 2019? Fermi Problems: A New Reasoning Challenge for AI. arXiv preprint arXiv:2110.14207 (2021)
2021 arXiv
-
[42]
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . https://www....
2020
-
[43]
Danupat Khamnuansin, Tawunrat Chalothorn, and Ekapol Chuangsuwanich
-
[44]
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al. 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310....
2023 arXiv
-
[45]
Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. arXiv:2004.12832 [cs.IR] https://arxiv.org/abs/2004.12832
2020 arXiv
-
[46]
Dongkyu Kim, Byoungwook Kim, Donggeon Han, and Matouš Eibich. 2024. AutoRAG: Automated Framework for optimization of Retrieval Augmented Generation Pipeline. arXiv:2410.20878 [cs.CL] https://arxiv.org/abs/2410.20878
2024 arXiv
-
[47]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Ken- ton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Sl...
2019 doi
-
[48]
Thiago Laitz, Konstantinos Papakostas, Roberto Lotufo, and Rodrigo Nogueira
-
[49]
arXiv preprint arXiv:2406.05733 (2024)
MrRank: Improving Question Answering Retrieval System through Multi- Result Ranking Model. arXiv preprint arXiv:2406.05733 (2024)
2024 arXiv
-
[50]
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero- Shot Relation Extraction via Reading Comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017) . Association for Computational Linguistics, Vancouver, Cana...
2017 doi
-
[51]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing...
2020
-
[52]
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. In Proceedings of the 44th Annual International ACM SIGIR Conf...
2021
-
[53]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measur- ing How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), Smaranda Muresan, Preslav Nakov, and Aline Villa...
2022 doi
-
[54]
Jerry Liu. 2022. LlamaIndex. https://doi.org/10.5281/zenodo.1234
2022 doi
-
[55]
arXiv:2401.06910 [cs.IR]
InRanker: Distilled Rankers for Zero-shot Information Retrieval. arXiv:2401.06910 [cs.IR]
-
[56]
Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent Retrieval for Weakly Supervised Open Domain Question Answering. In Proc. ACL. 6086– 6096
2019
-
[57]
Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma, Kaiyan Zhang, Bowen Zhou, and Jie Zhou. 2024. Generative multi-modal knowledge retrieval with large language models. In Proceedings of the AAAI Conference on Artificial Intel- ligence, Vol. 38. 18733–18741
2024
-
[58]
Xueguang Ma, Kai Sun, Ronak Pradeep, Minghan Li, and Jimmy Lin. 2022. Another Look at DPR: Reproduction of Training and Replication of Retrieval. In Advances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Stavanger, Norway, April 10–14, 2022, Pro...
2022
-
[59]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When Not to Trust Language Models: Investigating Effectiveness and Limitations of Parametric and Non-Parametric Memories. arXiv preprint (2022)
2022
-
[60]
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In EMNLP
2018
-
[61]
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering Ambiguous Open-domain Questions. In EMNLP
2020
-
[62]
Yaqi Liu, Xiaoyu Zhang, Xiaobin Zhu, Qingxiao Guan, and Xianfeng Zhao. 2017. Listnet-based object proposals ranking. Neurocomputing 267 (2017), 182–194
2017
-
[63]
Zheng Liu, Yujia Zhou, Yutao Zhu, Jianxun Lian, Chaozhuo Li, Zhicheng Dou, Defu Lian, and Jian-Yun Nie. 2024. Information Retrieval Meets Large Language Models. In Companion Proceedings of the ACM on Web Conference 2024 . 1586– 1589
2024
-
[64]
Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085 (2019)
2019 arXiv
-
[65]
Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. arXiv preprint arXiv:2003.06713 (2020)
2020 arXiv
-
[66]
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-stage document ranking with BERT. arXiv preprint arXiv:1910.14424 (2019)
2019 arXiv
-
[67]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu F...
2019
-
[68]
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Mail- lard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a Benchmark for Knowledge Intensive Language Ta...
2021
-
[69]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2017. MS MARCO: A Human-Generated MAchine Reading COmprehension Dataset. https://openreview.net/forum?id=Hk1iOLcle
2017
-
[70]
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers. arXiv preprint arXiv:2112.07899 (2021)
2021 arXiv
-
[71]
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! arXiv preprint arXiv:2312.02724 (2023)
2023 arXiv
-
[72]
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and Narrowing the Compositionality Gap in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.)....
2023 doi
-
[73]
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al. 2023. Large language models are effective text rankers with pairwise ranking prompting. arXiv preprint arXiv:2306.17563 (2023)
2023 arXiv
-
[74]
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2021 Conference of the North Amer...
2021 doi
-
[75]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67
2020
-
[76]
Bhawna Piryani, Jamshid Mozafari, and Adam Jatowt. 2024. Chroniclingamer- icaqa: A large-scale question answering dataset based on historical american newspaper pages. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retr...
2024
-
[77]
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models. arXiv:2309.15088 [cs.IR] https://arxiv.org/abs/2309.15088
2023 arXiv
-
[78]
Muhammad Shihab Rashid, Jannat Ara Meem, Yue Dong, and Vagelis Hristidis
-
[79]
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. Rocketqav2: A joint training method for dense passage retrieval and passage re-ranking. arXiv preprint arXiv:2110.07367 (2021)
2021 arXiv
-
[80]
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al . 1995. Okapi at TREC-3. Nist Special Publication Sp 109 SIGIR ’25, July 13–18, 2025, Padova, IT Abdallah et al. (1995), 109
1995
-
[81]
Tomáš Koˇ ciský, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Her- mann, Gábor Melis, and Edward Grefenstette. 2018. The NarrativeQA Reading Comprehension Challenge. Transactions of the Association for Computational Linguistics TBD (2018), TBD. https://TBD
2018
-
[82]
Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen- tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving Passage Retrieval with Zero-Shot Question Generation. (2022). https://arxiv.org/abs/2204.07496
2022 arXiv
-
[83]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Asso...
2016 doi
-
[84]
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. 2023. In-Context Retrieval-Augmented Lan- guage Models. Transactions of the Association for Computational Linguistics 11 (2023), 1316–1331. https://doi.org/10.1162/tacl_a_00605
2023 doi
-
[85]
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021. Sim- ple Entity-centric Questions Challenge Dense Retrievers. In Empirical Methods in Natural Language Processing (EMNLP)
2021
-
[86]
arXiv preprint arXiv:2402.10866 (2024)
EcoRank: Budget-Constrained Text Re-ranking Using Large Language Models. arXiv preprint arXiv:2402.10866 (2024)
2024 arXiv
-
[87]
Nilanjan Sinhababu, Andrew Parry, Debasis Ganguly, Debasis Samanta, and Pabitra Mitra. 2024. Few-shot Prompting for Pairwise Ranking: An Effective Non-Parametric Retrieval Model. arXiv preprint arXiv:2409.17745 (2024)
2024 arXiv
-
[88]
Raphaël Sourty, Jose G Moreno, Lynda Tamine, and François-Paul Servant. 2022. Cherche: A new tool to rapidly implement pipelines in information retrieval. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3283–3288
2022
-
[89]
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid Questions Meet Long-Form Answers. In Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association f...
2022
-
[90]
Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agent. ArXiv abs/2304.09542 (2023)
2023 arXiv
-
[91]
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021. Colbertv2: Effective and efficient retrieval via lightweight late interaction. arXiv preprint arXiv:2112.01488 (2021)
2021 arXiv
-
[92]
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi
-
[93]
Manveer Singh Tamber, Ronak Pradeep, and Jimmy Lin. 2023. Scaling Down, LiT- ting Up: Efficient Zero-Shot Listwise Reranking with Seq2seq Encoder-Decoder Models. arXiv preprint arXiv: 2312.16098 (2023)
2023 arXiv
-
[94]
Simone Tedeschi, Simone Conia, Francesco Cecconi, and Roberto Navigli. 2021. Named Entity Recognition for Entity Linking: What Works and What’s Next. In Findings of the Association for Computational Linguistics: EMNLP 2021 . Associa- tion for Computational Linguistics, Punta C...
2021
-
[95]
Amit Singhal et al. 2001. Modern information retrieval: A brief overview. IEEE Data Eng. Bull. 24, 4 (2001), 35–43
2001
-
[96]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[97]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal
-
[98]
Jiajia Wang, Jimmy Xiangji Huang, Xinhui Tu, Junmei Wang, Angela Jennifer Huang, Md Tahmid Rahman Laskar, and Amran Bhuiyan. 2024. Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges. Comput. Surveys 56, 7 (2024), 1–33
2024
-
[99]
Jiexin Wang, Adam Jatowt, and Masatoshi Yoshikawa. 2022. Archivalqa: A large- scale benchmark dataset for open-domain question answering over historical news collections. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information R...
2022
-
[100]
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019. QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions. In Proceed- ings of the 2019 Conference on Empirical Methods in Natural Language Process- ing and the 9th International Joint Conference on Nat...
2019 doi
-
[101]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019 doi
-
[102]
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv:2309.07597 [cs.CL]
2023 arXiv
-
[103]
Haoyi Xiong, Jiang Bian, Yuchen Li, Xuhong Li, Mengnan Du, Shuaiqiang Wang, Dawei Yin, and Sumi Helal. 2024. When search engine services meet large language models: visions and challenges. IEEE Transactions on Services Computing (2024)
2024
-
[104]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal
-
[105]
Ikuya Yamada, Akari Asai, and Hannaneh Hajishirzi. 2021. Efficient Passage Retrieval with Hashing for Open-domain Question Answering. In ACL
2021
-
[106]
Liu Yang, Junjie Hu, Minghui Qiu, Chen Qu, Jianfeng Gao, W Bruce Croft, Xi- aodong Liu, Yelong Shen, and Jingjing Liu. 2019. A hybrid retrieval-generation neural conversation model. In Proceedings of the 28th ACM international confer- ence on information and knowledge manageme...
2019
-
[107]
Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Proceedings of the 2015 Con- ference on Empirical Methods in Natural Language Processing , Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association ...
2015 doi
-
[108]
Transactions of the Association for Computational Linguistics (2022)
MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics (2022)
2022
-
[109]
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021. Pretrained transformers for text ranking: BERT and beyond. In Proceedings of the 14th ACM International Conference on web search and data mining . 1154–1156
2021
-
[110]
Soyoung Yoon, Eunbi Choi, Jiyeon Kim, Hyeongu Yun, Yireun Kim, and Seung won Hwang. 2024. ListT5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot Retrieval. arXiv:2402.15838 [cs.IR] https://arxiv.org/abs/2402.15838
2024 arXiv
-
[111]
Wenhui Wang, Hangbo Bao, Shaohan Huang, Li Dong, and Furu Wei. 2020. Minilmv2: Multi-head self-attention relation distillation for compressing pre- trained transformers. arXiv preprint arXiv:2012.15828 (2020)
2020 arXiv
-
[112]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171 (2022)
2022 arXiv
-
[113]
Hansi Zeng, Hamed Zamani, and Vishwa Vinay. 2022. Curriculum Learning for Dense Retrieval Distillation. In Proc. SIGIR. 1979–1983
2022
-
[114]
Nengjun Zhu, Jian Cao, Xinjiang Lu, and Qi Gu. 2021. Leveraging pointwise prediction with learning to rank for top-N recommendation. World Wide Web 24 (2021), 375–396
2021
-
[115]
Bennett, Junaid Ahmed, and Arnold Overwijk
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In Proceedings of the 9th International Conference on Learning Representa...
2021
-
[119]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...
2018
-
[122]
Xiao Yu, Yunan Lu, and Zhou Yu. 2024. LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems. arXiv preprint arXiv:2403.00982 (2024)
2024 arXiv
-
[123]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
2019
-
[126]
Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023. Rankt5: Fine-tuning t5 for text ranking with ranking losses. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inf...
2023
-
[2011]
In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing , Regina Barzilay and Mark Johnson (Eds.)
Robust Disambiguation of Named Entities in Text. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing , Regina Barzilay and Mark Johnson (Eds.). Association for Computational Linguistics, Edinburgh, Scotland, UK., 782–792. https://aclanthol...
2011
-
[2018]
In NAACL-HLT
FEVER: a Large-scale Dataset for Fact Extraction and VERification. In NAACL-HLT
-
[2019]
Social IQa: Commonsense Reasoning about Social Interactions. In Pro- ceedings of the 2019 Conference on Empirical Methods in Natural Language Pro- cessing and the 9th International Joint Conference on Natural Language Process- ing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vinc...
2019 doi
-
[2020]
spaCy: Industrial-strength natural language processing in python. (2020)
2020
-
[2022]
In European Conference on Information Retrieval
Out-of-domain semantics to the rescue! zero-shot hybrid retrieval models. In European Conference on Information Retrieval . Springer, 95–110
-
[2024]
arXiv preprint arXiv:2405.20978 (2024)
Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training. arXiv preprint arXiv:2405.20978 (2024)
2024 arXiv
-
[2025]
arXiv:2501.15245 [cs.CL] https://arxiv.org/abs/2501.15245
ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Re- trieval. arXiv:2501.15245 [cs.CL] https://arxiv.org/abs/2501.15245
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.