REVIEW 3 major objections 5 minor 69 references
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Conventional InfoNCE contrastive fine-tuning often reduces retrieval effectiveness of state-of-the-art dense retrievers; cross-encoder listwise distillation with diverse synthetic queries improves it more consistently.
desk verdict Useful empirical study with a real finding on contrastive fine-tuning, but the listwise-distillation claim is undercut by an omitted control. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Cross-encoder listwise distillation is the load-bearing mechanism: for each synthetic query the student retrieves the top 20 passages, the teacher cross-encoder scores all of them, and the student minimizes the KL divergence between the teacher's normalized score distribution and its own cosine-similarity distribution over the same list, learning graded relevance rather than binary positive/negative labels. The InfoNCE contrastive loss is retained only as a small auxiliary objective (weight 0.1). The supporting machinery is a fully synthetic data pipeline: Llama-3.1 8B generates six query types, a two-stage filter keeps only queries whose passage is in the retriever's top 20 and ranked first by the teacher, and teacher scores are min-max normalized for both distillation and hard-negative filtering.
What would settle it
Take several strong embedding models, fine-tune each on one held-out corpus using only the InfoNCE contrastive loss with hard negatives and the paper's teacher-based negative filtering, and measure NDCG@10 against the unfine-tuned model. If contrastive-only fine-tuning improves effectiveness in more than a minority of cases, the paper's central claim fails. A second check: train the same student on a single best synthetic query type and on the full six-type mixture at equal data sizes; if the mixture does not beat the single type, the diversity claim fails.
Extended reading notes
Core claim
The paper's central discovery is that the conventional InfoNCE contrastive loss is often the wrong tool for fine-tuning a strong embedding model on a new corpus: across BGE, GTE, Arctic, and the unsupervised E5 variant, contrastive-only fine-tuning generally reduced effectiveness even after passage de-duplication, hard-negative mining, and teacher-based negative filtering. The authors find that cross-encoder listwise distillation, which trains the student to match the teacher's graded relevance distribution over a list of candidate passages, consistently outperforms contrastive-only fine-tuning, and that combining a small contrastive term (weight 0.1) with listwise distillation works best. They further discover that diversifying synthetic query generation across six query types produces stronger training data than any single type in isolation, and that synthetic queries match human-written queries in training utility. Using synthetic queries and teacher scores from RankT5-3B and a Gemma-based reranker, they train a BERT-base embedding model and report state-of-the-art effectiveness among BERT embedding models on the evaluated benchmarks.
Load-bearing premise
The load-bearing premise is that the cross-encoder teachers' relevance scores are reliable enough to filter training queries and to act as distillation targets on target corpora, including corpora the teacher was never trained on; if the teacher misranks a passage, the student is trained on that error.
Editorial extensions
If this is right
- Fine-tuning a dense retriever on a new corpus with only InfoNCE contrastive loss should no longer be the default; a teacher-based listwise objective is the more reliable choice according to these results.
- Corpus-specific adaptation can be done entirely with synthetic data, because synthetic queries are as effective as human-written queries for training.
- Training on several synthetic query types at once is more robust than tailoring generation to the evaluation query format, so diversity is a training resource.
- The released BERT-base model, trained with synthetic queries and two cross-encoder teachers, reaches state-of-the-art effectiveness among BERT embedding models on the evaluated retrieval benchmarks.
- The recipe transfers across four base models, so it is not tied to a single architecture and can be applied to other BERT-base embedding models.
Reading between the lines
- Editorial inference: if contrastive fine-tuning degrades strong models even with denoised negatives, part of what is often credited to harder negatives and larger batches may actually be the richer graded signal from a teacher; a controlled test would fix data and vary only the loss.
- Editorial inference: the datasets where fine-tuning underperformed (FEVER, Climate-FEVER, SCIDOCS, ArguAna) are exactly those where the RankT5 teacher's reranking was weaker, suggesting teacher reliability is the main transfer bottleneck; a natural extension is per-dataset teacher weighting or ensembling.
- Editorial inference: if the synthetic-vs-human query equivalence generalizes, the cost of adapting retrieval to specialized corpora drops sharply, since query collection and annotation are replaced by LLM generation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies corpus-specific fine-tuning of BERT-base dense retrievers using synthetic queries generated by Llama-3.1. The authors report that fine-tuning with the standard InfoNCE contrastive loss often degrades the retrieval effectiveness of strong embedding models (BGE, GTE, Arctic), whereas combining cross-encoder listwise distillation with a small contrastive component improves effectiveness on a focused set of datasets. They also compare synthetic query types, finding that a mixture of query types with no downsampling outperforms single types, and that synthetic queries are competitive with human-written queries. Finally, they train a general-purpose model, cadet-embed-base-v1, using synthetic data and teacher distillation, and report strong BEIR results.
Significance. The paper addresses a practically important question: whether the dominant InfoNCE contrastive fine-tuning recipe is robust for corpus-specific adaptation of strong dense retrievers. If the main findings hold, the work would be a useful corrective to the field's reliance on contrastive learning and would strengthen the case for listwise distillation and synthetic-query pipelines. The paper has notable strengths: it evaluates multiple embedding models, uses query-level significance tests with Holm-Bonferroni correction, includes a careful discussion of training-data contamination in Appendix A, and releases the model, query-generation code, and training code. The main weaknesses are attribution gaps: the combined-loss comparison confounds distillation with a tenfold reduction in contrastive weight, the query-diversity claim confounds diversity with data quantity, and the claimed consistency of gains is limited by teacher reliability. These issues are addressable and do not invalidate the reported experimental observations, but they currently prevent the paper's strongest conclusions from being fully supported.
major comments (3)
- [Section 3.2.4, Table 2] The central comparison changes two variables at once. The '+ FT with contrastive loss' rows train with InfoNCE at full weight (Eq. 1), while the '+ FT with combined loss' rows use a 0.1 weight on the contrastive loss plus the listwise distillation loss. Because the contrastive gradient is reduced tenfold, the observed improvement could be due to gentler optimization rather than to the added distillation signal. The Introduction's claim that the results demonstrate 'the necessity of cross-encoder listwise distillation' is therefore not supported. A control that fine-tunes with the same 0.1-weight contrastive loss and no distillation, on the same filtered data and hyperparameters, is required. This concern is independent of teacher quality: even with a perfect teacher, the active ingredient of the reported gain would remain unidentified.
- [Section 4.3.1, Table 4] The claim that diverse query types yield greater effectiveness than any single query type is supported only in the 'All Query Types (No Downsampling)' row, which also has much more training data. Under the controlled downsampled comparison, the all-types mixture never beats the best single query type on any of the six datasets (e.g., DL19: all-types 72.6 vs zero-shot 73.0; SciFact: all-types 73.5 vs zero-shot 75.5). The abstract's statement about diverse query types is therefore overstated. The authors should either compare at equal total query budgets or explicitly rephrase the finding as a data-quantity effect.
- [Sections 3.3.1 and 4.2, Tables 1 and 3] The promised 'more consistent' improvement is bounded by teacher reliability. Table 1 shows that RankT5 reranking reduces NDCG@10 relative to the BGE retriever on FEVER, Climate-FEVER, ArguAna, and SCIDOCS, and Table 3 shows that the combined-loss fine-tuning degrades on exactly those datasets. Because Section 4.1's focused comparison excludes these datasets, the claim of consistent gains is established only where the teacher is strong. The authors acknowledge this in Sections 3.3.1 and 4.2, but the boundary condition should be stated as a qualification of the main claim, and ideally tested by comparing the combined loss against the 0.1-weight contrastive-only control on at least one teacher-failure dataset.
minor comments (5)
- [Section 3.3 and references] There are several typographical issues, including 'V oorhees' in the TREC-COVID citation and section text; please fix these and normalize the capitalization of CLIMATE-FEVER.
- [Table 4] Unlike Tables 2 and 3, Table 4 does not report significance tests; given the query-level testing used elsewhere, adding such tests would materially strengthen the query-type comparison.
- [Section 3.2.4] The hard-negative threshold (0.6), distillation temperatures, and contrastive weight (0.1) are tuned on E5-unsupervised with MSMARCO queries evaluated on DL19/DL20, then applied to all other models and datasets; a short sensitivity analysis, or an explicit statement that these parameters were not re-tuned per model, would help readers assess generality.
- [Section 4.4] The concluding discussion describes the pipeline as 'purely synthetic, from generated queries to relevance signals', but the few-shot user queries are generated using examples from the MSMARCO training set, and the static examples are produced by GPT4o; the text should clarify that human-written examples are used only as prompts, not as training labels.
- [Section 4.3.2, Table 5] The comparison of human-written and synthetic queries uses a 56K subset of synthetic queries 'aligned as closely as possible' to the same passages; please report the exact passage-overlap count so that the fairness of the comparison can be verified.
Circularity Check
No significant circularity: the training pipeline is self-contained and the main claims are not equivalent to their inputs; the flagged issues are benchmark-tuning and teacher-overlap confounds, not circular reductions.
full rationale
The claimed derivation chain is: sample corpus passages, generate six synthetic query types with Llama-3.1, filter queries with the retriever and a cross-encoder teacher, train the bi-encoder with a 0.1-weighted InfoNCE term plus KL listwise distillation against teacher scores, and evaluate on BEIR/TREC-DL. None of these steps defines the output in terms of the input by construction: the student embeddings are not the teacher scores, and the evaluation queries are not the synthetic training queries. The self-citations in the paper (e.g., Tamber et al. 2023b for RankT5 out-of-domain reliability, Tamber et al. 2024 for embedding distillation, Lin et al. 2020 for TCT-ColBERT, Pradeep et al. 2024 for passage deduplication) are contextual and none is load-bearing; the RankT5 reliability claim is also supported by external citations (Zhuang et al. 2023; Qin et al. 2023; Yoon et al. 2024). I weighed two reviewer concerns and they are correctness/attribution issues rather than circularity. First, Section 3.2.4 selects hyperparameters, including the 0.1 contrastive weight, on DL19/DL20, so the DL19/DL20 rows of Tables 2 and 4 are not fully independent confirmations, and the '+ FT with contrastive loss' baseline uses full-weight InfoNCE while the combined arm uses 0.1 weight, leaving the contribution of distillation per se unidentified without a 0.1-weight contrastive-only control; this affects causal attribution, not the logical equivalence of the derivation. Second, Appendix A acknowledges that the Gemma teacher was trained on MSMARCO, NQ, ArguAna, HotpotQA, FEVER, and PubMedQA, so some Table 6 benchmark scores partly inherit teacher exposure; the paper's BEIR(13*) average drops the main overlapping datasets and shows the non-overlapping results, so the central effectiveness claim has independent content. Because no prediction reduces by construction and no load-bearing step is forced by a self-citation chain, the circularity score is 2 rather than higher.
Assumptions & free parameters
free parameters (4)
- Hard-negative filter threshold =
0.6 (60% of positive score)
- Distillation temperatures =
0.05 for student, 0.3 for teacher
- Contrastive loss weight in combined loss =
0.1
- RankT5/Gemma teacher weight ratio =
0.25 / 0.75
assumptions (5)
- domain assumption Synthetic queries generated by Llama-3.1 8B from sampled passages are sufficiently representative of real user queries for the target corpus.
- domain assumption Cross-encoder teacher scores (RankT5-3B, Gemma reranker) provide reliable relevance supervision on target corpora, including out-of-domain ones.
- domain assumption Discarding queries whose passage is not ranked first among top-20 by the retriever after teacher reranking does not bias the comparison between loss functions.
- domain assumption Evaluation datasets used for the 'unseen corpus' claim are not substantially represented in base model training.
- standard math The one-sided paired t-test with Holm-Bonferroni correction across queries is a valid significance model for retrieval metric differences.
Cite this review
Pith. "Pith review of Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data." pith.science (2026). https://pith.science/paper/WQ4OR5WG
@misc{pith2026250519274,
author = {Pith},
title = {Pith review of: Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/WQ4OR5WG}},
note = {Machine review of arXiv:2505.19274}
}
read the original abstract
We investigate improving the retrieval effectiveness of embedding models through the lens of corpus-specific fine-tuning. Prior work has shown that fine-tuning with queries generated using a dataset's retrieval corpus can boost retrieval effectiveness for the dataset. However, we find that surprisingly, fine-tuning using the conventional InfoNCE contrastive loss often reduces effectiveness in state-of-the-art models. To overcome this, we revisit cross-encoder listwise distillation and demonstrate that, unlike using contrastive learning alone, listwise distillation can help more consistently improve retrieval effectiveness across multiple datasets. Additionally, we show that synthesizing more training data using diverse query types (such as claims, keywords, and questions) yields greater effectiveness than using any single query type alone, regardless of the query type used in evaluation. Our findings further indicate that synthetic queries offer comparable utility to human-written queries for training. We use our approach to train an embedding model that achieves state-of-the-art effectiveness among BERT embedding models. We release our model and both query generation and training code to facilitate further research.
Figures
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. GPT-4 Technical Report . arXiv:2303.08774
arXiv 2023
-
[4]
Negar Arabzadeh, Alexandra Vtyurina, Xinyi Yan, and Charles L. A. Clarke. 2022. Shallow pooling for sparse labels. Inf. Retr., 25(4):365–385
work page 2022
-
[5]
Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2023. Task-aware Retrieval with Instructions . In Findings of the Association for Computational Linguistics: ACL 2023, pages 3650--3675, Toronto, Canada. Association for Computational Linguistics
work page 2023
-
[6]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset . arXiv:1611.09268v3
arXiv 2016
-
[7]
Alexander Bondarenko, Maik Fr \"o be, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, and Matthias Hagen. 2020. Overview of Touch \'e 2020: Argument Retrieval . In Experimental IR Meets Multilinguality, Multimodality, and Interaction, pages 384--395, Cham. Springer Internat...
work page 2020
-
[8]
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. InPars: Data Augmentation for Information Retrieval using Large Language Models . arXiv:2202.05144
arXiv 2022
Show all 69 references
-
[9]
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A Full-Text Learning to Rank Dataset for Medical Information Retrieval . In Advances in Information Retrieval, pages 716--722, Cham. Springer International Publishing
2016
-
[10]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...
2020
-
[11]
Chanyeol Choi, Junseong Kim, Seolhwa Lee, Jihoon Kwon, Sangmo Gu, Yejin Kim, Minkyung Cho, and Jy-yong Sohn. 2024. Linq-Embed-Mistral Technical Report . arXiv:2412.03223
2024 arXiv
-
[12]
Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinso...
2024
-
[13]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel Weld. 2020. SPECTER : Document-level Representation Learning using Citation-informed Transformers . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2270--2282, On...
2020
-
[14]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020. Overview of the TREC 2020 Deep Learning Track . In Proceedings of the Twenty-Ninth Text REtrieval Conference Proceedings (TREC 2020), Gaithersburg, Maryland
2020
-
[15]
Voorhees
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M. Voorhees. 2019. Overview of the TREC 2019 Deep Learning Track . In Proceedings of the Twenty-Eighth Text REtrieval Conference Proceedings (TREC 2019), Gaithersburg, Maryland
2019
-
[16]
Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot Dense Retrieval From 8 Examples . In The Eleventh International Conference on Learning Representations
2023
-
[17]
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims . arXiv:2012.00614
2020 arXiv
-
[18]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models . arXiv:2407.21783
2024 arXiv
-
[19]
Luyu Gao, Yunyi Zhang, Jiawei Han, and Jamie Callan. 2021. Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup . In Proceedings of the 6th Workshop on Representation Learning for NLP
2021
-
[20]
Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. DBpedia-Entity v2: A Test Collection for Entity Search . In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in ...
2017
-
[21]
a tter, Sophia Althammer, Michael Schr \
Sebastian Hofst \"a tter, Sophia Althammer, Michael Schr \"o der, Mete Sertkan, and Allan Hanbury. 2020. Improving efficient neural ranking models with cross-architecture knowledge distillation . arXiv:2010.02666
2020 arXiv
-
[22]
Sebastian Hofst\" a tter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021. Efficiently teaching an effective dense retriever with balanced topic aware sampling. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in In...
2021
-
[23]
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023. InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval . arXiv:2301.01820
2023 arXiv
-
[24]
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference o...
2019
-
[25]
Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT . In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '20, page 39–48, New ...
2020
-
[26]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019
-
[27]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention . In Proceedings of the ACM SIGOPS 29th Symposium on Operating ...
2023
-
[28]
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models . arXiv:2405.17428
2024 arXiv
-
[29]
Chaofan Li, Minghao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Yingxia Shao, Defu Lian, and Zheng Liu. 2025. Making Text Embedders Few-Shot Learners . In The Thirteenth International Conference on Learning Representations
2025
-
[30]
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards General Text Embeddings with Multi-stage Contrastive Learning . arXiv:2308.03281
2023 arXiv
-
[31]
Davis Liang, Peng Xu, Siamak Shakeri, Cicero Nogueira dos Santos, Ramesh Nallapati, Zhiheng Huang, and Bing Xiang. 2020. Embedding-based Zero-shot Retrieval through Query Generation . arXiv:2009.10270
2020 arXiv
-
[32]
How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense Retrieval
Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. "How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense Retrieval" . In Findings of the Association for Computational Linguistics: EMNLP 2023,...
2023
-
[33]
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020. Distilling dense representations for ranking using tightly-coupled teachers . arXiv:2010.11386
2020 arXiv
-
[34]
Ji Ma, Ivan Korotkov, Yinfei Yang, Keith Hall, and Ryan McDonald. 2021. Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation . In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main V...
2021
-
[35]
Macedo Maia, Siegfried Handschuh, Andr\' e Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW'18 Open Challenge: Financial Opinion Mining and Question Answering . In Companion Proceedings of the The Web Conference 2018, WWW '18, page 1941–194...
2018
-
[36]
Aditya Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim, Sashank Reddi, and Sanjiv Kumar. 2022. In defense of dual-encoders for neural ranking . In International Conference on Machine Learning, pages 15376--15400. PMLR
2022
-
[37]
Luke Merrick, Danmei Xu, Gaurav Nuti, and Daniel Campos. 2024. Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models . arXiv:2405.05374
2024 arXiv
-
[38]
Gabriel de Souza P Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge. 2025. NV-Retriever: Improving text embedding models with effective hard-negative mining . arXiv:2407.15831
2025 arXiv
-
[39]
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. MTEB : Massive Text Embedding Benchmark . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2014--2037, Dubrovnik, Croatia. Association fo...
2023
-
[40]
Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to docTTTTTquery . Online preprint
2019
-
[41]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding . arXiv:1807.03748
2018 arXiv
-
[42]
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! arXiv:2312.02724
2023 arXiv
-
[43]
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Ragnar\"ok: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track . arXiv:2406.16828
2024 arXiv
-
[44]
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, et al. 2023. Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting . arXiv:2306.17563
2023 arXiv
-
[45]
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering . arXiv:2010.08191
2021 arXiv
-
[46]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer . The Journal of Machine Learning Research, 21(1):5485--5551
2020
-
[47]
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, QiaoQiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2021. R ocket QA v2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking . In Proceedings of the 2021 Conference on Empirical Methods in Natural Langua...
2021
-
[48]
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, and Christopher Potts. 2023. UDAPDR : Unsupervised domain adaptation via LLM prompting and distillation of rerankers. In Proceedings of the 2023 Conference on Empi...
2023
-
[49]
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. C ol BERT v2: Effective and Efficient Retrieval via Lightweight Late Interaction . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computatio...
2022
-
[50]
Smith, Luke Zettlemoyer, and Tao Yu
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. One Embedder, Any Task: Instruction-Finetuned Text Embeddings . In Findings of the Association for Computational Linguistics: ACL 2023, ...
2023
-
[51]
Manveer Singh Tamber, Ronak Pradeep, and Jimmy Lin. 2023 a . Pre-processing Matters! Improved Wikipedia Corpora for Open-Domain Question Answering . In European Conference on Information Retrieval, pages 163--176
2023
-
[52]
Manveer Singh Tamber, Ronak Pradeep, and Jimmy Lin. 2023 b . Scaling Down, LiTting Up: Efficient Zero-Shot Listwise Reranking with Seq2seq Encoder-Decoder Models . arXiv:2312.16098
2023 arXiv
-
[53]
Manveer Singh Tamber, Jasper Xian, and Jimmy Lin. 2024. Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models . arXiv:2406.09355
2024 arXiv
-
[54]
Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models . arXiv:2104.08663
2021 arXiv
-
[55]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER : a Large-scale Dataset for Fact Extraction and VER ification . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Huma...
2018
-
[56]
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. TREC-COVID: constructing a pandemic information retrieval test collection . In ACM SIGIR Forum, volume 54, pages 1--12. ACM New York...
2021
-
[57]
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the Best Counterargument without Prior Topic Knowledge . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 241--251, Melbourne, Australi...
2018
-
[58]
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534--7...
2020
-
[59]
Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
2021
-
[60]
Kexin Wang, Nandan Thakur, Nils Reimers, and Iryna Gurevych. 2022 a . GPL : Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense Retrieval . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: ...
2022
-
[61]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022 b . Text embeddings by weakly-supervised contrastive pre-training. arXiv:2212.03533
2022 arXiv
-
[62]
Benjamin Warner, Antoine Chaffin, Benjamin Clavi \'e , Orion Weller, Oskar Hallstr \"o m, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, et al. 2024. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and lo...
2024 arXiv
-
[63]
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. C-Pack: Packed Resources For General Chinese Embeddings . In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '24...
2024
-
[64]
Sohee Yang and Minjoon Seo. 2020. Is Retriever Merely an Approximator of Reader? arXiv:2010.10999
2020 arXiv
-
[65]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering . In Proceedings of the 2018 Conference on Empirical Methods in Natural Lang...
2018
-
[66]
Soyoung Yoon, Eunbi Choi, Jiyeon Kim, Hyeongu Yun, Yireun Kim, and Seung-won Hwang. 2024. L ist T 5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot Retrieval . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: L...
2024
-
[67]
Dun Zhang, Jiacheng Li, Ziyang Zeng, and Fulong Wang. 2024. Jasper and Stella: distillation of SOTA embedding models . arXiv:2412.19048
2024 arXiv
-
[68]
Xinyu Zhang, Sebastian Hofst \"a tter, Patrick Lewis, Raphael Tang, and Jimmy Lin. 2023. Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models . arXiv:2312.02969
2023 arXiv
-
[69]
Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023. RankT5: Fine-Tuning T5 for Text Ranking with Ranking Losses . In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in In...
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.