REVIEW 5 major objections 8 minor 69 references
ICLERB: In-Context Learning Embedding and Reranker Benchmark
T0 review · 5 major / 8 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Retrieval for in-context learning should be scored by how much a document improves the LLM's answer, and a 335M-parameter reranker trained with 10,000 LLM queries outperforms models over twenty times larger on ICLERB.
desk verdict Useful benchmark idea, but the leaderboard and the RLRAIF claim both rest on an unvalidated DPO proxy; worth reviewing, needs empirical anchoring. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the DPO metric, defined for a query $q$, a document $d$, a correct response $r$, and an optional incorrect response $\bar r$ as $\mathrm{DPO}(q,d)=\log\sigma\left(\log\frac{p_M(r\mid q,d)}{p_M(r\mid q)}-\log\frac{p_M(\bar r\mid q,d)}{p_M(\bar r\mid q)}\right)$, which measures how much adding $d$ to a one-shot prompt increases the LLM's relative log-probability of the correct answer. ICLERB uses this value as the ground-truth relevance label for ranking documents, and RLRAIF uses it as the reward signal in an active-learning loop that selects $(q,d)$ pairs by balancing high expected reward, high uncertainty, and batch diversity. The fine-tuning step is a pairwise ranking loss applied to the resulting reward comparisons, updating a small adapter on top of pre-trained embeddings. This single metric therefore carries both the benchmark's definition of document utility and the training signal that lets a small retriever outperform much larger models.
What would settle it
Take a held-out set of queries with known answers, compute DPO scores for all documents, and also measure actual answer accuracy when each document is used as the one-shot demonstration; if documents with higher DPO do not systematically yield higher accuracy, or if a retriever trained to maximize DPO does not improve end-task accuracy, the benchmark and the training algorithm lose their justification.
Extended reading notes
Core claim
The core claim is that retrieval for in-context learning should be reframed from a search problem to a recommendation problem: given a query, the retriever should rank documents by their utility in improving the LLM's response, and that utility can be measured by the DPO metric. ICLERB operationalizes this by building ground-truth relevance labels from DPO scores across multiple multiple-choice datasets and open-weight LLMs, then ranks embedding models and rerankers with nDCG@10 and nDCG@50. The same DPO signal is used as the reward in RLRAIF, which treats data acquisition as a contextual bandit problem that balances exploration and exploitation in both the query and document spaces, and trains a small non-linear adapter on top of a frozen embedding model. With roughly 10,000 DPO evaluations, the fine-tuned 335M-parameter reranker reaches the top of the ICLERB leaderboard, beating models an order of magnitude larger, which the paper takes as evidence that alignment with ICL utility matters more than raw model capacity.
Load-bearing premise
Everything rests on the assumption that the DPO score of a single document in a one-shot prompt is a faithful measure of how much that document improves the LLM's true end-task accuracy.
Editorial extensions
If this is right
- Retriever rankings produced by ICLERB disagree with rankings from semantic-similarity benchmarks on several models, so RAG system builders should expect different component choices when the goal is in-context learning rather than search.
- A 335M-parameter model fine-tuned with roughly 10,000 DPO evaluations can beat retrieval models more than twenty times larger, indicating that aligning the training signal with ICL utility can matter more than model capacity.
- Search-optimized rerankers can rank below their own embedding counterparts on ICLERB, suggesting that fine-tuning for semantic relevance can actively hurt retrieval for ICL.
- Because RLRAIF requires only log-probability access to an open-weight LLM, the same recipe can be applied to any domain with a query set and response labels without building a dedicated retrieval training dataset.
- Future ICLERB releases with more datasets and LLMs may change the leaderboard, and RLRAIF can be applied to other base models to test whether the reported gains persist across architectures.
Reading between the lines
- Editorial inference: the single-document, one-shot DPO measure ignores interactions among multiple demonstrations, so a multi-document extension that scores sets of documents might rank retrievers differently and could change what RLRAIF optimizes.
- Editorial inference: because ICLERB uses one fixed prompt template and one demonstration, the stability of its rankings under prompt phrasing is untested; a prompt-variation study would clarify how much of the leaderboard reflects retrieval quality rather than template effects.
- Editorial inference: the benchmark and training loop depend on open-weight LLM log probabilities, so extending the approach to proprietary or closed LLMs would require an alternative feedback signal, such as sampled-completion accuracy.
- Editorial inference: the RLRAIF acquisition strategy could be tested on other ranking models and other base embeddings to see whether the dual exploration-exploitation trade-off gives the reported gains beyond the single model and datasets used here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes reframing retrieval for in-context learning as a recommendation problem rather than a search problem, and introduces ICLERB, a benchmark that ranks embedding models and rerankers by how well the documents they retrieve increase the probability an LLM assigns to the correct answer, as quantified by a DPO-based metric (Eq. 7). The benchmark is instantiated on three MCQ-style datasets (TruthfulQA, Emotion, ProductER) and three open-weight LLMs, and the resulting leaderboard (Table 2) is compared against MTEB. The paper then introduces RLRAIF, an active-learning/contextual-bandit procedure that fine-tunes a small cross-encoder adapter using DPO rewards, and reports that the resulting model, cm-rerank-mxbai-rlaif-v0.1, outperforms much larger existing retrieval models on ICLERB with a budget of only 10k DPO evaluations.
Significance. If the central claims hold, ICLERB provides a utility-oriented complement to semantic-similarity benchmarks such as MTEB, and the RLRAIF result is a practically valuable demonstration that a small, cheaply fine-tuned retriever can be competitive with much larger models for retrieval-augmented ICL. The paper is also careful to split queries into train/test sets, to average results over repeated splits, and to report per-dataset and per-LLM breakdowns, which is a strength. However, the significance of both contributions is conditional on the validity of the DPO metric as a proxy for actual ICL accuracy; the manuscript does not yet supply that validation, and the RLRAIF method is described at a level that prevents independent reproduction.
major comments (5)
- [§3.1, Eq. (7)] The benchmark's ground truth is the DPO metric, and Section 3.1 asserts that DPO has 'desirable properties' citing [11], but the paper provides no empirical evidence that documents ranked higher by DPO actually improve the LLM's end-task accuracy. This is load-bearing because Table 2 ranks all models by DPO-based nDCG and because RLRAIF (Section 5.2) uses DPO as its reward. Please add a direct validation study: for held-out queries, compute the correlation (or rank agreement) between DPO scores and the LLM's task accuracy when the top document is used as a 1-shot demonstration, and probe sensitivity to the fixed prompt template, the choice of the incorrect answer r̄, and the use of a single demonstration. Without such evidence, ICLERB measures agreement with a specific surrogate, not utility for ICL.
- [§3.2, Table 2] The paper states that dataset splits and experiments are repeated and averaged, but no standard deviations, confidence intervals, or significance tests are reported anywhere. The difference between the first and second rows of Table 2 is 0.0047 nDCG@10 (0.7238 vs. 0.7191), and a large cluster of models lies within 0.01 of each other, so the ranking order may be within noise. Please report paired bootstrap confidence intervals over test queries (and per dataset/LLM) or equivalent significance tests to support the claim that RLRAIF's model significantly outperforms the baselines.
- [§5.2, §5.3] The RLRAIF acquisition function is described only as a qualitative bullet list ('exploitation in the document space,' 'information for ranking loss,' 'exploration,' 'diversity in the batch'), with no equations or hyperparameter values. Since RLRAIF is a major contribution and the empirical result depends on its specific exploration/exploitation trade-off, the description is not reproducible; the reader cannot tell what was actually optimized. Please provide the exact acquisition function, its coefficients, batch sizes, number of acquisition rounds, the training hyperparameters, and ideally release the code at submission.
- [§3.3.1, Table 1] ProductER is a new dataset introduced in this paper, described as 'manually curated' but in fact generated with OpenAI o1-preview, with no human validation or inter-annotator statistics reported. It is also not released ('we aim to release ... in the near future'), so the benchmark cannot currently be reproduced on one of its three datasets. Please describe the generation and validation protocol in detail, report label-quality checks, and release the dataset or clearly mark ProductER results as preliminary and separate from the main leaderboard.
- [§5.2, §3.2] Because RLRAIF is trained to maximize the DPO reward and ICLERB evaluates retrievers using the same DPO-based nDCG, the reported superiority of cm-rerank-mxbai-rlaif-v0.1 is partly a measure of alignment between the training objective and the evaluation metric. The test queries are held out, so this is not formally circular, but it does mean that the headline claim ('small models fine-tuned with our RLRAIF algorithm outperform large state-of-the-art retrieval models') is only as strong as the validity of DPO as a proxy for ICL accuracy. I therefore see this as connected to the validation requested in the first major comment, and recommend that the paper either provide that validation or temper the abstract and conclusion claims until it is available.
minor comments (8)
- [§2.4, Eq. (7)] The sentence 'the DPO metric ... is defined as the negative of the DPO loss [52]' is confusing because the original DPO loss is a training loss to be minimized; clarify the sign conventions and define σ explicitly.
- [§3.3.2] There is a typo in 'inital version' (should be 'initial').
- [Table 9] The table is labeled as 'fully extracted from the MTEB benchmark on November 27th, 2024' in the caption, but the body text says 'Table 9 summarizes the performance ... as reported by MTEB' without a date; please ensure the access date is stated consistently.
- [§5.2, Eq. (10)] Equation (10) is the standard logistic pairwise ranking loss; please state this explicitly and cite [6] in the text near the equation rather than only in the reference list.
- [§3.1] The claim that DPO is 'additive over independent queries' is used to justify aggregate evaluation, but the additive quantity is not written out; a brief statement of how DCG aggregations are pooled across queries would help.
- [§3.2] The text says test queries and documents are randomly subsampled, while §3.3.1 says the corpus of documents is fixed to be the training-set ground-truth responses; clarify how documents from the test split, if any, are treated to avoid ambiguity.
- [Appendix C, Figure 2] Figure 2 is not referenced in the main text; add a reference in Section 5.1 or 5.3.
- [Table 2] Models with undisclosed sizes are marked with '–'; for clarity, consider labeling them as 'not disclosed' or 'API' in the table caption.
Circularity Check
No construction-level circularity: RLRAIF is evaluated on held-out DPO labels, and the DPO metric is fully defined in the paper; the main weakness is the unvalidated DPO-to-accuracy proxy, which is a correctness concern rather than a circular derivation.
full rationale
The paper's derivation chain is not circular in the formal sense. ICLERB defines its ground truth via the DPO metric in Eq. 7, and the paper states this definition explicitly rather than importing an opaque oracle. RLRAIF is trained on DPO rewards and evaluated on held-out test queries: Section 3.2 describes an 80/20 train/test split, repeated over trials, and Section 5.3 reports the adapter was trained with a budget of 10k DPO values and then scored on the ICLERB benchmark. Its top nDCG@10 result in Table 2 is therefore a genuine generalization result, not a value equal to its training target by construction. The evaluation metric (Eqs. 8-9) is a standard rank-agreement metric with DPO-derived gains, and no equation in the paper reduces the claimed outperformance to the RLRAIF training loss (Eq. 10). The self-citation to [11], co-authored by present author Emile Contal, is the origin of the DPO metric, but because Eq. 7 and the listed properties are restated in the paper and are mathematical properties of the definition, the citation is not load-bearing for the derivations. The genuinely weak point is external validity: the paper equates DPO with enhancing LLM accuracy without reporting any correlation between DPO and true end-task accuracy, and it does not ablate the prompt template, the single-demonstration setting, or the choice of the incorrect answer r-bar. That is a correctness or validation gap, not a circularity, because DPO is an input assumption rather than a quantity derived from the things it is used to predict. Under the stated rules, an unvalidated proxy does not constitute circularity, so the score remains low despite the reviewer's concern.
Assumptions & free parameters
free parameters (4)
- adapter inner dimension d1 =
100
- DPO query budget =
10k DPO values (~5M tokens)
- test-set subsample size =
average 10^6 (q,d) pairs per dataset
- acquisition-function coefficients =
not specified
assumptions (5)
- domain assumption DPO(q,d) as defined in Eq. 7 faithfully measures how much a document improves ICL utility
- domain assumption One retrieved demonstration (1-shot) suffices to evaluate ICL retrieval utility
- domain assumption Open-weight LLM log probabilities are reliable and accessible for reward computation
- domain assumption The warm-start split (same documents, unseen queries) matches the target deployment
- domain assumption nDCG with gain 2^DPO is a valid ranking metric for negative-valued utilities
invented entities (2)
-
ProductER dataset
-
cm-rerank-mxbai-rlaif-v0.1
Cite this review
Pith. "Pith review of ICLERB: In-Context Learning Embedding and Reranker Benchmark." pith.science (2026). https://pith.science/paper/KZBC6YG7
@misc{pith2026241118947,
author = {Pith},
title = {Pith review of: ICLERB: In-Context Learning Embedding and Reranker Benchmark},
year = {2026},
howpublished = {\url{https://pith.science/paper/KZBC6YG7}},
note = {Machine review of arXiv:2411.18947}
}
read the original abstract
In-Context Learning (ICL) enables Large Language Models (LLMs) to perform new tasks by conditioning on prompts with relevant information. Retrieval-Augmented Generation (RAG) enhances ICL by incorporating retrieved documents into the LLM's context at query time. However, traditional retrieval methods focus on semantic relevance, treating retrieval as a search problem. In this paper, we propose reframing retrieval for ICL as a recommendation problem, aiming to select documents that maximize utility in ICL tasks. We introduce the In-Context Learning Embedding and Reranker Benchmark (ICLERB), a novel evaluation framework that compares retrievers based on their ability to enhance LLM accuracy in ICL settings. Additionally, we propose a novel Reinforcement Learning-to-Rank from AI Feedback (RLRAIF) algorithm, designed to fine-tune retrieval models using minimal feedback from the LLM. Our experimental results reveal notable differences between ICLERB and existing benchmarks, and demonstrate that small models fine-tuned with our RLRAIF algorithm outperform large state-of-the-art retrieval models. These findings highlight the limitations of existing evaluation methods and the need for specialized benchmarks and training strategies adapted to ICL.
Figures
Reference graph
Works this paper leans on
-
[11]
RAGSys: Item-Cold-Start Recommender as RAG System
Emile Contal and Garrin McGoldrick. RAGSys: Item-Cold-Start Recommender as RAG System. IR-RAG @ SIGIR24, 2024
work page 2024
-
[1]
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 57125–57211. Curran Associates, Inc., 2023
work page 2023
-
[2]
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset, 2018
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset, 2018
work page 2018
-
[3]
A Test Collection for Entity Search in DBpedia
Krisztian Balog and Robert Neumayer. A Test Collection for Entity Search in DBpedia. In Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’13, page 737–740, New York, NY , USA, 2013. Association for Computing Machinery
work page 2013
-
[4]
Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A Large Annotated Corpus for Learning Natural Language Inference. In Lluís Márquez, Chris Callison- Burch, Jian Su, Daniele Pighin, and Yuval Marton, editors, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 632–642. Associa...
work page 2015
-
[5]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwi...
work page 2020
-
[6]
Learning to Rank with Nonsmooth Cost Functions
Christopher Burges, Robert Ragno, and Quoc Le. Learning to Rank with Nonsmooth Cost Functions. In B. Schölkopf, J. Platt, and T. Hoffman, editors, Advances in Neural Information Processing Systems, volume 19. MIT Press, 2006
work page 2006
-
[7]
Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar
Wei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. Pre-training Tasks for Embedding-based Large-scale Retrieval. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020) . OpenReview.net, 2020
work page 2020
Show all 69 references
-
[8]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A Simple Frame- work for Contrastive Learning of Visual Representations. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of ...
2020
-
[9]
Cohere Embed Model
Cohere. Cohere Embed Model. https://cohere.com/blog/introducing-embed-v3 , 2023
2023
-
[10]
Cohere Rerank Model
Cohere. Cohere Rerank Model. https://cohere.com/blog/rerank-3, 2024
2024
-
[12]
Zeta-Alpha-E5-Mistral
Arthur Câmara, Dinos Papakostas, Mathias Parisot, Fernando Rejon Barrera, and Jakub Zavrel. Zeta-Alpha-E5-Mistral. https://huggingface.co/zeta-alpha-ai/ Zeta-Alpha-E5-Mistral , 2024
2024
-
[13]
Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge
Gabriel de Souza P. Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge. NV-Retriever: Improving text embedding models with effective hard- negative mining, 2024
2024
-
[14]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of...
2019
-
[15]
A Survey for In-context Learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. A Survey for In-context Learning. CoRR, abs/2301.00234, 2023
2023 arXiv
-
[16]
The faiss library
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre- Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library. arXiv preprint arXiv:2401.08281, 2024
2024 arXiv
-
[17]
The Llama 3 Herd of Models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[18]
SimCSE: Simple Contrastive Learning of Sentence Embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP...
2021
-
[19]
Universal Language Model Fine-tuning for Text Classifi- cation, 2018
Jeremy Howard and Sebastian Ruder. Universal Language Model Fine-tuning for Text Classifi- cation, 2018
2018
-
[20]
Unsupervised Dense Information Retrieval with Contrastive Learning
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised Dense Information Retrieval with Contrastive Learning. Transactions on Machine Learning Research, 2023
2023
-
[21]
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
Gautier Izacard and Edouard Grave. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computationa...
2021
-
[22]
Atlas: Few-shot Learning with Retrieval Augmented Language Models, 2022
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. Atlas: Few-shot Learning with Retrieval Augmented Language Models, 2022
2022
-
[23]
Cumulated Gain-Based Evaluation of IR Techniques
Kalervo Järvelin and Jaana Kekäläinen. Cumulated Gain-Based Evaluation of IR Techniques. ACM Transactions on Information Systems, 20(4), 2002
2002
-
[24]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023
-
[25]
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense Passage Retrieval for Open-Domain Question Answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on E...
2020
-
[26]
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab and Matei Zaharia. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , SIGIR ’20, page 39–48, New York, ...
2020
-
[27]
Linq-Embed-Mistral:Elevating Text Retrieval with Improved GPT Data Through Task-Specific Control and Quality Refinement
Junseong Kim, Seolhwa Lee, Jihoon Kwon, Sangmo Gu, Yejin Kim, Minkyung Cho, Jy yong Sohn, and Chanyeol Choi. Linq-Embed-Mistral:Elevating Text Retrieval with Improved GPT Data Through Task-Specific Control and Quality Refinement. Linq AI Research Blog, 2024
2024
-
[28]
Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming Catas- trophic Forgetting...
2017
-
[29]
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv preprint arXiv:2405.17428, 2024
2024 arXiv
-
[30]
Open Source Strikes Bread - New Fluffy Embeddings Model, 2024
Sean Lee, Aamir Shakir, Darius Koenig, and Julius Lipp. Open Source Strikes Bread - New Fluffy Embeddings Model, 2024
2024
-
[31]
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2021
2021
-
[32]
Making Text Embedders Few-Shot Learners, 2024
Chaofan Li, MingHao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Yingxia Shao, Defu Lian, and Zheng Liu. Making Text Embedders Few-Shot Learners, 2024
2024
-
[33]
Schapire
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. A Contextual-Bandit Approach to Personalized News Article Recommendation. In Proceedings of the 19th International Conference on World Wide Web, pages 661–670, New York, NY , USA, 2010. ACM
2010
-
[34]
Angle-optimized text embeddings
Xianming Li and Jing Li. Angle-optimized text embeddings. arXiv preprint arXiv:2309.12871, 2023
2023 arXiv
-
[35]
Unified Demonstration Retriever for In-Context Learning
Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu. Unified Demonstration Retriever for In-Context Learning. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Ass...
2023
-
[36]
To- wards General Text Embeddings with Multi-stage Contrastive Learning.CoRR, abs/2308.03281, aug 2023
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. To- wards General Text Embeddings with Multi-stage Contrastive Learning.CoRR, abs/2308.03281, aug 2023
2023 arXiv
-
[37]
Lin, Jacob Hilton, and Owain Evans
Stephanie C. Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Annual Meeting of the Association for Computational Linguistics , 2021
2021
-
[38]
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What Makes Good In-Context Examples for GPT-3? In Eneko Agirre, Marianna Apidianaki, and Ivan Vuli´c, editors, Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowl...
2022
-
[39]
Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums
Xueqing Liu, Chi Wang, Yue Leng, and ChengXiang Zhai. Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums. In Proceedings of the 4th ACM SIGSOFT International Workshop on NLP for Software Engineering , pages 2–5, 2018
2018
-
[40]
Active Learning for Ranking through Expected Loss Optimization
Bo Long, Olivier Chapelle, Ya Zhang, Yi Chang, Zhaohui Zheng, and Belle Tseng. Active Learning for Ranking through Expected Loss Optimization. In Proceedings of the 33rd interna- tional ACM SIGIR conference on Research and development in information retrieval , pages 267–274, 2010
2010
-
[41]
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, 2022
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, 2022
2022
-
[42]
Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024
2024
-
[43]
Embedding And Clustering Your Data Can Improve Contrastive Pretraining, 2024
Luke Merrick. Embedding And Clustering Your Data Can Improve Contrastive Pretraining, 2024
2024
-
[44]
Distributed Representations of Words and Phrases and their Compositionality, 2013
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Distributed Representations of Words and Phrases and their Compositionality, 2013. 19
2013
-
[45]
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference...
2022
-
[46]
The bayesian approach to global optimization
Jonas Mockus. The bayesian approach to global optimization. In System Modeling and Optimization: Proceedings of the 10th IFIP Conference New York City, USA, August 31– September 4, 1981, pages 473–481. Springer, 2005
1981
-
[47]
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. MTEB: Massive Text Embedding Benchmark. In Andreas Vlachos and Isabelle Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2023)...
2023
-
[48]
Text and Code Embeddings by Contrastive Pre-Training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nkoulou, Girish Sastry, Gretchen Krueger, David Schnurr, Felipe Petroski ...
2022 arXiv
-
[49]
Zhao, Yi Luan, Keith B
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y . Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. Large Dual Encoders Are Generalizable Retrievers, 2021
2021
-
[50]
New embedding models and API updates
OpenAI. New embedding models and API updates. https://openai.com/index/ new-embedding-models-and-api-updates/ , 2024
2024
-
[51]
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs. In Yaser Al-Onaizan, Mohit Bansal, and Yun- Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processin...
2024
-
[52]
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in N...
2023
-
[53]
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21:1–67, 2020
2020
-
[54]
In-Context Retrieval-Augmented Language Models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. In-Context Retrieval-Augmented Language Models. Transactions of the Association for Computational Linguistics , 11:1316–1331, 11 2023
2023
-
[55]
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019
2019
-
[56]
Learning To Retrieve Prompts for In- Context Learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning To Retrieve Prompts for In- Context Learning. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conference of the North American Chapter of the Association fo...
2022
-
[57]
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. CARER: Contextualized affect representations for emotion recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3687–3697, Brussels, Belgium, O...
2018
-
[58]
Active learning literature survey
Burr Settles. Active learning literature survey. 2009
2009
-
[59]
Algorithms for reinforcement learning
Csaba Szepesvári. Algorithms for reinforcement learning. Springer nature, 2022
2022
-
[60]
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, 2021
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, 2021
2021
-
[61]
V oyage AI Rerankers
Voyage AI. V oyage AI Rerankers. https://blog.voyageai.com/2024/09/30/ rerank-2/, 2024
2024
-
[62]
Voyage AI Text Embeddings
Voyage AI. Voyage AI Text Embeddings. https://blog.voyageai.com/2024/09/18/ voyage-3/, 2024
2024
-
[63]
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv preprint arXiv:2212.03533, 2022
2022 arXiv
-
[64]
Learning to Retrieve In-Context Examples for Large Language Models
Liang Wang, Nan Yang, and Furu Wei. Learning to Retrieve In-Context Examples for Large Language Models. In Yvette Graham and Matthew Purver, editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (V olume 1: Long Pa...
2024
-
[65]
Adina Williams, Nikita Nangia, and Samuel R. Bowman. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Marilyn A. Walker, Heng Ji, and Amanda Stent, editors, Proceedings of the 2018 Conference of the North American Chapter of the Association fo...
2018
-
[66]
Mind: A Large-Scale Dataset for News Recom- mendation
Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, et al. Mind: A Large-Scale Dataset for News Recom- mendation. In Proceedings of the 58th annual meeting of the association for computational linguistics, pa...
2020
-
[67]
C-Pack: Packed Resources for General Chinese Embeddings
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. C-Pack: Packed Resources for General Chinese Embeddings. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang, editors, Proceedings of the 47th International...
2024
-
[68]
Compositional Exem- plars for In-context Learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. Compositional Exem- plars for In-context Learning. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Confe...
2023
-
[69]
Average Retrieval
Dun Zhang. stella_en_1.5B_v5. https://huggingface.co/dunzhang/stella_en_1.5B_ v5, 2023. 21 A Appendix: Supplementary Results from ICLERB A.1 ICLERB Results per Dataset A.1.1 TruthfulQA Table 3: ICLERB results for TruthfulQA [37], sorted by nDCG@10 Organization Model Name nDCG@...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.