Pith. sign in

REVIEW 5 major objections 8 minor 69 references

ICLERB: In-Context Learning Embedding and Reranker Benchmark

T0 review · 5 major / 8 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Retrieval for in-context learning should be scored by how much a document improves the LLM's answer, and a 335M-parameter reranker trained with 10,000 LLM queries outperforms models over twenty times larger on ICLERB.

desk verdict Useful benchmark idea, but the leaderboard and the RLRAIF claim both rest on an unvalidated DPO proxy; worth reviewing, needs empirical anchoring. read the letter →

arxiv 2411.18947 v1 pith:KZBC6YG7 submitted 2024-11-28 cs.LG cs.IR

classification cs.LGcs.IR
keywords in-contextlearningretrieval-augmentedgenerationretrievalasrecommendationtorankdirectpreferenceoptimizationactiveLLMfeedbackbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that when retrieval supplies examples to an LLM's prompt, the retriever should be judged by how much it improves the LLM's answers, not by semantic similarity to the query. It introduces ICLERB, a benchmark that scores embedding models and rerankers with the DPO metric, which measures the change in the LLM's log-probability of the correct answer when a document is added to a one-shot prompt. It also introduces RLRAIF, a reinforcement learning-to-rank algorithm that actively chooses which query-document pairs to score and fine-tunes a reranker using about 10,000 such scores. On ICLERB, the resulting 335M-parameter reranker ranks first, ahead of state-of-the-art retrieval models that are more than twenty times larger. The paper concludes that retrieval for ICL is a recommendation problem and that training retrievers against LLM utility rather than search relevance can compensate for much smaller model size.

What carries the argument

The central object is the DPO metric, defined for a query $q$, a document $d$, a correct response $r$, and an optional incorrect response $\bar r$ as $\mathrm{DPO}(q,d)=\log\sigma\left(\log\frac{p_M(r\mid q,d)}{p_M(r\mid q)}-\log\frac{p_M(\bar r\mid q,d)}{p_M(\bar r\mid q)}\right)$, which measures how much adding $d$ to a one-shot prompt increases the LLM's relative log-probability of the correct answer. ICLERB uses this value as the ground-truth relevance label for ranking documents, and RLRAIF uses it as the reward signal in an active-learning loop that selects $(q,d)$ pairs by balancing high expected reward, high uncertainty, and batch diversity. The fine-tuning step is a pairwise ranking loss applied to the resulting reward comparisons, updating a small adapter on top of pre-trained embeddings. This single metric therefore carries both the benchmark's definition of document utility and the training signal that lets a small retriever outperform much larger models.

What would settle it

Take a held-out set of queries with known answers, compute DPO scores for all documents, and also measure actual answer accuracy when each document is used as the one-shot demonstration; if documents with higher DPO do not systematically yield higher accuracy, or if a retriever trained to maximize DPO does not improve end-task accuracy, the benchmark and the training algorithm lose their justification.

Watch

Extended reading notes

Core claim

The core claim is that retrieval for in-context learning should be reframed from a search problem to a recommendation problem: given a query, the retriever should rank documents by their utility in improving the LLM's response, and that utility can be measured by the DPO metric. ICLERB operationalizes this by building ground-truth relevance labels from DPO scores across multiple multiple-choice datasets and open-weight LLMs, then ranks embedding models and rerankers with nDCG@10 and nDCG@50. The same DPO signal is used as the reward in RLRAIF, which treats data acquisition as a contextual bandit problem that balances exploration and exploitation in both the query and document spaces, and trains a small non-linear adapter on top of a frozen embedding model. With roughly 10,000 DPO evaluations, the fine-tuned 335M-parameter reranker reaches the top of the ICLERB leaderboard, beating models an order of magnitude larger, which the paper takes as evidence that alignment with ICL utility matters more than raw model capacity.

Load-bearing premise

Everything rests on the assumption that the DPO score of a single document in a one-shot prompt is a faithful measure of how much that document improves the LLM's true end-task accuracy.

Editorial extensions

If this is right

  • Retriever rankings produced by ICLERB disagree with rankings from semantic-similarity benchmarks on several models, so RAG system builders should expect different component choices when the goal is in-context learning rather than search.
  • A 335M-parameter model fine-tuned with roughly 10,000 DPO evaluations can beat retrieval models more than twenty times larger, indicating that aligning the training signal with ICL utility can matter more than model capacity.
  • Search-optimized rerankers can rank below their own embedding counterparts on ICLERB, suggesting that fine-tuning for semantic relevance can actively hurt retrieval for ICL.
  • Because RLRAIF requires only log-probability access to an open-weight LLM, the same recipe can be applied to any domain with a query set and response labels without building a dedicated retrieval training dataset.
  • Future ICLERB releases with more datasets and LLMs may change the leaderboard, and RLRAIF can be applied to other base models to test whether the reported gains persist across architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the single-document, one-shot DPO measure ignores interactions among multiple demonstrations, so a multi-document extension that scores sets of documents might rank retrievers differently and could change what RLRAIF optimizes.
  • Editorial inference: because ICLERB uses one fixed prompt template and one demonstration, the stability of its rankings under prompt phrasing is untested; a prompt-variation study would clarify how much of the leaderboard reflects retrieval quality rather than template effects.
  • Editorial inference: the benchmark and training loop depend on open-weight LLM log probabilities, so extending the approach to proprietary or closed LLMs would require an alternative feedback signal, such as sampled-completion accuracy.
  • Editorial inference: the RLRAIF acquisition strategy could be tested on other ranking models and other base embeddings to see whether the dual exploration-exploitation trade-off gives the reported gains beyond the single model and datasets used here.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 8 minor

Summary. The paper proposes reframing retrieval for in-context learning as a recommendation problem rather than a search problem, and introduces ICLERB, a benchmark that ranks embedding models and rerankers by how well the documents they retrieve increase the probability an LLM assigns to the correct answer, as quantified by a DPO-based metric (Eq. 7). The benchmark is instantiated on three MCQ-style datasets (TruthfulQA, Emotion, ProductER) and three open-weight LLMs, and the resulting leaderboard (Table 2) is compared against MTEB. The paper then introduces RLRAIF, an active-learning/contextual-bandit procedure that fine-tunes a small cross-encoder adapter using DPO rewards, and reports that the resulting model, cm-rerank-mxbai-rlaif-v0.1, outperforms much larger existing retrieval models on ICLERB with a budget of only 10k DPO evaluations.

Significance. If the central claims hold, ICLERB provides a utility-oriented complement to semantic-similarity benchmarks such as MTEB, and the RLRAIF result is a practically valuable demonstration that a small, cheaply fine-tuned retriever can be competitive with much larger models for retrieval-augmented ICL. The paper is also careful to split queries into train/test sets, to average results over repeated splits, and to report per-dataset and per-LLM breakdowns, which is a strength. However, the significance of both contributions is conditional on the validity of the DPO metric as a proxy for actual ICL accuracy; the manuscript does not yet supply that validation, and the RLRAIF method is described at a level that prevents independent reproduction.

major comments (5)
  1. [§3.1, Eq. (7)] The benchmark's ground truth is the DPO metric, and Section 3.1 asserts that DPO has 'desirable properties' citing [11], but the paper provides no empirical evidence that documents ranked higher by DPO actually improve the LLM's end-task accuracy. This is load-bearing because Table 2 ranks all models by DPO-based nDCG and because RLRAIF (Section 5.2) uses DPO as its reward. Please add a direct validation study: for held-out queries, compute the correlation (or rank agreement) between DPO scores and the LLM's task accuracy when the top document is used as a 1-shot demonstration, and probe sensitivity to the fixed prompt template, the choice of the incorrect answer r̄, and the use of a single demonstration. Without such evidence, ICLERB measures agreement with a specific surrogate, not utility for ICL.
  2. [§3.2, Table 2] The paper states that dataset splits and experiments are repeated and averaged, but no standard deviations, confidence intervals, or significance tests are reported anywhere. The difference between the first and second rows of Table 2 is 0.0047 nDCG@10 (0.7238 vs. 0.7191), and a large cluster of models lies within 0.01 of each other, so the ranking order may be within noise. Please report paired bootstrap confidence intervals over test queries (and per dataset/LLM) or equivalent significance tests to support the claim that RLRAIF's model significantly outperforms the baselines.
  3. [§5.2, §5.3] The RLRAIF acquisition function is described only as a qualitative bullet list ('exploitation in the document space,' 'information for ranking loss,' 'exploration,' 'diversity in the batch'), with no equations or hyperparameter values. Since RLRAIF is a major contribution and the empirical result depends on its specific exploration/exploitation trade-off, the description is not reproducible; the reader cannot tell what was actually optimized. Please provide the exact acquisition function, its coefficients, batch sizes, number of acquisition rounds, the training hyperparameters, and ideally release the code at submission.
  4. [§3.3.1, Table 1] ProductER is a new dataset introduced in this paper, described as 'manually curated' but in fact generated with OpenAI o1-preview, with no human validation or inter-annotator statistics reported. It is also not released ('we aim to release ... in the near future'), so the benchmark cannot currently be reproduced on one of its three datasets. Please describe the generation and validation protocol in detail, report label-quality checks, and release the dataset or clearly mark ProductER results as preliminary and separate from the main leaderboard.
  5. [§5.2, §3.2] Because RLRAIF is trained to maximize the DPO reward and ICLERB evaluates retrievers using the same DPO-based nDCG, the reported superiority of cm-rerank-mxbai-rlaif-v0.1 is partly a measure of alignment between the training objective and the evaluation metric. The test queries are held out, so this is not formally circular, but it does mean that the headline claim ('small models fine-tuned with our RLRAIF algorithm outperform large state-of-the-art retrieval models') is only as strong as the validity of DPO as a proxy for ICL accuracy. I therefore see this as connected to the validation requested in the first major comment, and recommend that the paper either provide that validation or temper the abstract and conclusion claims until it is available.
minor comments (8)
  1. [§2.4, Eq. (7)] The sentence 'the DPO metric ... is defined as the negative of the DPO loss [52]' is confusing because the original DPO loss is a training loss to be minimized; clarify the sign conventions and define σ explicitly.
  2. [§3.3.2] There is a typo in 'inital version' (should be 'initial').
  3. [Table 9] The table is labeled as 'fully extracted from the MTEB benchmark on November 27th, 2024' in the caption, but the body text says 'Table 9 summarizes the performance ... as reported by MTEB' without a date; please ensure the access date is stated consistently.
  4. [§5.2, Eq. (10)] Equation (10) is the standard logistic pairwise ranking loss; please state this explicitly and cite [6] in the text near the equation rather than only in the reference list.
  5. [§3.1] The claim that DPO is 'additive over independent queries' is used to justify aggregate evaluation, but the additive quantity is not written out; a brief statement of how DCG aggregations are pooled across queries would help.
  6. [§3.2] The text says test queries and documents are randomly subsampled, while §3.3.1 says the corpus of documents is fixed to be the training-set ground-truth responses; clarify how documents from the test split, if any, are treated to avoid ambiguity.
  7. [Appendix C, Figure 2] Figure 2 is not referenced in the main text; add a reference in Section 5.1 or 5.3.
  8. [Table 2] Models with undisclosed sizes are marked with '–'; for clarity, consider labeling them as 'not disclosed' or 'API' in the table caption.

Circularity Check

0 steps flagged · score 2.0 of 10

No construction-level circularity: RLRAIF is evaluated on held-out DPO labels, and the DPO metric is fully defined in the paper; the main weakness is the unvalidated DPO-to-accuracy proxy, which is a correctness concern rather than a circular derivation.

full rationale

The paper's derivation chain is not circular in the formal sense. ICLERB defines its ground truth via the DPO metric in Eq. 7, and the paper states this definition explicitly rather than importing an opaque oracle. RLRAIF is trained on DPO rewards and evaluated on held-out test queries: Section 3.2 describes an 80/20 train/test split, repeated over trials, and Section 5.3 reports the adapter was trained with a budget of 10k DPO values and then scored on the ICLERB benchmark. Its top nDCG@10 result in Table 2 is therefore a genuine generalization result, not a value equal to its training target by construction. The evaluation metric (Eqs. 8-9) is a standard rank-agreement metric with DPO-derived gains, and no equation in the paper reduces the claimed outperformance to the RLRAIF training loss (Eq. 10). The self-citation to [11], co-authored by present author Emile Contal, is the origin of the DPO metric, but because Eq. 7 and the listed properties are restated in the paper and are mathematical properties of the definition, the citation is not load-bearing for the derivations. The genuinely weak point is external validity: the paper equates DPO with enhancing LLM accuracy without reporting any correlation between DPO and true end-task accuracy, and it does not ablate the prompt template, the single-demonstration setting, or the choice of the incorrect answer r-bar. That is a correctness or validation gap, not a circularity, because DPO is an input assumption rather than a quantity derived from the things it is used to predict. Under the stated rules, an unvalidated proxy does not constitute circularity, so the score remains low despite the reviewer's concern.

Assumptions & free parameters 4 free parameters · 5 assumptions · 2 invented entities

The ICLERB ground truth depends on the DPO proxy, which is assumed to reflect ICL utility; RLRAIF adds hand-chosen hyperparameters (d1, budget, acquisition weights) and the authors introduce a new dataset (ProductER) that is not yet available. These items constitute the main unverified inputs to the central claims.

free parameters (4)
  • adapter inner dimension d1 = 100
    Set by hand in Section 5.2 for the non-linear adapter trained by RLRAIF; the paper does not justify this choice or report sensitivity.
  • DPO query budget = 10k DPO values (~5M tokens)
    Chosen budget in Section 5.3; no ablation on budget size in the text.
  • test-set subsample size = average 10^6 (q,d) pairs per dataset
    Chosen in Section 3.2 citing computational constraints; the subsample may bias which query-document pairs are scored.
  • acquisition-function coefficients = not specified
    The components (exploitation, information, exploration, diversity) are listed in Section 5.2 but the balance weights are not given, so the method as described has implicit free parameters.
assumptions (5)
  • domain assumption DPO(q,d) as defined in Eq. 7 faithfully measures how much a document improves ICL utility
    Adopted from the authors' RAGSys paper (Ref [11]); ICLERB ground truth relevance is DPO, so any mismatch between DPO and true ICL accuracy invalidates the benchmark.
  • domain assumption One retrieved demonstration (1-shot) suffices to evaluate ICL retrieval utility
    Section 3.1 explicitly limits ICLERB to single-document DPO; multi-document effects such as redundancy are out of scope, but typical RAG uses several documents.
  • domain assumption Open-weight LLM log probabilities are reliable and accessible for reward computation
    Section 3.3.2 and 2.4 require token-probability access; closed models are excluded, and logprob calibration is assumed.
  • domain assumption The warm-start split (same documents, unseen queries) matches the target deployment
    Section 3.2 states this emulates the typical use case; cold-start documents are not evaluated, limiting the benchmark's scope.
  • domain assumption nDCG with gain 2^DPO is a valid ranking metric for negative-valued utilities
    Eq. 8 modifies standard DCG to handle negative DPO values; this choice affects all leaderboard scores.
invented entities (2)
  • ProductER dataset
    purpose: Test retrieval for product entity resolution (is product A the same as product B)
    Created by the authors using OpenAI's o1-preview and not yet released (Section 3.3.1), so external validation is impossible.
  • cm-rerank-mxbai-rlaif-v0.1
    purpose: Fine-tuned reranker meant to outperform larger retrieval models on ICL
    Evaluated only on the ICLERB benchmark in this paper; no external ICL datasets or ablations back the generalization claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ICLERB: In-Context Learning Embedding and Reranker Benchmark." pith.science (2026). https://pith.science/paper/KZBC6YG7

@misc{pith2026241118947,
  author       = {Pith},
  title        = {Pith review of: ICLERB: In-Context Learning Embedding and Reranker Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KZBC6YG7}},
  note         = {Machine review of arXiv:2411.18947}
}
read the original abstract

In-Context Learning (ICL) enables Large Language Models (LLMs) to perform new tasks by conditioning on prompts with relevant information. Retrieval-Augmented Generation (RAG) enhances ICL by incorporating retrieved documents into the LLM's context at query time. However, traditional retrieval methods focus on semantic relevance, treating retrieval as a search problem. In this paper, we propose reframing retrieval for ICL as a recommendation problem, aiming to select documents that maximize utility in ICL tasks. We introduce the In-Context Learning Embedding and Reranker Benchmark (ICLERB), a novel evaluation framework that compares retrievers based on their ability to enhance LLM accuracy in ICL settings. Additionally, we propose a novel Reinforcement Learning-to-Rank from AI Feedback (RLRAIF) algorithm, designed to fine-tune retrieval models using minimal feedback from the LLM. Our experimental results reveal notable differences between ICLERB and existing benchmarks, and demonstrate that small models fine-tuned with our RLRAIF algorithm outperform large state-of-the-art retrieval models. These findings highlight the limitations of existing evaluation methods and the need for specialized benchmarks and training strategies adapted to ICL.

Figures

Figures reproduced from arXiv: 2411.18947 by the authors.

Figure 1
Figure 1. Retrieval performance for the task of ICL, as measured by nDCG@10 in the In-Context [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. ICLERB rank reached with different DPO budgets, for RLRAIF and an approach randomly [PITH_FULL_IMAGE:figures/full_fig_p029_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 59 canonical work pages

  1. [11]

    RAGSys: Item-Cold-Start Recommender as RAG System

    Emile Contal and Garrin McGoldrick. RAGSys: Item-Cold-Start Recommender as RAG System. IR-RAG @ SIGIR24, 2024

  2. [1]

    Transformers as statisticians: Provable in-context learning with in-context algorithm selection

    Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 57125–57211. Curran Associates, Inc., 2023

  3. [2]

    MS MARCO: A Human Generated MAchine Reading COmprehension Dataset, 2018

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset, 2018

  4. [3]

    A Test Collection for Entity Search in DBpedia

    Krisztian Balog and Robert Neumayer. A Test Collection for Entity Search in DBpedia. In Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’13, page 737–740, New York, NY , USA, 2013. Association for Computing Machinery

  5. [4]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A Large Annotated Corpus for Learning Natural Language Inference. In Lluís Márquez, Chris Callison- Burch, Jian Su, Daniele Pighin, and Yuval Marton, editors, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 632–642. Associa...

  6. [5]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwi...

  7. [6]

    Learning to Rank with Nonsmooth Cost Functions

    Christopher Burges, Robert Ragno, and Quoc Le. Learning to Rank with Nonsmooth Cost Functions. In B. Schölkopf, J. Platt, and T. Hoffman, editors, Advances in Neural Information Processing Systems, volume 19. MIT Press, 2006

  8. [7]

    Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar

    Wei-Cheng Chang, Felix X. Yu, Yin-Wen Chang, Yiming Yang, and Sanjiv Kumar. Pre-training Tasks for Embedding-based Large-scale Retrieval. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020) . OpenReview.net, 2020

Show all 69 references
  1. [8]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A Simple Frame- work for Contrastive Learning of Visual Representations. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of ...

  2. [9]

    Cohere Embed Model

    Cohere. Cohere Embed Model. https://cohere.com/blog/introducing-embed-v3 , 2023

  3. [10]

    Cohere Rerank Model

    Cohere. Cohere Rerank Model. https://cohere.com/blog/rerank-3, 2024

  4. [12]

    Zeta-Alpha-E5-Mistral

    Arthur Câmara, Dinos Papakostas, Mathias Parisot, Fernando Rejon Barrera, and Jakub Zavrel. Zeta-Alpha-E5-Mistral. https://huggingface.co/zeta-alpha-ai/ Zeta-Alpha-E5-Mistral , 2024

  5. [13]

    Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge

    Gabriel de Souza P. Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge. NV-Retriever: Improving text embedding models with effective hard- negative mining, 2024

  6. [14]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of...

  7. [15]

    A Survey for In-context Learning

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. A Survey for In-context Learning. CoRR, abs/2301.00234, 2023

  8. [16]

    The faiss library

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre- Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. The faiss library. arXiv preprint arXiv:2401.08281, 2024

  9. [17]

    The Llama 3 Herd of Models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783, 2024

  10. [18]

    SimCSE: Simple Contrastive Learning of Sentence Embeddings

    Tianyu Gao, Xingcheng Yao, and Danqi Chen. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP...

  11. [19]

    Universal Language Model Fine-tuning for Text Classifi- cation, 2018

    Jeremy Howard and Sebastian Ruder. Universal Language Model Fine-tuning for Text Classifi- cation, 2018

  12. [20]

    Unsupervised Dense Information Retrieval with Contrastive Learning

    Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. Unsupervised Dense Information Retrieval with Contrastive Learning. Transactions on Machine Learning Research, 2023

  13. [21]

    Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering

    Gautier Izacard and Edouard Grave. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty, editors, Proceedings of the 16th Conference of the European Chapter of the Association for Computationa...

  14. [22]

    Atlas: Few-shot Learning with Retrieval Augmented Language Models, 2022

    Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. Atlas: Few-shot Learning with Retrieval Augmented Language Models, 2022

  15. [23]

    Cumulated Gain-Based Evaluation of IR Techniques

    Kalervo Järvelin and Jaana Kekäläinen. Cumulated Gain-Based Evaluation of IR Techniques. ACM Transactions on Information Systems, 20(4), 2002

  16. [24]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  17. [25]

    Dense Passage Retrieval for Open-Domain Question Answering

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. Dense Passage Retrieval for Open-Domain Question Answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on E...

  18. [26]

    ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT

    Omar Khattab and Matei Zaharia. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , SIGIR ’20, page 39–48, New York, ...

  19. [27]

    Linq-Embed-Mistral:Elevating Text Retrieval with Improved GPT Data Through Task-Specific Control and Quality Refinement

    Junseong Kim, Seolhwa Lee, Jihoon Kwon, Sangmo Gu, Yejin Kim, Minkyung Cho, Jy yong Sohn, and Chanyeol Choi. Linq-Embed-Mistral:Elevating Text Retrieval with Improved GPT Data Through Task-Specific Control and Quality Refinement. Linq AI Research Blog, 2024

  20. [28]

    Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell

    James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming Catas- trophic Forgetting...

  21. [29]

    NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

    Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv preprint arXiv:2405.17428, 2024

  22. [30]

    Open Source Strikes Bread - New Fluffy Embeddings Model, 2024

    Sean Lee, Aamir Shakir, Darius Koenig, and Julius Lipp. Open Source Strikes Bread - New Fluffy Embeddings Model, 2024

  23. [31]

    Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2021

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2021

  24. [32]

    Making Text Embedders Few-Shot Learners, 2024

    Chaofan Li, MingHao Qin, Shitao Xiao, Jianlyu Chen, Kun Luo, Yingxia Shao, Defu Lian, and Zheng Liu. Making Text Embedders Few-Shot Learners, 2024

  25. [33]

    Schapire

    Lihong Li, Wei Chu, John Langford, and Robert E. Schapire. A Contextual-Bandit Approach to Personalized News Article Recommendation. In Proceedings of the 19th International Conference on World Wide Web, pages 661–670, New York, NY , USA, 2010. ACM

  26. [34]

    Angle-optimized text embeddings

    Xianming Li and Jing Li. Angle-optimized text embeddings. arXiv preprint arXiv:2309.12871, 2023

  27. [35]

    Unified Demonstration Retriever for In-Context Learning

    Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu. Unified Demonstration Retriever for In-Context Learning. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Ass...

  28. [36]

    To- wards General Text Embeddings with Multi-stage Contrastive Learning.CoRR, abs/2308.03281, aug 2023

    Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. To- wards General Text Embeddings with Multi-stage Contrastive Learning.CoRR, abs/2308.03281, aug 2023

  29. [37]

    Lin, Jacob Hilton, and Owain Evans

    Stephanie C. Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Annual Meeting of the Association for Computational Linguistics , 2021

  30. [38]

    Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. What Makes Good In-Context Examples for GPT-3? In Eneko Agirre, Marianna Apidianaki, and Ivan Vuli´c, editors, Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowl...

  31. [39]

    Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums

    Xueqing Liu, Chi Wang, Yue Leng, and ChengXiang Zhai. Linkso: a dataset for learning to retrieve similar question answer pairs on software development forums. In Proceedings of the 4th ACM SIGSOFT International Workshop on NLP for Software Engineering , pages 2–5, 2018

  32. [40]

    Active Learning for Ranking through Expected Loss Optimization

    Bo Long, Olivier Chapelle, Ya Zhang, Yi Chang, Zhaohui Zheng, and Belle Tseng. Active Learning for Ranking through Expected Loss Optimization. In Proceedings of the 33rd interna- tional ACM SIGIR conference on Research and development in information retrieval , pages 267–274, 2010

  33. [41]

    Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, 2022

    Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, 2022

  34. [42]

    Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024

    Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. Sfr-embedding-2: Advanced text embedding with multi-stage training, 2024

  35. [43]

    Embedding And Clustering Your Data Can Improve Contrastive Pretraining, 2024

    Luke Merrick. Embedding And Clustering Your Data Can Improve Contrastive Pretraining, 2024

  36. [44]

    Distributed Representations of Words and Phrases and their Compositionality, 2013

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. Distributed Representations of Words and Phrases and their Compositionality, 2013. 19

  37. [45]

    Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference...

  38. [46]

    The bayesian approach to global optimization

    Jonas Mockus. The bayesian approach to global optimization. In System Modeling and Optimization: Proceedings of the 10th IFIP Conference New York City, USA, August 31– September 4, 1981, pages 473–481. Springer, 2005

  39. [47]

    MTEB: Massive Text Embedding Benchmark

    Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers. MTEB: Massive Text Embedding Benchmark. In Andreas Vlachos and Isabelle Augenstein, editors, Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2023)...

  40. [48]

    Text and Code Embeddings by Contrastive Pre-Training

    Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nkoulou, Girish Sastry, Gretchen Krueger, David Schnurr, Felipe Petroski ...

  41. [49]

    Zhao, Yi Luan, Keith B

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y . Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. Large Dual Encoders Are Generalizable Retrievers, 2021

  42. [50]

    New embedding models and API updates

    OpenAI. New embedding models and API updates. https://openai.com/index/ new-embedding-models-and-api-updates/ , 2024

  43. [51]

    Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs

    Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs. In Yaser Al-Onaizan, Mohit Bansal, and Yun- Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processin...

  44. [52]

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in N...

  45. [53]

    Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research, 21:1–67, 2020

  46. [54]

    In-Context Retrieval-Augmented Language Models

    Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. In-Context Retrieval-Augmented Language Models. Transactions of the Association for Computational Linguistics , 11:1316–1331, 11 2023

  47. [55]

    Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

    Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019

  48. [56]

    Learning To Retrieve Prompts for In- Context Learning

    Ohad Rubin, Jonathan Herzig, and Jonathan Berant. Learning To Retrieve Prompts for In- Context Learning. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conference of the North American Chapter of the Association fo...

  49. [57]

    CARER: Contextualized affect representations for emotion recognition

    Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. CARER: Contextualized affect representations for emotion recognition. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3687–3697, Brussels, Belgium, O...

  50. [58]

    Active learning literature survey

    Burr Settles. Active learning literature survey. 2009

  51. [59]

    Algorithms for reinforcement learning

    Csaba Szepesvári. Algorithms for reinforcement learning. Springer nature, 2022

  52. [60]

    BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, 2021

    Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models, 2021

  53. [61]

    V oyage AI Rerankers

    Voyage AI. V oyage AI Rerankers. https://blog.voyageai.com/2024/09/30/ rerank-2/, 2024

  54. [62]

    Voyage AI Text Embeddings

    Voyage AI. Voyage AI Text Embeddings. https://blog.voyageai.com/2024/09/18/ voyage-3/, 2024

  55. [63]

    Text Embeddings by Weakly-Supervised Contrastive Pre-training

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text Embeddings by Weakly-Supervised Contrastive Pre-training. arXiv preprint arXiv:2212.03533, 2022

  56. [64]

    Learning to Retrieve In-Context Examples for Large Language Models

    Liang Wang, Nan Yang, and Furu Wei. Learning to Retrieve In-Context Examples for Large Language Models. In Yvette Graham and Matthew Purver, editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (V olume 1: Long Pa...

  57. [65]

    Adina Williams, Nikita Nangia, and Samuel R. Bowman. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Marilyn A. Walker, Heng Ji, and Amanda Stent, editors, Proceedings of the 2018 Conference of the North American Chapter of the Association fo...

  58. [66]

    Mind: A Large-Scale Dataset for News Recom- mendation

    Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, et al. Mind: A Large-Scale Dataset for News Recom- mendation. In Proceedings of the 58th annual meeting of the association for computational linguistics, pa...

  59. [67]

    C-Pack: Packed Resources for General Chinese Embeddings

    Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. C-Pack: Packed Resources for General Chinese Embeddings. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang, editors, Proceedings of the 47th International...

  60. [68]

    Compositional Exem- plars for In-context Learning

    Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. Compositional Exem- plars for In-context Learning. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Confe...

  61. [69]

    Average Retrieval

    Dun Zhang. stella_en_1.5B_v5. https://huggingface.co/dunzhang/stella_en_1.5B_ v5, 2023. 21 A Appendix: Supplementary Results from ICLERB A.1 ICLERB Results per Dataset A.1.1 TruthfulQA Table 3: ICLERB results for TruthfulQA [37], sorted by nDCG@10 Organization Model Name nDCG@...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.