Pith. sign in

REVIEW 19 cited by

Fine-Tuning LLaMA for Multi-Stage Text Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08319 v1 pith:KJ4KQ32N submitted 2023-10-12 cs.IR

classification cs.IR
keywords modelsretrievaleffectivenesslanguagellmsdemonstratefine-tuninglarge
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The effectiveness of multi-stage text retrieval has been solidly demonstrated since before the era of pre-trained language models. However, most existing studies utilize models that predate recent advances in large language models (LLMs). This study seeks to explore potential improvements that state-of-the-art LLMs can bring. We conduct a comprehensive study, fine-tuning the latest LLaMA model both as a dense retriever (RepLLaMA) and as a pointwise reranker (RankLLaMA) for both passage retrieval and document retrieval using the MS MARCO datasets. Our findings demonstrate that the effectiveness of large language models indeed surpasses that of smaller models. Additionally, since LLMs can inherently handle longer contexts, they can represent entire documents holistically, obviating the need for traditional segmenting and pooling strategies. Furthermore, evaluations on BEIR demonstrate that our RepLLaMA-RankLLaMA pipeline exhibits strong zero-shot effectiveness. Model checkpoints from this study are available on HuggingFace.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

    cs.IR 2024-12 conditional novelty 7.0 of 10

    CIR-LVLM fine-tunes Qwen-VL-Chat with LoRA and hybrid task and instance-specific prompts to produce query and target embeddings, achieving new state-of-the-art recall on Fashion-IQ, Shoes, and CIRR.

  2. SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

    cs.IR 2026-08 conditional novelty 6.0 of 10

    SPEAR, a PDN-style framework with gradient-isolated embeddings, multiplicative rewrite gating, and a dynamic rewrite selector, reports large offline and online gains over Dewu's production search baseline.

  3. Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

    cs.CL 2025-09 conditional novelty 6.0 of 10

    In multilingual retrieval-augmented generation, models cite English evidence more accurately than translated evidence, and this language preference can outweigh document relevance.

  4. MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers

    cs.IR 2025-06 conditional novelty 6.0 of 10

    A zero-shot mixture of sparse, dense, and simulated human retrievers, weighted by pre- and post-retrieval geometry signals, beats individual small retrievers and 7B LLM retrievers on four scientific retrieval benchmarks.

  5. RaDeR: Reasoning-aware Dense Retrieval Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A math-trained dense retriever and reranker, built from MCTS reasoning trajectories and self-reflection, outperforms strong baselines on reasoning-intensive retrieval benchmarks and beats BM25 on chain-of-thought queries.

  6. Reranking with Compressed Document Representation

    cs.IR 2025-05 conditional novelty 6.0 of 10

    A reranker trained on 8-token PISCO document embeddings plus a short query achieves near-identical nDCG@10 to full-text rerankers on BeIR and TREC-DL while running up to 16x faster.

  7. QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A fine-tuned ensemble of a generative small language model and an embedding model outperformed zero-shot LLMs on Korean search relevance labeling, with reported Cohen's kappa of 0.646 versus 0.387 and 60x lower latency.

  8. CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment

    cs.IR 2025-02 conditional novelty 6.0 of 10

    CliniQ is a public EHR retrieval benchmark with 77,206 LLM-annotated relevance judgments, showing that BM25 is a strong baseline and that semantic matches drive dense-retriever gains.

  9. mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval

    cs.IR 2025-01 conditional novelty 6.0 of 10

    The paper introduces mFollowIR, a multilingual instruction-following retrieval benchmark across Russian, Chinese, and Persian, and finds that English instruction-trained models transfer cross-lingually but struggle in...

  10. Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A single LLM re-ranker can be configured at runtime to different depths and widths, with training tricks that keep compressed variants close to full-scale accuracy.

  11. PaSa: An LLM Agent for Comprehensive Academic Paper Search

    cs.IR 2025-01 conditional novelty 6.0 of 10

    PaSa, a two-agent LLM system trained with session-level RL, reports substantially higher recall than existing academic search baselines on complex paper-finding queries.

  12. AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark

    cs.IR 2024-12 conditional novelty 6.0 of 10

    AIR-Bench uses LLMs to generate retrieval test data over 69 datasets, 9 domains and 13 languages, and reports a 0.82 rank correlation between model rankings on its generated MS MARCO set and human-labeled MS MARCO.

  13. Boosting LLM-based Relevance Modeling with Distribution-Aware Robust Learning

    cs.IR 2024-12 conditional novelty 6.0 of 10

    DaRL improves LLM relevance ranking by augmenting training data with OOD-detected samples, applying multi-stage fine-tuning, and calibrating overconfident predictions.

  14. Drowning in Documents: Consequences of Scaling Reranker Inference

    cs.IR 2024-11 conditional novelty 6.0 of 10

    Modern cross-encoder rerankers improve Recall@10 at small candidate counts but degrade when reranking thousands of documents, sometimes falling below first-stage dense retrieval alone.

  15. GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

    cs.CL 2026-08 conditional novelty 5.0 of 10

    GEM unifies text generation and retrieval embeddings in one model, generating a reasoning analysis before encoding the query, and improves retrieval on reasoning-intensive and instruction-following benchmarks.

  16. ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval

    cs.CL 2025-01 conditional novelty 5.0 of 10

    ASRank re-ranks retrieved documents by scoring how well each document supports a zero-shot answer scent generated by a large LLM, beating UPR and RankGPT on several QA datasets.

  17. State Space Models are Strong Text Rerankers

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Mamba-1 and Mamba-2 rerankers match comparably sized transformers on ranking accuracy but are less efficient in training and inference, with Mamba-2 improving on both fronts.

  18. Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense Retrieval

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Reasoning-infused text embedding, which prepends LLM-generated reasoning to queries before embedding, improves zero-shot dense retrieval on BRIGHT.

  19. LineRetriever: Planning-Aware Observation Reduction for Web Agents

    cs.CL 2025-06 conditional novelty 4.0 of 10

    LineRetriever uses a small LM to select relevant lines from web page observations, cutting context by up to 73% with only small success-rate drops on web agent benchmarks.

Pith tools