hub Canonical reference

Passage Re-ranking with BERT

Rodrigo Nogueira, Kyunghyun Cho · 2019 · cs.IR · arXiv 1901.04085

Canonical reference. 88% of citing Pith papers cite this work as background.

82 Pith papers citing it

Background 88% of classified citations

open full Pith review browse 82 citing papers arXiv PDF

abstract

Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural language processing tasks such as question-answering and natural language inference. In this paper, we describe a simple re-implementation of BERT for query-based passage re-ranking. Our system is the state of the art on the TREC-CAR dataset and the top entry in the leaderboard of the MS MARCO passage retrieval task, outperforming the previous state of the art by 27% (relative) in MRR@10. The code to reproduce our results is available at https://github.com/nyu-dl/dl4marco-bert

hub tools

JSON dossier citing papers JSON arXiv source

citation-role summary

background 15 baseline 1

citation-polarity summary

background 14 baseline 1 support 1

claims ledger

abstract Recently, neural models pretrained on a language modeling task, such as ELMo (Peters et al., 2017), OpenAI GPT (Radford et al., 2018), and BERT (Devlin et al., 2018), have achieved impressive results on various natural language processing tasks such as question-answering and natural language inference. In this paper, we describe a simple re-implementation of BERT for query-based passage re-ranking. Our system is the state of the art on the TREC-CAR dataset and the top entry in the leaderboard of the MS MARCO passage retrieval task, outperforming the previous state of the art by 27% (relative)

co-cited works

representative citing papers

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

cs.CL · 2026-06-15 · unverdicted · novelty 8.0

MetaSyn benchmark shows LLM pipelines recover at most 52.7% of ground-truth included studies due to screening failures on PI/ECO eligibility, despite 90.9% retrieval recall at K=200.

From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability

cs.IR · 2026-06-06 · unverdicted · novelty 8.0

A benchmark and ontology-driven framework links 434 cardiovascular devices to patents at 91.6% recall, producing 6.8M high-confidence links for regulatory-IP integration.

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

cs.CR · 2026-04-07 · unverdicted · novelty 8.0

The first SoK on LLM-based AutoPT frameworks provides a six-dimension taxonomy of agent designs and a unified empirical benchmark evaluating 15 frameworks via over 10 billion tokens and 1,500 manually reviewed logs.

Learning to Unscramble Feynman Loop Integrals with SAILIR

hep-ph · 2026-04-06 · unverdicted · novelty 8.0

A self-supervised transformer learns to unscramble Feynman integrals for online IBP reduction, delivering bounded memory use on complex two-loop topologies while matching Kira's speed on the hardest cases tested.

BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

cs.IR · 2021-04-17 · accept · novelty 8.0

BEIR is a heterogeneous zero-shot benchmark showing BM25 as a robust baseline while re-ranking and late-interaction models perform best on average at higher cost, with dense and sparse models lagging in generalization.

Dense Passage Retrieval for Open-Domain Question Answering

cs.CL · 2020-04-10 · accept · novelty 8.0

Dense dual-encoder retrievers outperform BM25 by 9-19% absolute in top-20 passage retrieval accuracy across open-domain QA datasets and enable new state-of-the-art end-to-end QA results.

ContextNest: Verifiable Context Governance for Autonomous AI Agent

cs.AI · 2026-07-02 · unverdicted · novelty 7.0

ContextNest formalizes context governance for AI agents using hash-chained documents and deterministic selectors, with experiments showing higher answer quality and perfect determinism versus standard retrieval.

Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method

cs.DB · 2026-06-06 · unverdicted · novelty 7.0

An adaptive two-phase semantic filter using clustering then a hybrid proxy trained on LLM confidence achieves 1.6-2.0x speedup over prior methods at 90% accuracy on 10K document corpora.

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

cs.CL · 2026-06-02 · unverdicted · novelty 7.0

Re-ranking retrieval candidates via a cross-encoder trained on continuous perturbation-based attribution scores improves citation faithfulness and gold-answer alignment in legal QA over semantic similarity.

Test-Time Training for Zero-Resource Dense Retrieval Reranking

cs.IR · 2026-05-31 · unverdicted · novelty 7.0

DART adapts a scoring matrix at inference time via gradient updates on pseudo-labels from top/bottom documents to gain +2.1% mean NDCG@10 on six BEIR benchmarks with under 10ms added latency.

SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning

cs.CR · 2026-05-27 · unverdicted · novelty 7.0

SilentRetrieval is a data poisoning attack achieving 84.6% HR@10 and 57.5% ASR-LLM on Natural Questions via coordinated beam search and trigger fusion while preserving document fluency.

Layer-wise Token Compression for Efficient Document Reranking

cs.IR · 2026-05-20 · unverdicted · novelty 7.0 · 2 refs

Layer-wise Token Compression applies adaptive token pooling at middle transformer layers for cross-encoder rerankers, preserving MS MARCO ranking quality while raising QPS up to 25% on passages and 116% on documents, with added gains on listwise LLM rerankers and a regularizer effect for long inputs

Very Efficient Listwise Multimodal Reranking for Long Documents

cs.IR · 2026-05-12 · unverdicted · novelty 7.0

ZipRerank delivers state-of-the-art multimodal listwise reranking accuracy for long documents at up to 10x lower latency via early interaction and single-pass scoring.

Hypothesis-Driven Deep Research with Large Language Models: A Structured Methodology for Automated Knowledge Discovery

cs.AI · 2026-05-11 · unverdicted · novelty 7.0

HDRI is a six-principle eight-stage framework for hypothesis-organized LLM research featuring gap-driven iteration, traceable fact reasoning, and subject locking, realized in INFOMINER with reported gains in fact density and completeness.

Prism-Reranker: Beyond Relevance Scoring -- Jointly Producing Contributions and Evidence for Agentic Retrieval

cs.IR · 2026-04-26 · accept · novelty 7.0

Prism-Reranker models output relevance, contribution statements, and evidence passages to support agentic retrieval beyond scalar scoring.

Bayesian Active Learning with Gaussian Processes Guided by LLM Relevance Scoring for Dense Passage Retrieval

cs.IR · 2026-04-20 · unverdicted · novelty 7.0

BAGEL is a Bayesian active learning framework that uses Gaussian Processes to propagate LLM relevance signals across embedding space and guide global exploration, outperforming standard LLM reranking under identical budgets on four retrieval benchmarks.

KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

cs.CV · 2026-04-18 · unverdicted · novelty 7.0

KIRA is a unified architecture for visual RAG that reports 0.97 retrieval precision, 1.0 grounding, and 0.707 domain correctness across medical, circuit, satellite, and histopathology domains via hierarchical chunking, dual-path retrieval, and evidence-conditioned generation.

Scaling Laws for Cross-Encoder Reranking

cs.IR · 2026-03-05 · unverdicted · novelty 7.0

Cross-encoder reranker performance scales predictably via power laws with model size and training exposure, allowing accurate forecasts for 400M and 1B models and data-heavy compute allocation.

SPIRE: Structure-Preserving Interpretable Retrieval of Evidence

cs.IR · 2026-02-12 · unverdicted · novelty 7.0

SPIRE presents a tree-structured retrieval method using subdocuments, paths, and dual contextualization that produces higher-quality and more diverse citations than passage-based baselines on HTML QA benchmarks.

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

cs.CL · 2020-05-22 · accept · novelty 7.0

RAG models set new state-of-the-art results on open-domain QA by retrieving Wikipedia passages and conditioning a generative model on them, while also producing more factual text than parametric baselines.

Relevance Is Not Permission: Warranted Attention for Value Contributions

cs.AI · 2026-06-29 · unverdicted · novelty 6.0 · 2 refs

Warrant adds a query-item permission gate g_ij to attention value terms, improving primary metrics in 27 of 32 comparisons across CTDG, MTPP, RAG, STPP, and TKG tasks.

AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering

cs.CL · 2026-06-27 · unverdicted · novelty 6.0

AB-RAG adaptively budgets retrieval in RAG by combining three confidence signals to decide when to stop or fetch more evidence, separating correct from incorrect answers at 57.6% vs 0% exact match on a factoid dataset.

Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation

cs.LG · 2026-06-27 · unverdicted · novelty 6.0

Presents a WildChat-derived benchmark for multi-agent routing as set-valued prediction and reports that supervised methods outperform nearest-neighbor and zero-shot LLM baselines in both unconstrained accuracy and constrained cost settings.

HistoRAG: Embedding Historical Methodology in Retrieval-Augmented Generation Through Critical Technical Practice

cs.CL · 2026-06-16 · unverdicted · novelty 6.0

HistoRAG embeds historiographical principles into RAG via temporal windowing, decoupled retrieval, and contestable LLM relevance judgments, evaluated on 102k Der Spiegel articles from 1950-1979.

citing papers explorer

Showing 33 of 33 citing papers after filters.

From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability cs.IR · 2026-06-06 · unverdicted · none · ref 2 · internal anchor
A benchmark and ontology-driven framework links 434 cardiovascular devices to patents at 91.6% recall, producing 6.8M high-confidence links for regulatory-IP integration.
Test-Time Training for Zero-Resource Dense Retrieval Reranking cs.IR · 2026-05-31 · unverdicted · none · ref 6 · internal anchor
DART adapts a scoring matrix at inference time via gradient updates on pseudo-labels from top/bottom documents to gain +2.1% mean NDCG@10 on six BEIR benchmarks with under 10ms added latency.
Layer-wise Token Compression for Efficient Document Reranking cs.IR · 2026-05-20 · unverdicted · none · ref 32 · 2 links · internal anchor
Layer-wise Token Compression applies adaptive token pooling at middle transformer layers for cross-encoder rerankers, preserving MS MARCO ranking quality while raising QPS up to 25% on passages and 116% on documents, with added gains on listwise LLM rerankers and a regularizer effect for long inputs
Very Efficient Listwise Multimodal Reranking for Long Documents cs.IR · 2026-05-12 · unverdicted · none · ref 46 · internal anchor
ZipRerank delivers state-of-the-art multimodal listwise reranking accuracy for long documents at up to 10x lower latency via early interaction and single-pass scoring.
Bayesian Active Learning with Gaussian Processes Guided by LLM Relevance Scoring for Dense Passage Retrieval cs.IR · 2026-04-20 · unverdicted · none · ref 27 · internal anchor
BAGEL is a Bayesian active learning framework that uses Gaussian Processes to propagate LLM relevance signals across embedding space and guide global exploration, outperforming standard LLM reranking under identical budgets on four retrieval benchmarks.
Scaling Laws for Cross-Encoder Reranking cs.IR · 2026-03-05 · unverdicted · none · ref 32 · internal anchor
Cross-encoder reranker performance scales predictably via power laws with model size and training exposure, allowing accurate forecasts for 400M and 1B models and data-heavy compute allocation.
SPIRE: Structure-Preserving Interpretable Retrieval of Evidence cs.IR · 2026-02-12 · unverdicted · none · ref 17 · internal anchor
SPIRE presents a tree-structured retrieval method using subdocuments, paths, and dual contextualization that produces higher-quality and more diverse citations than passage-based baselines on HTML QA benchmarks.
CoDeR: Local Constraint-Compatible Retrieval Beyond Semantic Similarity cs.IR · 2026-06-11 · unverdicted · none · ref 25 · internal anchor
CoDeR augments standard topical dense retrieval with a bi-encoder compatibility scorer trained via contrastive lexical-polarity supervision to reduce early exposure to constraint-violating documents.
STORM: Stepwise Token Optimization with Reward-Guided Beam Search cs.IR · 2026-06-09 · unverdicted · none · ref 38 · internal anchor
STORM trains lexical query rewriters via reward-guided beam search that converts retrieval metrics into stepwise token signals, enabling 0.6B-8B models to rival dense retrievers on TREC, BEIR and MIRACL without index changes.
Interactive Multi-Turn Retrieval for Health Videos cs.IR · 2026-05-02 · unverdicted · none · ref 25 · internal anchor
DATR combines coarse CLIP-based retrieval with multi-turn query fusion and cross-encoder re-ranking to improve health video retrieval, supported by the new MHVRC corpus.
Beyond Single-Score Ranking: Facet-Aware Reranking for Controllable Diversity in Paper Recommendation cs.IR · 2026-03-11 · unverdicted · none · ref 2 · internal anchor
SciFACE improves facet-specific paper ranking NDCG scores by training separate cross-encoders for Background and Method similarity on 5,891 GPT-4o-mini labeled pairs, outperforming SPECTER by up to 31 points.
Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems cs.IR · 2026-01-08 · unverdicted · none · ref 10 · internal anchor
W-RAC decouples extraction from semantic planning via structured units and LLM grouping to match traditional retrieval performance at roughly 10x lower LLM token cost.
ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking cs.IR · 2025-06-04 · unverdicted · none · ref 3 · internal anchor
ProRank uses RL-based prompt warmup and fine-grained scoring to train small language models that surpass LLM rerankers on BEIR.
Unsupervised Dense Information Retrieval with Contrastive Learning cs.IR · 2021-12-16 · unverdicted · none · ref 163 · internal anchor
Contrastive learning trains unsupervised dense retrievers that beat BM25 on most BEIR datasets and support cross-lingual retrieval across scripts.
The Crowded Embedding Space: A Mean-Field Mechanism for Emergent Marginalization in Retrieval-Augmented Agents cs.IR · 2026-06-01 · unverdicted · none · ref 12 · internal anchor
A mean-field analysis of embedding-space crowding shows a phase transition and Fokker-Planck dynamics that drive retrieval-augmented agents to self-organize toward exclusive service of majority interests.
LRanker: LLM Ranker for Massive Candidates cs.IR · 2026-05-27 · unverdicted · none · ref 16 · internal anchor
LRanker combines K-means candidate aggregation with graph-partitioned ensemble of query embeddings to improve LLM ranking accuracy and scalability on massive candidate pools, reporting 3-30% gains on RBench tasks up to 6.8M candidates.
Lost in the Evidence? Reproducing Document Position and Context Size Effects in RAG cs.IR · 2026-05-26 · unverdicted · none · ref 17 · internal anchor
Reproducibility study shows position and context size effects in RAG depend on topic sampling and retrieval quality, proposes calibration for stable trends, and releases code after finding discrepancies with prior industry work.
CALMem : Application-Layer Dual Memory for Conversational AI cs.IR · 2026-05-20 · unverdicted · none · ref 7 · internal anchor
CALMem delivers virtually unbounded effective context for LLM conversations via an application-layer dual memory architecture with intra-session retrieval and token-adaptive injection.
KG-First, LLM-Fallback: A Hybrid Microservice for Grounded Skill Search and Explanation cs.IR · 2026-05-02 · unverdicted · none · ref 12 · internal anchor
SkillGraph-Service builds a provenance-preserving knowledge graph from multiple competency frameworks and achieves nDCG@5 above 0.94 with sub-200 ms latency via KG-first hybrid retrieval and constrained LLM explanations.
Efficient Listwise Reranking with Compressed Document Representations cs.IR · 2026-04-29 · unverdicted · none · ref 23 · internal anchor
RRK compresses documents to multi-token embeddings for efficient listwise reranking, enabling an 8B model to achieve 3x-18x speedups over smaller models with comparable or better effectiveness.
Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation for Dense Retrieval cs.IR · 2026-04-06 · unverdicted · none · ref 21 · internal anchor
Stratified sampling preserving teacher score distribution outperforms hard-negative mining as a robust baseline for knowledge distillation in dense retrieval.
The Role of Vocabularies in Learning Sparse Representations for Ranking cs.IR · 2025-09-20 · unverdicted · none · ref 13 · internal anchor
Larger 100K vocabularies in SPLADE models, especially those initialized with ESPLADE pretraining, improve retrieval effectiveness after pruning compared to 32K baselines while keeping similar efficiency.
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey cs.IR · 2025-09-09 · unverdicted · none · ref 81 · internal anchor
A comprehensive survey that organizes query expansion methods in the PLM/LLM era along four design dimensions, synthesizes application patterns, and outlines future directions.
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval cs.IR · 2025-04-20 · unverdicted · none · ref 11 · internal anchor
LLM-generated synthetic hard negatives for training dense retrievers consistently underperform corpus-mined negatives from BM25 and cross-encoders across 10 BEIR datasets, with non-monotonic gains from scaling the generator from 4B to 30B parameters.
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs cs.IR · 2024-06-17 · unverdicted · none · ref 4 · internal anchor
ITEM is a new iterative utility judgment loop for RAG that maps Schutz's three levels of relevance to retrieval, utility scoring, and generation, yielding measured gains on TREC DL, WebAP, GTI-NQ, and NQ.
DSIRM: Learning Query-Bridged Discrete Semantic Identifiers for E-commerce Relevance Modeling cs.IR · 2026-06-03 · unverdicted · none · ref 18 · internal anchor
DSIRM uses query-bridged contrastive quantization and generative LLMs to create relevance-aware discrete semantic identifiers, reporting +1.54% offline AUC and online lifts on Tmall production data.
RAG-Match: Retrieval-Augmented Knowledge Injection and Hierarchical Reasoning for Calibrated Semantic Relevance cs.IR · 2026-05-25 · unverdicted · none · ref 23 · internal anchor
RAG-Match is a three-stage framework for semantic relevance modeling that integrates knowledge-augmented pretraining, hierarchical reasoning alignment, and preference-based decision calibration, outperforming LLM baselines on a search benchmark.
SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents cs.IR · 2026-05-08 · unverdicted · none · ref 11 · internal anchor
SkillSelect-Serve improves same-budget bundle recall and mean utility for LLM agent skill selection over fixed top-k retrieval by using structured Skill Services, a Micro-Agent Requirement Planner, and dual-granularity utility modeling on 35,353 skills and 586 queries.
LLM-Oriented Information Retrieval: A Denoising-First Perspective cs.IR · 2026-05-01 · unverdicted · none · ref 135 · 2 links · internal anchor
Argues for a denoising-first paradigm in LLM-oriented information retrieval, framing challenges via a four-stage progression and providing a taxonomy of signal-to-noise optimization techniques across the pipeline.
FRAGATA: Semantic Retrieval of HPC Support Tickets via Hybrid RAG over 20 Years of Request Tracker History cs.IR · 2026-04-15 · unverdicted · none · ref 4 · internal anchor
Fragata applies hybrid RAG to enable semantic retrieval of HPC support tickets across 20 years of history, handling language differences, typos, and varied wording better than traditional keyword search.
RAGe: A Retrieval-Augmented Generation Evaluation Framework cs.IR · 2026-05-23 · unverdicted · none · ref 25 · internal anchor
RAGe is a modular evaluation framework that correlates retrieval and generation quality with hardware constraints to recommend optimal RAG components for specific datasets.
A Case-Driven Multi-Agent Framework for E-Commerce Search Relevance cs.IR · 2026-05-07 · unverdicted · none · ref 24 · internal anchor
A case-driven multi-agent system automates the full pipeline of bad-case detection, annotation, and resolution for e-commerce search relevance using Annotator, Optimizer, and User agents plus supporting components.
Let's measure run time! Extending the IR replicability infrastructure to include performance aspects cs.IR · 2019-07-10 · unverdicted · none · ref 13 · internal anchor
Position paper proposing to extend the OSIRRC replicability infrastructure with two performance benchmark scenarios, backed by a case study on neural re-ranking model runtimes.

Passage Re-ranking with BERT

hub tools

citation-role summary

citation-polarity summary

claims ledger

co-cited works

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer