Pith. sign in

REVIEW 2 cited by

RepBERT: Contextualized Text Embeddings for First-Stage Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.15498 v2 pith:DRGTAI53 submitted 2020-06-28 cs.IR

classification cs.IR
keywords embeddingsrepbertretrievalcontextualizeddocumentsfirst-stagequeriesachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although exact term match between queries and documents is the dominant method to perform first-stage retrieval, we propose a different approach, called RepBERT, to represent documents and queries with fixed-length contextualized embeddings. The inner products of query and document embeddings are regarded as relevance scores. On MS MARCO Passage Ranking task, RepBERT achieves state-of-the-art results among all initial retrieval techniques. And its efficiency is comparable to bag-of-words methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PARK: Personalized academic retrieval with knowledge-graphs

    cs.IR 2025-07 conditional novelty 6.0 of 10

    PARK personalizes academic search by embedding a citation-derived knowledge graph into the same vector space as a neural retrieval model, beating baselines in three of four domains.

  2. UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval

    cs.AI 2026-08 conditional novelty 5.0 of 10

    UniGD couples generative retrieval with explicit relevance scoring in one model, reporting +5.78% ad revenue, 33.1% lower latency at Kuaishou, and improved Recall@10 on NQ320K and MS300K.

Pith tools