Pith. sign in

REVIEW 5 cited by

Understanding the Behaviors of BERT in Ranking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.07531 v4 pith:OR24DKIG submitted 2019-04-16 cs.IR cs.CL

classification cs.IRcs.CL
keywords bertrankingtasksbehaviorsdocumentexperimentalmarcopassage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper studies the performances and behaviors of BERT in ranking tasks. We explore several different ways to leverage the pre-trained BERT and fine-tune it on two ranking tasks: MS MARCO passage reranking and TREC Web Track ad hoc document ranking. Experimental results on MS MARCO demonstrate the strong effectiveness of BERT in question-answering focused passage ranking tasks, as well as the fact that BERT is a strong interaction-based seq2seq matching model. Experimental results on TREC show the gaps between the BERT pre-trained on surrounding contexts and the needs of ad hoc document ranking. Analyses illustrate how BERT allocates its attentions between query-document tokens in its Transformer layers, how it prefers semantic matches between paraphrase tokens, and how that differs with the soft match patterns learned by a click-trained neural ranker.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning an Effective Premise Retrieval Model for Efficient Mathematical Formalization

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A compact BERT retriever for Lean premises, trained with a Lean-specific tokenizer and fine-grained similarity plus re-ranking, outperforms prior retrievers on most splits and lifts MiniF2F pass@1.

  2. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A survey organizes current research on trustworthy RAG into six pillars, reliability, privacy, safety, fairness, explainability, and accountability, and maps methods, metrics, and open problems for each.

  3. HNCSE: Advancing Sentence Embeddings via Hybrid Contrastive Learning with Hard Negatives

    cs.CL 2024-11 reject novelty 4.0 of 10

    HNCSE reports 2-point average STS gains over SimCSE using positive mixing and hard-negative mixing, but the method is under-specified and unverified.

  4. Revisiting Semantic Representation and Tree Search for Similar Question Retrieval

    cs.CL 2019-08 conditional novelty 4.0 of 10

    Similar-question retrieval using BERT embeddings can be accelerated by a k-means tree with beam search, at a cost of about 0.04 MAP on a Quora Question Pairs ranking task.

  5. A Study of BERT for Non-Factoid Question-Answering under Passage Length Constraints

    cs.IR 2019-08 conditional novelty 4.0 of 10

    BERT fine-tuning substantially improves non-factoid passage re-ranking over prior baselines, with a 256-token input window performing best and chunking providing a workaround for longer passages.

Pith tools