Pith. sign in

REVIEW 4 cited by

Context-Aware Sentence/Passage Term Importance Estimation For First Stage Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.10687 v2 pith:JJMB4STN submitted 2019-10-23 cs.IR

classification cs.IR
keywords termretrievaltextcontextualizeddeepquerywhenalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Term frequency is a common method for identifying the importance of a term in a query or document. But it is a weak signal, especially when the frequency distribution is flat, such as in long queries or short documents where the text is of sentence/passage-length. This paper proposes a Deep Contextualized Term Weighting framework that learns to map BERT's contextualized text representations to context-aware term weights for sentences and passages. When applied to passages, DeepCT-Index produces term weights that can be stored in an ordinary inverted index for passage retrieval. When applied to query text, DeepCT-Query generates a weighted bag-of-words query. Both types of term weight can be used directly by typical first-stage retrieval algorithms. This is novel because most deep neural network based ranking models have higher computational costs, and thus are restricted to later-stage rankers. Experiments on four datasets demonstrate that DeepCT's deep contextualized text understanding greatly improves the accuracy of first-stage retrieval algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A shared integrated teacher improves both sparse and dense text-image retrieval, letting a sparse retriever match or beat dense baselines on MSCOCO and several Flickr30k settings.

  2. HyReC: Exploring Hybrid-based Retriever for Chinese

    cs.IR 2025-06 conditional novelty 6.0 of 10

    HyReC unifies dense, lexicon, and learned word-segment retrieval into one model and reports improved C-MTEB retrieval scores for Chinese.

  3. ERU-KG: Efficient Reference-aligned Unsupervised Keyphrase Generation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ERU-KG uses reference-trained SPLADE term importances plus neighbor-document noun phrases to generate present and absent keyphrases without keyphrase labels, and reports strong benchmark and retrieval results.

  4. A Multi-Task Evaluation of LLMs' Processing of Academic Text Input

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    The abstract reports Gemini underperforms on four academic text tasks, but the attached full text is an unrelated biomedical retrieval paper, leaving the claims unverifiable.

Pith tools