Pith. sign in

REVIEW 1 cited by

From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.04733 v2 pith:T2OH6PM2 submitted 2022-05-10 cs.IR cs.CL

classification cs.IRcs.CL
keywords modelsdistillationsparsetrainingbeendenseneuralsame
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural retrievers based on dense representations combined with Approximate Nearest Neighbors search have recently received a lot of attention, owing their success to distillation and/or better sampling of examples for training -- while still relying on the same backbone architecture. In the meantime, sparse representation learning fueled by traditional inverted indexing techniques has seen a growing interest, inheriting from desirable IR priors such as explicit lexical matching. While some architectural variants have been proposed, a lesser effort has been put in the training of such models. In this work, we build on SPLADE -- a sparse expansion-based retriever -- and show to which extent it is able to benefit from the same training improvements as dense models, by studying the effect of distillation, hard-negative mining as well as the Pre-trained Language Model initialization. We furthermore study the link between effectiveness and efficiency, on in-domain and zero-shot settings, leading to state-of-the-art results in both scenarios for sufficiently expressive models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comparative Study of Text Retrieval Models on DaReCzech

    cs.IR 2024-11 conditional novelty 4.0 of 10

    A benchmark on the Czech DaReCzech dataset finds Gemma2 most accurate, Contriever least accurate, and SPLADE/PLAID the best efficiency-quality trade-off.

Pith tools