REVIEW 2 cited by
The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Information Retrieval using dense low-dimensional representations recently became popular and showed out-performance to traditional sparse-representations like BM25. However, no previous work investigated how dense representations perform with large index sizes. We show theoretically and empirically that the performance for dense representations decreases quicker than sparse representations for increasing index sizes. In extreme cases, this can even lead to a tipping point where at a certain index size sparse representations outperform dense representations. We show that this behavior is tightly connected to the number of dimensions of the representations: The lower the dimension, the higher the chance for false positives, i.e. returning irrelevant documents.
Forward citations
Cited by 2 Pith papers
-
Quantum-inspired Embeddings Projection and Similarity Metrics for Representation Learning
A quantum-inspired, parameter-light projection head compressing BERT embeddings to 256 dimensions matches a classical dense head on TREC passage reranking and improves on small training sets.
-
MST-R: Multi-Stage Tuning for Retrieval Systems and Metric Evaluation
A multi-stage retrieval system (MST-R) improves Recall@10 from 0.78 to 0.87 on the ObliQA regulatory dataset, and a passage-concatenation baseline inflates the RePASs answer metric to 0.95.
Discussion (0). Continue with ORCID to comment.