Pith. sign in

REVIEW 2 cited by

The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.14210 v2 pith:3N2URKYF submitted 2020-12-28 cs.IR cs.CL

classification cs.IRcs.CL
keywords representationsdenseindexsizesinformationlargelow-dimensionalretrieval
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Information Retrieval using dense low-dimensional representations recently became popular and showed out-performance to traditional sparse-representations like BM25. However, no previous work investigated how dense representations perform with large index sizes. We show theoretically and empirically that the performance for dense representations decreases quicker than sparse representations for increasing index sizes. In extreme cases, this can even lead to a tipping point where at a certain index size sparse representations outperform dense representations. We show that this behavior is tightly connected to the number of dimensions of the representations: The lower the dimension, the higher the chance for false positives, i.e. returning irrelevant documents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantum-inspired Embeddings Projection and Similarity Metrics for Representation Learning

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A quantum-inspired, parameter-light projection head compressing BERT embeddings to 256 dimensions matches a classical dense head on TREC passage reranking and improves on small training sets.

  2. MST-R: Multi-Stage Tuning for Retrieval Systems and Metric Evaluation

    cs.IR 2024-12 conditional novelty 5.0 of 10

    A multi-stage retrieval system (MST-R) improves Recall@10 from 0.78 to 0.87 on the ObliQA regulatory dataset, and a passage-concatenation baseline inflates the RePASs answer metric to 0.95.

Pith tools