Pith. sign in

REVIEW 1 cited by

Investigating the Scalability of Approximate Sparse Retrieval Algorithms to Massive Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11628 v1 pith:JDOQK6OP submitted 2025-01-20 cs.IR

classification cs.IR
keywords retrievaldatasetsalgorithmsembeddingssparseapproximatedistributionaleffectiveness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learned sparse text embeddings have gained popularity due to their effectiveness in top-k retrieval and inherent interpretability. Their distributional idiosyncrasies, however, have long hindered their use in real-world retrieval systems. That changed with the recent development of approximate algorithms that leverage the distributional properties of sparse embeddings to speed up retrieval. Nonetheless, in much of the existing literature, evaluation has been limited to datasets with only a few million documents such as MSMARCO. It remains unclear how these systems behave on much larger datasets and what challenges lurk in larger scales. To bridge that gap, we investigate the behavior of state-of-the-art retrieval algorithms on massive datasets. We compare and contrast the recently-proposed Seismic and graph-based solutions adapted from dense retrieval. We extensively evaluate Splade embeddings of 138M passages from MsMarco-v2 and report indexing time and other efficiency and effectiveness metrics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?

    cs.CY 2025-08 unverdicted novelty 4.0 of 10

    The manuscript is internally inconsistent: the abstract claims VLM art-attribution experiments, while the full text is an unrelated hybrid-search benchmark paper.

Pith tools