Pith. sign in

REVIEW 5 cited by

A Survey on Locality Sensitive Hashing Algorithms and their Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.08942 v1 pith:QR5AVGQY submitted 2021-02-17 cs.DB

A Survey on Locality Sensitive Hashing Algorithms and their Applications

classification cs.DB
keywords hashinglocalitysensitivesurveyapplicationdomainsfindinghigh-dimensional
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Finding nearest neighbors in high-dimensional spaces is a fundamental operation in many diverse application domains. Locality Sensitive Hashing (LSH) is one of the most popular techniques for finding approximate nearest neighbor searches in high-dimensional spaces. The main benefits of LSH are its sub-linear query performance and theoretical guarantees on the query accuracy. In this survey paper, we provide a review of state-of-the-art LSH and Distributed LSH techniques. Most importantly, unlike any other prior survey, we present how Locality Sensitive Hashing is utilized in different application domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ASH: Asymmetric Scalar Hashing With Learned Dimensionality Reduction for High-Fidelity Vector Quantization

    cs.IR 2026-06 unverdicted novelty 7.0

    ASH achieves state-of-the-art ANN recall and speed across compression levels by learning an orthonormal projection for dimensionality reduction followed by scalar quantization in an asymmetric encoder-decoder setup.

  2. H3D: Benchmarking Unsupervised Text Hashing for Fine-Grained Document Deduplication

    cs.IR 2026-07 conditional novelty 6.0

    Lexical non-learning hashes match near-duplicates well, while BGE-based quantized embeddings better preserve rewritten scientific similarity, under a shared ranking protocol on CSFCube and RELISH.

  3. When More Cores Hurts: The Vector Database Scaling Paradox in HPC

    cs.DC 2026-06 unverdicted novelty 6.0

    Large-scale HPC evaluation of Qdrant, Milvus, and Weaviate reveals that workload patterns limit scaling and extra cores can reduce throughput, exposing a cloud-to-HPC design mismatch.

  4. RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching

    cs.DC 2026-05 unverdicted novelty 5.0

    RcLLM accelerates generative recommendation inference by 1.31x-9.51x in TTFT through beyond-prefix KV caching, replicated user caches, sharded item caches, affinity scheduling, and selective attention with negligible ...

  5. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets

    cs.CV 2023-04 unverdicted novelty 5.0

    The AICrowd dataset has 90% training duplicates and 93% validation-to-training leakage; a perceptual hashing pipeline detects and mitigates these issues.