Pith. sign in

GPU-Accelerated ANNS: Quantized for Speed, Built for Change

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promising path to high-performance ANNS: they provide massive parallelism for distance computations, are readily available, and can co-locate with downstream applications. Despite these advantages, current GPU-accelerated ANNS systems face three key limitations. First, real-world applications operate on evolving datasets that require fast batch updates, yet most GPU indices must be rebuilt from scratch when new data arrives. Second, high-dimensional vectors strain memory bandwidth, but current GPU systems lack efficient quantization techniques that reduce data movement without introducing costly random memory accesses. Third, the data-dependent memory accesses inherent to greedy search make overlapping compute and memory difficult, leading to reduced performance. We present Jasper, a GPU-native ANNS system with both high query throughput and updatability. Jasper builds on the Vamana graph index and overcomes existing bottlenecks via three contributions: (1) a CUDA batch-parallel construction algorithm that enables lock-free streaming insertions, (2) a GPU-efficient implementation of RaBitQ quantization that reduces memory footprint up to 8x without the random access penalties, and (3) an optimized greedy search kernel that increases compute utilization, resulting in better latency hiding and higher throughput. Our evaluation across five datasets shows that Jasper achieves up to 1.84x higher query throughput than CAGRA and achieves up to 80% peak utilization as measured by the roofline model. Jasper's construction scales efficiently and constructs indices an average of 7x faster than CAGRA while providing updatability that CAGRA lacks. Compared to BANG, the previous fastest GPU Vamana implementation, Jasper delivers 10-74x faster queries.

fields

cs.DC 1

years

2026 1

verdicts

ACCEPT 1

representative citing papers

InferScale: GPU-Native KV Injection for Personalized LLM Serving

cs.DC · 2026-07-29 · accept · novelty 6.0

GPU-resident precomputed KV injection with Chunked RoPE and context-window encoding makes personalized LLM memory latency nearly independent of retrieval budget while nearly matching prompt-injection accuracy.

citing papers explorer

Showing 1 of 1 citing paper.

  • InferScale: GPU-Native KV Injection for Personalized LLM Serving cs.DC · 2026-07-29 · accept · none · ref 12 · internal anchor

    GPU-resident precomputed KV injection with Chunked RoPE and context-window encoding makes personalized LLM memory latency nearly independent of retrieval budget while nearly matching prompt-injection accuracy.