Pith. sign in

QJL: 1-bit quantized JL transform for KV cache quan- tization with zero overhead

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it

years

2026 8 2025 1

representative citing papers

Runtime-Certified Bounded-Error Quantized Attention

cs.LG · 2026-05-20 · unverdicted · novelty 6.0

A tiered KV cache architecture computes per-head per-step error bounds on quantized attention and uses adaptive fallback to guarantee bounded or exact outputs relative to FP16 reference.

AXELRAM: Quantize Once, Never Dequantize

cs.LG · 2026-04-03 · conditional · novelty 6.0

AXELRAM performs attention on quantized KV cache using a fixed orthogonal-transform codebook, reducing multiplications by 102.4x and fixing sign-sensitivity spikes via gradient-free calibration.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

cs.LG · 2025-04-28 · unverdicted · novelty 6.0

TurboQuant achieves near-optimal vector quantization distortion for both MSE and inner products via random rotation and per-coordinate scalar quantization, with a formal proof that it matches lower bounds within a factor of approximately 2.7.

UltraQuant: 4-bit KV Caching for Context-Heavy Agents

cs.LG · 2026-06-18 · unverdicted · novelty 4.0

UltraQuant applies 4-bit KV caching with TurboQuant-style rotation and custom AMD kernels to context-heavy agent workloads, delivering 3.47x P50 TTFT reduction in cache-pressured rounds and 1.63x throughput gain over FP8.

citing papers explorer

Showing 9 of 9 citing papers.