Pith. sign in

Subgen: Token generation in sublinear time and memory

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.LG 3 cs.DS 1

years

2026 3 2025 1

verdicts

UNVERDICTED 4

representative citing papers

Nearly Optimal Attention Coresets

cs.DS · 2026-05-07 · unverdicted · novelty 8.0

ε-coresets for attention exist of size O(√d e^{ρ+o(ρ)}/ε) for unit-norm keys/values and queries of norm ≤ρ, nearly matching the Ω(√d e^ρ/ε) lower bound.

RoPE-Aware Bit Allocation for KV-Cache Quantization

cs.LG · 2026-06-23 · unverdicted · novelty 7.0

Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.

The risk of KV cache compression

cs.LG · 2026-07-01 · unverdicted · novelty 6.0

The paper derives a characterization of minimax risk for KV cache compression and maps it to practical design principles and an algorithm tested on LongBench.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

cs.LG · 2025-04-28 · unverdicted · novelty 6.0

TurboQuant achieves near-optimal vector quantization distortion for both MSE and inner products via random rotation and per-coordinate scalar quantization, with a formal proof that it matches lower bounds within a factor of approximately 2.7.

citing papers explorer

Showing 4 of 4 citing papers.

  • Nearly Optimal Attention Coresets cs.DS · 2026-05-07 · unverdicted · none · ref 52

    ε-coresets for attention exist of size O(√d e^{ρ+o(ρ)}/ε) for unit-norm keys/values and queries of norm ≤ρ, nearly matching the Ω(√d e^ρ/ε) lower bound.

  • RoPE-Aware Bit Allocation for KV-Cache Quantization cs.LG · 2026-06-23 · unverdicted · none · ref 42

    Block-GTQ performs RoPE-aware greedy bit allocation on KV caches using per-block energy scores, cutting logit MAE 32-80% versus uniform TQ-MSE and lifting long-context task scores substantially at 2-3 bits per dimension.

  • The risk of KV cache compression cs.LG · 2026-07-01 · unverdicted · none · ref 26

    The paper derives a characterization of minimax risk for KV cache compression and maps it to practical design principles and an algorithm tested on LongBench.

  • TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cs.LG · 2025-04-28 · unverdicted · none · ref 64

    TurboQuant achieves near-optimal vector quantization distortion for both MSE and inner products via random rotation and per-coordinate scalar quantization, with a formal proof that it matches lower bounds within a factor of approximately 2.7.