Pith. sign in

Paper Citation Record · LEDGER

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 7 inbound Pith citation observations for arXiv:2412.08890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08890 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:32:39.990624Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:26:55.395701Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.257305Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6996c122-ee75-4e9d-a5c2-89085e9018f3 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head check- points.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Gqa: Training generalized multi-query transformer models from multi-head check- points

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.921656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.921656Z digest=sha256:40748e92713c60bde9181fe451a06d61bf04be0045eb1695080adc92e5afddba

Observation 7304570a-1d63-4f0b-a935-fde0456f0668 · outbound

This paper cites Longformer: The Long-Document Transformer.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.928529Z digest=sha256:514012c0002017757025a6dffc5157f5ac54c835f27604a3251adeb8e78e05b7

Observation 16760b27-39cb-4865-a6f4-e8cb8057b340 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.931782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.931782Z digest=sha256:6f143dd1ee608b934a2f86689ede228cbbc2cc01d7a776507a1ab358e52a4f70

Observation 2a14adf7-5e1a-4605-a1b1-e1e792b0d052 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.944609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.944609Z digest=sha256:3a3661f35d7030e1f8d04ca5ae89dcade7e15637879b25981c47dcdd3e1c1f82

Observation ad012494-5137-4507-bdda-9643a54b41c4 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.947554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.947554Z digest=sha256:269bcd2581b367477aa6cfa321812b82de2f57ba4e00f08fcd0286a2b598b8cf

Observation 44398d71-fd32-41e3-987c-c6fb764c9fef · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.949962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.949962Z digest=sha256:6551711c08eb57c967004f15b2a05300ef0dff3d4571aeccf6378580224f4564

Observation a6e0556e-6d30-467b-90d9-eb08435f1879 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.952337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.952337Z digest=sha256:b34aaf1641b6114f97142c13efaa3d164764fe0bc110b741ef3240ae45b28aee

Observation 9a573994-344b-4f0a-9c48-6cd4df391568 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries SnapKV: LLM Knows What You are Looking for Before Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.954480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.954480Z digest=sha256:bbfbc59df3135bd31c53e6ac226134697d8449b6754ff6c73e0b4b23d9e96d83

Observation 3ff7a8d3-7e33-4302-a0d9-1ceac9c3347d · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.956930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.956930Z digest=sha256:468d21b35cb2ad9474e367a2f570f7ccc5a40982caccd833e72aa141d83d0296

Observation 940f14fa-18b7-4b36-ba3f-0d9e6e88aca2 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.959695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.959695Z digest=sha256:3ea627a1963c359b1677a7ed9b4b37871cb03e83721aa32f835a3e9ad73c3115

Observation 0325c185-17c6-478f-8446-98063aa181c8 · outbound

This paper cites k-Sparse Autoencoders.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries k-Sparse Autoencoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.962371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.962371Z digest=sha256:fb380e0432fb00161ae488b8958908ad0f949e90f051c6fac69206bd73fd02a9

Observation 04eba951-ee5d-4c55-9cf5-612fecbf5103 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.965032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.965032Z digest=sha256:ea41f3c9a0b355241458fe50e54d26de22999cc927176ffe0b7b5ac06e6b66e0

Observation de5dc1e6-9709-4b4c-aa49-54da6f37c8c2 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Fast Transformer Decoding: One Write-Head is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.967735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.967735Z digest=sha256:c334e7f9d5aad7217efd000f70e9f5bea0810398430662052b8e58b6a96ce9cf

Observation 9c290bd5-84e9-4430-ac12-2e7c84472eb9 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.973242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.973242Z digest=sha256:b939255a1c5c0bb6ea09972a4ac1e4b633732ea208258c8af692908f56dd2d9c

Observation 565271d7-35bb-44d4-841d-f0e363bb956c · outbound

This paper cites Attention is all you need.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Attention is all you need

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:32:40.151637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:32:39.976451Z digest=sha256:090cf7bb862717d355da145e5ff0da8649e8df5a3d0c9880d19709a2820ecab2

Observation 7704faf6-a812-4e49-99c2-178b456a92d9 · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.982079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.982079Z digest=sha256:e45f7c140c23b674e49cb3d8ae7a92a9f9f3e8cf6101eec59abf78b386663a42

Observation f8b489a9-e388-4c65-9016-444ae4d34c70 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.987580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.987580Z digest=sha256:2f0381d48294e3d83e0efb8cbe6677f774e103b550659ec899fabe3cdc8ee635

Observation b7a2ab87-7f4b-483f-8907-2113a51c0c9b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Training Verifiers to Solve Math Word Problems

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.935049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.935049Z digest=sha256:e65364a8436e7442624965277df8d6ee6c934796e7ed910cd6f3a18f058af553

Observation cebab9c8-3e35-4a9e-8c79-c79e017138a2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.979186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.979186Z digest=sha256:6a15e55acd145afe8d6b52e4de0e1f7a5ded35a0e36b15c0ff413d8490bb4efa

Observation 2fe16acb-c6c4-42fa-9764-69071f0ef38a · outbound

This paper cites Loki: Low-rank Keys for Efficient Sparse Attention.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Loki: Low-rank Keys for Efficient Sparse Attention

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.970666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.970666Z digest=sha256:0f775cd2dcdb1f5039f78b98e4ce059407d88e01d3611b92ff946a2491482e5b

Observation adb11b81-00c7-4e26-8c6f-2ee114404806 · outbound

This paper cites APPENDIX A I MPLEMENTATION DETAILS Algorithm 1 illustrates a naive implementation of OMP for understanding.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries APPENDIX A I MPLEMENTATION DETAILS Algorithm 1 illustrates a naive implementation of OMP for understanding

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:32:40.143729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T17:32:39.990624Z digest=sha256:7435646abe8a46244ef7914248af3129209caf756101e0967eda477cefcbd93b

Observation 7ca79eb1-e2b5-4e91-8f93-196b6704489d · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.938181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.938181Z digest=sha256:7aa71aedaa9029ca117aee36a43432a866e3234ad0fb0b88bec33a9f33b9dc3f

Observation 79b0f42e-f2ec-4eca-8a4e-b02232bc8564 · outbound

This paper cites Effectively Compress KV Heads for LLM.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries Effectively Compress KV Heads for LLM

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.984824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.984824Z digest=sha256:03bc1651c0ac93d507cb02f3fb10bae18725e1ea20b213bc5451443141549bb9

Observation 633141e6-7103-42f7-8519-27a5ba51c752 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.924932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.924932Z digest=sha256:8ce60fafa80fdd80b09301e1a942957f6e16b1e3441e752fdb8200bfb539b220

Observation dcea4b12-2e86-4870-a977-e26441bb3c7c · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:32:39.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:32:39.941220Z digest=sha256:0c6582be322ddef75e1939e1932632a3e775da165a449e84b155be576443540c

Pith citing papers

Observation e0d28393-71b5-46a6-a19d-bb518fb9b178 · inbound

PolarQuant: Quantizing KV Caches with Polar Transformation cites this paper.

PolarQuant: Quantizing KV Caches with Polar Transformation Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T13:26:55.395701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:26:55.395701Z digest=sha256:f99834d7ef05791662227182e2aaecfffa7c67d7aa0843bdccda20ac5a1c3251

Observation 624879cc-1d29-4d46-85a0-ca699718236c · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.270351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:79067b81da0d9fc61de1567ed8edb33624b1d541af86b49f8bbeaca9afbf5b37

Observation 1916c9f4-cb6c-49b4-b9f7-9d002151684e · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.436604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.436604Z digest=sha256:16972589229eca94ba2e658e66ea492fc40961327ab0ce3092c25019d39ae9f2

Observation d94ec039-2f70-4e0b-ae55-9c8fedb922fe · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:05.138038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:462a5d7256453eb7af9cf13db76e25ce0733f9416ffd6fccc69253ecb8dcd6ff

Observation 28d8da42-9c4a-465a-9895-a050ec8cdefa · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.258656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:63770399f1d2929eb183224c279f85126b88a103852ac2e5c7edb7a065873828

Observation 81229a86-7276-4c50-9de8-532cb996de2a · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.119424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:4c78b6cf0dd85b0f50f4c0139088025898379a53d24bfbcad8f1f3dbd9c805a2

Observation b3d8b8f7-8db9-4516-91fe-ce7f84b53e57 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:04.379792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:04.379792Z digest=sha256:0fbf212c6387d0511cb8ef8c7c7cb8afd0101acbac6ab21725f705c9843fc0ca