Pith. sign in

Paper Citation Record · LEDGER

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2506.04642.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04642 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:09.969345Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:47:31.653881Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T06:16:20.379382Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95ddf0a1-0c66-4313-9ee2-d46877579bad · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.814940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.814940Z digest=sha256:6ea23741805b48529f7d2f8a6af1d02ca16e0c3e9ba8dd2c711a68643e415bc0

Observation 269b8026-bb4a-49c8-85a8-2ccdac6e1ca7 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.820633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.820633Z digest=sha256:fbfc48d3441a140b32ffb75c41b0e3a78f458699e7a651858c5cdebd08e33440

Observation 484ef146-5291-4782-9dd0-fdc49ba24a4b · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Palu: Compressing KV-Cache with Low-Rank Projection

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.825323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.825323Z digest=sha256:7a62fe7039e86c67543bd7abae13d340ad3e5699c85f273c03d487b90d3f73ef

Observation 3ff40ad3-f55c-46db-926c-d51afbbd0020 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.832690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.832690Z digest=sha256:699fc904e206ea412a0e2adb8e71fa630e238d690d09a0211033100c832f8c48

Observation f11cb9aa-bcc5-4735-a342-3b24f0178fc2 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.837526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.837526Z digest=sha256:880a2dc3120f1d336c3709b49a3a28147cf0aa0f5b638af12f24ba049106ff3a

Observation 352a4bf6-f68e-4646-98a1-65937f420473 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.842117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.842117Z digest=sha256:9d1a580681d6d0bc5b909dc533f06c00ccb8948f38a71218819b8bbda0551e2f

Observation b3e9de94-c814-45eb-be2e-14625b18e067 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.847376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.847376Z digest=sha256:fb3c01a4efea83723347233f0083d221923390ec110e5f13dad1c1af4fd9b79f

Observation 5bd22f1a-af28-459a-9ec7-6ad28740fc42 · outbound

This paper cites The Llama 3 Herd of Models.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.852555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.852555Z digest=sha256:4f2bdd6687f09065eed8f17a800107e667ddc117b25f0c12506d1b2c9fff5363

Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.857737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.857737Z digest=sha256:4079e1be21fe909928f61351bb0421dd8df5c4deeaa5791b8eab92b399711454

Observation dec87567-bce7-4b01-9448-54d41b907068 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.862820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.862820Z digest=sha256:5b12572b890978d41923089efab91dd062e4aac3b9b4f3ea2f8b55d12385848e

Observation 2d192574-f02b-4608-8b1f-b3895ab9bcc1 · outbound

This paper cites Mistral 7B.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.868431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.868431Z digest=sha256:60382df1a521e81a5cab31bd1de9b7fa8964a95de4fee105f55f9f9a219c6387

Observation dac99d7f-cb4a-4bca-8a6d-92b8b0ae33c7 · outbound

This paper cites QCQA: Quality and Capacity-aware grouped Query Attention.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering QCQA: Quality and Capacity-aware grouped Query Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.873091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.873091Z digest=sha256:3e32c7adcd6d6679bd1fdd51afae39e4d2061a39a0a6956ce7a2cf203c6b4963

Observation d693ca11-8206-4880-b0da-41d86b5d3cf7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.879572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.879572Z digest=sha256:6848cd7b726cd087ba11fa1708b691e48bcacf4c28b02e49cc421cff642ebbb1

Observation e40f040f-1e6d-45c2-ab28-b6db27918ed3 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.757623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.884072Z digest=sha256:7e09d118410ae573cfa2e93f8fcdf9ea05ac365bc220f8760f8dd91d62801dc3

Observation 3b2606d6-5df9-419f-af99-9e6c5e097be2 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.733674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.889818Z digest=sha256:d4d4fff733c15db79b0ea75c62b740536f2084a9752f8c3a8f5f3c769c97a7c4

Observation 74362e38-9bcc-4d39-8626-bfe1b68a55ea · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.715024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.896238Z digest=sha256:9b22ab61433e96147dfa3b90e905e6b98af91337e50052b238e2a94f32971cee

Observation 5f34c911-4420-46cd-a5ae-21abb5158006 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Fast Transformer Decoding: One Write-Head is All You Need

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.901527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.901527Z digest=sha256:06f65b929785b2ad07554fc7efcaf88863382b9907c791088181c4117a7fae93

Observation c9696f4f-1dba-44c6-8b5a-95c3d630a868 · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.907446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.907446Z digest=sha256:35c7953af3d803462451342c817de1ae515c2c118b833bcd3acf2f9d6b6cb773

Observation 59a20d92-ce06-48fe-866a-4610cd1b7b11 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.914455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.914455Z digest=sha256:4a07ced9d2b57ccedb945b6706d3347463da553942a68893e88499906e66425c

Observation d6cb2726-e83b-48eb-b906-833c6a995ec5 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.920790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.920790Z digest=sha256:8c033442e63cf46d2be70d95fa8a4c171dfd1e31212da1a029a99490854085c7

Observation 1396b045-3cf4-47a9-a74b-d7326c79c876 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.926851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.926851Z digest=sha256:2628f1ce42864931037ab006d4c370acdc8e6c27500b5e6ed2f4fb793fe3a14c

Observation a773ade7-6a5e-4d4b-bb47-81844a9ece5c · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.695637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.933063Z digest=sha256:b276fb10d85525e828d616bd5d617489bb3fb73a0e876f2dd318b5c2715f6872

Observation 1a3c3d1d-0243-40c2-ab7c-11379f14c19f · outbound

This paper cites Attention Is All You Need.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Attention Is All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.938312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.938312Z digest=sha256:6aef50d8416c495efd3453454a90f9592b710f01e33e6fa1666b90fb60b08e0f

Observation bd5e42cf-18cb-4fea-9e28-bde12a9c36a8 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.943240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.943240Z digest=sha256:d8515186f93188f1f29b9e2b2a9d0794d79e60e0334ffd48e924c1f6dd05fe0d

Observation 0f93231f-9814-4dc5-85c4-ecb43a2b7e43 · outbound

This paper cites Effectively Compress KV Heads for LLM.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Effectively Compress KV Heads for LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.947823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.947823Z digest=sha256:545cbd5780486bfa3d2d9206c9e24d148fd993c13b208ec29d9307ba1ba72777

Observation 2c7daa70-4d90-4fdb-8a50-23d4b97d8245 · outbound

This paper cites Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.952447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.952447Z digest=sha256:68734dc7a7bb7c5eaa27c8d2601ecbee70478a5956d22502efc962f28280c389

Observation ee8404e7-937d-4390-a1da-bb201b4e0a22 · outbound

This paper cites an unresolved cited work.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:45:10.649641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:45:09.957404Z digest=sha256:1ee7d422390b46a0bfad689b245bb367e933e6fa29ecdf66af6ac196966c59dd

Observation 8a744468-b86b-4f68-96bc-c8900601386b · outbound

This paper cites online" 'onlinestring :=.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering online" 'onlinestring :=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.964123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.964123Z digest=sha256:e23c54b675235c2bc05175c3208d715e4a3cf12f59a309273dc679986d198fb3

Observation 1c9ebf99-05a5-4efc-9c46-b326a0197ddf · outbound

This paper cites write newline.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.969345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.969345Z digest=sha256:b8cda7ea31b3c6d199d6109310bcfad26c3decd10d0fdff7b5eec69400eae5bf

Pith citing papers

Observation 2d0a32dc-d528-42fe-83d6-2749609e6e03 · inbound

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching cites this paper.

S$^4$R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:47:32.072537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T00:47:31.653881Z digest=sha256:8023fbf3e018b0d93aeb09f626b82ebf772b1c862f0b8a0aade7add2c4c497fe