Pith. sign in

Paper Citation Record · LEDGER

CaliDrop: KV Cache Compression with Calibration

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.19906.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19906 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:15.844310Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7202c61d-e046-47d4-a4a6-dd9616e5842a · outbound

This paper cites GPT-4 Technical Report.

CaliDrop: KV Cache Compression with Calibration GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.646922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.646922Z digest=sha256:afce986d1e9211ecd628da413ee8bd10906283d23814baf1c237a0bd21577f93

Observation 2676b63f-3f03-4c4b-8fed-6fa6fcb6f796 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

CaliDrop: KV Cache Compression with Calibration Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:17.445507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.651831Z digest=sha256:730b4893a74cf4d6716da4093690e564e0aeae29aa7b70ca5296a27907135432

Observation f6b2c56f-f772-4ace-89cc-1e45982ccf17 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

CaliDrop: KV Cache Compression with Calibration Longbench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:17.244358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.656502Z digest=sha256:fecaf8b62fee0164250accbe72a3a711d9c1edb64ff617058eca692fb4fa8854

Observation 345bb7a4-d061-4786-9478-af7175703411 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

CaliDrop: KV Cache Compression with Calibration Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.661526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.661526Z digest=sha256:57fd160929978d811e7402cbe60ebd39fedc498904fa9e3027c79393b9a01908

Observation 71be8909-196f-489c-9806-77b6bc72c888 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

CaliDrop: KV Cache Compression with Calibration PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.666030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.666030Z digest=sha256:09edeb0d98014c10fa2c1f2f9fb9569a2c25998f9ab6976b0959d0b078c353ae

Observation 72c9b27c-2c3b-4a5d-87d3-bc691b0d4124 · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

CaliDrop: KV Cache Compression with Calibration Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.670799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.670799Z digest=sha256:26fa40b96ad9adeee046112a218c5c1444bcba54bfd9ee2764342fbbf4e1c810

Observation 57454bc2-d061-44aa-ae76-7f3e17e17fe3 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

CaliDrop: KV Cache Compression with Calibration SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.675892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.675892Z digest=sha256:04508635d131ddf88ed7cb05e53886ace1b7137cfcfd2a495624891c342b5979

Observation 7806b74f-bbdc-40fc-8703-d73765ce6bdd · outbound

This paper cites A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression.

CaliDrop: KV Cache Compression with Calibration A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.680442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.680442Z digest=sha256:1d3f591aee7a56f48b3e13655648717d4333c25ea07dc8d5feda74d3ebed59ab

Observation 11d233ba-8593-4977-af05-59f648ad1ee3 · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

CaliDrop: KV Cache Compression with Calibration QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.685067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.685067Z digest=sha256:9677015f5fa30d7464e09d956354c01a35d3eb0e377aefeefe2df7e2263b8488

Observation df3fa1f2-e40d-4e1b-a806-412d58cc275e · outbound

This paper cites The Llama 3 Herd of Models.

CaliDrop: KV Cache Compression with Calibration The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.689722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.689722Z digest=sha256:5550a5484aa0c8bb40b558bb8a82d728fbbb68d120a38668ffedbe143e6c57a7

Observation 96d90404-72c7-4a92-a26e-07c29d737a99 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

CaliDrop: KV Cache Compression with Calibration Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.694151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.694151Z digest=sha256:faddf155b26de3abaf3efac861a51f7a8aabf4b8b82746472bccd2945a44225b

Observation cf847818-5763-493b-a472-5af7c4a4b078 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.

CaliDrop: KV Cache Compression with Calibration Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.698374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.698374Z digest=sha256:adce1df8b4bec657712c460ab81a0a9fa03c58606dc311b01afae740bc9a9a3b

Observation f42b0cec-832c-44e0-ba4a-1db6ace0f765 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

CaliDrop: KV Cache Compression with Calibration Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.702165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.702165Z digest=sha256:6b150db7bce67aea7661a481c58c74438061b3f37a7625cf61b7382be85ab2b9

Observation 83187b8a-959e-40c0-a79a-35d4d6a8652f · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

CaliDrop: KV Cache Compression with Calibration ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.706240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.706240Z digest=sha256:7a2c854715abd6f5f0ebd0c2b2b810b23f9b5764669fd331fa57a13fe40ba2f9

Observation ffc9ce03-1f66-45f5-a56c-51d090804ebd · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

CaliDrop: KV Cache Compression with Calibration KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.710581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.710581Z digest=sha256:624b987a01301ae1ff0770bceb3510a3781396b4919b1caafb380acc611428a1

Observation 01f02dc7-00ff-4b0d-9a53-8034531de720 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

CaliDrop: KV Cache Compression with Calibration RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.714902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.714902Z digest=sha256:9290c85cf74a88eec17a932d409f7520e8356d19fcfede2717ccd6de6fc43d84

Observation 9f8d88bd-33bb-4e47-b3c8-d8853acda7ea · outbound

This paper cites Mistral 7B.

CaliDrop: KV Cache Compression with Calibration Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.719168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.719168Z digest=sha256:9815353d50b26368dbfc8daa3a635af7e7eabe1668ca9163a0cc26256c978b6d

Observation f92c6351-36f5-467d-8658-d8e1baa35229 · outbound

This paper cites Hydragen: High-Throughput LLM Inference with Shared Prefixes.

CaliDrop: KV Cache Compression with Calibration Hydragen: High-Throughput LLM Inference with Shared Prefixes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.723671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.723671Z digest=sha256:567848eb6285f650b30f21e58d4f736599e568970b222ddf2cbe109c9c29dcba

Observation 5310eae0-5bb4-4afe-ac14-c9176c428768 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

CaliDrop: KV Cache Compression with Calibration GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.727858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.727858Z digest=sha256:da10a67e98be16b9eb37f2fdf1d52b31098f01ac4793c1b51fcea34cca7a490c

Observation 30bf4c5f-1431-4f2c-b1cd-cc5c284d9e16 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

CaliDrop: KV Cache Compression with Calibration SnapKV: LLM Knows What You are Looking for Before Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.732475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.732475Z digest=sha256:cc6eed7d319e2223498ae19df1721928b5d03c047a6e78d4aee3c48fece53e3b

Observation c3d33bfc-f572-4d21-b961-be136cd0003b · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

CaliDrop: KV Cache Compression with Calibration DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.736969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.736969Z digest=sha256:b7eb9984fbb74e86af12d9892da8f958df4d8c0e16846770c2dee8015a6c350a

Observation 5fb05e49-b401-4fdc-b3a4-2db964c44424 · outbound

This paper cites DeepSeek-V3 Technical Report.

CaliDrop: KV Cache Compression with Calibration DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.741444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.741444Z digest=sha256:b51ef2e007a1818175736a8fdd341745df05f27925c00a31ef3362f43715a43b

Observation bcdcf966-4e0f-41c1-b5b8-effde72e5654 · outbound

This paper cites Lost in the middle: How language models use long contexts.

CaliDrop: KV Cache Compression with Calibration Lost in the middle: How language models use long contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.745777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.745777Z digest=sha256:0fb64cddaa56aa0b392d308150e01894cb0089a03c93fd698fb396188a0e4205

Observation b3a2c4d2-f9f9-4352-8645-e4e7df5738a0 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

CaliDrop: KV Cache Compression with Calibration Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:17.081334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.749791Z digest=sha256:6f8f9fa46c5c45a8de701e51692a49b1d97eebd65d6276945747ee3f035afd35

Observation 365ca0fe-707b-42f2-9793-ab56062917e7 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

CaliDrop: KV Cache Compression with Calibration KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.932174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.754094Z digest=sha256:7a9b030662c365412684304fe5e844efa276e98a6006855c675772ff33342a0b

Observation 597ffd37-bc5a-4e92-844c-e9e60014ebe3 · outbound

This paper cites CAKE: Cascading and adaptive KV cache eviction with layer preferences.

CaliDrop: KV Cache Compression with Calibration CAKE: Cascading and adaptive KV cache eviction with layer preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.797279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.758394Z digest=sha256:385124f89fa6f155fe58caf486c029b70cfe352d476714e42bc0d4d086416e21

Observation ba1b3470-96d1-46fd-a945-120164147543 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

CaliDrop: KV Cache Compression with Calibration Fast Transformer Decoding: One Write-Head is All You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.762573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.762573Z digest=sha256:bb6aa3809af82932588e071063e547de3fc7282212d54b0d4bc30b8528c00541

Observation 00e1e89f-cb37-4df4-a152-1064e9263afb · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

CaliDrop: KV Cache Compression with Calibration Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.766960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.766960Z digest=sha256:af61ae3c0ab2921c51f5055b582f5777189e9e83d03004fcf4ab2fc0f6fe7a6e

Observation 13304490-bca6-4a82-90c0-81c7b3e72b6d · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

CaliDrop: KV Cache Compression with Calibration You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.771372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.771372Z digest=sha256:dbdbcfaa2bcaa786b7d786b6bd053cba508971ef964d2720326d62d33894bda1

Observation dfe709df-f236-40e4-8208-1325c1a77790 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context llm inference.

CaliDrop: KV Cache Compression with Calibration Quest: Query-aware sparsity for efficient long-context llm inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.699882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.776238Z digest=sha256:7a0802f088213972fabd63383d8e2fe573b6a9fd1e9ec027205aed58f919f6ad

Observation d9e996b0-3d34-438f-8007-6825fdca0878 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CaliDrop: KV Cache Compression with Calibration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.780545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.780545Z digest=sha256:4e4822c24968f32d478bb416406bdd1d8c24e351b6d75ae047c57d0dfb147a39

Observation b0d8488a-6186-46fc-9703-dec6a9253b90 · outbound

This paper cites With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation.

CaliDrop: KV Cache Compression with Calibration With Greater Text Comes Greater Necessity: Inference-Time Training Helps Long Text Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.784636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.784636Z digest=sha256:8e46cb431e8f1f4b04336c3a0eac6e425f4e81ae9838e3e8d4cb5809c56893ea

Observation f55d36b7-9840-4ccd-a473-1fbad5488290 · outbound

This paper cites Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks.

CaliDrop: KV Cache Compression with Calibration Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.789011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.789011Z digest=sha256:f4643c48dc2e591c6e7f2c6a2db705c20e77f4a88f6c0e8050f37fde8f4e63d6

Observation 7b6f0485-9572-4977-8526-ce7db81918e2 · outbound

This paper cites Layer-Condensed KV Cache for Efficient Inference of Large Language Models.

CaliDrop: KV Cache Compression with Calibration Layer-Condensed KV Cache for Efficient Inference of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.793887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.793887Z digest=sha256:f27e54fecd3e7049337e8d2561726baefd02922ece278c52fe297fa3e6aec927

Observation 475fbee9-8551-418f-8854-3e867a0151ed · outbound

This paper cites Efficient streaming language models with attention sinks.

CaliDrop: KV Cache Compression with Calibration Efficient streaming language models with attention sinks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.660199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.798557Z digest=sha256:6db472ba8ed0d1322d1915ca04e84345adc430fd10ce0d2246c39114daaad44d

Observation 06bfe2da-1a58-40bf-9640-c9831ed69550 · outbound

This paper cites RefreshKV: Updating Small KV Cache During Long-form Generation.

CaliDrop: KV Cache Compression with Calibration RefreshKV: Updating Small KV Cache During Long-form Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.803004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.803004Z digest=sha256:de6da023a57f10a427eb7a5d0174bb8109a98f5158a145ed7a94edcd62ec60c4

Observation 1280fc6c-c0e7-48ed-90fa-0b80692b91d2 · outbound

This paper cites PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference.

CaliDrop: KV Cache Compression with Calibration PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.807556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.807556Z digest=sha256:2e8e83729597630490e9bc5a1cab4c32d91cffee81cffa42c770a07b349862fc

Observation 94bf7996-0ad1-4918-bb21-0b53445cdb60 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

CaliDrop: KV Cache Compression with Calibration No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.812442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.812442Z digest=sha256:0dcf2d7b6066af7b67051388e58cb7370d401e6c5055bd5abde03b5fd8cf5e04

Observation 3d124d0f-0688-4e16-8aea-961b5b574e41 · outbound

This paper cites Effectively Compress KV Heads for LLM.

CaliDrop: KV Cache Compression with Calibration Effectively Compress KV Heads for LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.817194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.817194Z digest=sha256:62f87e019b281deca8237867590fb5762d52b243fd02f0063d503feda7e09896

Observation 0aef9934-a407-4c49-8162-428eb3976373 · outbound

This paper cites KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization.

CaliDrop: KV Cache Compression with Calibration KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.821815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.821815Z digest=sha256:75b175bd3d126bb89c34c15196c494ebfcbdd46b383b6eaf6390efe38f990ef0

Observation 0600c39f-b547-4039-b16f-492813fa9ea3 · outbound

This paper cites DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction.

CaliDrop: KV Cache Compression with Calibration DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.826481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.826481Z digest=sha256:335336beb547eff35de016f157c1786227adff9892f98451a2a94a03b6eda686

Observation a025ba81-5030-40a3-aec2-5800b2d8dab2 · outbound

This paper cites Cam: Cache merging for memory-efficient llms inference.

CaliDrop: KV Cache Compression with Calibration Cam: Cache merging for memory-efficient llms inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.621659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.831200Z digest=sha256:f139bf812066f5848a72c87dd4f434da841bcd60de34e3435b63afe777786576

Observation f312100d-bf84-4f50-a921-f44980329ab3 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

CaliDrop: KV Cache Compression with Calibration Barrett, Zhangyang Wang, and Beidi Chen

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:57:16.584256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:57:15.835368Z digest=sha256:240cab397c7243aac9c281ca7c44ed6146fd4e20d34a495cacbcb3159b702878

Observation 20efee29-b8eb-44cc-a975-33186125f2c7 · outbound

This paper cites DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs.

CaliDrop: KV Cache Compression with Calibration DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:15.839816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.839816Z digest=sha256:f692c63928529bbb49ba6057a85a0627f8fb233e74f2f03b252091d0494be73b

Observation 8161cfc0-3de8-4c17-bc42-c96431256f5e · outbound

This paper cites RelayAttention for Efficient Large Language Model Serving with Long System Prompts.

CaliDrop: KV Cache Compression with Calibration RelayAttention for Efficient Large Language Model Serving with Long System Prompts

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-06T13:57:15.844310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:15.844310Z digest=sha256:ce95543fbacf019e541d222c2d2e1d9d38f36f1576d86125b63eea71fc0e6127

Pith citing papers

No inbound Pith citation observations are available.