Pith. sign in

Paper Citation Record · LEDGER

Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2407.18003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.18003 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:49:24.114689Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5219d9de-e731-44a0-90d5-02d903ff308e · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.467156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:9629f1641379f77f7df71fb428eb919d74b60f175cb60bc56c8d880f387835e9

Observation d5b3f2d9-d1dd-4bdf-a1ed-cc660c9f4577 · inbound

PoM: Efficient Image and Video Generation with the Polynomial Mixer cites this paper.

PoM: Efficient Image and Video Generation with the Polynomial Mixer Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:47.127086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:47.127086Z digest=sha256:1ec9c0ebd81a0824b6b04ee9b39c2f650326ec40902b2f5ba94ac4ea097ab4ee

Observation 85aaef17-8091-466c-b1e5-1ef66ac6ee03 · inbound

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification cites this paper.

Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:00.325854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:00.325854Z digest=sha256:30ecd45de4915fc93802078277472a60085d3fc789e34b079b787cc47daefec3

Observation b8c53615-97bd-437a-8068-f3f255f0c4e4 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.742080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.742080Z digest=sha256:60944b8029bdd6d11e36563c2ac145a68ac1f4a4a1b9bd5f65d3641875a40d8e

Observation b92d5962-84f4-4bbe-a72c-02ea9c1698b1 · inbound

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding cites this paper.

Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:16.660204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:16.660204Z digest=sha256:37a402cac76f57f07be919e302a23e8829ea83a5759fa5887188e8484d2356b3

Observation d9d6c010-c1f2-4d1a-ad8e-4885cbdcba68 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.420768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:62e6eddb51b008260d7422f20e9ad89ba490f969ad292214f49f439676d69073

Observation a6650d7d-366e-4a5c-96bd-68c48e7e7dc7 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.798694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:d7d4f88aca1b1fb7862f1bf2dad4dc5ba76683da9b74edd946fa5cad28aea075

Observation a27890fd-b34e-4647-ad95-b4702abfd465 · inbound

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models cites this paper.

Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:15.745107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:15.745107Z digest=sha256:722251adf7fd4d304c658c259b28c2186e92add1172fde924a18431ea17054a3

Observation 4897a087-dd91-41dd-a307-597c18fc83a6 · inbound

Semantic Scheduling for LLM Inference cites this paper.

Semantic Scheduling for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T01:09:21.238031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:09:21.238031Z digest=sha256:5426bc6d0ad8f769b19f3f4c22ec8635f80878e39baf0c26d652fe1bf19dcee8

Observation 788f684c-688f-4c76-a0d2-c0d1d8010cba · inbound

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference cites this paper.

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:20:42.997865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:20:42.997865Z digest=sha256:79447d69614632b9f5d5ba25d0ef23cede6a033a9e76ca573c8fa472e2626b7a

Observation 93a778ed-c57c-4e7a-ad1c-0000ebe9a475 · inbound

CommVQ: Commutative Vector Quantization for KV Cache Compression cites this paper.

CommVQ: Commutative Vector Quantization for KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:49:24.114689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:49:24.114689Z digest=sha256:f3d887a65f9e9f0dfbd8e46190e1970ff258caeb1fd50c33055f36ccf55cfbb9

Observation 461a5b00-b6bf-454a-900d-d65b113bee91 · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.768568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.768568Z digest=sha256:8d40fe1335d8b1acd434eeba655e4b75e97948f8705580717cbd26f68b05860b

Observation 34f513a7-45ee-4117-9607-8b2ff82b929e · inbound

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers cites this paper.

SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:01.272177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:10:01.272177Z digest=sha256:80ee991f4ef1d55c7c53672a41e061e9e0ceaf8363ad5186b8a304922111fd2d

Observation f37ec339-0b23-4a03-bf5b-9a5ccd011380 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.255004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.255004Z digest=sha256:25814d48c138a454c8b792bf9b191c1d08102435a787ac73ef2edd483611a5f1

Observation 2932188b-feb0-4505-b425-00706e65caa3 · inbound

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression cites this paper.

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:37.229366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:37.229366Z digest=sha256:0a6495c1315e3d155e8f2d47fbcc002a1d37468cc29cf9885d810798ae0b6d2b

Observation a39092e4-3a0b-44e8-ade1-6d67846053ab · inbound

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent cites this paper.

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.886353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:38:43.886353Z digest=sha256:38bd0dafe82c094a87ab00429a917f253cc84ff5c879b07efc1e34a22e71d303

Observation 94b1b7e0-2608-4c71-b194-af0574f455d9 · inbound

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery cites this paper.

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:40.748817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:40.748817Z digest=sha256:e0755ff6f6f862cb23494348a8ab928e052c32ddb5040a3b31ec62938e7c8f9e

Observation eafe9d99-b0e2-4280-91ad-76196c72b741 · inbound

A Distributed Learned Hash Table cites this paper.

A Distributed Learned Hash Table Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:59.019444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:46:59.019444Z digest=sha256:c11648d02b5f3ec93d1a5259a9cbe06af63a053c1fac310a129b5ce088e3f11f

Observation 491957c2-5a6c-460d-9d6c-9cc99aeca3ed · inbound

Adaptive KV-Cache Compression without Manually Setting Budget cites this paper.

Adaptive KV-Cache Compression without Manually Setting Budget Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:10:56.981798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:10:56.981798Z digest=sha256:ff70edb4a3c209612e85d51bb7278c87e6d03d7977889ca4bf280025c92afcb4

Observation d25a45b9-ab71-462e-8b35-d99d6b246da4 · inbound

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference cites this paper.

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:13.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:13.096313Z digest=sha256:621e686faf20c6a3254675e71d60bf3196c2e8a3724bcc5b644233927c11079b

Observation 64453fe3-4c28-4191-8b70-bc5a7136d4cb · inbound

EvolKV: Evolutionary KV Cache Compression for LLM Inference cites this paper.

EvolKV: Evolutionary KV Cache Compression for LLM Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:30.115688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:53:30.115688Z digest=sha256:f4aaa7146eb3a474f5e253f20c096bccf9d2592314125ebbbf3d6667b507f56c

Observation 2eb1b009-0208-437c-85a3-eaf280a9a3db · inbound

OjaKV: Context-Aware Online Low-Rank KV Cache Compression cites this paper.

OjaKV: Context-Aware Online Low-Rank KV Cache Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:26:24.702813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T13:26:02.980973Z digest=sha256:c1cf1962866f3fe45c30239709e7c877fdadf2e4363782920b182b3e3758cd2f

Observation b21af5fd-a71a-4e57-a297-453a0d1f8cd8 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.765255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:8b91651daa7d47a89ff9f3f9ca0ffe2292b8d5c81ccc754716fad77a08847c55

Observation ef827935-40c0-4108-87f3-cef628ce9116 · inbound

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction cites this paper.

EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:09:36.990898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T01:09:07.983785Z digest=sha256:a3bd1b71d4b28c8dfe9dc5a01d1a592f76051530282c97b8073d38abfc53c322

Observation f78f4f86-0f53-4092-a17d-fcd360ec2729 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.770772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:c99af1850365aa23814145d2f027de0340a54f2da99c53b8e6185bee88470b2f

Observation 842f3199-2da6-41b7-be1e-3e9db7153cfb · inbound

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression cites this paper.

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.203487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:55:45.739280Z digest=sha256:8bb33f79af5e18c78a6a8d4a27b774a5b7b814fa42ef3b4c8b8bd3f48f23e878

Observation 5da7ee58-3e80-436a-8159-41aac99f4f6c · inbound

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers cites this paper.

How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.191396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:17:52.344313Z digest=sha256:8c122926d76e8411da3f4de6cdf1a059feed3c8b573af43f37303205ec5eb805

Observation 8dce6e59-b5f3-4860-ba32-eb1112d25229 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T01:54:34.032488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:4d02e1b35e763248d2814f3b18534ee646d040fdfe8afd7943c8ca2603323fbe

Observation 2b58e1d1-177a-4766-91d7-f114ac55f69b · inbound

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference cites this paper.

On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:36:06.555949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T17:24:04.123827Z digest=sha256:f156486b45c1c6abd716412da97dd16a3765ab694510ec25d5e3f3967f3153ab

Observation 2d31888f-da9b-495c-a3fa-7b03b3942266 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.876717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:d077407624eb1d9d42488d54952bad70cc2b728f1ee4a578a1a4bb458f1a7122

Observation 7804a3b1-2b9e-4312-924b-2cf77aa08c07 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.969873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:a28a88ca9a92d83ed5f1950daf9030de7f8845ffa00492b321f04206deff1208

Observation 444c5df5-e6c5-478e-b03c-62bc57f60cfa · inbound

FlowNar: Scalable Streaming Narration for Long-Form Videos cites this paper.

FlowNar: Scalable Streaming Narration for Long-Form Videos Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.742615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T19:08:36.655886Z digest=sha256:40881fae87f7574730af355b04d00ee6164fb31b76633dd0ecfde748bae7e717

Observation 95bdb966-0c58-4681-837b-d792a80376a8 · inbound

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators cites this paper.

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T06:16:09.070414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:16:09.070414Z digest=sha256:82e684fe7b7857a97f95a7c2f380313ac2e5a985b9a39223179e39c53c60b0e5

Observation 6c6b5074-bdbc-4779-8613-04ebeaf226a6 · inbound

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents cites this paper.

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:51.051073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:50:51.051073Z digest=sha256:8d2511967af8feed5f6f32f1abafb7293643e1c70f138a92eaaf940bf9850e27

Observation 7f9d27ab-a6ec-4cb4-b4e3-a59308de6cc1 · inbound

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models cites this paper.

KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:52:00.794462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:52:00.794462Z digest=sha256:423fbf14296916c0d4b7882d9ac2c0644feb0cbccdeaa5455bc4d65eee8bbcba