Pith. sign in

Paper Citation Record · LEDGER

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity

As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2412.02252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02252 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:46:02.187649Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e527d8e1-16c0-4cf5-8462-f8d1a65b1883 · outbound

This paper cites GPT-4 Technical Report.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.946174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.946174Z digest=sha256:ca935eb4d989a08c43769f8fdabde5bac371455094ac51e0662d015c85fc3cd6

Observation 1dffe21b-8859-475d-94c4-5c36025c6c9a · outbound

This paper cites GQA : Training generalized multi-query transformer models from multi-head checkpoints.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity GQA : Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.951909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.951909Z digest=sha256:bc4fab496f01752f55789170e6e41595afee127ce3fe9a5ececc73a489a02bde

Observation af388563-0fd1-4914-b83f-c64befbb3d96 · outbound

This paper cites L -eval: Instituting standardized evaluation for long context language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L -eval: Instituting standardized evaluation for long context language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.957532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.957532Z digest=sha256:be06a91321521eba9bdb742083966c2a049ae7e887755035df9d9e92abdf5259

Observation 94bc9141-ae7c-4e15-8733-7cca1e4f12b0 · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.962711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.962711Z digest=sha256:4aa0633a42313316235f76830032d57139ebd16f84c5c9c3126a02dc5bd7cc41

Observation a6855b94-fefa-4307-8602-bd8b7a9b0790 · outbound

This paper cites Codeplan: Repository-level coding using llms and planning.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Codeplan: Repository-level coding using llms and planning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.969363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:01.967701Z digest=sha256:4c8a128fba504a312f1f54b9329238e3e9ae07857f9d3bc84e6f40508e6aab54

Observation b3fd7210-7bd1-4343-a18b-21d1373c918e · outbound

This paper cites Longformer: The Long-Document Transformer.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.972899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.972899Z digest=sha256:7400c8d205bd79370ceb24a170ac64d6c253329ce015e43d97903c93087b3378

Observation 8604c5cb-b478-4d7a-9a74-16ea1eda6e5c · outbound

This paper cites Leveraging redundancy in attention with Reuse Transformers.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Leveraging redundancy in attention with Reuse Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.978654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.978654Z digest=sha256:cff9a116069038cd8ad39d048b7d4e80be347188536da56cfcef11299ccc2141

Observation 87ff2fb3-099b-47c4-8c1a-56e45b31dc06 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.984234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.984234Z digest=sha256:5e0561a23ace693dc751f7baa85ef7006802232515268142f692e30a3bf536b2

Observation 98db3a0e-e6b3-4768-b801-d984748f683a · outbound

This paper cites an unresolved cited work.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.989136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.989136Z digest=sha256:05509084e28dca7ecd441216d8426a5d9e659a5f0779a744c8418bda53a1b054

Observation 648e1687-64f3-42c1-9dd0-29fe3dfb6c8e · outbound

This paper cites The Llama 3 Herd of Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.993734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.993734Z digest=sha256:d5ac2bd377e0eafaaecae34974873b1e02f090c864a8ad33bccb46c2b9b7f5b6

Observation d6cf978e-0eb9-44c4-9592-8fb91282da87 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:01.998574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:01.998574Z digest=sha256:bbee13fafa08835f7c495247ebee8fd13b91b51ad663d395025bb9fcf82d9299

Observation e9bd01b4-c9c0-4188-846e-c488d9910d89 · outbound

This paper cites Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Challenges in Deploying Long-Context Transformers: A Theoretical Peak Performance Analysis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.003540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.003540Z digest=sha256:6e9a0f4c9e3af5cdf1794cb471fbe0168185c1848dce4b508e39373d8f3999ce

Observation 97576a13-ca55-44f5-8fa1-338fa89e2dd6 · outbound

This paper cites How to train long-context language models (effectively), 2024.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity How to train long-context language models (effectively), 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.008707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.008707Z digest=sha256:08a8afe2e819ae60ceb027276b2d10487a5a76caba5becf677dc9026a0b59f26

Observation 6d91c8be-ec2f-49e7-a0dc-09b627cdd68e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.014246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.014246Z digest=sha256:b81530a86c15f0894413a3f614e1690614d0248c639068b039c6989d5c01522f

Observation c8416f7e-c70d-404e-8e23-107bbe694135 · outbound

This paper cites LM -infinite: Zero-shot extreme length generalization for large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LM -infinite: Zero-shot extreme length generalization for large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.019172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.019172Z digest=sha256:060bcc37648ff0be577590309476b2b30669c5b05423acba9ba582f81ccf1118

Observation 9ba4f3af-5ff8-430e-ba6d-531a423b4017 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.023498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.023498Z digest=sha256:b05a500b01c42dae7665129ec818e0a64249cad2b3e9c4db5b04cadf43090f15

Observation 726bed92-0016-47fd-bf64-5646b9ccce34 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.028093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.028093Z digest=sha256:1134c901f47dbbc898b4e8a7d20467fce87cadf5aaed4318717a94b958aeafa1

Observation 8abda245-0878-4b14-a047-106a40160817 · outbound

This paper cites L ong LLML ingua: Accelerating and enhancing LLM s in long context scenarios via prompt compression.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity L ong LLML ingua: Accelerating and enhancing LLM s in long context scenarios via prompt compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.032467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.032467Z digest=sha256:7de0b141063ec1d154044dfb30815470e2f80cf107b0b7ca00e74aaaa4cd42b0

Observation 2c712d5a-b9bc-4d6b-910d-28e376a5e614 · outbound

This paper cites Compressing context to enhance inference efficiency of large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Compressing context to enhance inference efficiency of large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.037136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.037136Z digest=sha256:02fa7165dcf733063aa622a04aed881a98bcb26483eecb3db5266f4632650fb0

Observation d53162a3-2441-42cc-b969-463bd21dfebe · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity SnapKV: LLM Knows What You are Looking for Before Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.041282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.041282Z digest=sha256:47fd9fbc9b5f23ceb61d06323a25a6754b6f73f15d102292531037e0d752b7c8

Observation d3c1ed64-65ff-48c9-b87f-ccf31b814ee6 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.945100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.046056Z digest=sha256:5cf0d834f2a697454b524e9cc380af56b9406d2023c0f8b9421bf89203dbd45e

Observation f33968a4-7f52-4a65-8c53-7783935bb369 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.050262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.050262Z digest=sha256:9ae7d354298b970a223b8d270c48f0375a336a42a7bfe3d3357e14768c7a98c7

Observation ab5fa997-8eea-4f97-b365-6962333cce6e · outbound

This paper cites MiniCache: KV Cache Compression in Depth Dimension for Large Language Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MiniCache: KV Cache Compression in Depth Dimension for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.056231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.056231Z digest=sha256:017fe122dfcda4449e02c27d4026392540bd593878c6553ddc102e68a4f25304

Observation 93b7599b-0ce8-460a-9224-ba235a9dcbb1 · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.061116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.061116Z digest=sha256:6bf3e44c3e6d36c8dd6d8f4c7a10fbd1bd71df83af9969b1167eba9957841150

Observation b59d75df-2bc9-45be-ba7a-312c438fabb4 · outbound

This paper cites QLLM : Accurate and efficient low-bitwidth quantization for large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QLLM : Accurate and efficient low-bitwidth quantization for large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.928658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.066259Z digest=sha256:bc4fc4f0808c880f6bbd115553cb0d6cab7a9e4710ee9049127d4acaf5b31370

Observation c89effdd-193b-4566-b17e-66c90c46a7be · outbound

This paper cites and Liu, B.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Liu, B

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.914397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.071944Z digest=sha256:f8d1dec0220aec19a49d70d8215c58f08a19f083adac80d4ba6f14783a0f3607

Observation 0b4de720-5d2c-4470-9e81-52b027644318 · outbound

This paper cites The jensen-shannon divergence.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity The jensen-shannon divergence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.077749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.077749Z digest=sha256:cc8cb6d399b3ea1d35f9999b6f6f78eb2e73c1740d3bb398bf8e0c0226303176

Observation 3619e8c9-1702-46f2-b142-ebb2aedf3158 · outbound

This paper cites V., Qiu, L., and Zhang, D.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity V., Qiu, L., and Zhang, D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.082175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.082175Z digest=sha256:d8531a05a72dffc8ab438784f00b0e2dfdeb4e10dbe7b1cb6ca3be76cfd446ae

Observation 81d264c3-548c-4428-b888-beaf6af0b3d0 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Pytorch: An imperative style, high-performance deep learning library

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.086882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.086882Z digest=sha256:d088799900dd0212fc656f3cfd73ed5ac54bed89ea23f46edd868bdd0708d485

Observation b050983d-e7b2-410d-9863-633e7927f744 · outbound

This paper cites Efficiently scaling transformer inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficiently scaling transformer inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.091264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.091264Z digest=sha256:eedf23118c570d64277189e1257bbbde2aa39fe1d63211e175d3405be198bb90

Observation 3e497bba-cdb3-4c0f-91b4-4d73d0e00615 · outbound

This paper cites Zero: memory optimizations toward training trillion parameter models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Zero: memory optimizations toward training trillion parameter models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.095682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.095682Z digest=sha256:aca8a59e62df45003f6400175cc43386176100bd14d03754e08e879b35fddd69

Observation 45c9968c-51dc-4ab1-86fc-214c971c4686 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.100260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.100260Z digest=sha256:575c8e837141c70b52de93a1f140d404203774353ecde76c71d9b6832089ec12

Observation 35ded138-c09d-4515-9700-f63b0c235c4a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.104917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.104917Z digest=sha256:ba72723378b3eef58b387c3f44a02da24947535f2f59802fc3f665d008d2ef4e

Observation 158d2ced-a283-4de7-8ea4-2364c600a9f8 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Fast Transformer Decoding: One Write-Head is All You Need

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.109644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.109644Z digest=sha256:4350a9deec461022ffed891d32829a51ab95978b7182d3e19a4e4d7ded5fda0d

Observation f5459b65-f3c2-4d98-bd24-5c01e8d3d8cf · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Flexgen: high-throughput generative inference of large language models with a single gpu

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.114792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.114792Z digest=sha256:1c2d1acd2e45b189c178e73626852875c9243bd7b26ccab36109091e1475a5c0

Observation d672a08f-02e0-405a-9bc9-9f60643d96df · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.119255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.119255Z digest=sha256:fe262dadadbfc21a5c98991da890c2d5fc2d6292e8d951b5f4b5d3cce88cf922

Observation ecaf3672-7514-4d70-8c26-d46b392151e2 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.123980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.123980Z digest=sha256:81d21e6724dd14303af1afa4811075e11a0be1fa185a061215c3fe3a43e30e1b

Observation 39cb8619-c1c9-4898-9287-a94f10a8ec0a · outbound

This paper cites QUEST : Query-aware sparsity for efficient long-context LLM inference.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity QUEST : Query-aware sparsity for efficient long-context LLM inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.852007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.129198Z digest=sha256:c2162aa896e8758f7087ea7ed61cc0896851f7e9b529293db8ca400d60a1e5e3

Observation 8bd4e694-b592-4ecf-84b5-5cef6d10d38e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.133622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.133622Z digest=sha256:5541e6efb125ed11c1de1e24545eca48a0972f4f7e309f209c8c6c42a926e3a4

Observation 9abecf7b-6a21-4e70-9c1e-f0a740fd5e5f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.139083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.139083Z digest=sha256:34177bc6e90a4c3c22de05882865b0f767d00553f10a94335763c3fb012a1817

Observation eb66d2cf-2166-4468-b377-69189d136794 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.143383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.143383Z digest=sha256:d70535842b94284c589b41739fa55d6361abbdb02c6e6f9ed3691db20fcf9ae0

Observation 61f7bfbc-d96c-4484-86fb-599962c76105 · outbound

This paper cites N., Kaiser, L.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity N., Kaiser, L

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.147598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.147598Z digest=sha256:1c80ea42d17c460b117f3b63d07d9b254a34e76b283ce356d5239b204fa0914d

Observation 6599008c-d467-400c-8a03-864bac047884 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Transformers: State-of-the-art natural language processing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.151899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.151899Z digest=sha256:3370f0a77bb4f432971fe8e38fcdee37a66aba2b91d486a9dddaf8c0d11f13fd

Observation 637c17ae-8609-4060-bd5e-ee9e6e7a49c8 · outbound

This paper cites and Tu, K.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity and Tu, K

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.156015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.156015Z digest=sha256:b0df54998e431f85eee72b057683dd460715e0eaa0ad25eea5d46909daa1fe6a

Observation 066b066f-a3f8-4d08-913b-dcc30296097c · outbound

This paper cites Efficient streaming language models with attention sinks.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Efficient streaming language models with attention sinks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.160318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.160318Z digest=sha256:716bf3bf39e333c31798166283038e014b0c14a663b7f81efcc753c033300565

Observation bee1dc97-650a-4ec2-b3ee-320486b3b22d · outbound

This paper cites Sharing attention weights for fast transformer.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity Sharing attention weights for fast transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.164871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.164871Z digest=sha256:091cc3734f37044e5ef034426d56446760bcf30b3f2b5ea58db4aeddcb8591e5

Observation 7d73e4f0-05ec-4fff-8adc-4ef91e667dfa · outbound

This paper cites A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.169436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.169436Z digest=sha256:3328ca2bb70505a09f73a469e2523e1d359116c81bbb2a74464baa309478a14c

Observation 3547b98f-4cdc-4e26-9900-7e289f56daf2 · outbound

This paper cites B ench: Extending long context evaluation beyond 100 K tokens.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity B ench: Extending long context evaluation beyond 100 K tokens

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.174036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.174036Z digest=sha256:4333401e718873eb3766a7181072602bd364bafabdc37fa605b8dd697b171ed0

Observation 3dd0b3a7-e200-4e85-8a80-b2346180c1a2 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.178568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.178568Z digest=sha256:7805f7202421bdde5d26f7bac95eba88137a06c711beb89837678c2f77714c1a

Observation 72a49ccc-cccd-4788-a785-97374f2dc53e · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:46:02.798121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-11T23:46:02.183128Z digest=sha256:d433cee19c2c5f644ec2be17c262533f23f10e5f3e5f0a9a400da3d2a52eab18

Observation 1359955e-27e9-4492-9b9e-c2e319126b9c · outbound

This paper cites write newline.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.187649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.187649Z digest=sha256:295d9c066ca6e267f7e83acea7a225ffbcd1a8a33d6dcb854912cebd82fe845f

Pith citing papers

No inbound Pith citation observations are available.