Pith. sign in

Paper Citation Record · LEDGER

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2606.24467.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.24467 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T00:11:57.763089Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T12:56:45.291959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact12
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed88dafd-6b39-4038-90b5-adccec79cdad · outbound

This paper cites GPT-4 Technical Report.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.020178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:50b5442e28a07e9c9c65adb4bb7bbeecc804162cb894850585aa1291aca87405

Observation 1da52462-0a7a-4175-a718-a484a8f40fc7 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:ee664fc5e2361c1107a8e661b5bf6cfada0cd4860e24875222dbee4a877d5899

Observation 161d8000-00ce-4020-bce7-9a781805fa75 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:0d7965339585edeff6bd929336e330baf0e818ebc5cd8709f76f87d149e7c79a

Observation 082c550d-f1f2-4e4b-9086-4876c57c6822 · outbound

This paper cites The Llama 3 Herd of Models.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference The Llama 3 Herd of Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.017771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:8be156b5d6a4db06abe9c5199627b6a5b2142d7334794110d43d93abb8414866

Observation ba327b48-32b6-4ba4-a725-b7086922418b · outbound

This paper cites Qwen2.5 Technical Report.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Qwen2.5 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.044274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:9331ad29ef0dd0ce2925b70a6d19a275e9982481988231f904b14e9dd586e1c1

Observation 840454fa-54ab-4dfd-8a6c-e695e68eabcb · outbound

This paper cites Wang, Y.-G.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Wang, Y.-G

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:64f8934f200c2e40ea962d70ec176f30372b06db7c05eb3b6368f1f0d241edbd

Observation db03baaf-a8e7-4ab6-ad70-a8f34aa04018 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:7a8a46d331fcdd63467a4f93c07f0288c696e2ae6fde7965b9e4d6544a673248

Observation 95bb6d88-1f3e-47bd-8e44-4d6174afc6c8 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.047166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:c3b8e156acaf3a201e583e0c36383a67ff61d689dce12fc41899e328e133c277

Observation 176e4426-a23b-4b4d-96b4-7f985a40e75f · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:61db344e97b48a8bcf9c12105abf74058ddbe78fcb50cfa28aecd3216aa8e77f

Observation e45f06e1-7dfe-43e0-bd94-2fe77032d648 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:b405f4dd4ec6acaa50c84e2427e9e2203955f276ebc8eec9b431b864d868def7

Observation 59c55722-3d3e-4f8b-b4c8-15fa7c2c5b6e · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:708894460fc7ef9b109b7727f5c5969c0b09d896c306c9fc1ad07bfec31c3f7b

Observation 57f7f1be-1426-4318-968a-5fbbcaeee64d · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:b3c250b73b146f164cfa3c5d3e1b6f5c01f97cd73245eb70bf130b05d4defc11

Observation a59fb7dd-2e04-460d-8b4d-4e58d22fb382 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.037140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:32f7e4f36497f823760b8ad5f00275f694c5b71d6810ea41d3c97d0f946b0c3e

Observation b0c13995-5a76-4ded-9274-2fb9a43db4c8 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:9851a1956db46f17e830d10658a839bdc7dae7a73bbff62abd567c45063add03

Observation ae98eed7-e0a9-484e-b8d9-a1fe2adcd90c · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:672050689b9cd5e1200ba499ff3c0473f451e2cff97c45bb5a4aa49144b56bc7

Observation 6a4aec25-f8a9-4260-be27-fa2d8910f6e0 · outbound

This paper cites Ainslie, J.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Ainslie, J

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:051e900cd789b802211dc50cc94bd8775cd7733d96f738df442638320274d02b

Observation 6ca51b5c-94c1-4a51-bab3-4691d41f8f00 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:c2894aa357d4cff4c696ba089e8efdf518fc02f790f5ce225d5d2c414c6d384d

Observation 7f0c9e18-9979-4692-9235-4dffb5f47504 · outbound

This paper cites Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Not all heads matter: A head-level KV cache compression method with integrated retrieval and reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.042075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:4a9417bade446d52f09aa31a3a3c1639fa195b8ea6963af22e8075efb61bb5ff

Observation cdcca5e3-b201-4475-84b2-8babd76b69a8 · outbound

This paper cites Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.039510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:a4092965a31d84adb688d23afc09a6e9c115785eda86a77047b70fbf1f65d92f

Observation 2d926237-75ef-4cc0-9cd3-eaec92c1b960 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:9c231b0ea49cd3589540f3d6db587e5e6bf5a7b272db1e149b91c91c47b41cfd

Observation 4d5d0be6-e42c-4357-9e61-444de75be77d · outbound

This paper cites Zhang, Y.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Zhang, Y

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:61202c45d1f0267dbc2a1e2502bc8fd14a43ccd67d6243473ef912e40d4a6d20

Observation 33567cdf-0ff6-4191-9970-d799ef3abda9 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:e048f88134a14dc7d93a6db1d5d1bac18fae1e6cec1e19ac0e1dc1767aafcedb

Observation 81338119-39bd-47fc-b16b-1fce231afc7f · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:19d21e77fad251db29bd53c4b772121375746f51a279d77d0dff0000e8654ec2

Observation 13e049c5-6e4f-4053-8fa5-69be06fac3b9 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:f35a4cc191eabd08701ba928fd48959f6853f37ecfa84b262a4aef9ed2201b72

Observation 24e356cb-8a28-4b84-8ee9-47151dc07cd3 · outbound

This paper cites In-context Learning and Induction Heads.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference In-context Learning and Induction Heads

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.031461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:58d5b3e5e902d415ec7906f37d1683fb920c3edcbd2468b6a21e3bcc0b7dcfcb

Observation 7c08a395-bf1a-49ed-838e-735220d663a0 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:86888e30d9eee8475f7bceb605aed497e94be3399a38aab0b83a22edd5bc6426

Observation cc7873b9-9a69-4ffb-8bf2-7ca248dcc20f · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Attention Heads of Large Language Models: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.033929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:3995b6e701b349d90db54b1a4874adf922e8aadad57c3666c1e3754b417d1d5a

Observation a08b16b2-a9a9-4867-a995-00fcfbe4a800 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:7fd0386628088a3b20cb4cd7fca636dba813ece81f7320d81fc03493cb4da06f

Observation 8729ed96-5f0a-4d9b-a027-454555b9498e · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:355b33ee440b8a7f394350467f359633d4aec1f722de10aef796b74b9cfefdfd

Observation 3443cbb6-d9b8-46e5-923a-70d22af1dae9 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:57b848c966d0f30cf02f0240ab7e28c0fbbdf22a3ddb36b0d995a492c191979d

Observation 2f9276c2-1721-4d88-bea3-3c653ebbe4fe · outbound

This paper cites Which Attention Heads Matter for In-Context Learning?.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Which Attention Heads Matter for In-Context Learning?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.030660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:b0f15f16a650e14d0db889b057cce90808eb3c26533e9328b798c048c0bd8646

Observation da331ebc-b840-425a-af55-3348867563e3 · outbound

This paper cites RazorAttention: Efficient KV Cache Compression Through Retrieval Heads.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.028076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:fa84475678347a7079e538d5931de58acd236012526608e65c5574298ad88d57

Observation a31b896d-ed4e-4428-8bb3-3714b9e5800b · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:58.026479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:4e7c901146f43838c1a2b149688e13044f6e8c4ddff746b968b35ed34e1edf4c

Observation 83395ff6-9e1a-478e-ad13-8686005f2d84 · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:6fb87fa1a954054f8f0c93bb85cb5daf3bb52fb8a1df029a4214fe8c0e0f5074

Observation 18816ac6-2b77-4c89-9f59-bcd679dc6ff6 · outbound

This paper cites Kamradt, Needleinahaystack, https://github.com/gkamradt/LLMTest_NeedleInAHaystack, 2023.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Kamradt, Needleinahaystack, https://github.com/gkamradt/LLMTest_NeedleInAHaystack, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:a9272d7934f3c8e518df57a043024888c62f287e5aa82e979e9e73da5ecd6b5b

Observation 5787d02d-c279-4a55-9d16-fb257ad763a3 · outbound

This paper cites Dao, Flashattention-2: Faster attention with better parallelism and work partitioning, in: The Twelfth International Conference on Learning Representations, 2024.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Dao, Flashattention-2: Faster attention with better parallelism and work partitioning, in: The Twelfth International Conference on Learning Representations, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:70d334f823ebdefbfa0bcd4399248fc985f7527931e27817c32b2eb0fa5b49b5

Observation 5f655f36-06c5-4b15-895a-e09ed15f658f · outbound

This paper cites Jiang, Y.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Jiang, Y

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:b6183b9c1cfca4130232796e49e2a5b9305c69d90a3440d9669d8dc55fbeea57

Observation 38333a1e-4bf2-4285-8410-eb6961e9fada · outbound

This paper cites an unresolved cited work.

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T00:11:57.763089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T00:11:57.763089Z digest=sha256:0625038800757ba99a819840ed8777efc8783fe338147c4f179d20beedb51e93

Pith citing papers

Observation 22bdc162-20b6-4d40-8af9-6cef3be9f150 · inbound

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory cites this paper.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T12:56:45.291959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:56:45.291959Z digest=sha256:f894decc268b88040da50d8d1161342719d88a786172a6a54084ac4f9e2affe0