Pith. sign in

Paper Citation Record · LEDGER

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2607.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.01237 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T17:54:40.241144Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80d8136a-b205-4969-8543-214a27830729 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:4a3beac3196e296b5213da3938b9edebd894aa827684d2c729151a7f871751dd

Observation 8fec53c5-7593-415f-8f71-ab3edcf77ef2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Chain-of-thought prompting elicits reasoning in large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:330d6cc2a83e35fcf3429f8eb51bfc297ca760dfb38c47569fb987623bfe2b5f

Observation 6e18fe63-e85f-441f-9519-2d6d0825cebf · outbound

This paper cites Qwen3 Technical Report.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Qwen3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:8e2be829583295acc767e0ca7bb580d097e0914adc1b5c92288199d295fea12c

Observation 0b339e05-75aa-4724-ab84-c3047cb81653 · outbound

This paper cites CRANE: Reasoning with constrained LLM generation.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression CRANE: Reasoning with constrained LLM generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:aad72f6749fbd9fc2a2ae718d450d22a29d919239c0209a3b826e6700fe49630

Observation 3d897440-0980-48d5-b75b-aac80246a55d · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:7079f94b00c6f8cdd09430bae21551d424ccbcab464c0af2ed4221f2a771f923

Observation 335dbaa4-29f2-468a-a406-8b2e36a447af · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:96f546b5de85e5a91b50221be8d527682282137b820a1dfdef4a7664b6b6d9a1

Observation d93f3c0a-8555-47ff-9b4b-af02c91d07da · outbound

This paper cites R-KV: Redundancy-aware KV cache compression for reasoning models.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression R-KV: Redundancy-aware KV cache compression for reasoning models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:356b5260a2f36a6a22024e45448260561a562e99b9820fd4fa33b5852fb2ab65

Observation d2a3af9d-b160-416e-868d-1f48e120884f · outbound

This paper cites Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:b5572acccd92a9a61ffdb970d8f750fa40f680788210163e3da88e7c049ef859

Observation ed5bee5c-3c90-41b8-8291-083765bb628b · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression SnapKV: LLM knows what you are looking for before generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:19199ff8ffd3bc42d029452edc29bca2fe1812f503b60615a8be3ec4b043137f

Observation 63c4c100-f793-4e9f-af04-8c88ccf4e9cc · outbound

This paper cites KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:0373a9df657f4fe8604eafc49a4c9d012bc129b979c80d924c0b6937d1b73223

Observation 3345f60f-9dc5-4f5d-b6ab-8de28f5de93c · outbound

This paper cites Efficient streaming language models with attention sinks.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient streaming language models with attention sinks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:dc9f19c2ee9c2a444bb39d064c7b5cb2f6ca00be4d38cc366688220497893bb9

Observation 1e478993-cd24-4aeb-8541-ed3bd8c5c5a2 · outbound

This paper cites Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:de980b35bed6685251d6ac7f6a0adbd168b3b3f99eeb47851d769191c98f0b3e

Observation 5116eaeb-bf7c-4fb6-a92e-60434d4dc6db · outbound

This paper cites Inference-time hyper-scaling with KV cache compression.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Inference-time hyper-scaling with KV cache compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:f78a43369f5f17f62bbedb564e5e650325f9fbcbb3f4e8a3b9ce284c9bff4a87

Observation 15b06e5f-d426-4701-94ad-eeb15ebfaef3 · outbound

This paper cites Lee, Sangdoo Yun, and Hyun Oh Song.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Lee, Sangdoo Yun, and Hyun Oh Song

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:73faf399fce85a255119e8fa67099b5983d77b6da3882bb9c3403399d353a40c

Observation 82ceb0a5-06db-4e0d-ac3e-37b2dc3541f5 · outbound

This paper cites Expected attention: Kv cache compres- sion by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636, 2025.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Expected attention: Kv cache compres- sion by estimating attention from future queries distribution.arXiv preprint arXiv:2510.00636, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:171b43792317bb5f9a6342b0979b4a53679927d06450b9e33643b103165ccfb0

Observation dbbab5dc-b0d4-4e5d-b221-7b9f76b6d62d · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:88a9d2945f87105ddcc4de479112615455ef18786fa7de90ebb9001de50de856

Observation 0bfc3dce-0fec-4c44-8810-353f7a78a7c8 · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Criticbench: Benchmarking llms for critique-correct reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:248409c8150659cd5efe65309fa00c7422cdf20cf18f1e1ba016f7400014813a

Observation 5d8bc600-dc96-494d-a2ee-bfcfe26830aa · outbound

This paper cites The Llama 3 Herd of Models.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:850c41999e0bda36873437cb1bfd63971da9c564b1e41d6b1fa497d661d49e10

Observation 5fb4f415-7e74-4c7d-a47f-539e634a7b93 · outbound

This paper cites ChunkKV: Semantic-preserving KV cache compression for efficient long-context LLM inference.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression ChunkKV: Semantic-preserving KV cache compression for efficient long-context LLM inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:a21f595f95769f001f25e699d62c61fdbf135a35ac1e371e5f6a72c05ea1b039

Observation e30c151d-eb63-4b81-ac4c-b5860959317a · outbound

This paper cites Efficient many- shot in-context learning with dynamic block-sparse attention.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient many- shot in-context learning with dynamic block-sparse attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:fdcc89eb80af8d91e18f339cc7ccc7c5aa2ca095adaf48cc354c9a88be29e49a

Observation c9e6893b-1cfb-4058-aa25-7a2a7c835586 · outbound

This paper cites ThinKV: Thought-adaptive KV cache compression for efficient reasoning models.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression ThinKV: Thought-adaptive KV cache compression for efficient reasoning models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:ed5a0232579828a09150b074d798ce31da28aa9bb7e66390197c7c5294c77ac7

Observation 750fa320-e1a8-478d-8bcc-4bb9ce9bc88a · outbound

This paper cites QuoKA: Query-oriented KV selection for efficient LLM prefill.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression QuoKA: Query-oriented KV selection for efficient LLM prefill

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:841943fde044966baadcfdf4f18c06e969da78d430f364b18c13c67834d17864

Observation 79c718dc-7f54-45a4-9361-af01cab2e35c · outbound

This paper cites Icecache: Memory-efficient KV-cache management for long-sequence LLMs.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Icecache: Memory-efficient KV-cache management for long-sequence LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:4b262a7357668ea55fa34a8aee925752f06df9ea8d6172daf5f33980d589a610

Observation b5a24779-c8b7-4b78-82f7-9a24bd5cb16e · outbound

This paper cites Cache what lasts: Token retention for memory-bounded KV cache in LLMs.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Cache what lasts: Token retention for memory-bounded KV cache in LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:b3d7ffe11392a67e34d5eb01e481745df3ec9d0b5e3a15e628f6edbcc447d9a3

Observation 999036f9-037c-43e6-8e96-5112f799346d · outbound

This paper cites Keydiff: Key similarity-based KV cache eviction for long-context LLM inference in resource-constrained environments.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Keydiff: Key similarity-based KV cache eviction for long-context LLM inference in resource-constrained environments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:664e8007926566056e205f66da2b36fd5267c54af391b5eb1e1b95e950ee28cb

Observation 3e383c10-4093-4d2f-8011-0cbb4b6de5a3 · outbound

This paper cites Spargeattention: Accurate and training-free sparse attention accelerating any model inference.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Spargeattention: Accurate and training-free sparse attention accelerating any model inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:23e39a386f6ff7ed7c4ec164f1965512c8438950a6cbc6eace04c829fd0d71ee

Observation 677ab63c-0793-4954-88c1-c2af34493308 · outbound

This paper cites MoBA: Mixture of block attention for long-context LLMs.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression MoBA: Mixture of block attention for long-context LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:1a451ef2bdbb0eeb763c638fc680308fce9b5589b3c605337a1b3cb73e43ccef

Observation c0eeae7f-99e4-49bf-9d91-04e44b6701e9 · outbound

This paper cites Twilight: Adaptive attention sparsity with hierarchical top-$p$ pruning.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Twilight: Adaptive attention sparsity with hierarchical top-$p$ pruning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:fb91650c20ba72822ef9049f00663ef62a0cc7fa17055d62b6081d8b3f9a5bea

Observation 118e17f8-1133-4b77-82f2-038bbc8cab5a · outbound

This paper cites Xattention: Block sparse attention with antidiagonal scoring.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Xattention: Block sparse attention with antidiagonal scoring

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:2b1936ae7ab2e57199dabb47bd172b9d5b957ef0272e248affccb6aa9ba0119f

Observation ee3d6372-d41e-40d7-9753-1b6458120aa0 · outbound

This paper cites Efficient attention mechanisms for large language models: A survey.arXiv preprint arXiv:2507.19595, 2025.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Efficient attention mechanisms for large language models: A survey.arXiv preprint arXiv:2507.19595, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:b94f276af322532cdc25060e19fb70a7e0817ad83f056c296306c4b5f59e13bc

Observation 76cad784-ca42-43f8-9b26-23cd6d04ee87 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:54a417389b3a4bd8e4997cb20a1c8866020c17a26e83d5cd0ae2dfeb169e4495

Observation 028e36d6-ed91-4bc3-bfb8-bbc1187b1239 · outbound

This paper cites Kwai Summary Attention Technical Report.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Kwai Summary Attention Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:9bd0619a0182e4f0153750187355115b86fbef44b043f86c641507cf5cfb51eb

Observation b312c64b-aaae-4a59-baf1-cb57768ac2b7 · outbound

This paper cites Where does in-context learning \\ happen in large language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Where does in-context learning \\ happen in large language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:50657d281b970bb37ef3a0542f1338b9719fe6348033918a53ef5aa8b1a56af6

Observation 10ac2138-1f68-4805-a52e-61ab2333a2f0 · outbound

This paper cites GPT-4 Technical Report.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:53420c4fb275af7a953e9458ee62dc838962c266409aa60fe13fbe06de712047

Observation d48e8749-ea35-4244-969d-1251a3e85bb6 · outbound

This paper cites A survey on large language model acceleration based on KV cache management.Transactions on Machine Learning Research, 2025.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression A survey on large language model acceleration based on KV cache management.Transactions on Machine Learning Research, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:3df9dc74037f35ecb0daee2d482ee863300a6c1b20a7473532341ffe260db89c

Observation a7cc8594-c288-4b44-9823-d51196e2b317 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Gonzalez, Clark Barrett, and Ying Sheng

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:dff2e239179a1191504c9fc2bbe565dbb574020cd09393a78c67b17d27374e5b

Observation dd657f61-0eb9-4b77-b823-fa4fdc01e8bc · outbound

This paper cites an unresolved cited work.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:d11c4e50dc33adfb80b749a7d2997e4beef58f9e82da6888e753c83b324a3761

Observation d94c1bf0-fc2c-4c66-9078-7f105845093d · outbound

This paper cites CAKE: Cascading and adaptive KV cache eviction with layer preferences.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression CAKE: Cascading and adaptive KV cache eviction with layer preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:3d0922159be8db0c4528d6df2e1db73cf938574e9cbcdb361a7985e9dea809ff

Observation 4c5d13d9-9675-4c0a-b820-c72e1fc49d62 · outbound

This paper cites American invitational mathematics examination (aime) 2024, 2024.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression American invitational mathematics examination (aime) 2024, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:2cc614a2a70b817799bd9abda0d62271ccfaa8f151193a48783bc50e80bf09d1

Observation 1a37132e-3a93-4810-8d07-c4668696b1d5 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:895434fc126fdf39c6b14eee87acfefee6aeec80fdace33ff820f44edebb85fb

Observation 2edd8f80-9b6d-4e67-86f7-278ed53eb693 · outbound

This paper cites Data engineering for scaling language models to 128k context.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression Data engineering for scaling language models to 128k context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:6f56b32c2eda465996868189877b535aa03c0acb7c9e410d7efa3d29b6d07775

Pith citing papers

No inbound Pith citation observations are available.