Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:18:18.542070Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.05896.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:18:18.542070Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c055c66-937b-4714-b416-2a93fc7c64c3 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b9240e-b607-4c10-96c7-951783dffdde · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb2e7ac-7654-44df-bd4e-50c5a5bd1c4f · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Code Llama: Open Foundation Models for Code
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a7a6b8-2d7e-4dd0-85b2-208b50436ba1 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0f265e-8a7e-4b8f-9fc3-b56ab7a8531e · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference BigTranslate: Augmenting Large Language Models with Multilingual Translation Capability over 100 Languages
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4c3da6-16b0-4a75-8e3c-aefbc8fcd193 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Attention is all you need,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f9e62c-f1e5-42cd-a751-790c9f667659 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Language models are unsupervised multitask learners,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8d2ff1-9305-44d5-8d28-3e113cc26bf4 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e865ae17-693f-497c-bc1c-2f350d3c17b2 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 635d9776-ddcd-4f9a-abd7-b52a8ce58373 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc970ca-837e-4ab2-907a-dc13a6579332 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f949eb0b-9063-46f8-8d69-5c8ca568bc7b · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Cam: Cache merging for memory-efficient LLMs inference,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4403e995-db45-424f-b38b-3cc8f5c321be · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01665fc2-6ef2-4415-bf6a-6d14cac1887e · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78aba339-3138-4ca1-84c4-ff1e468612f8 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad1737c-78b2-4c46-a9ee-7431ac55c420 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Efficient streaming language models with attention sinks,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 42f33cf2-bfd5-4168-9fb0-4c80c50091a4 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da5dd4a-5a53-4efc-b9ac-9d29aff8cf88 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference SnapKV: LLM Knows What You are Looking for Before Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ccda45-dabe-436d-ace9-c95026b75da0 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PyramidInfer: Pyramid KV cache compression for high-throughput LLM inference,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe9687a2-803f-4f7a-96c7-78bec9e8fd1d · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 63227244-9140-4c27-846e-1fd5a6032cb9 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Model tells you what to discard: Adaptive KV cache compression for LLMs,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e3e2cf7-8c4b-4ae0-95cb-10a70f904c40 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference LooGLE: Can Long-Context Language Models Understand Long Contexts?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a6bf00-40b7-4191-9f24-2baea217dc2f · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d453edd-8979-4d3e-82b0-9f9cf16794cf · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Tsplit: Fine-grained gpu memory management for efficient dnn training via tensor splitting,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c231ff2d-aca3-4261-9acb-1ec24f10313d · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Het-gmp: A graph-based system approach to scaling large embedding model training,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d6c90f88-3dd7-4017-902f-92f4384eae7b · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Platod2gl: An efficient dynamic deep graph learning system for graph neural net- work training on billion-scale graphs,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 726d85d1-ff7f-45fb-b65e-30e9558d788f · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Optimizing tensor programs on flexible storage,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46a7c86d-7356-41e5-a6ff-ba62edc91568 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference You Only Cache Once: Decoder-Decoder Architectures for Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f659a8-005c-48d9-b736-fb859617fdb4 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cfb6ec-bace-40b4-bb9a-039f587b4091 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f9eca48-fff4-4f70-95ef-c7730d1660cf · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Flexgen: High-throughput generative inference of large language models with a single gpu,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdf17bf-0f5e-4a83-8e81-3bfc63d2eb51 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Efficient sparse attention needs adaptive token release,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24931133-797e-4db4-92b1-5103e501c2a3 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d4850b-4452-45df-91fa-515eae97b1f6 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a767261c-f5ef-4981-b34d-80cdd15589ab · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2098b003-d575-429f-a8aa-8c95cf730341 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference Lost in the middle: How language models use long contexts,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bcb1c91-78dc-451c-a2f1-327f1cdbc760 · outbound
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference LongBench: A bilingual, multitask benchmark for long context understanding,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.