Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T04:58:48.072701Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.22389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T04:58:48.072701Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a5eb13f-1a44-446a-a3a8-4ad8c73ae0c6 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mistral 7B
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d393a459-3d3a-45d5-8569-c0bd97478b75 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Qwen2 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b2bd9c-f749-42eb-9cb7-980212cc8ceb · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llama 3 model card,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da7909b2-9490-49a6-86de-26876e8dbc64 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding How long can context length of open-source llms truly promise?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa6817ee-8830-46ac-a1cf-e0de2eed627d · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Language models are few-shot learners,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d9bc7c-6939-467d-bd4f-49f6c3eb6c83 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265aae74-fe63-4ce2-bff8-76c421346c67 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longbench: A bilingual, multitask benchmark for long context understanding,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66a1dd3-0fe1-46a0-9064-8ec084fbd602 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding P3-llm: An integrated npu-pim accelerator for llm inference using hybrid numerical formats,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64458efa-d583-4833-b270-826b924a6558 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A survey on large language model acceleration based on kv cache management,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1efbd9f7-260f-488c-b7d1-d0c26ae5ee0b · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Codec: Prefix-shared decoding kernel for llms,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ddd9b6-e43d-45d6-aeea-f49e8bbffe2f · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Orca: A distributed serving system for transformer- based generative models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bc9d2b-61b4-461b-8350-8052822ef91a · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Splitwise: Efficient generative llm inference using phase splitting,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91226077-408b-4a62-8895-c59e6dead0e7 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kivi: A tuning-free asymmetric 2bit quantization for kv cache,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9444d6e0-37ff-4da2-b119-2d5ac9ffbcad · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient streaming language models with attention sinks,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b35529-fa0f-45f7-9700-bc509221e412 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Duoattention: Efficient long-context llm inference with retrieval and streaming heads,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e9a2ab-f619-4ccc-8439-672908bf54cf · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Snapkv: Llm knows what you are looking for before gener- ation,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f01aef9d-f544-4f9e-81b1-22137343fa9d · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sepllm: Accelerate large language models by com- pressing one segment into one separator,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c03b91-62dc-4df3-88b6-452aaa2d9d7b · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Lm-infinite: Zero-shot extreme length generalization for large language models,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f39ce39-6cc0-495b-966b-d541fcf165f6 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Longformer: The Long-Document Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a69917-216f-40a3-b323-06d044da6719 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54487552-cbc4-45c4-a1d4-2f33166ce04a · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318a4ef6-4f36-4e23-939a-687ecb7d8edd · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kvo-llm: Boosting long-context generation throughput for batched llm inference,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e2cabfc-b120-4daf-86f8-207a0a322d8b · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 313f3b18-dce6-4810-b5ce-f1a393fa2957 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Alisa: Accelerating large language model inference via sparsity-aware kv caching,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d78b5e8-4f14-4873-93fa-b52b7a00c4d3 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Mata: A memory-efficient attention accelerator for llms exploiting look-back kv cache pruning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88bcf029-6fc7-4492-8f9d-7280cd5933e4 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Unicaim: A unified cam/cim architecture with static- dynamic kv cache pruning for efficient long-context llm inference,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d746782c-237a-4ea6-850b-e105d2e2bfc5 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Token-picker: Accelerating attention in text generation with minimized memory transfer via probability estimation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79c477fa-51e5-4b0a-8746-14a8ec912734 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dias: Distance-based attention sparsity for ultra-long- sequence transformer with tree-like processing-in-memory architecture,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5d4989-9e38-45d6-93e6-795bbd388645 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Veda: Efficient llm generation through voting-based kv cache eviction and dataflow-flexible accelerator,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e531de-bec9-4f9a-b671-a9e308b1d187 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv-cache oriented query-aware sparse attention accelerator with cross-stage precision-configurable digital cim,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0f6cc6-865f-4b67-af00-7818900e3b05 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding End-to-end acceleration of generative models with runtime regularized kv cache management,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deeb2018-df6f-49d8-af1b-701c3b66cca0 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Edgellm: A highly efficient cpu-fpga heterogeneous edge accelerator for large language models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082819a2-cf6e-4748-b2e9-7663597bb079 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63bd56cf-985d-4172-ba8b-b72d0e54149a · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Flightllm: Efficient large language model inference with a complete mapping flow on fpgas,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f0df5c5-732d-48a8-902d-ee7f519581f4 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ofq-llm: Outlier-flexing quantization for efficient low- bit large language model acceleration,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca253343-8b28-448c-901b-7b5abf8a5582 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20f6a49-0c49-43c5-94ad-f2b9c415ab98 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Apt-llm: Exploiting arbitrary-precision tensor core comput- ing for llm acceleration,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc3d7e8-46f9-48a6-b05e-c1b19e0be6fb · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding A Survey on Efficient Inference for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a0208c-2ad0-46de-b0b8-7c0713d15844 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding When to stop? towards efficient code generation in llms with excess token prevention,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ec201e-38ec-4d68-a31d-44fe07b3fea5 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Llmcompass: Enabling efficient hardware design for large language model inference,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604e267d-01bf-49e1-805d-3ab2733c5df8 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Energy cost modelling for optimizing large language model inference on hardware accelerators,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9168e695-d063-4ad4-9634-4fa85f8a9393 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89491f21-4613-47d2-8be2-e7c1ee4649b6 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Skipkv: Selective skipping of kv generation and storage for efficient inference with large reasoning models,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a0e496a-1570-4ff2-900b-8d6699310dce · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Titanus: Enabling kv cache pruning and quantization on-the-fly for llm acceleration,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605bc04d-cf93-41d7-a727-1a6322b69f05 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Infinigen: Efficient generative inference of large language models with dynamic kv cache management,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe095db-cf41-4bfe-b822-d71c0a6c13f6 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sparq attention: Bandwidth-efficient llm inference,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 610a51a3-5140-4cf8-873c-e171f7a54f2f · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Algorithm 232: Heapsort,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3832966-788f-4656-a365-4e0ff33f09a0 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Sorting networks and their applications,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da2e8d6-b1c9-4d46-ab3b-8d346b62474f · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Efficient memory management for large language model serving with PagedAttention,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbae7ab-3070-43ee-8ecf-b4e80e042fe3 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6216b1-3f70-4659-8ed8-230610832737 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ten lessons from three generations shaped google’s tpuv4i: Industrial product,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6c5676-5725-4df3-8714-ccfd8c9d5462 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dramsim3: A cycle-accurate, thermal-capable dram sim- ulator,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d45b6d6-2d5b-49e7-83f1-77b9cf4008f3 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 469204c9-e339-46d2-a2b8-a54f42bcba56 · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Dynamickv: Task-aware adaptive kv cache compression for long context llms,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9345b92f-f78d-4469-8e47-d447e96ef55d · outbound
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding DeepScaleTool: A tool for the accurate estimation of technology scaling in the deep-submicron era,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.