Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:11.429005Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2507.03153.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:11.429005Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T07:06:53.318182Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T08:55:35.129021Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 283541c0-7eae-45bd-94b7-ec057d567d2a · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b596e2d0-6399-4cf9-a0de-0c5b61b62bd6 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d9f2aa4-a207-4c12-9005-4d935434b456 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Longformer: The Long-Document Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc6f1ec-f1bd-48fc-909b-5a8d15a18bba · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c16654e-8ea8-45e1-93e7-fe4bfd9029f7 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b58bdb54-51f7-452a-9701-6a74709b43ac · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d79de82-e945-491d-b9ae-c7fa6d4d4158 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation baaaabff-b972-4e5e-b815-5aec8ec9e06b · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a3bcea62-8d5c-458c-942f-7e698426e0d6 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference 2024.{Cost- Efficient} large language model serving for multi-turn conversations with{CachedAttention}
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4f8178d-b21c-4288-bb66-550d6a64f86f · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51290079-dd5d-49e2-becb-197b6672b7bb · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Reformer: The Efficient Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f005e9bb-00fd-43ba-ae2e-564e9bfdc5a8 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673f8366-0a7c-4916-8c3e-704d82699165 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8e17e11e-569d-4aa1-8c22-918e25924b4a · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 821624bc-6404-49cd-beb8-71168644d3a4 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88891cc-7e72-4ad8-b2bc-9b95e0e3169c · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f95e3c7-396e-45e7-a18c-415c57366bd0 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0ddab3eb-4b04-4d4b-86dd-934d3b0a6c6a · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bde7af9-e12f-464e-940f-a70578758c78 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Advances in Neural Information Processing Systems 37 (2024), 22947– 22970
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4cfc5cd7-ebf8-4bf4-91fb-20b43ebd3780 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd570bcb-6f0c-45d5-9dd8-8bc3a5fab274 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference s1: Simple test-time scaling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a83586-5834-405a-a6b6-f5e81933883a · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b00083-ebe0-4ad6-b566-9c8e9e4d6279 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cdbeddd-2d21-4162-aa5f-5982cab441b1 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb6a93f-78c8-48a2-af27-994cb7722caf · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cdcfa3c6-f8c5-4f3a-b2cf-7f7d0504f755 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Code Llama: Open Foundation Models for Code
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37b0486f-032e-4737-905f-db21fa971db7 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8d5ada3d-7956-4b8f-a0e9-bab06d28c0a3 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fa13dd56-c45f-45b3-9124-c34322e448b1 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87760f9-9b4d-4089-8ba3-851c6ef17fc6 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f4e39e-5f34-4aae-a0e3-33f962a8c334 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02aaafc-504c-4ac7-a1b7-fdee5b54e97c · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d142924-a365-45e9-9f08-d8cdae1f92f4 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 30d3ef68-2644-49ff-86db-2a3f9113d671 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Efficient Streaming Language Models with Attention Sinks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca89c058-e067-484c-a479-72a3b1eb8b75 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference XAttention: Block Sparse Attention with Antidiagonal Scoring
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce532b92-7fce-410b-84e2-38f491b37009 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1b89ed75-3b16-4829-bdb5-889e414fcf96 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 30b42c94-a547-4810-941f-147bf261f7ce · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23595a6e-6812-41cf-bb8e-053175fa418d · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 379f6667-7c1b-4af3-a2aa-8d206fe92ad7 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bd63e885-b224-42e4-9093-0fa42a612a38 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a4b9e8-a88a-4978-85f7-292e48558915 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81360729-f072-4c54-b3ad-1f54066d5c5c · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 312b2387-692d-428f-b045-4841fb73ad65 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference OPT: Open Pre-trained Transformer Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbd7307-683a-477e-b16c-06b410c55bbe · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2ca5ed73-8e11-4321-8050-b223b8889972 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8aa3366e-4b03-4a9e-97e7-c694e98953dd · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72cd1495-ea5b-4234-864d-b2433786a270 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d750438-4284-444a-9063-5f37783b2aa0 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference 2024.{DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 845365b8-9b6d-474c-b622-dd86c4a0ac78 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Pointer Sentinel Mixture Models
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad134e0-f174-484e-9c75-853194db01f2 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference Advances in Neural Information Processing Systems 35 (2022), 16344–16359
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a831765-b1c5-4afe-8fc2-00a250eba7be · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference In Proceedings of the 29th Symposium on Operating Systems Principles
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30a9b1a4-667b-4e1b-be27-98305084f558 · outbound
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference A Complete Survey on LLM-based AI Chatbots
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b19ae0-e370-40bd-b243-539118cd83c6 · inbound
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c2597739-d9ed-4c33-b905-30feae5dc31f · inbound
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f74a384a-6f6a-4d7f-82a1-563ce4dd2dd2 · inbound
Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 306398cd-4251-4152-b2c4-6f9056b11667 · inbound
Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.