Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T02:48:18.330339Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2604.18529.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T02:48:18.330339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:09.632043Z
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 62b20ab4-46a3-40fa-861d-1e253f799089 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd8c002f-7794-4580-8d2f-66e13e415024 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ec778ea7-a6a6-43fa-9307-44b1d5cadad4 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 016aa677-6fac-4a7c-9cf6-8e4825e56472 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1641f35-4272-4b6b-b1c2-61c12af56784 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(Rotterdam, Netherlands)(ASPLOS ’25)
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59921c60-abf9-4b0d-bba5-f0aeb786217a · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed November 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac04a6fe-30bb-47a7-a0f8-22093f9bc7a9 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f8d7267a-7304-4f84-a113-e6452afaff3c · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fcca7652-81b8-4f7a-93df-e7b63bf2dd9d · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e1d3ab66-5641-4e35-8342-dfea167095bb · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 323bb551-8f17-4304-aab3-5d2798c161d7 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ff7deec7-ca88-4bf7-a39c-3b413d63759b · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f106ff4f-bcb9-4f75-9f4e-4a6a431f32bb · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0f1fb9a6-851b-4c62-96f1-32b588ba5bf6 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 882b027b-c9f7-41bf-a994-039e64c4a6e9 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed November 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7de55fd5-b176-40e1-8d20-47371fba73f5 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e5feed70-adc8-4ebb-80cc-2e99a763f8f1 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d9530f04-112c-4495-adc2-b78c5472185e · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing In Findings of the Association for Computational Linguistics: ACL 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ecc2fabf-db43-4187-a15a-d92d7f8a6c85 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8bb6983d-d5e9-40cc-9c36-c311745f903f · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Nyx: Virtualizing dataflow execution on shared FPGA platforms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5492f8f5-a6b3-47a4-ae72-08e2fe89a96e · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5500f534-b187-4d50-8a4b-aeb9f96a01d8 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 64c44172-09ec-4f5d-8c1e-d8d8867c4aab · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6330ada9-fe4e-4cdf-a550-65b534718baf · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 34e174e2-5332-47db-a5ff-f86ac7bffd48 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing 2025.A CXL progress report: The elephant is learning to dance.https://www.eejournal.com/article/a-cxl-progress-report- the-elephant-is-learning-to-dance/
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6a5ddbcd-6e9e-4e04-8fc0-7cdf005f5eca · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c48b33e6-e6ba-4e4a-b662-be8f827eccd5 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Advances in Neural Information Processing Systems 37 (2024), 22947– 22970
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0eedc597-4243-4447-b206-9debfddf297b · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f972df56-d797-4761-8e44-b17c62ac90ff · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6831bc56-b83e-488a-89be-31953da0acb2 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3e8927d5-bccf-42e2-95c2-c5b0b6ec3c67 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Advances in Neural Information Processing Systems 36 (2023), 52342–52364
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aca5ac1d-3565-461b-96e6-6f7e7c5cd0e8 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4e8fa6ac-1daf-4cd0-8240-740f34e9a3db · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 867aa5ff-5f4f-4dcd-b112-6c5025056ff2 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 793f5eb2-ca87-4108-8151-06de6d20d91a · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dec2ac56-a880-43d1-946b-356688b77134 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 01e6bcd7-d73b-497b-a57f-430a6ba92712 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing RayN: Ray Tracing Acceleration with Near-memory Computing
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aafd7be7-f2d8-4ef0-bd22-6d24e65299af · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59189d4e-dbbd-4eff-ae6d-5c132ad8076c · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation be3ceb72-7492-41ae-b023-dbc65e57a7d6 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Grape: Practical and efficient graphed execution for dynamic deep neural networks on gpus
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 251a5aa6-532b-4cce-95ed-26d1fc718aa6 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Accessed April 2026
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f236e617-6e9d-4948-b24b-2bb99cb6c303 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 802373b7-a209-42e4-b052-e6e81f382f25 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8d626a5b-09bc-4ed6-b148-e9878a391591 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d0990dfc-0785-447b-9775-cac6128e78d7 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bfbc2e60-2ebd-4fab-a5fd-63c7411d8440 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Efficient Streaming Language Models with Attention Sinks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 313561c4-ef2e-4dea-9525-bb4b9c48d236 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8218bed6-b7a5-4a03-ba3c-2136cb234631 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing LazyFormer: Self Attention with Lazy Update
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9bee5ae2-23a3-4263-90d3-5e91ea0247a1 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing arXiv:2512.18194 [cs.DC] https://arxiv.org/abs/2512.18194
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 22205822-6f34-413c-8064-63367be331b8 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 01a377fa-b53f-4dfe-94f0-f9ed7a7bc7b7 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f79e7d39-93d7-4662-8af7-3c8f32826329 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Jenga: Effective Memory Management for Serving LLM with Heterogeneity
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac968331-dc79-4d57-bea6-729aa4e6c1aa · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing OPT: Open Pre-trained Transformer Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 430a27a3-436f-49f6-bf9f-022b8b1c1f66 · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c71ebe57-ba35-4ff3-a9be-db911743dcbe · outbound
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ef6d700-32a1-4f43-a1a4-14841684c2f8 · inbound
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.