Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T15:03:07.651839Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 100 inbound Pith citation observations for arXiv:2309.06180.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T15:03:07.651839Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:17:55.992637Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 08dccdd5-98d0-49e9-8fc2-68424e18710b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ea5d1a3a-06f2-4b50-92b3-4ff93d6f7d9b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Layer Normalization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 020a3db1-7f86-418e-a32c-f59844f53d99 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a93fed71-3f98-4bbe-830a-fcc14605e2ba · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbca0eb2-521c-44bd-b4ae-1ac2f4b43420 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa3293c6-1d77-49e6-addc-9ad3407bc999 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 824e324e-c329-43f7-a910-2250e37ef738 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Training Deep Nets with Sublinear Memory Cost
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a908b3f5-513c-4a51-bfce-6ff7685937dc · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Gonzalez, Ion Stoica, and Eric P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fc5054a-6c19-48f0-ba31-5b0a7cb2c84a · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention PaLM: Scaling Language Modeling with Pathways
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5b99f57-bbde-4dcd-ba2a-b0c62cecee45 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c795a00-0234-4a5e-9838-52853962a2c9 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f60ccf6b-6748-4095-a8ea-2770d22cd497 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5421875f-a657-44c1-bfaf-46a7ab549b3f · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8475f1f3-5c14-4de7-be9d-5744bad783c3 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Advances in Neural Information Processing Systems 35 (2022), 16344–16359
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22f35f2e-51c1-411f-8b07-44a03537b375 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0b92cc6-7b6f-4c43-ab9a-64c35c09c07d · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ecbb2d0a-e0ea-41db-8bac-7d9947ce92af · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7f0b4fcf-41f2-4301-87d0-aded4ffa0da1 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47d9e4d5-8cb7-4862-9c75-82ccc985a1c0 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 649cd9d0-3c15-4f92-8c73-69860ecf4dac · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7171c841-29bd-4898-8258-adc70234a42c · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41b994ad-0cee-4d6e-b3ac-4f54a8316759 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69f0cf92-1c2e-4a5d-8b51-f3002f882278 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention In 16th USENIX Symposium on Oper- ating Systems Design and Implementation (OSDI 22)
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 585c1e67-09ac-4976-8be7-93d98b0e0134 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b4f3895-665d-4efb-8785-9c4633f6c278 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5db4699b-9297-40fb-b135-f11ec619aa4e · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf0d9868-991c-4e70-97a0-d86b8857242b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4734b47f-db5b-4741-bfc8-23bc9e216ee7 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 090d8bc3-9a77-4681-8587-00d9448e823b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42994f68-102c-4f9c-a533-fa2ce75b0a5c · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6611d062-ea7b-4b1a-8e1c-a04ea637c734 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d738a63-1597-49db-8fff-96a626340e43 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 20103cf5-3f6f-4484-8635-05f8794c2528 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ecff823f-3c41-4263-bcf4-36e0aadf93f2 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d242a37f-13c8-4bf9-9225-05daabbd8863 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6573c59c-e979-4709-aa4d-49841a90a965 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention TensorFlow-Serving: Flexible, High-Performance ML Serving
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe748ac2-8082-48f9-af31-e99b892a8716 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6970e59a-e3bd-45f9-a0e1-1ca53edf0330 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f9a882f-fecc-46b0-880b-1058a20b94bc · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 30a47c21-5c3f-4249-bff7-a8cc348693f6 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention GPT-4 Technical Report
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a1cc0d46-d309-403d-82ff-3d711eff3e3a · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd199c6e-1b18-4516-8610-557303632b5a · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eefaaf41-bbe0-44e1-a186-ceee8ffa759b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 659423ce-ded9-4c05-970b-b38ca297a6d8 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Efficiently Scaling Transformer Inference
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 17795cef-1a39-41a4-b047-882332367102 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 464d2b8d-cfa6-4cba-b873-82cf65375e79 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention In USENIX Annual Technical Conference
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e37e6ef-7bb2-4691-816c-8a6faf31ca45 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 27a18ba0-8337-42ad-8a2c-12987079e10f · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04f25914-1bb7-4542-b0fe-bc53824a95b6 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b514fe48-8fb9-4ec3-9311-fb3b4b968682 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cb0556b-ffb3-46a6-a7b9-f38e2f5426be · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 453495f7-8cb5-4a78-9245-7bd54665c69b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention OLLA: Optimizing the Lifetime and Location of Arrays to Reduce the Memory Usage of Neural Networks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6272fdd7-76f8-4f80-8260-be9929539786 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18a53911-7127-4192-a747-a445c79dc7b9 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Hashimoto
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fbc8220c-365a-45f1-892c-eb1ab0919310 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd745b65-5172-4485-b2fd-d57ee7c8f11e · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention LLaMA: Open and Efficient Foundation Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df58395b-c622-403e-a2a9-7aa76c595b9c · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e39f4097-4adb-442f-a6bb-c8bbc759b2d2 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75222c7e-88ef-4087-bdfc-ce459d8a17e0 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06833f31-c947-4493-95c3-cd1f6d8c7e6b · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e8f5a54-6b76-468d-8568-6e67e32a93a4 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies: Industry Papers
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 53f81328-4859-432f-9bf9-3a24ff5e5060 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c21b7371-c2d6-4eed-8a18-faf9cb7ef895 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0bfcdb19-46a3-40da-b3db-a74e46841850 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ff26025-541d-4a6a-90c8-18935974fb7a · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84b51dc0-35c3-45c4-b7bc-e799058e8007 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75626fe5-d11e-4f77-8cd3-5785faa1cee0 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention OPT: Open Pre-trained Transformer Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cfb028ce-306c-4577-a080-51df01dc9fe1 · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d18360e-836f-41ae-9d50-4b3b6f8916ca · outbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a09cee7c-e6fb-4a8d-8a7f-8e6db35e1ce7 · inbound
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 44a29f19-7905-4a89-a7ad-7673f5708cf5 · inbound
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e90940b9-37a1-46b0-bbfb-bbd9e07b8577 · inbound
Yi: Open Foundation Models by 01.AI Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd045068-0ad5-4337-b55a-975fd7ff9702 · inbound
Seed1.5-VL Technical Report Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84d3b738-2ec1-488c-a7a9-80e4e4f191b3 · inbound
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a5eeee8-d283-4b6c-88b8-539ab86f7d08 · inbound
YAC: Bridging Natural Language and Interactive Visual Exploration with Generative AI for Biomedical Data Discovery Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e306f6b-cea3-4bfa-824d-2702ab97836a · inbound
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8352fb9f-2190-44d3-9802-aee7ad2f0fb8 · inbound
Future Policy Approximation for Offline Reinforcement Learning Improves Mathematical Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16ce4b77-a270-4a88-b99e-4a8ff69b7ace · inbound
SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa27d06a-24dd-487a-a6f2-55b48744a6e2 · inbound
SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5828a3-4a97-4d7e-acc2-6b66459574f5 · inbound
OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b09c3e4c-bd94-49fa-a538-a7ac835e08a2 · inbound
Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fdd69aa-b0f5-4dbb-93be-cdb0a736e4fc · inbound
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7b06a5f-0f9a-4511-bc40-5fe4f38a3c81 · inbound
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c3ed04-8613-45e9-bef5-9f8ce13c9ed3 · inbound
Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76bc791-9b01-41e9-a8b3-5726bb1c4118 · inbound
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e12f028e-a27a-49da-8839-6559f4736ef6 · inbound
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a04c1cb-01ee-4d22-8017-2e79f859b5a6 · inbound
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7bae5c0c-e4bc-491b-952a-6a0e910ed617 · inbound
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5a381d-ad51-42b3-9a13-c7d69b8bf89b · inbound
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ffd6fc1-8044-4e95-8afa-9db274f6991d · inbound
Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21438ac9-8943-42c3-a3a1-ee107a5e15fb · inbound
NVIDIA Nemotron 3: Efficient and Open Intelligence Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dabb644d-35e5-417d-9992-56045faa4ed3 · inbound
EVE: A Generator-Verifier System for Generative Policies Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f47a5b-8cb9-4c34-b86d-fe721771d66b · inbound
Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32ca7d8-4c0b-4a10-93a1-8148a62a2950 · inbound
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79106028-0630-4467-b74d-e580f9316f61 · inbound
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 55908e28-d2b4-4fda-8dfa-a68a54216ec3 · inbound
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0145f447-cc1b-4b41-8253-502b0bac7606 · inbound
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a51cf2-28e8-489c-860c-e93b170f367c · inbound
Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a6b8c7-9528-42b9-8dcf-2db0fd1a2a66 · inbound
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5e16b7c-324e-4620-9fcc-84330660ff6a · inbound
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c365dc1-ac4e-4c10-9549-d0a342db390b · inbound
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 714
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd911f9d-4ae1-4fbf-8312-e4656226fca4 · inbound
SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ac49f6-aba6-43bb-97c0-4fc05963c61f · inbound
MemFactory: Unified Inference & Training Framework for Agent Memory Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00658988-ccf1-4cf2-8028-d7dca793c69e · inbound
Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4a8daeb-d2d5-4605-9713-562afeb2418b · inbound
Reduced-Mass Orbital AI Inference via Integrated Solar, Compute, and Radiator Panels Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e52651ee-5d58-4c47-aa5c-48ab9dbc99c4 · inbound
Benchmarking Compound AI Applications for Hardware-Software Co-Design Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6d71ee04-a1e3-418d-8710-12d1b98f8dc7 · inbound
SLM Finetuning for Natural Language to Domain Specific Code Generation in Production Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10a927ef-35f0-4dea-bde0-019c641a9532 · inbound
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e61aa8b3-c2d6-468f-a5e1-0615aa80d029 · inbound
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a87505b2-e5e1-4561-9e1c-d1794343e3d4 · inbound
Local Rules for Directing the Emergence of Global Properties in Complex Structures Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a964f57f-d5ad-4310-aeaa-d56d086dd49e · inbound
Hierarchical vs. Flat Iteration in Shared-Weight Transformers Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bf9c7bb6-5a24-4be6-a233-0cba20a5f1e3 · inbound
LLM4Log: A Systematic Review of Large Language Model-based Log Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8510faa-f26c-49eb-9d44-bcd0b9754c9e · inbound
LLM4Log: A Systematic Review of Large Language Model-based Log Analysis Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49a2065e-77cc-4752-894b-5225250cbd7d · inbound
EasyVideoR1: Easier RL for Video Understanding Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4806eeca-cb61-45cc-9895-cf72e050fb07 · inbound
Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 969bc451-96bf-41e1-85d8-f8af4cfde9ca · inbound
Neural Garbage Collection: Learning to Forget while Learning to Reason Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3a847f2-9f36-4042-bdbf-bb6d3ec55701 · inbound
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 147a6f06-0d99-4e6c-af08-006dcfa5ea24 · inbound
Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 56f30735-a94e-4f8b-83f8-bf582ad11a9b · inbound
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e876114-b1b1-4689-b641-77fa36e076c3 · inbound
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88eabef9-6c84-43c2-bf68-82201e3f6ab9 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5cc1126-612b-47d4-950e-78cad4bfa8b0 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e09ce1cb-c9df-48ce-bcec-b1abddffe5b7 · inbound
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0185c61-96ba-4647-8b6e-0a23c9aab31a · inbound
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a75bdfc5-750c-4315-8524-cfeecea13224 · inbound
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6cb5c210-50d6-49c3-a800-a33efa8ae065 · inbound
How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c992578f-207e-46e3-8e13-f0010e0a8f38 · inbound
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32b523ad-def9-4fe4-950d-d00400fcfe55 · inbound
KL for a KL: On-Policy Distillation with Control Variate Baseline Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4817d65-5487-43c3-916f-1ff8151a4cdc · inbound
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49752ddc-d2a5-441d-a809-c9c8cb5f4677 · inbound
Slipstream: Trajectory-Grounded Compaction Validation for Long-Horizon Agents Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c2d228a4-6dfa-4e57-9ad0-9d888623c70c · inbound
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91304a7c-a0e0-47fa-9aa1-2eab54439058 · inbound
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4764dc5-d6b6-4fad-808f-2a81a08cff98 · inbound
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22f72be2-9076-4ace-9a53-8d280b79fc15 · inbound
An Executable Benchmarking Suite for Tool-Using Agents Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 43461c72-10b5-4330-91ed-4783fa6158c8 · inbound
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f701fc2-3453-4ce1-866f-64a16063b7a8 · inbound
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98e77faf-e7b7-497a-ab03-9963f6cb71b2 · inbound
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69df9bfd-99df-488c-954f-f41289fc3d83 · inbound
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fa28c05-8db8-4bc8-a1ad-319aee585b18 · inbound
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c54eed09-b2f9-45a1-a00a-7ce456243ab3 · inbound
Generative Floor Plan Design with LLMs via Reinforcement Learning with Verifiable Rewards Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96c141c2-fe94-4275-b0e6-076a59498710 · inbound
Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2402af29-3a3a-43ca-968a-235d89c6b7e9 · inbound
MeMo: Memory as a Model Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e84f747f-2ce1-48b9-87ea-656b6fdaa540 · inbound
MeMo: Memory as a Model Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e058c1c4-0739-4b7b-9e41-db8e0a35eb58 · inbound
Policy-Grounded Dynamic Facet Suggestions for Job Search Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3d04e3f3-a7b8-40a4-895a-56ecab78a467 · inbound
R2V Agent: Teaching SLMs When to Ask for Help Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 09ddf539-c9e4-441e-b92e-8413f0c31e30 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7672d4f9-ed35-468c-9705-6137f79c72c3 · inbound
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b99c7520-b32f-4b93-a5ca-c30ad5988397 · inbound
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72746384-2afe-4780-a416-fdbe9be3c996 · inbound
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03917308-d875-440b-9684-266e62329fc9 · inbound
KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14b8c744-e225-4851-b42e-957d4467c921 · inbound
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b011f7fa-9874-453b-8649-d5dc50d75c65 · inbound
LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2846e575-b3af-49bc-85bc-7c86f40ffed2 · inbound
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad17b4d6-3f03-4b2e-a6e9-6cd28c9f670e · inbound
Runtime-Certified Bounded-Error Quantized Attention Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ccdd2310-86f0-4a47-a25e-e61b8922d932 · inbound
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 85dcd48c-f689-411c-b247-c166ab2f5352 · inbound
TO-Agents: A Multi-Agent AI Pipeline for Preference-Guided Topology Optimization Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5bdee5d-6b54-491d-9c74-b70c16387b58 · inbound
Asymmetric Virtual Memory Paging for Hybrid Mamba-Transformer Inference Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 34c4f42d-43e7-4577-8b93-866d8b07cbc4 · inbound
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd290409-6fe2-431b-8ffb-8ac5e58023bb · inbound
Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d7ed762-0779-442a-b35b-a5737d33461a · inbound
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0465413f-89c2-4459-a945-f5b3dacbc004 · inbound
The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0df5a6f-056a-48b6-bf7a-342314744147 · inbound
ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8e49efcf-4d58-4f29-b07a-0610410dff25 · inbound
A Unified Structured Query Understanding Framework for Industrial Semantic Search Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 428c5f72-fa48-40ad-a416-05461800ea45 · inbound
Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2524a5f8-f395-47b8-8e88-026ff4d2f44c · inbound
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4fc5ae9e-3a64-42ef-95cb-bd2983a67c28 · inbound
MRMMIA: Membership Inference Attacks on Memory in Chat Agents Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 86cf3189-d1ac-40de-9ed0-67bd3350a318 · inbound
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eb8ab6dd-fd67-44ed-bb74-56c90c3eace4 · inbound
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88cc7254-93c5-475d-bf14-610287411264 · inbound
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.