Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.140836Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 1 inbound Pith citation observation for arXiv:2502.08182.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:10:23.140836Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T15:26:01.283448Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T15:30:17.957796Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dfd3e3b5-9f1e-48c8-bd5b-a897ed731d4c · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://huggingface.co/ docs/accelerate/index
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0366ba-6bb8-4a30-9625-dce2bdc4c27f · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees URL https://sharegpt.com/
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c809b2a0-2bd0-4602-9c5b-dbf0129b93cd · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Acharya, B
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aea5f611-4591-4f01-8d04-77904f1a748f · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SYMPHONY: Improving Memory Management for LLM Inference Workloads
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b16632af-cffe-4664-a125-b340c71fcf6f · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73bab42-7498-48f5-bf9d-7c33882b4062 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf37bd70-da3b-4841-97b7-fd4028c59005 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Alomari, N
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d1de85-13b0-46c8-b1a0-474069fc9f6b · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Neutrino Production via $e^-e^+$ Collision at $Z$-boson Peak
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd779b54-fef4-4499-b02c-d0d4df4e3041 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c060a7-b00f-4a3d-be13-bf570ed796c5 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e3514d0-fad9-46f4-b93b-1a6e6adea279 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Evaluating Large Language Models Trained on Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fa23cf-a362-426c-bd5e-0524c2d95eaa · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec6d3e6-3e3c-4a2c-8d75-392c61dbe549 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Cheng, Y
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 361882a4-e08b-49c1-b7f5-b9e10f42196d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Choquette, W
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a217b87c-3eb6-4156-8223-f9e31b9318d4 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, X
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb26b63-d62b-433a-b9dd-49bd525b9a5e · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Crankshaw, G.-E
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8864959f-64f8-4c43-8edb-a01c5363656d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d267163-9e1d-4904-b99b-ae00d8b36c70 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c055ce8a-9ab6-43ec-9bda-0a119794ef09 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 868c2bb5-b7aa-4d78-b8f9-516d52af42b3 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Improving LLM Abilities in Idiomatic Translation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d1cb3d-82fc-453c-be89-0574bc0e903d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2081ed-afb5-4bab-9329-e09f717e306c · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Elliott, M
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f513ccb5-2799-4452-84e7-1bf801a016da · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gambhir and V
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aab287a1-5331-43ed-9816-6b4d6f824e8a · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05f2565b-fce5-4450-aa48-bc54c5e0504e · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast State Restoration in LLM Serving with HCache
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0875813d-2c66-4ca6-bea5-74fc4b222a31 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079c3fe8-7d92-4e73-9b8f-29ff9a898f21 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e9542ac-3d38-458a-9453-dbf985ba63f4 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c2f1dc1-ed41-4cce-8198-907b1da0d872 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Gujarati, R
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bedb766-c3d8-4009-a6f4-8cabf5d05c7c · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4102a7aa-35ba-4dfa-8a29-3e9d8efcfe92 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Data Interpreter: An LLM Agent For Data Science
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6ccbed-f64c-4d5b-8fd7-4062c9ac36e2 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16fe9b7-a4e6-4618-80d9-5f3d3eb0c748 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LoRA: Low-Rank Adaptation of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 080b3edf-e083-4039-befa-b6418483dec4 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Jayaram Subramanya, D
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1d0193d-511a-475f-be6f-ed6b8ca82988 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8bef64-1fbb-4748-884f-0e9ed2d915f8 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A System for Microserving of LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98edf3ec-67d3-4150-bca6-d89a37b78ddb · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees P/D-Serve: Serving Disaggregated Large Language Model at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ab360a-47e1-4829-a2cc-787f9493e12a · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Kasner and O
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551ac547-c551-4136-985d-96db252a50a6 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees TransLLaMa: LLM-based Simultaneous Translation System
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6999be8-0db9-4b49-9d4b-110ca6713b84 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Koziolek, S
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a8b374e-b10d-4e9a-9594-7daa28bf3f7d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 843c7021-a90c-4e97-831c-8d9dcc4d6a57 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3bd7c19c-e72d-450c-8ce4-93341d3e4ac1 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c8af5f9-0bca-478e-a65f-338ecd210b1d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Li and Y
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00cfd601-3e0b-45c4-bde4-0a203f62de22 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400705793
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67037a46-e217-4161-9602-77ce5d054a37 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78c06f67-3fea-4181-9542-0d3614066805 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees JarviX: A LLM No code Platform for Tabular Data Analysis and Optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933940f3-c2ba-45f7-842e-3de9c9c88baa · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 978-1-939133-40-3
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ee6f993-8a04-4e54-b4ac-517074a9be2b · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1869ca4a-04c4-436d-80ab-e9a07dd08f58 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patel, E
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 23f0e975-cf2d-4376-aefd-7619d2c91c38 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3080e40-7188-4226-b8df-5d76f9baf677 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2edf95f3-8250-47f2-b3ef-aa5241ad9e9d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859140e1-5240-453b-bb61-92f8c698617f · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Radford, J
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89689631-2df4-43fd-acbb-f8bd492b7d58 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLaMAX: Scaling Linguistic Horizons of LLM by Enhancing Translation Capabilities Beyond 100 Languages
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ab214a-be66-4b90-be9b-270b12727c8d · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee04131e-c1a4-4b82-a082-b917a0c239fc · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sheng, L
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1865967-9974-40fd-9d08-a0f2b497d8ef · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Patke, D
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9dd51793-6362-43ee-bb19-d1a6b6bfa105 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9798400712869
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 662abd6e-5878-4a70-b9fb-c2d4c8dd6cf4 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a99035e2-d60e-4602-a506-c98c8c941309 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39c68a13-1fb8-4926-a96e-7bd0d068d316 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Taori, I
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94410835-311f-4f22-8d66-a29178cc0ccd · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57493d46-6d0c-4de4-b325-1d93336b9913 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees SynCode: LLM Generation with Grammar Augmentation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c171c9a-a435-4661-abb7-c565c78cbb53 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Attention Is All You Need
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d609aa-25e5-40d3-9b0b-0a22d781f7de · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c88a061-ac65-4cc4-9197-81ade5109dc1 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Sivakumar
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04c34b7b-ae75-4b9c-b79b-59d040fa0e2e · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Efficient Streaming Language Models with Attention Sinks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f6e1bd-4e2e-472b-ad3c-45f13b31fea5 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 702c9932-288c-49ce-8502-41fabbff8863 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ef46cc2-c744-4699-ab3a-3650b3babf5b · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c879ee-4105-496f-9fea-418b7edf8bb9 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6ed6c039-3b37-495c-8cbe-b021ae7d85a0 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a06af6-94c4-4c7d-8068-6b465a061033 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e23aa1-de2f-45ca-a08b-7f8daffda8d9 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9e39e33-03ee-45d4-9459-ef5051e604bf · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Fast Distributed Inference Serving for Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb437f5-8eb6-4b00-9e21-9ba31727476f · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhong, S
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d24709d-5531-4835-a6a8-ae200dff1cbb · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Enhancing LLM with Evolutionary Fine Tuning for News Summary Generation
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49859a21-b5f5-4422-afb2-cbf81ec3b713 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0698959-7015-43d3-ab8a-17a0f7d293ca · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees A Comparative Study of Offline Models and Online LLMs in Fake News Detection
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 684d01be-6bba-4e3a-b5c0-df6027b1f96b · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Zhang, Y
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d4a23382-f0b9-483a-aa37-fb196362d0f2 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees OPT: Open Pre-trained Transformer Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d5ee91-2e11-40cb-943f-04393324bc8c · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78ff80cd-7750-4a06-bf4b-db00c4e2132a · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM-Enhanced Data Management
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7f6cc5-0cdd-4cf1-83c4-8069fea15ff8 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees ISBN 9781450381376
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 863f88bd-2095-4e28-8bc3-5619ec192cf1 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees doi: https://doi.org/10.1016/j.csl.2021
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c827061c-605c-455c-b76a-ab5a0c378816 · outbound
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f7bd89d-1367-495f-859a-f4031d98415d · inbound
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.