Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.29575.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1283f76c-0bbf-4207-80e6-721c88cf4e48 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The rapid adoption of generative ai,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71750f8e-e65d-4dfc-8fc7-5c8e831faba4 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The adoption of chatgpt,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3aec83-1a24-4388-adde-a9bb58b1fdf0 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Quantifying large language model usage in scientific papers,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a31b72-eb9d-4ed6-8138-bb0287ef2850 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llumnix: Dynamic scheduling for large language model serving,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a627426c-a4f4-4206-a394-04d805de12ae · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0fee93-9e4d-46c1-8b28-cfa5f828b8ae · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7e53ee-93d8-47cb-a734-76e6e4e5c880 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficiently scaling transformer inference,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd5fb7e-2239-4859-89da-77361ba9beb1 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8401ac48-13a1-42e6-96f7-d75a09241a32 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9be6525-2ed8-4deb-baa1-20852599966e · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmvisor: A real-time latency attribution model for multi-tenant llm serving,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 169dd382-9cd3-48bf-8c43-d32d50983289 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Predicting llm inference latency: A roofline-driven ml method,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa6fa12-8937-465d-87bf-2c10a68eb48a · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Language mod- els are few-shot learners,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06875560-041e-47bd-bca5-4e0a529eaf44 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLaMA: Open and Efficient Foundation Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab551368-fb6b-46ed-b8d7-789c77352651 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Attention is all you need,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2897af-d4a1-4f69-8e86-b899daadc481 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Orca: A distributed serving system for transformer-based generative models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f5eeae-c9e5-433a-b6c4-22b4f42723c9 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient memory management for large language model serving with pagedattention,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aa6c78b-0554-4beb-a73e-df1f537fc2e2 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving TensorRT-LLM,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94854dc3-b780-4314-b2a6-ef0ecc871228 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving DeepSpeed-MII,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f142fcee-c690-41b7-a722-d974ca81131e · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Slora: Scalable serving of thousands of lora adapters,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ded418-0c69-4282-b2fd-5ca8d11f8edc · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e556d816-ff08-40d6-835e-f20286af23c8 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast inference from transform- ers via speculative decoding,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11947e0b-a77a-40a4-9b7f-51eeb4fa4fc2 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c8d29b-ce4a-4423-b2d6-fc9afdaae304 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast Transformer Decoding: One Write-Head is All You Need
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fff4c06-15c0-4d6c-82a7-e8a0c5e61c97 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9d34a5-5e25-41b3-b4f1-f9dde91eb23f · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f35459a-2052-4824-bc59-7fc008a098fc · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec390e4f-1a26-45ca-b8ef-c96e3928b659 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Vidur: A large-scale simulation framework for llm inference,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a09dfd-6bd0-4c5c-b9c0-dad1f0f4702a · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmcompass: Enabling efficient hardware design for large language model inference,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d956b91-494c-4617-83d1-4b357831ad48 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Amali: An analytical model for accurately modeling llm inference on modern gpus,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46db8890-3e1f-4963-9f18-6cc14c0f8fef · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fairness in serving large language models,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1c6471-f486-4f1d-a3b2-dc42413c0c53 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Clean sharegpt dataset,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b314849-88b5-42a6-910a-381b7b4b23f0 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mistral 7B
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a6e250-6339-4122-b15f-b4bac850f715 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Granite Code Models: A Family of Open Foundation Models for Code Intelligence
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab58eaea-b3aa-4790-bff5-1bc003501ea7 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Opt: Open pre-trained transformer language models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543d0be1-ba38-436c-9a08-fd300a83dc8a · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Qwen2 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c589949a-a181-4c18-b42c-a44f50ab1084 · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Ai and memory wall,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa8d4de-c4d2-4dc7-919c-4814191e17ec · outbound
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving OPT: Open Pre-trained Transformer Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.