Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:29:34.814139Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2501.05460.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:29:34.814139Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T00:31:12.965178Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T10:27:56.708993Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6a4cebb-9521-4820-ac94-9c94739b3328 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c278c7-4849-44d0-9d21-2051357c10b3 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159e985a-df5e-4fb8-895a-24eba915a341 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d368e730-0f3a-4b4b-bf29-ddbf99e9d34c · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Bayesian performance analysis for black-box optimization benchmarking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7115aa6f-977d-4886-a4fa-bf59f60a44de · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation A survey on evaluation of large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b5c4a0e-3eb6-4663-9f5d-4bcc72741d2a · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4bfc5aba-9f8e-4f16-807a-02fa3f311093 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0117dca-f92c-49d7-9e15-5433c53e1873 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation P/D-Serve: Serving Disaggregated Large Language Model at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2377e117-9098-4867-96d4-78956e29de2a · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15dd31be-cc95-4cb3-ba54-d9daf26363bd · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Gonzalez, Hao Zhang, and Ion Stoica
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a7ee87-fa7b-44fd-a5b2-3782d8931485 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation SnapKV: LLM Knows What You are Looking for Before Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c28bde-86c0-41e9-9b92-df4fa936a110 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Visual instruction tuning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6228cef4-4dd2-4c59-a563-b9bd40c58774 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation A Survey of Resource-efficient LLM and Multimodal Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff5c09c-b35e-4453-a966-7649deab5b34 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e71c408-6d18-48c6-b9ce-04c7d489c5c6 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Splitwise: Efficient generative llm inference using phase splitting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 40f4ea5e-993e-4f6c-a998-6dfd33b497e4 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c49337-4b17-45cb-943c-27d8b7d00f4a · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Déjàvu: Kv-cache streaming for fast, fault-tolerant generative llm serving
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 338d2da4-22cc-4777-afe7-9b661fd0bbc2 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Multimodal large language models: A survey
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd044171-a942-48a2-b84e-255b46599e40 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Next-qa: Next phase of question-answering to explaining temporal actions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e4a0f9a-7f97-42a2-893c-ebb612184070 · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b3ed10-eb80-449a-9dbf-e10719527c3a · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Orca: A distributed serving system for Transformer-Based generative models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22e01768-2e30-4e87-ac5c-33606952ea2c · outbound
Efficiently Serving Large Multimodal Models Using EPD Disaggregation Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b31d7bd5-83c8-41cd-9bf8-4f1d994fd4cf · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fbf9efa4-beef-43dd-a7d5-9b06a76d13ae · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f6663c9-ffdd-45ef-81fa-24d7aa46c2c0 · inbound
RTP-LLM: High-Performance Alibaba LLM Inference Engine Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6dfd23c0-4950-4a8d-8547-2c2653179600 · inbound
M*: A Modular, Extensible, Serving System for Multimodal Models Efficiently Serving Large Multimodal Models Using EPD Disaggregation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.