Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T06:31:47.330226Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 60 inbound Pith citation observations for arXiv:2308.16369.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T06:31:47.330226Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:21.474198Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
48 of 48 outbound references displayed
External citation measurements
15
pith, observed 2026-08-05T02:28:24.338817Z
Observation a105e427-4160-4241-a2a1-abea7e724f96 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://aws.amazon.com/ codewhisperer/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e902f87a-52a2-477f-b2db-bf0cb17a17d1 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://claude.ai
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5824472-cfc1-4eaa-a535-a0a4fd72e88c · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.bing.com/chat
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef6b5aba-023a-4817-bfab-f635250e345c · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://character.ai
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56523f58-fab0-4040-aba0-86d409455c8e · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://chat.openai.com
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e3dd7e90-3bd7-42c4-aba1-130564e9b81b · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/NVIDIA/ FasterTransformer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43b77621-ebaf-4d1e-b36f-2f25293ab284 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/features/ copilot
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0d36354-b63a-49a9-9f1d-e2ba054a5726 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://bard.google.com
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f3d80ce-3263-4912-b024-4ea163211f2d · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://komo.ai/
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39d7275e-1c3c-40f4-891d-4f19b6eb04f9 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://huggingface.co/ decapoda-research/llama-13b-hf
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a093a949-6477-47fb-a696-b07e3dd97d9e · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://docs.nvidia.com/deeplearning/ performance/dl-performance-matrix- multiplication/index.html
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bde7ef6-0a05-4405-9d02-3e1d2b17f4d0 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/karpathy/nanoGPT
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 232a8aaf-560b-433e-a12c-5b90447c3ab3 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //developer.nvidia.com/nvidia-triton- inference-server
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 881f4fa7-d0ce-40f2-81cd-3e103c0997b8 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.theaidream.com/post/openai-gpt- 3-understanding-the-architecture
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf913a27-8cc5-4926-9738-1286a09b353f · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.perplexity.ai/
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8beb3fb5-dab8-4427-abcc-cd05394948b4 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://replit.com/site/ ghostwriter
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 766a7a7a-b3a2-44bc-abe2-146c818423af · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://huggingface.co/ text-generation-inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d844abbd-a60c-4928-953c-92d11bd7164b · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //blog.gopenai.com/how-to-speed-up-llms- and-use-100k-context-window-all-tricks-in- one-place-ffd40577b4c
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70ef3112-31bb-427e-a177-33cf0f40ad31 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //core.vmware.com/blog/using-nvidias-aiml- frameworks-generative-ai-vmware-vsphere
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e290db03-dd13-44ba-83d3-f8f0d22b44a2 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/vllm-project/vllm
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57426900-144d-47fd-b70b-4beaf116d805 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://facebookresearch.github.io/xformers/ components/ops.html
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3108b96-2be2-4ef8-a367-7d3d33348065 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://you.com/
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46f32074-a200-4aa0-b324-b8e31b72980c · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Ef- ficient large scale language modeling with mixtures of experts
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 369ca656-70c0-469a-98cc-3e165495f53a · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Varuna: scal- able, low-cost training of massive deep learning models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c134673a-3b18-404f-bd34-597431574682 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Language models are few-shot learn- ers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b517b334-edd3-480e-992c-a8a2bb97ca17 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills PaLM: Scaling Language Modeling with Pathways
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b9d845d-20b8-4a40-a068-9ae8477833d3 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Clipper: 15 A {Low-Latency} online prediction serving system
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bffd7178-3a30-4336-84cb-641bfa3001d5 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78cb58de-9125-440b-8f0a-143734629bf6 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49d17720-2838-4863-96b5-484def4aff56 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Llm.int8(): 8-bit matrix multiplication for transformers at scale
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95821179-3f0a-46cf-99c3-c3627a8a6e56 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Qlora: Efficient finetuning of quan- tized llms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7dc33df-c59a-41fb-b718-2479d2455d15 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Gptq: Accurate post-training quantization for generative pre-trained transformers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c1d80fe-473a-49b0-b9fa-b0bb66b0011c · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Lee, Anjali Sridhar, Shruti Bhosale, Carole-Jean Wu, and Benjamin Lee
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3aaddcb9-7463-4584-b7fa-8b24980c6c5c · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Gpipe: Effi- cient training of giant neural networks using pipeline parallelism
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df99e933-060a-4123-968b-bfe415ee2feb · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Scaling Laws for Neural Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91b23fbd-e8cd-4d7b-bc13-55e0957c0490 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Accelerating distributed MoE training and inference with lina
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4998e0ee-cb84-48de-ac81-d8c86618bc94 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Pipedream: gen- eralized pipeline parallelism for dnn training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e7aea7c-03f5-435c-a20b-bb02b85b1657 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills GPT-4 Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5acd7e4e-86ab-4722-8564-226b18a2a283 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Efficiently scaling transformer inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation be8945db-5e77-470a-89e5-db6a777c31e8 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Rabe and Charles Staats
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 217f0a45-d75c-46a4-82a5-905704a3d4f9 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fast transformer decoding: One write- head is all you need
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 072d2831-7e23-465d-abb9-25b1180f7b16 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8d9d635-39ef-4241-8838-2d62a9677aad · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30a76471-d0e8-45a1-871f-a6cae359a10e · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Retentive network: A successor to transformer for large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d805c9c-26f5-4ccb-9f51-85326a5dead2 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f7c027a-6cd8-4c75-b390-5e21be4f88f1 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fast distributed inference serving for large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88eb4b57-e784-4861-8c1a-1ad18faf17fe · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Smoothquant: Accu- rate and efficient post-training quantization for large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ea16e24-5c31-490d-9cd4-83eb87925b95 · outbound
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Orca: A distributed serving system for Transformer-Based generative mod- els
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fe6461ac-5e7b-4f5d-bc39-59168f834719 · inbound
A Survey on Efficient Inference for Large Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 282
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4abe331b-dea7-429e-ba62-883728ee7af3 · inbound
HybridFlow: A Flexible and Efficient RLHF Framework SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03427540-7973-4bf9-81b7-c76b81325c5a · inbound
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46754dae-cb19-485e-a3fa-b617fed68ded · inbound
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69c6f004-7f74-445a-aa37-fe42a90f72e9 · inbound
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b6323c5-3711-4110-b424-4813dec83e4f · inbound
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 99981bc5-d0a4-49de-a7c1-f551b1b376a8 · inbound
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b847bca7-d887-4b86-a732-357b7bab6771 · inbound
ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 477b55fb-b7f4-4549-b48c-4db1c93ae349 · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1fe7abec-fc54-4930-b95b-570e5b69389c · inbound
MineDraft: A Framework for Batch Parallel Speculative Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb2bc04-fe28-4b80-ac28-0de489d8a815 · inbound
MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2095a72b-29ea-4b2a-9214-62c9df2099f3 · inbound
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d19e4f0-00d1-4c70-adb2-c60c3e58c9b7 · inbound
Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b5e4f32-6891-40e8-b599-9ed2b0bdcda3 · inbound
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f552e6cc-291e-4fd6-b5a1-1cb3ebfb1402 · inbound
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c7d2a04-1531-4ce8-bb39-8401e855bb25 · inbound
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57459ce9-f5b2-40e0-8e67-b098b2e4c155 · inbound
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45f9f4c5-2df1-4438-9432-efc6ee4104f3 · inbound
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2be23e24-8a6f-4e5d-9e05-c97a6778dd31 · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12f4c3e5-f983-4d7e-a3d1-950e5b34a66a · inbound
EdgeFM: Efficient Edge Inference for Vision-Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f32c318-668d-4544-9bcd-c918306c2685 · inbound
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 477ce940-9db6-4f33-af43-ec239089882b · inbound
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1963e390-410a-4633-9172-f1778c1d7c3f · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5c63d26-c383-43fd-a980-76a81b75c7af · inbound
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0f9de6d-25c9-47db-8e77-3506007e33a8 · inbound
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f64a633-45d3-4dc7-92ba-1a9c74883330 · inbound
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 294494ee-41a1-4a85-aeab-d53556d90933 · inbound
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1672a77c-5603-4e69-9190-7b6a1395e35c · inbound
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2168dfa0-e6e7-4c2b-beb1-4b679bfe853d · inbound
Training-Inference Consistent Segmented Execution for Long-Context LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21e964eb-0864-46e1-b933-fce79a301279 · inbound
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb7488f5-483e-447d-a260-5b81a734ce55 · inbound
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 760b1ed1-c8fb-41bc-8920-71fee3715f0e · inbound
Beyond Scaling: Agents Are Heading to the Edge SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3aa6fc2c-e0f9-40a7-ba07-69dfe47f7249 · inbound
KVBuffer: IO-aware Serving for Linear Attention SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78ccea88-cc51-4bc3-adca-e216a2848772 · inbound
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2dd0a23f-066c-4141-9f14-2f9b8061721e · inbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1eb6d41-9712-4c4a-bf3a-cb16e797ba18 · inbound
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f312682-f682-4d2d-8c83-9582019b93c4 · inbound
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25b7cead-120f-477d-99f3-08adeda338f8 · inbound
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eed0d6e2-6f1f-4ac5-abb2-b116cc224555 · inbound
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 748ada5c-771b-459a-a0ab-0341f53d841b · inbound
DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6fd6099-1cf8-45f5-820c-ab3e9b1da9c1 · inbound
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 093eaa09-3de9-45f5-b0a4-dcbc4333307b · inbound
Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b62544c9-918c-4d02-a3a5-d65f9598d3cc · inbound
Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a4533d2-5f58-4f84-aa5c-5b6e1c69650a · inbound
ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 001e5b19-1864-4616-932c-50058cf2494e · inbound
LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55d8ed2b-9a86-44fa-8e1e-ed2a07997247 · inbound
PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91a132f0-dacf-4838-a14c-0b6a4a24dd4d · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67be4816-2fe7-469e-b273-0698ee5d3ed6 · inbound
KernelSight-LM: A Kernel-Level LLM Inference Simulator SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db4971cb-75b1-46cc-a5bc-9efe313ef085 · inbound
Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c91eeed-246f-4938-89dd-1552cefc34ba · inbound
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673c8a16-7e9e-488c-813f-c4e0314ff1de · inbound
CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4ad8f6e-ec23-4eec-be11-31e37ec220d3 · inbound
BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd6711a-345d-4153-b0ce-deb89e82792a · inbound
[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce8cbfb-3ea3-4862-a784-c002cdd0a97c · inbound
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25433c1c-42f0-4738-9c5a-b4d897fd6b82 · inbound
A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986a2d1d-7550-43df-a407-ef37ae9c12e2 · inbound
DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b83278f4-402a-4a2b-864e-341bc3b30f79 · inbound
Request-Level Energy Attribution for Batched LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab8ae17b-813b-4f18-8c37-91a3224871ba · inbound
Action Chunk Scheduling for Batched Robot Policy Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6ae754-fee4-4a2b-a942-2cc2210b866c · inbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33809743-9fe6-4fe0-afd9-442265d89329 · inbound
Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.