Pith. sign in

Paper Citation Record · LEDGER

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

As of 5 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 60 inbound Pith citation observations for arXiv:2308.16369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.16369 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:31:47.330226Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:21.474198Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact3
  • verified fuzzy43
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

15
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a105e427-4160-4241-a2a1-abea7e724f96 · outbound

This paper cites https://aws.amazon.com/ codewhisperer/.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://aws.amazon.com/ codewhisperer/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.522365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:8abb4a80618130394b691d572cbf32dc559b14a104172153bc18e6bc29254268

Observation e902f87a-52a2-477f-b2db-bf0cb17a17d1 · outbound

This paper cites https://claude.ai.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://claude.ai

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.528671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:3b536621f6e8d76df2d07efc2a9c4bf27d733b1babebc9a8d35976e3b71710ca

Observation f5824472-cfc1-4eaa-a535-a0a4fd72e88c · outbound

This paper cites https://www.bing.com/chat.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.bing.com/chat

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.531607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:75f9361632c975d92d99114b01a9199df8c83b06db30b59f5f23b86d135af7d9

Observation ef6b5aba-023a-4817-bfab-f635250e345c · outbound

This paper cites https://character.ai.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://character.ai

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.534417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:de50396613a850b45d92f3ffb67f1a0bbf164afc04001742b06045aaf88a717a

Observation 56523f58-fab0-4040-aba0-86d409455c8e · outbound

This paper cites https://chat.openai.com.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://chat.openai.com

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.387231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:1bd23e0404c68db5446259ea30430f7d2c22a0c2921ba0cae523a37c2e9baac6

Observation e3dd7e90-3bd7-42c4-aba1-130564e9b81b · outbound

This paper cites https://github.com/NVIDIA/ FasterTransformer.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/NVIDIA/ FasterTransformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.390775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:e986627a6a54fb3ee1cd5e5fda17db2421ca09d5c03e6a63db02f4174600921e

Observation 43b77621-ebaf-4d1e-b36f-2f25293ab284 · outbound

This paper cites https://github.com/features/ copilot.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/features/ copilot

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.393917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:6c2678189d9931f4fa95b1b000b9415269ef5234596e7136954234ed45e51c19

Observation a0d36354-b63a-49a9-9f1d-e2ba054a5726 · outbound

This paper cites https://bard.google.com.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://bard.google.com

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.396924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:15ba9db2f1774a44ba5009cd2a835e604c879efc267c1550a4fc4709f975b6f1

Observation 8f3d80ce-3263-4912-b024-4ea163211f2d · outbound

This paper cites https://komo.ai/.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://komo.ai/

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T06:31:47.400069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:0a35003a0c1e84c9e7542e00533da75909ff9d5713e52c38ff643ef0784d4ea5

Observation 39d7275e-1c3c-40f4-891d-4f19b6eb04f9 · outbound

This paper cites https://huggingface.co/ decapoda-research/llama-13b-hf.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://huggingface.co/ decapoda-research/llama-13b-hf

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.403590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:ecd67b707a36cd5cfe35795c9ad29b63b96d40d0db5574c4c071e7fe5056db6b

Observation a093a949-6477-47fb-a696-b07e3dd97d9e · outbound

This paper cites https://docs.nvidia.com/deeplearning/ performance/dl-performance-matrix- multiplication/index.html.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://docs.nvidia.com/deeplearning/ performance/dl-performance-matrix- multiplication/index.html

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.407072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:cc0ab43fef7d2d44030c9176fd0d23b679eb177df80a68cd907f674f4c0539c5

Observation 6bde7ef6-0a05-4405-9d02-3e1d2b17f4d0 · outbound

This paper cites https://github.com/karpathy/nanoGPT.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/karpathy/nanoGPT

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.410143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:3226811f5a92fc05b3412133cdf8896f825ff3cb448f5c48eb831a3d0971e2fe

Observation 232a8aaf-560b-433e-a12c-5b90447c3ab3 · outbound

This paper cites https: //developer.nvidia.com/nvidia-triton- inference-server.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //developer.nvidia.com/nvidia-triton- inference-server

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.413152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:2ebe725ec1d767c61e7152dce330c478286ebb567fb06f944a1aa1505feb4983

Observation 881f4fa7-d0ce-40f2-81cd-3e103c0997b8 · outbound

This paper cites https://www.theaidream.com/post/openai-gpt- 3-understanding-the-architecture.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.theaidream.com/post/openai-gpt- 3-understanding-the-architecture

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.416734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:65b15f5c792415abb9517bcf378b78b0caa46693416ac1f1b92b2c211da0b90b

Observation bf913a27-8cc5-4926-9738-1286a09b353f · outbound

This paper cites https://www.perplexity.ai/.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://www.perplexity.ai/

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.420385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:24ee4beaf0cedd98f69ce2d9102d1b0ff3e81c468942b972666214e66b41247f

Observation 8beb3fb5-dab8-4427-abcc-cd05394948b4 · outbound

This paper cites https://replit.com/site/ ghostwriter.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://replit.com/site/ ghostwriter

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.425685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:9d20581d1d56c37bf98767925da377a6b4d956ef17c495d4b304ec3f78d61c90

Observation 766a7a7a-b3a2-44bc-abe2-146c818423af · outbound

This paper cites https://huggingface.co/ text-generation-inference.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://huggingface.co/ text-generation-inference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.428812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:1d77b9977dc4953244aa307b45fd6bc933120501a279086d31b395b9783e0292

Observation d844abbd-a60c-4928-953c-92d11bd7164b · outbound

This paper cites https: //blog.gopenai.com/how-to-speed-up-llms- and-use-100k-context-window-all-tricks-in- one-place-ffd40577b4c.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //blog.gopenai.com/how-to-speed-up-llms- and-use-100k-context-window-all-tricks-in- one-place-ffd40577b4c

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.432185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:7435a6206a50bc6f2f54f5519a55506a527288dd853d9fdcaf890a3c70f5b29e

Observation 70ef3112-31bb-427e-a177-33cf0f40ad31 · outbound

This paper cites https: //core.vmware.com/blog/using-nvidias-aiml- frameworks-generative-ai-vmware-vsphere.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https: //core.vmware.com/blog/using-nvidias-aiml- frameworks-generative-ai-vmware-vsphere

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.435797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:baac8367217517004b384cefd575706e6042cefb6d9af940dd6a7996bb854712

Observation e290db03-dd13-44ba-83d3-f8f0d22b44a2 · outbound

This paper cites https://github.com/vllm-project/vllm.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://github.com/vllm-project/vllm

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.439266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:473f412b94d61eae59ef416faccaec1e4ad51d75d90a4270b4a075c8eefc140e

Observation 57426900-144d-47fd-b70b-4beaf116d805 · outbound

This paper cites https://facebookresearch.github.io/xformers/ components/ops.html.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://facebookresearch.github.io/xformers/ components/ops.html

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.442729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:ff6f534ae82d0181edcd7fa2c04f834e4cb794cd6a373091609dae60d3c289b4

Observation b3108b96-2be2-4ef8-a367-7d3d33348065 · outbound

This paper cites https://you.com/.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills https://you.com/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.445810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:cd851c507990576585e49cd32d01cf2e867997b297c1307ff22fe00b2c69d1e9

Observation 46f32074-a200-4aa0-b324-b8e31b72980c · outbound

This paper cites Ef- ficient large scale language modeling with mixtures of experts.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Ef- ficient large scale language modeling with mixtures of experts

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.448872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:bb672c795e151706511417b05fb3692e0a6e2d1abfae6e827d88c59d9b054652

Observation 369ca656-70c0-469a-98cc-3e165495f53a · outbound

This paper cites Varuna: scal- able, low-cost training of massive deep learning models.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Varuna: scal- able, low-cost training of massive deep learning models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.451841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:876102dd0e8d5f534751fc615d53f45810305e16b044f92535b3ccf8d5dd7ac3

Observation c134673a-3b18-404f-bd34-597431574682 · outbound

This paper cites Language models are few-shot learn- ers.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Language models are few-shot learn- ers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.455157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:b45b5940c30577a0e427e1aec26449dff56a6deddbbe1ae2651aa7a04e484ef1

Observation b517b334-edd3-480e-992c-a8a2bb97ca17 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills PaLM: Scaling Language Modeling with Pathways

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:31:47.370660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:9545bb61d2afef2399e5cd493edcae9e8c6e16d90d501f34b73b3b9f74d29b22

Observation 1b9d845d-20b8-4a40-a068-9ae8477833d3 · outbound

This paper cites Clipper: 15 A {Low-Latency} online prediction serving system.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Clipper: 15 A {Low-Latency} online prediction serving system

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.458092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:001fc65983d9a5a67b4bbeafde5a7728d68b8ea10dbad73559c5798783aafc06

Observation bffd7178-3a30-4336-84cb-641bfa3001d5 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.461983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:2cf65c5f36e1aed2118de2f27b911d7d06a74a00e98115e968b36763effaa320

Observation 78cb58de-9125-440b-8f0a-143734629bf6 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.465251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:5cdd44739d7448d8697cb12b281b75145b6d83ceb5685856471bb0c1ee1c1b5f

Observation 49d17720-2838-4863-96b5-484def4aff56 · outbound

This paper cites Llm.int8(): 8-bit matrix multiplication for transformers at scale.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Llm.int8(): 8-bit matrix multiplication for transformers at scale

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.469018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:a7c1e8019e3566dfd1116876dc68e3d4889a82f8564d0f676290120291dd648c

Observation 95821179-3f0a-46cf-99c3-c3627a8a6e56 · outbound

This paper cites Qlora: Efficient finetuning of quan- tized llms.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Qlora: Efficient finetuning of quan- tized llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.472500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:198fb10fe788b9824e594aa4836f385de313ad9f9daa423ef783b403f0737159

Observation f7dc33df-c59a-41fb-b718-2479d2455d15 · outbound

This paper cites Gptq: Accurate post-training quantization for generative pre-trained transformers.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Gptq: Accurate post-training quantization for generative pre-trained transformers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.476366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:6643ceb63215626b6980e9a8c0988bc7f03de2753c3d853abf501922eb0bc810

Observation 7c1d80fe-473a-49b0-b9fa-b0bb66b0011c · outbound

This paper cites Lee, Anjali Sridhar, Shruti Bhosale, Carole-Jean Wu, and Benjamin Lee.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Lee, Anjali Sridhar, Shruti Bhosale, Carole-Jean Wu, and Benjamin Lee

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.479244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:2ce93a611a1b119d21bb8529cc1ac332b1423c1bfa67c5f79f75e3192d14f22c

Observation 3aaddcb9-7463-4584-b7fa-8b24980c6c5c · outbound

This paper cites Gpipe: Effi- cient training of giant neural networks using pipeline parallelism.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Gpipe: Effi- cient training of giant neural networks using pipeline parallelism

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.483062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:92e6999a9346f190fb43417a769d003632a9ddaeff3c2931013ca8253f5fa6b6

Observation df99e933-060a-4123-968b-bfe415ee2feb · outbound

This paper cites Scaling Laws for Neural Language Models.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Scaling Laws for Neural Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:31:47.383653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:f434d1a90a982322af7d906f379d1c250e7386df374d671fd11b43dc45ff1036

Observation 91b23fbd-e8cd-4d7b-bc13-55e0957c0490 · outbound

This paper cites Accelerating distributed MoE training and inference with lina.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Accelerating distributed MoE training and inference with lina

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.486920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:7999ac01106e37cee3aa98b679e4a6400d2c03044b15ac7cfb11ca30e9d05bcc

Observation 4998e0ee-cb84-48de-ac81-d8c86618bc94 · outbound

This paper cites Pipedream: gen- eralized pipeline parallelism for dnn training.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Pipedream: gen- eralized pipeline parallelism for dnn training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.490548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:41e5d388a63a99a9017a7c374d0589a2f4cb254aa26accfc9d99b14c04402dae

Observation 3e7aea7c-03f5-435c-a20b-bb02b85b1657 · outbound

This paper cites GPT-4 Technical Report.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills GPT-4 Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:31:47.363642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:c2dc84dc7d0676822aac027995c9a1870fa97803108f76410d1b8e8834929231

Observation 5acd7e4e-86ab-4722-8564-226b18a2a283 · outbound

This paper cites Efficiently scaling transformer inference.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Efficiently scaling transformer inference

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.493754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:c8e84956fb160ffcbf8b347fe8c70f0aa3beecc688ae87d09280f1ce77146272

Observation be8945db-5e77-470a-89e5-db6a777c31e8 · outbound

This paper cites Rabe and Charles Staats.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Rabe and Charles Staats

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.497115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:942b1283b39295d777bdd3d35c83b2a22c347939ef3e6f1a2c3a05c211cf005c

Observation 217f0a45-d75c-46a4-82a5-905704a3d4f9 · outbound

This paper cites Fast transformer decoding: One write- head is all you need.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fast transformer decoding: One write- head is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.500267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:1ce113a52d70bfdd900281ba722389df5d7c1b9692077a5bab1a1d147c3c94fc

Observation 072d2831-7e23-465d-abb9-25b1180f7b16 · outbound

This paper cites Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.504655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:2637d1c3cc5beeda2c5034baa868e431105e6dbadef4c1dff22b660e89d8d33a

Observation a8d9d635-39ef-4241-8838-2d62a9677aad · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:31:47.377211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:f5f9a18febfb89fa24da675c4433bfce36a7dea7307ffffb36a310ec74867697

Observation 30a76471-d0e8-45a1-871f-a6cae359a10e · outbound

This paper cites Retentive network: A successor to transformer for large language models.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Retentive network: A successor to transformer for large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.509432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:88cca76c25ffb541f9ab665126adf470fd7dbc8e1076f503166f099ba5071515

Observation 6d805c9c-26f5-4ccb-9f51-85326a5dead2 · outbound

This paper cites Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Chi, Tat- sunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.512757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:4b7c4ed78491b6244c184d1295c8c085ee7e8912f012cda31ce58006fa3ee5cb

Observation 9f7c027a-6cd8-4c75-b390-5e21be4f88f1 · outbound

This paper cites Fast distributed inference serving for large language models.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Fast distributed inference serving for large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.515978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:b00e16e080aea65e79d8c84a3b7643392ce3f112e6c66cdccf3e2326185377bc

Observation 88eb4b57-e784-4861-8c1a-1ad18faf17fe · outbound

This paper cites Smoothquant: Accu- rate and efficient post-training quantization for large language models.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Smoothquant: Accu- rate and efficient post-training quantization for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.519160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:e841ca2d8b7b5b1ff008f095ebb07a79a2ebef159c577deecba1464395b178a0

Observation 4ea16e24-5c31-490d-9cd4-83eb87925b95 · outbound

This paper cites Orca: A distributed serving system for Transformer-Based generative mod- els.

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills Orca: A distributed serving system for Transformer-Based generative mod- els

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:31:47.525359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:31:47.330226Z digest=sha256:5bd28f132df46f89134ebb9ced26765bb259d7443a68f8124ae2b9dc9d960f4e

Pith citing papers

Observation fe6461ac-5e7b-4f5d-bc39-59168f834719 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:c2d60ad48c8bfd37672a7fa64ba4b541b18b1c90205e1316cd939e388126d4a9

Observation 4abe331b-dea7-429e-ba62-883728ee7af3 · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:18552536f3f1a1bf2afb0a7c967f0bd4bfc9b9ecca401e51fc3ae19dbb6a25fc

Observation 03427540-7973-4bf9-81b7-c76b81325c5a · inbound

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads cites this paper.

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T11:49:16.707196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T11:49:16.654836Z digest=sha256:4a7723f819102516bab0ce841ac57c8ad069f7465e771f2415a786e318ca71f1

Observation 46754dae-cb19-485e-a3fa-b617fed68ded · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T16:58:11.909937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:51c501f83cd53c5c9d5e33ad03c11dcde27c786242c2f4d31b25543b47dbb547

Observation 69c6f004-7f74-445a-aa37-fe42a90f72e9 · inbound

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints cites this paper.

Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:45:03.952532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:44:35.756018Z digest=sha256:84350f8011c2b917cc9b0f534f52afef0b223afd5ec4e3b5de26fae3a13723e1

Observation 2b6323c5-3711-4110-b424-4813dec83e4f · inbound

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference cites this paper.

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:59:04.724780Z digest=sha256:4c6a559b7f95a02a097100ee9a87ffee6f6c8d5205150ab8891bf56f47d5e508

Observation 99981bc5-d0a4-49de-a7c1-f551b1b376a8 · inbound

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators cites this paper.

SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:00:39.372114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T01:59:30.583027Z digest=sha256:7fefd887ea8d8f8ca85b2efc4c013b54a27c74fe0fb9775012ff1a146eea5f3d

Observation b847bca7-d887-4b86-a732-357b7bab6771 · inbound

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators cites this paper.

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:58:42.971189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:54:08.052482Z digest=sha256:175f3b83702b632730d06646c02d0786c515929fba139fd5a474259b90450b19

Observation 477b55fb-b7f4-4549-b48c-4db1c93ae349 · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T12:15:06.813295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:f02edead355bdad90c877ed1fe5f04ae2107a966b5fc3b341d25abe9d3cd12a0

Observation 1fe7abec-fc54-4930-b95b-570e5b69389c · inbound

MineDraft: A Framework for Batch Parallel Speculative Decoding cites this paper.

MineDraft: A Framework for Batch Parallel Speculative Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:14:46.669096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:14:46.669096Z digest=sha256:0e6c14f7ae4c788c9110e1e0581dbe5161bcaa1be7837a13b6d07f2928101cdf

Observation 6fb2bc04-fe28-4b80-ac28-0de489d8a815 · inbound

MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems cites this paper.

MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T10:30:21.369025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:30:21.369025Z digest=sha256:d44992f8bb203464c04b45443f4dd0c57d73ba0ff3c672b78ff27bda111161f2

Observation 2095a72b-29ea-4b2a-9214-62c9df2099f3 · inbound

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate cites this paper.

Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:05:08.438659Z digest=sha256:e21ff848153ab5fe9d88feb440f89272a8c0636d081da45bc25bf74d8639927a

Observation 1d19e4f0-00d1-4c70-adb2-c60c3e58c9b7 · inbound

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees cites this paper.

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:51:11.343200Z digest=sha256:f2dd52d8a227aa438dd394a4ff3d6f97bf98c89c3be2977d6bbb59972874979b

Observation 5b5e4f32-6891-40e8-b599-9ed2b0bdcda3 · inbound

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs cites this paper.

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:31:25.826205Z digest=sha256:ced0226de895dff903179188796b117cff526b221a46a06a8841e13a54e915cd

Observation f552e6cc-291e-4fd6-b5a1-1cb3ebfb1402 · inbound

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding cites this paper.

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T17:56:39.124969Z digest=sha256:e9635a0d252fe293ced01e32e2605e77f1a08a3f5057bc1db00ac06b0354f0a4

Observation 4c7d2a04-1531-4ce8-bb39-8401e855bb25 · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:b8fa3edec8931b19fefe3635f87f6021065bf86587d765f9437a3f1399e17c6e

Observation 57459ce9-f5b2-40e0-8e67-b098b2e4c155 · inbound

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving cites this paper.

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T13:15:21.201950Z digest=sha256:cbb68fc738b7e45a648e709e76d4e3704868273844f7a572667c62301489c61f

Observation 45f9f4c5-2df1-4438-9432-efc6ee4104f3 · inbound

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems cites this paper.

MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:30:56.899306Z digest=sha256:2d0429588f6cd5b7d51aff2d07a0a8af46e5b66a18683c1a0e1226d62b357826

Observation 2be23e24-8a6f-4e5d-9e05-c97a6778dd31 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:47f52383bcae605be341996d810a51171558d50bd7063edf9e1234bac34f3e8b

Observation 12f4c3e5-f983-4d7e-a3d1-950e5b34a66a · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:55:34.732042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:b532b3ca97562069b4022c8d2fbe84bc728d5fe63e214cce2624938788dd2a36

Observation 1f32c318-668d-4544-9bcd-c918306c2685 · inbound

GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving cites this paper.

GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:37:08.671539Z digest=sha256:307bf1ae10211dfe6e5de28dbee1191f91dd1073755bb551c36899cb809707dc

Observation 477ce940-9db6-4f33-af43-ec239089882b · inbound

PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers cites this paper.

PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:49:56.357400Z digest=sha256:d877d239b6e89b17260a2ebb9763170bf32c41402608af4be478725877463b23

Observation 1963e390-410a-4633-9172-f1778c1d7c3f · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:27:13.781231Z digest=sha256:d7f01820755d2a855f02b64120847eabc51cf953a789cae0a7083c5fbfa1cc7c

Observation a5c63d26-c383-43fd-a980-76a81b75c7af · inbound

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference cites this paper.

Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:35:10.235272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T00:31:12.965178Z digest=sha256:44c771548554690aa9b6274677024eb450bab6347eb5666b307ea6079f655aba

Observation e0f9de6d-25c9-47db-8e77-3506007e33a8 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:29:28.916831Z digest=sha256:72cc3e47a52a1bf9bb34dd1cce9af1891653cc4b32953ee762cb76e3ba470371

Observation 9f64a633-45d3-4dc7-92ba-1a9c74883330 · inbound

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving cites this paper.

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:47:42.025621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T17:43:56.767714Z digest=sha256:517330b17b88877308b8851b9fd495a2854ad0f87e03d95d9d84bda03a3d908d

Observation 294494ee-41a1-4a85-aeab-d53556d90933 · inbound

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference cites this paper.

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T00:56:41.542160Z digest=sha256:4b85a1e4a95fb59079eb8a9acd71f95aee969b1f7e38f69ca1453172e3b9dfaa

Observation 1672a77c-5603-4e69-9190-7b6a1395e35c · inbound

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference cites this paper.

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:02:34.494214Z digest=sha256:0bee60b4eed1b58283f83ae5e316966a8f7bd18c7f8bf1962aea32c0e6f341a3

Observation 2168dfa0-e6e7-4c2b-beb1-4b679bfe853d · inbound

Training-Inference Consistent Segmented Execution for Long-Context LLMs cites this paper.

Training-Inference Consistent Segmented Execution for Long-Context LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T06:52:51.852156Z digest=sha256:07d34282eae3362e4812be4e8beb4e5ea98ad97c6365893857495b961d991775

Observation 21e964eb-0864-46e1-b933-fce79a301279 · inbound

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures cites this paper.

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:31:47.535686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:48:46.006687Z digest=sha256:ee90f971148f0be96f042388a66974fb75c5fce016de627596715fcf89474fd8

Observation fb7488f5-483e-447d-a260-5b81a734ce55 · inbound

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection cites this paper.

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:47.960862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T21:19:31.263068Z digest=sha256:6c5f689814c4bc7333135ac3d33d8572b197460e4ddd66d297edd7651b50f051

Observation 760b1ed1-c8fb-41bc-8920-71fee3715f0e · inbound

Beyond Scaling: Agents Are Heading to the Edge cites this paper.

Beyond Scaling: Agents Are Heading to the Edge SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:53:14.935042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:51:14.999341Z digest=sha256:5afc588ca0fe411ac5e24b7179a18547d670e155e3749970354d49221be3148e

Observation 3aa6fc2c-e0f9-40a7-ba07-69dfe47f7249 · inbound

KVBuffer: IO-aware Serving for Linear Attention cites this paper.

KVBuffer: IO-aware Serving for Linear Attention SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:03:15.260675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:00:06.171816Z digest=sha256:38f54804aaa598161a022ba81bc5d86904ae34b7eab76051600235b53f59da51

Observation 78ccea88-cc51-4bc3-adca-e216a2848772 · inbound

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption cites this paper.

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:58:04.862805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:57:59.624400Z digest=sha256:fadeff026151b3dd25eb336b927b8cde0c2fe66d8d5c240203d9f38b7dca3766

Observation 2dd0a23f-066c-4141-9f14-2f9b8061721e · inbound

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles cites this paper.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.271581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:461973463f5f4a0c1fc7636410a2737de5173898ec07d5530d4277c1bb927e62

Observation d1eb6d41-9712-4c4a-bf3a-cb16e797ba18 · inbound

Frontier: Towards Comprehensive and Accurate LLM Inference Simulation cites this paper.

Frontier: Towards Comprehensive and Accurate LLM Inference Simulation SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T03:49:31.093477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T03:47:49.835773Z digest=sha256:9934a2294aab37dbb2e257adaadde6a5ce59d11b24863a5d67827ab769ff987d

Observation 4f312682-f682-4d2d-8c83-9582019b93c4 · inbound

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System cites this paper.

AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-25T03:15:17.303901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T03:12:49.028342Z digest=sha256:17df356b879bc5a0526f3b4030c1f8ecd9595a9b3da0e342867acd45d8e478ed

Observation 25b7cead-120f-477d-99f3-08adeda338f8 · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:44:45.607766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:865d8e35259f3bff11a134b66874ef0a9e3c8977b0a727fdbbe3dc3b77778b1d

Observation eed0d6e2-6f1f-4ac5-abb2-b116cc224555 · inbound

How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving cites this paper.

How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:29.987690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T14:11:24.646931Z digest=sha256:081f9a3aa9961fee06f7a551e2c3266cb091be9a5228334ca245478d4d8aef56

Observation 748ada5c-771b-459a-a0ab-0341f53d841b · inbound

DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference cites this paper.

DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:40.708633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:47:26.074448Z digest=sha256:7bd07e91118455536642e66c8930ac3eb46a4a07ddb92ef3444c5c62688c5d78

Observation e6fd6099-1cf8-45f5-820c-ab3e9b1da9c1 · inbound

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference cites this paper.

Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:27:06.046624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T23:49:28.318260Z digest=sha256:0f7582b0ba4c5cd35df2b2c64172ba93095af71c171f6e8bec9572e75f192967

Observation 093eaa09-3de9-45f5-b0a4-dcbc4333307b · inbound

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving cites this paper.

Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.124678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:25:46.234576Z digest=sha256:8aa29ca744d8321bb19ab6b2377a831f6d1fd9d8adb404a05b3f3830701cd398

Observation b62544c9-918c-4d02-a3a5-d65f9598d3cc · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-03T07:27:44.463957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:82a85855ec74becafc0c9cd26d963db5eb4b69831114ef764ade67723f0d3980

Observation 5a4533d2-5f58-4f84-aa5c-5b6e1c69650a · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.785034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:ec0125bd4751540f5aff64b92082447bd7ba5723919dd26a1ef4d52f6600a6b7

Observation 001e5b19-1864-4616-932c-50058cf2494e · inbound

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs cites this paper.

LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:59:50.605397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T07:26:07.356352Z digest=sha256:9cbe885d4e6e9da7e3b84c6e05717cfaa8780d60c5ae2886dbb64b1008b698a4

Observation 55d8ed2b-9a86-44fa-8e1e-ed2a07997247 · inbound

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs cites this paper.

PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:07:23.425299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T21:04:43.874528Z digest=sha256:c57b051c7aa147eccd4e92191f2178085c3052c4bac9a8802366845783702869

Observation 91a132f0-dacf-4838-a14c-0b6a4a24dd4d · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T16:15:49.765804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T00:48:19.207465Z digest=sha256:b9c9a5c63de336aeb659cc96cd218d60165c41ab3433b9617353788485d4941e

Observation 67be4816-2fe7-469e-b273-0698ee5d3ed6 · inbound

KernelSight-LM: A Kernel-Level LLM Inference Simulator cites this paper.

KernelSight-LM: A Kernel-Level LLM Inference Simulator SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:19:02.751862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T23:09:38.092583Z digest=sha256:0adda0f06be5cad52e05998c93d517062c8c3ca25123ba8421e39631f8ed6cbd

Observation db4971cb-75b1-46cc-a5bc-9efe313ef085 · inbound

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting cites this paper.

Edge-Deployable LLM Fine-Tuning on a Single GPU for Telecom Network Troubleshooting SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T17:31:15.972423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:31:15.972423Z digest=sha256:e834ee91bd091afa1ea8fa753844f3f2f019d7b14be58c51de1139f25670a471

Observation 4c91eeed-246f-4938-89dd-1552cefc34ba · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:def9f7c36fa4d972b64d64d570a03638c527fc19af144d3d8ee7b063ff97385e

Observation 673c8a16-7e9e-488c-813f-c4e0314ff1de · inbound

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems cites this paper.

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:37:23.050625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T16:30:09.561233Z digest=sha256:5868ffc5863b76af501b698b604239f2bdcc857117c52ab2c9fc8aa32bd8618d

Observation d4ad8f6e-ec23-4eec-be11-31e37ec220d3 · inbound

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving cites this paper.

BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T05:43:12.690359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:43:12.690359Z digest=sha256:6c22590549c2799f46745641bf6b5cbac414c1d4c2f7d60154a0e992e9a92de2

Observation dfd6711a-345d-4153-b0ce-deb89e82792a · inbound

[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows cites this paper.

[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T07:51:15.549754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:51:15.549754Z digest=sha256:447e692294beec8229cb7413c7b9eed360a1d6fd2202404446abd54f810e01da

Observation 7ce8cbfb-3ea3-4862-a784-c002cdd0a97c · inbound

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models cites this paper.

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:32:01.677558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:32:01.677558Z digest=sha256:40d9e20a82c066d80ec4b841d4d95fda0702f95f06bf0b5303ecfc8479631bab

Observation 25433c1c-42f0-4738-9c5a-b4d897fd6b82 · inbound

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing cites this paper.

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:12.115153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:12.115153Z digest=sha256:8eaf12f4bad14a19bc838668ce9947f0cb6e8769369c2b805b536924a3ddf52e

Observation 986a2d1d-7550-43df-a407-ef37ae9c12e2 · inbound

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs cites this paper.

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T01:35:53.234326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:35:53.234326Z digest=sha256:e8330c3025c0f27ff1c9277de051aecaed84d7ddc8f80c33e2785abf8c3ac055

Observation b83278f4-402a-4a2b-864e-341bc3b30f79 · inbound

Request-Level Energy Attribution for Batched LLM Serving cites this paper.

Request-Level Energy Attribution for Batched LLM Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:44:54.104824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:44:54.104824Z digest=sha256:84b48862072da00f4d97b8c2f24e5ae7a488b34368506071a7678e84ce8e0f28

Observation ab8ae17b-813b-4f18-8c37-91a3224871ba · inbound

Action Chunk Scheduling for Batched Robot Policy Serving cites this paper.

Action Chunk Scheduling for Batched Robot Policy Serving SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:44:24.913037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:44:24.913037Z digest=sha256:dcce4c736cd73e314c3e06b36d6a995c34c029f81ee9ea05b39b2bb59e523b92

Observation 6f6ae754-fee4-4a2b-a942-2cc2210b866c · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.474198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.474198Z digest=sha256:8b83f7bd97b13d89a7d6b5dfcf00b199778aa27655e243d23810fef1e65b5a56

Observation 33809743-9fe6-4fe0-afd9-442265d89329 · inbound

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling cites this paper.

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:54:00.298222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:54:00.298222Z digest=sha256:52eb6a0ad130afdc1aea32d3155d91171c3c1dc80c83e9b3b117a9d963799e40