Pith. sign in

Paper Citation Record · LEDGER

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels

As of 11 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.18106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18106 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:05:22.596292Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:37:20.774251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:36:15.531550Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47cd4147-f3cc-47dd-ac60-4559b03ac1b3 · outbound

This paper cites https://pytorch.org/blog/acceleratin g-llama3/?hss_channel=lcp-78618366/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/acceleratin g-llama3/?hss_channel=lcp-78618366/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.331066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.392066Z digest=sha256:7660a2ab94f9934a8f3af03db7c07636d7d120354304722bdd4b03ee6c6988dc

Observation 83a42cb8-8a3e-474b-a87f-8c1bf358989e · outbound

This paper cites https://flashinfer.ai/ 2024/02/02/introduce-flashinfer.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://flashinfer.ai/ 2024/02/02/introduce-flashinfer.html

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.319515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.396611Z digest=sha256:c683584b89e864c1fc81375ee7b3387d90a4010aa8b9744ea738663ac9d19dc3

Observation 22947872-6e5a-40b3-a0f1-7ad9f52a1758 · outbound

This paper cites https://pytorch.org/blog/cutlass-p ing-pong-gemm-kernel/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/cutlass-p ing-pong-gemm-kernel/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.305150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.401104Z digest=sha256:3a837ac27298fecbfbce4b30381ac20313b3a5e104c311294fea94fd9d45cb03

Observation a18affcc-3f87-4cba-b693-1485f0841d30 · outbound

This paper cites https://huggingface.co/blog/layerskip.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://huggingface.co/blog/layerskip

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.291470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.405940Z digest=sha256:4f14f1005ccc44bbd192de0161a5280d9b842746a3145ba8a21a2ce1b4a5c83b

Observation 452c3868-d977-4ffc-a866-f90d86387830 · outbound

This paper cites https: //pytorch.org/blog/flash-decoding/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https: //pytorch.org/blog/flash-decoding/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.277652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.410870Z digest=sha256:da856d9d670469e463d860483fe77b230fc5f66c8da39ae0e7106d1787f6c6b2

Observation 64bc80c1-1cfc-46b9-aa7a-daf962464996 · outbound

This paper cites Mirror of https://gitee.com/ascend/pytorch.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mirror of https://gitee.com/ascend/pytorch

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.263812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.415024Z digest=sha256:0f9dde9d5d6d97945b89a0f59d600d5814f700821a6e5e33daa952e037f15bee

Observation dd28661e-e65c-4227-9cc1-8c37acf99bae · outbound

This paper cites https://github.com/p ybind/pybind11.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/p ybind/pybind11

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.250528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.419398Z digest=sha256:38efd7d92855f5c0216530aeede1ce741f2851e709d9718447ae500142687ac0

Observation 5edfda8b-cdd9-4bac-a27d-9644b9eb8474 · outbound

This paper cites https://docs.vllm.ai/e n/latest/automatic_prefix_caching/apc.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://docs.vllm.ai/e n/latest/automatic_prefix_caching/apc.html

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.232533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.423136Z digest=sha256:b9f5643891b38445e04de36794ef531c67e43ba23a06f92dd772553defce1d7b

Observation 83cefb15-b204-4231-81d8-6abdf8fcb6ae · outbound

This paper cites https://www.hiascend.com/docum ent/detail/en/canncommercial/700/modeldevpt /ptmigr/ptaoplist_000006.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/docum ent/detail/en/canncommercial/700/modeldevpt /ptmigr/ptaoplist_000006.html

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.219157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.426276Z digest=sha256:4e9e8a46c9908b07f5e1fb8a3cb14df6f7a548bac211324293d54b5382bc40a8

Observation 111339f5-494f-42d3-99f4-e97c7f379713 · outbound

This paper cites https://www.hiascend.com/doc_center/source /zh/Pytorch/60RC2/apiref/apilist/ptaoplist _000787.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/source /zh/Pytorch/60RC2/apiref/apilist/ptaoplist _000787.html

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.206340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.429593Z digest=sha256:816e2013d45317bb910cb6b5baebb0f51afe0b25e512bdc0b44be23a261cb759

Observation 9e869776-9d22-46e9-b9b5-733fc3a691ac · outbound

This paper cites https://www.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.194461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.434016Z digest=sha256:60ba539e412ecf2f9997808d7c33988add23160f53ebadf5334ea7ffb454d351

Observation f37ccd31-1e6d-4ad5-a445-3b98d387190f · outbound

This paper cites https://www.hiascend.com/doc_center/sour ce/zh/CANNCommunityEdition/80RC1alpha001/ap iref/fmkadptapi/ptaoplist_000142.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/sour ce/zh/CANNCommunityEdition/80RC1alpha001/ap iref/fmkadptapi/ptaoplist_000142.html

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.183067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.437889Z digest=sha256:b45d26931f8e6eba627a86a71f344bafd57368280fb63c4db4be4be86e586080

Observation 0bbb9e4b-c871-4b17-9262-7d6c6e8aa49e · outbound

This paper cites https://github.com/v llm-project/vllm/tree/main/.buildkite/nig htly-benchmarks.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/v llm-project/vllm/tree/main/.buildkite/nig htly-benchmarks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.169997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.441379Z digest=sha256:9e3da7df1271950b75913f4427476dffe3f1481c675893b17e3501b09a7fe03d

Observation f099ea11-6cec-4a5b-a15d-cbcf46a99a58 · outbound

This paper cites https://github.com /vllm-project/vllm/pull/8054.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com /vllm-project/vllm/pull/8054

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.155090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.445080Z digest=sha256:c142273787c3e147d6b6d58aa0f8174ec338e9d23cb03ab564810035d0129350

Observation 6e76a983-1087-48a1-bdbd-d0c58b75f3c2 · outbound

This paper cites https://developer.nvidia.com/blog/optimi zing-compute-shaders-for-l2-locality-using -thread-group-id-swizzling/ , July 2020.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://developer.nvidia.com/blog/optimi zing-compute-shaders-for-l2-locality-using -thread-group-id-swizzling/ , July 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.139298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.448986Z digest=sha256:5ddd46ec31769671bbe4ec48c920eb84f290521a609ea1b27305832be79657cd

Observation 6b08a2df-99af-4b74-a82c-04bb52b56086 · outbound

This paper cites Mnemosyne: Parallelization strategies for efficiently serving multi-million context length llm in- ference requests without approximations.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mnemosyne: Parallelization strategies for efficiently serving multi-million context length llm in- ference requests without approximations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.453803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.453803Z digest=sha256:01c3caa1931c71196a81ce3400bb34910a7337417b3d1422a9ad870964c08237

Observation 2d69371e-fbdb-400e-bda8-e7a15a12b8f2 · outbound

This paper cites Taming throughput- latency tradeoff in llm inference with sarathi-serve.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Taming throughput- latency tradeoff in llm inference with sarathi-serve

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.125471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.458077Z digest=sha256:e381c68c25b1619c29731777247cff9b168255e8e9d4eea1450f1db1440e0497

Observation 36bab499-f5d5-491e-8201-f8db57f5b7b7 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.462957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.462957Z digest=sha256:acc840190ebe11e262d34d19b34137fba75b68f9f9058b1193a40ac2ea918898

Observation 0d67471d-a83d-4795-9981-14a0fe5af7a9 · outbound

This paper cites an unresolved cited work.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:05:23.112597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.468538Z digest=sha256:1160810aacffa4b22a291261e343cd276e2127f7165e05ab5bf9d5046666b71f

Observation 75bfec98-92e3-4eb8-8d29-2d43e65f6b38 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.473831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.473831Z digest=sha256:7bc3ee2216900aa1a0d8c8fbd366addca946cd1053647057e3427426c7cf3549

Observation 5e8a7e0a-e400-4e19-943b-a7ad5eea69ee · outbound

This paper cites End-to-end object detection with transform- ers.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels End-to-end object detection with transform- ers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.478313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.478313Z digest=sha256:9984ec651d3da585ca587827d3c904213448c229a0801d17b65654c38a6f43c0

Observation 733cd9e4-6420-4366-b0d2-338805b786a5 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.482679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.482679Z digest=sha256:bca01f100173cab092b96355ad815aecb18004e9fd79a58a95ae976788903bd9

Observation d0b6a347-0352-4d1d-8e03-58225f182926 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.486482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.486482Z digest=sha256:cb78e2a94645996b1336a2cd0e9cb103224ab47de951d649debd0a74cc31e208

Observation be8e9371-9f9c-4480-b7d6-75c73f083ede · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Flashattention: Fast and memory- efficient exact attention with io-awareness

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.091824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.490605Z digest=sha256:75828f19aeea4e15c45d8d109106b66f6292bfb87aab8cbfd81253cdd3de1abe

Observation a0bb6f57-3fb0-440d-a1ef-f3fa95263ba5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels An image is worth 16x16 words: Transformers for image recognition at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.078223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.494447Z digest=sha256:aded18f96a80e02faff2c12d9514d0340f4ff2061c412e9b704f8a07c6a62490

Observation 5c526a61-35fc-4073-9207-c18d2d746ee0 · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.498348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.498348Z digest=sha256:596f3072775d1d2073af229d443dda26e87f54db7b62cd37df9e2d7ac2bfc4fa

Observation ac97fea9-8f50-4cd6-a9b7-c55ae0c2bb51 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.503397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.503397Z digest=sha256:eddf61f6ca81387341ed4ea47d65decebbbe714d0377a348f224b79c92fd3b17

Observation 9857202e-31fb-4f66-baee-4a8f4072702d · outbound

This paper cites P/D-Serve: Serving Disaggregated Large Language Model at Scale.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.508469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.508469Z digest=sha256:9968a719accada9bc0d5491e4e2601c472b59c395fafcf3c0e709bf8f30c988a

Observation b7800d6c-7386-4453-a1c9-cd79dd7fe34b · outbound

This paper cites POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:05:22.775442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.514760Z digest=sha256:3b2c3dbc30d03f33dbd63e6ab39dcc021741a6a77d29c390188814290b37ccd9

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.519070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.519070Z digest=sha256:8e602e07b875246ecbdb08413f7357d652c79d7fb62b8ade31299bd9f7a50a98

Observation bb016a02-9b7e-42f5-84f8-f03d8a495ba2 · outbound

This paper cites Efficient memory man- agement for large language model serving with page- dattention.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Efficient memory man- agement for large language model serving with page- dattention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.523327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.523327Z digest=sha256:36685ce7cdb9b16cb5da54dbf0532775fb12c898f438b7b43fd078bccef30654

Observation 81fc6725-0bf6-436b-b36f-2d510aa3f4b4 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Fast inference from transformers via speculative decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.527297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.527297Z digest=sha256:767d70e9f007e24901929df75e58173052b12db51d4e5e8201f4d2918c075a0f

Observation b1aec2a8-ca3b-4f6d-9fb9-70510bc93d1e · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.531342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.531342Z digest=sha256:129b29a181cca5ae6ab7910622c02d048173cccd0d875b7934e2496f07f200c6

Observation 106ad724-c428-45ed-a10e-76a854dcaa36 · outbound

This paper cites Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.047071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.535557Z digest=sha256:a39e0c96571ce1442bda158417bfd997ebe68245bed8ab7c8c99d361d051dd5d

Observation 07c0b696-aace-4bae-87a8-e6a60384a8ef · outbound

This paper cites Davinci: A scalable architecture for neural network com- puting.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Davinci: A scalable architecture for neural network com- puting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.033583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.539706Z digest=sha256:205a75789b2cbca49b5c5c12406abbfa6cb41406dad1a2ed474da9e2a0b4842f

Observation bfec7400-f702-40ed-9ca1-0d7ab176d340 · outbound

This paper cites FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.543764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.543764Z digest=sha256:be4e4ad266e22db9612a2a91764b64feb1d435c3f0e0016cd1dea76fff7e0ade

Observation bd49f4a0-a2f4-4f8f-9570-d8fc9bfaabf2 · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.548112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.548112Z digest=sha256:8f3c70c08903dfd2b1c1843bf9e032f8446bca114a0ed69f9ea885621c84dd6a

Observation e8651dbe-77c5-4aa3-9ffd-a011c769b9e4 · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Specinfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.018073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.553178Z digest=sha256:ead3dfa7bd81c0e1ac363c2c6bb4485c9ad371d90c0c4f349b5a8121add2011c

Observation 3fb187e3-ee08-4c7d-8e13-2b351b54ce6c · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.557261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.557261Z digest=sha256:329605c625c78b886f40064e4977d5ce678151743ebd7ed8c161ef9d48367d2b

Observation b8c36028-98ec-4624-a5cc-0d7b426b321d · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.561973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.561973Z digest=sha256:51d1e1f673d1cede845d3f749c5492e686b6021cde6d0ff04b256e5a3f4058e2

Observation 1f0dc711-fbd7-432d-be49-72d79524c17a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.567647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.567647Z digest=sha256:11f300ae525e5bafc9459e9e29b92377b3253d7745f377a8bffe9737b4fd490b

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.572111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.572111Z digest=sha256:b19470b6874303d25e9c9fe61a31722cb6b100dd5323857a6037f69ac44aec49

Observation 76612cb5-28c6-48c6-951e-65d09d21bbce · outbound

This paper cites ChunkAt- tention: Efficient self-attention with prefix-aware KV cache and two-phase partition.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels ChunkAt- tention: Efficient self-attention with prefix-aware KV cache and two-phase partition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.002281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.576332Z digest=sha256:fa37bf17be4e94db8bcfbde09e68b4f0b1b2e65271de67c36623860d0ee74de5

Observation d24a8bee-8606-495d-9c90-1a43c4577a8b · outbound

This paper cites Orca: A distributed serving system for transformer-based generative mod- els.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Orca: A distributed serving system for transformer-based generative mod- els

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.987665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.580400Z digest=sha256:281502a02900fab22371f708d544fbf20ea3b25764e8201ec9901f461235bf85

Observation 8058edb3-d53b-41ee-a168-34bf101eebb0 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.584734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.584734Z digest=sha256:78681911d0cd16cc75e978e52ca0a738244e4076ffc0f345df83dc535560df31

Observation 7878538c-d35e-4ba6-93b8-b14f9a978001 · outbound

This paper cites Lookahead: An inference acceleration frame- work for large language model with lossless generation accuracy.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Lookahead: An inference acceleration frame- work for large language model with lossless generation accuracy

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.973937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.589029Z digest=sha256:60dc3c07da7b11916e49829cadaf85e592c242a3ef783669149b9008a9b06714

Observation 794db7ca-1f52-48d4-af1f-04c218142f6b · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SGLang: Efficient Execution of Structured Language Model Programs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.592491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.592491Z digest=sha256:fdfb63acf2882c7f42749fe298f45cd0dfe496eb196740c295c1dd161d80921f

Observation c41de5c9-3c40-480d-8740-319d8f0be7ee · outbound

This paper cites Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.960244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T05:05:22.596292Z digest=sha256:3e8f14b7cfca5d20a05bde52af7076e05eea8c641c87b4acb26c607c68b15836

Pith citing papers

Observation 948b4573-1367-45f1-adf1-ed2a3aaa1f87 · inbound

AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments cites this paper.

AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.533193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T16:37:20.774251Z digest=sha256:1b804578da425c7e53eba263e3f449876b98ff646f8228cde4bfd2ab25230b05