Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:05:22.596292Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.18106.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:05:22.596292Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T16:37:20.774251Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T21:36:15.531550Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47cd4147-f3cc-47dd-ac60-4559b03ac1b3 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/acceleratin g-llama3/?hss_channel=lcp-78618366/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83a42cb8-8a3e-474b-a87f-8c1bf358989e · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://flashinfer.ai/ 2024/02/02/introduce-flashinfer.html
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22947872-6e5a-40b3-a0f1-7ad9f52a1758 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/cutlass-p ing-pong-gemm-kernel/
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a18affcc-3f87-4cba-b693-1485f0841d30 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://huggingface.co/blog/layerskip
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 452c3868-d977-4ffc-a866-f90d86387830 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https: //pytorch.org/blog/flash-decoding/
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64bc80c1-1cfc-46b9-aa7a-daf962464996 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mirror of https://gitee.com/ascend/pytorch
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd28661e-e65c-4227-9cc1-8c37acf99bae · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/p ybind/pybind11
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5edfda8b-cdd9-4bac-a27d-9644b9eb8474 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://docs.vllm.ai/e n/latest/automatic_prefix_caching/apc.html
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83cefb15-b204-4231-81d8-6abdf8fcb6ae · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/docum ent/detail/en/canncommercial/700/modeldevpt /ptmigr/ptaoplist_000006.html
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 111339f5-494f-42d3-99f4-e97c7f379713 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/source /zh/Pytorch/60RC2/apiref/apilist/ptaoplist _000787.html
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e869776-9d22-46e9-b9b5-733fc3a691ac · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f37ccd31-1e6d-4ad5-a445-3b98d387190f · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/sour ce/zh/CANNCommunityEdition/80RC1alpha001/ap iref/fmkadptapi/ptaoplist_000142.html
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0bbb9e4b-c871-4b17-9262-7d6c6e8aa49e · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/v llm-project/vllm/tree/main/.buildkite/nig htly-benchmarks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f099ea11-6cec-4a5b-a15d-cbcf46a99a58 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com /vllm-project/vllm/pull/8054
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e76a983-1087-48a1-bdbd-d0c58b75f3c2 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://developer.nvidia.com/blog/optimi zing-compute-shaders-for-l2-locality-using -thread-group-id-swizzling/ , July 2020
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b08a2df-99af-4b74-a82c-04bb52b56086 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mnemosyne: Parallelization strategies for efficiently serving multi-million context length llm in- ference requests without approximations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d69371e-fbdb-400e-bda8-e7a15a12b8f2 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Taming throughput- latency tradeoff in llm inference with sarathi-serve
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 36bab499-f5d5-491e-8201-f8db57f5b7b7 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d67471d-a83d-4795-9981-14a0fe5af7a9 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 75bfec98-92e3-4eb8-8d29-2d43e65f6b38 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8a7e0a-e400-4e19-943b-a7ad5eea69ee · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels End-to-end object detection with transform- ers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 733cd9e4-6420-4366-b0d2-338805b786a5 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Accelerating Large Language Model Decoding with Speculative Sampling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b6a347-0352-4d1d-8e03-58225f182926 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8e9371-9f9c-4480-b7d6-75c73f083ede · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Flashattention: Fast and memory- efficient exact attention with io-awareness
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0bb6f57-3fb0-440d-a1ef-f3fa95263ba5 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels An image is worth 16x16 words: Transformers for image recognition at scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5c526a61-35fc-4073-9207-c18d2d746ee0 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac97fea9-8f50-4cd6-a9b7-c55ae0c2bb51 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9857202e-31fb-4f66-baee-4a8f4072702d · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels P/D-Serve: Serving Disaggregated Large Language Model at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7800d6c-7386-4453-a1c9-cd79dd7fe34b · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 27cd03d3-196d-403b-b467-27393fa9a00e · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Scaling Laws for Neural Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb016a02-9b7e-42f5-84f8-f03d8a495ba2 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Efficient memory man- agement for large language model serving with page- dattention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81fc6725-0bf6-436b-b36f-2d510aa3f4b4 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Fast inference from transformers via speculative decoding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1aec2a8-ca3b-4f6d-9fb9-70510bc93d1e · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106ad724-c428-45ed-a10e-76a854dcaa36 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 07c0b696-aace-4bae-87a8-e6a60384a8ef · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Davinci: A scalable architecture for neural network com- puting
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bfec7400-f702-40ed-9ca1-0d7ab176d340 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd49f4a0-a2f4-4f8f-9570-d8fc9bfaabf2 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8651dbe-77c5-4aa3-9ffd-a011c769b9e4 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fb187e3-ee08-4c7d-8e13-2b351b54ce6c · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c36028-98ec-4624-a5cc-0d7b426b321d · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0dc711-fbd7-432d-be49-72d79524c17a · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8170ae-176f-48f0-9769-8c7a0df895a6 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Qwen2 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76612cb5-28c6-48c6-951e-65d09d21bbce · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels ChunkAt- tention: Efficient self-attention with prefix-aware KV cache and two-phase partition
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d24a8bee-8606-495d-9c90-1a43c4577a8b · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Orca: A distributed serving system for transformer-based generative mod- els
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8058edb3-d53b-41ee-a168-34bf101eebb0 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7878538c-d35e-4ba6-93b8-b14f9a978001 · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Lookahead: An inference acceleration frame- work for large language model with lossless generation accuracy
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 794db7ca-1f52-48d4-af1f-04c218142f6b · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SGLang: Efficient Execution of Structured Language Model Programs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41de5c9-3c40-480d-8740-319d8f0be7ee · outbound
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 948b4573-1367-45f1-adf1-ed2a3aaa1f87 · inbound
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.