Pith. sign in

Paper Citation Record · LEDGER

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill

As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2606.22541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22541 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T09:42:39.573568Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4be7d489-df1a-4ba7-b089-ea4b2a5d56a9 · outbound

This paper cites https://www.mindspore.cn/tutorials/experts /en/r2.3.1/operation/op_custom_ascendc.htm l, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill https://www.mindspore.cn/tutorials/experts /en/r2.3.1/operation/op_custom_ascendc.htm l, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:cfc1373c55ba5b6dc9ee99ce0f00ce06ad599db6f74ef2db1f797942b694a961

Observation 46952001-3ba0-4468-8b77-6914942ff88a · outbound

This paper cites https://www.hiascend.com/cann, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill https://www.hiascend.com/cann, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:69cc901e9dceace82bed88878324dc4b5f6a7ca61671277c9fe8ede19416c761

Observation 35ecc77b-37c4-4f03-a616-725d857c5d5b · outbound

This paper cites https://www.hiasce nd.com/document/detail/zh/canncommercial/8 3RC1/API/ascendcopapi/atlasascendc_api_07_ 0102.html, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill https://www.hiasce nd.com/document/detail/zh/canncommercial/8 3RC1/API/ascendcopapi/atlasascendc_api_07_ 0102.html, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:dc7f183d49c2fcd3f8d08eecbb60af92b16fb3278ccbbf53abb1d37cdc5816aa

Observation 3b596b1b-d4c9-4107-bf5b-a26d3232bc64 · outbound

This paper cites https://www.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill https://www

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:3018d49011575e6623977fcef52760e6802abacd2e6febc6b688209d7ab59a0a

Observation 5a4533d2-5f58-4f84-aa5c-5b6e1c69650a · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.785034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:2dfaad7553433d122b2ddb041c01f04a57cb808db72d6b1ad55e94109a057ce4

Observation 7583702a-b018-4649-b582-ea1fa8804ad2 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.802503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:30e24763af8acab934c58d6f9eced9790949522d280f8358a3d80c200fddfd30

Observation 76cdb926-82c6-488c-8fd9-d9ef18c7b3f9 · outbound

This paper cites Characterizing cloud-native llm inference at bytedance and exposing optimization challenges and opportunities for future ai accelerators.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Characterizing cloud-native llm inference at bytedance and exposing optimization challenges and opportunities for future ai accelerators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:2c5d101c3faa1364e0d861d3919ad82e29b8ee54f62082b40500b86a5a0cb1de

Observation b9ece214-3270-4622-80a8-f4a374c28a19 · outbound

This paper cites Amali: An analytical model for accurately modeling llm inference on modern gpus.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Amali: An analytical model for accurately modeling llm inference on modern gpus

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:0ba586def199f047c411b54d7bb4bb3f257ab4144099f46f91aeaa5b8620e4bc

Observation b7d6f262-50ac-4bf9-8823-25dafbf881aa · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Generating Long Sequences with Sparse Transformers

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.829437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:48b0a508e00f4c2490173a104b221b11a266e03c9a9782f3c4037572619c0901

Observation d8d69d94-1d59-4810-8994-04f8cff8dafe · outbound

This paper cites Lazy batching: An sla-aware batching system for cloud ma- chine learning inference.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Lazy batching: An sla-aware batching system for cloud ma- chine learning inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:79b0f28f6a8d97f875b910643bbabd2e5078295d6b94f49f5f7068199990d4b1

Observation 0acf6493-f839-4647-ab64-2e53476e2ff9 · outbound

This paper cites Prediction is all moe needs: Expert load distribution goes from fluctuating to stabilizing, 2024.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Prediction is all moe needs: Expert load distribution goes from fluctuating to stabilizing, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:583aa119a268b1914682107fec1cf291204dec26129b57daa670173494e48ed0

Observation 3f72e47b-de4b-4814-af1f-b1db31a7ade4 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:39:46.819780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:23cee8fc0774e6228649da6462f72152b78f86b2a0ae2255680514f900274332

Observation cb4737a2-5bb0-45e9-b72a-4fad63616e2e · outbound

This paper cites Deepseek-v3 technical report, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Deepseek-v3 technical report, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:a9e830cc55a5ed132af127291148fbc3e35954870d853752d5a31307a5a08a22

Observation 703825ae-cb2d-46d9-8535-07ccb01863ce · outbound

This paper cites Longnet: Scaling transformers to 1,000,000,000 tokens, 2023.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Longnet: Scaling transformers to 1,000,000,000 tokens, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:d84fc2ffb46fd46fed60c39e2940a7ee772f479a1615432f8ecbda5eeea4bbe5

Observation 3e2be5df-f02b-4fcb-84cb-790b57ddc27c · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.809933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:71ecca72a1c1a8ad63d955437014064cb3bf28df5143ed8252714db5b8d238fa

Observation edf51e26-9bc1-4bf6-b721-ba265cb41aeb · outbound

This paper cites Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.J.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Switch transformers: scaling to trillion parameter models with simple and efficient sparsity.J

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:ae0c5196c772aad65c2c02aee09ee1cf35bd8e29fdf6b1703c2dd686eeb20a88

Observation 57bb45f1-1410-4d86-96ba-796e23276792 · outbound

This paper cites Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.812502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:3ed8c9a9576380fdbfe7784d42d0e315c4405e827af3ecec1f5d7d95b4e57e24

Observation 8fdd3057-98ec-426d-bf98-3a00b1799f2d · outbound

This paper cites TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.822145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:d14ed9859b056cded9c596ec0cfce82e36e6d9de332fdf1cec822cc74bccd23c

Observation cb28a707-1303-41c5-8403-1b2d4f7a93ed · outbound

This paper cites Past-future scheduler for llm serving under sla guar- antees.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Past-future scheduler for llm serving under sla guar- antees

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:06bfd69431bbce7ab1e9b35b3a26169131e4f18222e852cc4e87e917a79e7d68

Observation 50cd2131-44bf-4af8-9404-d4bb2279ba9f · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.800083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:f97e5e2a158e60bd0804edaa5cab853e3185ac02610972cf9b15bf1323aa306c

Observation 116ba9ea-3c2a-4e3b-8b0d-22aa9632bda0 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.805074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:32cb6b6e71555f5fc75e43b7323f24cd7ac80af7f3fdbb25b5caa0a00b2be6a1

Observation 4530887f-bca3-46b8-9a18-a835f7d32302 · outbound

This paper cites Deepserve: serverless large language model serving at scale.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Deepserve: serverless large language model serving at scale

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:571b3f4f72afb7c3c430541fafb0c832b39283647ccc8dac046f5e5f95acf8c9

Observation 99a5961b-40c0-4a6a-9690-de00f01a1bd8 · outbound

This paper cites Ad- vancing transformer architecture in long-context large language models: A comprehensive survey, 2024.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Ad- vancing transformer architecture in long-context large language models: A comprehensive survey, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:5b829004e106946049692f5e7cba91a1e02cafd86c8925ca933a30e547886f81

Observation 5ebf2254-cd6c-4997-a6d5-6e83ad6d8079 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:cf81ee5d9da80829ce4bf855cf3b77af15091ca46c5fb9fc16279f5541016c56

Observation 1e972d62-124e-40de-a8b9-15103e96bfc5 · outbound

This paper cites Efficient memory man- agement for large language model serving with page- dattention.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Efficient memory man- agement for large language model serving with page- dattention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:853fe61f96bd69750f633bac2179fb661ad602634d11a3dcb75633f2a8551d20

Observation 204977dd-947f-455d-8da1-0a766ba8ca82 · outbound

This paper cites Lightseq: : Sequence level parallelism for distributed training of long context transformers.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Lightseq: : Sequence level parallelism for distributed training of long context transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:80c0baddb687d7cc3f9d5c6c5b431af32bcb166739fde8fb36400b8879d4eace

Observation a02cb5f7-b5af-431a-a4ad-2b877e49476e · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.807624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:357c7fee49516ba6014ec82a80736570fa08409c7e510f0598c4cad3fb3d83fa

Observation 805686ed-6088-4442-a425-2d0dba339105 · outbound

This paper cites Ub-mesh: A hierarchically localized nd-fullmesh data center network architecture.IEEE Micro, 45(5):20–29, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Ub-mesh: A hierarchically localized nd-fullmesh data center network architecture.IEEE Micro, 45(5):20–29, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:a387a13835905e8bb3b35e5dd8eb693655de0a27bd5cc922a4f2b3c772194073

Observation 5178e071-885e-44f4-8fed-b50993fa7318 · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.797595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:b79841de86dd3c554afd52cd118373cdabfed7642186b1265d265e109c30f1ff

Observation fa2864c0-9492-4731-a80c-3510d15ea2a7 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.787355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:8451ea73fa6d76320a70030cf5ec1bd1a80e83bfd5d1ac830f4a370ee8a1dc9a

Observation 20c1ff80-c833-4ead-a6a1-44a0ceb41827 · outbound

This paper cites Revealing the challenges of attention-ffn disaggregation for modern moe models and hardware systems.arXiv preprint arXiv:2602.09721, 2026.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Revealing the challenges of attention-ffn disaggregation for modern moe models and hardware systems.arXiv preprint arXiv:2602.09721, 2026

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.789810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:6ee3f83cb662b0196d8570e9e049adc619c9797ae7cbcd7cf3f15a161d96df82

Observation 0a329dab-8576-44f0-a9e7-cfadcdc99fbe · outbound

This paper cites Ring at- tention with blockwise transformers for near-infinite context.arXiv preprint arXiv:310.01889, 2023.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Ring at- tention with blockwise transformers for near-infinite context.arXiv preprint arXiv:310.01889, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:e12c82f45eb41b279a1b3e75049f960b7598fbf10589617b62068b01ba2c269c

Observation 0f5b23ae-5053-4db6-8d7e-746955fdcb68 · outbound

This paper cites 2025.Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill 2025.Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.792483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:aedadf1b35f7249e1ecceb4e2a9925fb1232356730e8d4d24e46c43e9d0ca0c0

Observation 03c1d626-c481-4ac6-af08-f10b825c6b8d · outbound

This paper cites Moe-gps: Guid- lines for prediction strategy for dynamic expert duplica- tion in moe load balancing, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Moe-gps: Guid- lines for prediction strategy for dynamic expert duplica- tion in moe load balancing, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:7522ba011f381cc8940b5be1727305a057e6a0df69b5d06b7032a5458a4f4789

Observation 9be696ba-5347-4158-9e67-559b2a5bff77 · outbound

This paper cites arXiv preprint arXiv:2503.07137 , year=.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill arXiv preprint arXiv:2503.07137 , year=

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:39:46.815037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:825cee40ca5e13d506901925f0d2ad77ee25e7c748c5ce33e0549eb1502af66f

Observation 937fdc4d-98d1-4638-9c22-aa213a2491ed · outbound

This paper cites Mar- coni: Prefix caching for the era of hybrid llms.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Mar- coni: Prefix caching for the era of hybrid llms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:150b66c9f255c531d118a900a773eeeb2e05fe0a23781272bf0325cf539b24ab

Observation 5dd4124e-755e-4399-9a2f-e562c7b2b729 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Splitwise: Efficient generative llm inference using phase splitting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:81f407c926bc8d768ca9bb5fc4f2f3e1bcb80693089a16b68b10c7f4b84fcdec

Observation 1ccbe570-388c-4600-8b89-ddc718145fc8 · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:b10fb6b6b0aaa7aceec8c6a37be4a8d1523249e4f51450f5691d806c9937e564

Observation bc7cd991-c42b-453e-a73c-73526dfc7ad6 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:98ec195b286fa155d394e7d5d0aca51c54bbcade5b22eed40faa866bc4a954c2

Observation 67934f4f-057e-4cfe-a62e-336c97447677 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.827089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:6f1749d3d394100fff9706fba565d611f7ae3b957be97ace7b92b2ab31a78659

Observation a43b99f4-1f84-4397-8ddb-1be63e5f172c · outbound

This paper cites Pangu pro moe: Mixture of grouped experts for efficient sparsity.arXiv preprint arXiv:505.21411, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Pangu pro moe: Mixture of grouped experts for efficient sparsity.arXiv preprint arXiv:505.21411, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:491a7743ba6e3c317b081a104607847ca42b0d5bd8d18cd34de188be00608cfd

Observation cc4d953c-4c6a-4efb-8ef7-09eec864a669 · outbound

This paper cites Kimi k2: Open agentic intelligence, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Kimi k2: Open agentic intelligence, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:ca90750315f351731b31fb2c49020cea06b548bd35bdf5c42f689c5b06e19840

Observation a5f27282-9de6-4598-8372-eb78e266ecd4 · outbound

This paper cites Thompson.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Thompson

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:52fd028efa5a1a2012bf9e062051ffa4f5bdf1929b1c9a9aea80ee27f2300363

Observation dec2331c-7653-4b01-badd-d2332eab8495 · outbound

This paper cites Step-3 is large yet affordable: Model-system co-design for cost-effective decoding.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Step-3 is large yet affordable: Model-system co-design for cost-effective decoding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.824708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:91fc4c09f2804cc068daf6864b8ec64a179adee0c60dcb6fab069cb67fd106b9

Observation e5d240f6-2937-4b2a-b7b8-7aa5bde81c8f · outbound

This paper cites Flexsp: Accelerating large language model training via flexible sequence parallelism.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Flexsp: Accelerating large language model training via flexible sequence parallelism

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:2660fdc7344caceccaaa51e3d284752fd58d00a2bda2c4c7df1b42e86e936c14

Observation ff62471a-f5a9-4a92-8796-c1da20797600 · outbound

This paper cites {WLB-LLM}:{Workload-Balanced} 4d parallelism for large language model training.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill {WLB-LLM}:{Workload-Balanced} 4d parallelism for large language model training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:eb6792e56311adceef1f10f26e4009abb3f25341ad4e531ba16b2f15d11922c1

Observation efbc5919-349b-4111-a6f4-99ccb0e1dbf7 · outbound

This paper cites Loongserve: Efficiently serving long-context large language models with elas- tic sequence parallelism.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Loongserve: Efficiently serving long-context large language models with elas- tic sequence parallelism

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:575288d0ca3afdff1896db8b49da50067e6ff0dedf91aab93c787fa41c57931e

Observation 49da9fe7-4f1a-49a3-be4e-3b78367201ee · outbound

This paper cites an unresolved cited work.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:5ddb7f01aff6bffa86839e67fdfb9b2130096554e2559050b2da1f53a5a1ed24

Observation 886bb5a6-d07d-4679-86a9-c50c313822ca · outbound

This paper cites Aegaeon: Effective gpu pooling for concurrent llm serving on the market.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Aegaeon: Effective gpu pooling for concurrent llm serving on the market

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:8198f9059d4ca72e07d32ade9e68bb43e7c19a358b33878e7c67171456b1a17c

Observation 1c195a77-a5ef-4408-8357-4db76eeabaa6 · outbound

This paper cites ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:39:46.831727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:1743b18ec2ee55671771e436c1b446b5450934907938dc923104bbcf5ebdac24

Observation bf273d27-d1ec-4a3c-9a26-d35ce2ddbcc3 · outbound

This paper cites Qwen3 technical report, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Qwen3 technical report, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:0e2e5c3c0f2b771bcebde15724b451dc74c18f14a391bee078a59e7904059050

Observation 57407930-470f-4da7-9341-e30756cfe0af · outbound

This paper cites Gonzalez, Clark Bar- rett, and Ying Sheng.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Gonzalez, Clark Bar- rett, and Ying Sheng

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:039eaef765743daf0931cfd94d01424c587281a9bb2d84f30a36b88b92c5b469

Observation bb309625-d9bd-4e83-93d6-440d9cdf4534 · outbound

This paper cites Bucketserve: Bucket-based dynamic batching for smart and efficient llm inference serving.arXiv preprint arXiv:2507.17120, 2025.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Bucketserve: Bucket-based dynamic batching for smart and efficient llm inference serving.arXiv preprint arXiv:2507.17120, 2025

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.794999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:71397be5127662f19473c29816da307b878751c7c996ae8a11cf07a9f911024e

Observation f14d1e4d-49aa-47f7-85db-bf224efb8ad1 · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:c6281c1159d2e8235ebed223e62d001aecb6302b48d5633e7548cedc67235636

Observation 02da2d4b-9545-4a50-a8b0-d8587fa031ee · outbound

This paper cites Sampleattention: Near- lossless acceleration of long context llm inference with adaptive structured sparse attention.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Sampleattention: Near- lossless acceleration of long context llm inference with adaptive structured sparse attention

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:35220497beffbf719ad1e55c7c34e376129cd92e422f4f87dab39805bcc994f2

Observation 6337a5f0-1d91-46be-a122-6f9dae8f542b · outbound

This paper cites Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-26T09:42:39.573568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:68ec8850e78bb671917930064ec23b5022a8077df68162d8fc25aad9c301d2c5

Observation 18a90714-7375-4a52-a134-3460c469c5e5 · outbound

This paper cites Serving Large Language Models on Huawei CloudMatrix384.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Serving Large Language Models on Huawei CloudMatrix384

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.817497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:f023a755093190edb8ca91fb6430d7c91399e4da8783aaadd88da8c9878cdb1a

Pith citing papers

No inbound Pith citation observations are available.