Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T22:52:39.416575Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2508.12851.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T22:52:39.416575Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:56:19.766584Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:07:44.131237Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4fb10cc8-a015-4a03-974c-0d137f463bd2 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a87e8e74-f757-4436-a62d-6ad928348af8 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mixtral of Experts
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2d376eaf-8462-4e61-b09d-5c46021642ce · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V3 Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 23bff876-6a79-4ee9-bd75-5de700e664aa · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Geforce rtx 40 series
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 19c09be9-0738-434a-94fc-4d3c69197c24 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Gpunion: Autonomous gpu sharing on campus
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 79250885-432f-42c5-807f-b29e73cd6a49 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5a23cd60-650e-4d0c-a848-4dd8e1d088cf · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 28f25519-5370-4395-94e5-818027fc61df · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 352ef6f0-0f0e-41ad-aa33-87cb1bd15739 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Expert Parallelism Load Balancer (EPLB)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 70e646ec-bad1-4f76-ad27-0e0e41e1e4ce · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 42834221-0771-4a8d-a968-82bfcca28b9d · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Moe-infinity: Efficient moe inference on personal machines with sparsity-aware expert cache
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 88c9b676-31a2-46a5-a55c-19bee0e0cc4c · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b5426018-7783-4318-9668-586020acbabd · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f80b0c76-ef3c-4d15-9b6a-9029202c8e7d · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 33b53e28-96ef-4d18-8419-7549eb4b4098 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 1653027f-19ef-47dd-a00b-ef4e2d8659da · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pointer sentinel mixture models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 162d2908-54e7-4457-b0be-4ee977d78c01 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement TACO: Topics in Algorithmic COde generation dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 7ff6cb8c-caa1-41f1-add2-96631a92d043 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 27729f97-78dd-44f1-a1cf-f9b3cba67c3c · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Joint application placement and request routing optimization for dynamic edge computing service management
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 46ec477a-dcc3-4e53-aa45-6acc8f0b077d · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Task placement and resource allocation for edge machine learning: A gnn- based multi-agent reinforcement learning paradigm
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 94421f3a-aee4-4696-bc92-8a53e4996ba2 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a4d1369b-13fd-4c1f-bf09-04ef30249892 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 295ed844-d8fd-44c3-abae-acb58ffc8c93 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation a59f0663-8ab0-4e3e-9bba-eacf0386e657 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Prophet: Fine-grained load balancing for parallel training of large- scale moe models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 2bbbc822-bdac-496b-b4ae-c8a673107c8d · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8223bee1-211b-4122-962f-abcb7fa044a7 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 189a01be-0957-4ce8-acbf-ca8f9b7ad920 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Accelerating distributed {MoE} training and inference with lina
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation c8181ca5-3127-45e8-8eeb-d2c31f7b72f3 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 91c32161-0484-4cd7-88be-11e5b78f1e3f · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 63ab3982-2d7b-4a8d-aac7-51adcfe8dbe5 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0a824cf5-fbcf-4467-af21-b6c80fe85591 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 8926653f-8995-4258-9dae-2495de2d0545 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation df5f0a61-7f5b-4617-b094-43cf6ee7c39e · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pipemoe: Accelerating mixture- of-experts through adaptive pipelining
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 6b75f761-5b45-4cd4-bc95-07f70bce0efb · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 5a664ba4-c476-4d49-b517-72779e4d8c19 · outbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tutel: Adaptive mixture-of-experts at scale
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation b7337072-9053-4129-8709-020fe34ac8a7 · inbound
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 67f9a1aa-d74f-4562-b519-63ae177f5500 · inbound
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f0ba5b1d-6ddb-4152-bc87-f0d5573af3fa · inbound
OrderMoE: An expert similarity driven distributed edge MoE inference Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.