Pith. sign in

Paper Citation Record · LEDGER

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

As of 1 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2508.12851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12851 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T22:52:39.416575Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:56:19.766584Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:07:44.131237Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact13
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fb10cc8-a015-4a03-974c-0d137f463bd2 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.049412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:411d0c0352a53679419e1785338041ddaa2ed29d24e9ca6d9841ef943aa067af

Observation a87e8e74-f757-4436-a62d-6ad928348af8 · outbound

This paper cites Mixtral of Experts.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mixtral of Experts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.955108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:eb69336d0faf2c0a75f62edfcb63c178b455d27a5c948e981e9bef7fa9dc46b4

Observation 2d376eaf-8462-4e61-b09d-5c46021642ce · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.936457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:1a0e0051b4c31f0be51b0686c79eb8175123312a50899a084c4c3bb54b6776c2

Observation 23bff876-6a79-4ee9-bd75-5de700e664aa · outbound

This paper cites Geforce rtx 40 series.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Geforce rtx 40 series

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.985010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ae8b22a537c0ad64e60a7416fc35fd429bbb755a3afd1624461c1ee34c67bbe5

Observation 19c09be9-0738-434a-94fc-4d3c69197c24 · outbound

This paper cites Gpunion: Autonomous gpu sharing on campus.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Gpunion: Autonomous gpu sharing on campus

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.950085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ca3b6721de6a36f1ec41b8d4f99e0bbe97876beedfa3489c5e94d30d6a478d8e

Observation 79250885-432f-42c5-807f-b29e73cd6a49 · outbound

This paper cites Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.874632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:7872025954eb0491167102133a4d437c3f8e44a6eec8a79f93ca181a5fc09b43

Observation 5a23cd60-650e-4d0c-a848-4dd8e1d088cf · outbound

This paper cites MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.924894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:c4434513432e601cd5b2ef37337228a87d78bb795c2408b6c4682b6c2bb6fb1a

Observation 28f25519-5370-4395-94e5-818027fc61df · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.942278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:475faafe18cfdf72e228d655e16ec4f1c2c2fdf546c99a8a335e63f375c192d0

Observation 352ef6f0-0f0e-41ad-aa33-87cb1bd15739 · outbound

This paper cites Expert Parallelism Load Balancer (EPLB).

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Expert Parallelism Load Balancer (EPLB)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.017792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:400926b75fec5d283b0a666385a76bc5eab148a9ece4d499b80953962b54007a

Observation 70e646ec-bad1-4f76-ad27-0e0e41e1e4ce · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.995958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:209337b0f1a73f2a9735de13c474efd97734843027f709d0a118c080ef038a20

Observation 42834221-0771-4a8d-a968-82bfcca28b9d · outbound

This paper cites Moe-infinity: Efficient moe inference on personal machines with sparsity-aware expert cache.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Moe-infinity: Efficient moe inference on personal machines with sparsity-aware expert cache

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.028077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:82985584d22384afee843855d3c7bd2ddd28030da52e9ef663211cd012f85535

Observation 88c9b676-31a2-46a5-a55c-19bee0e0cc4c · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.981756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:10ea161e5f188db160f1f7a15563dee319e3669bb44980a8d981546cee20a8c9

Observation b5426018-7783-4318-9668-586020acbabd · outbound

This paper cites an unresolved cited work.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-18T22:52:51.968571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:f4658b22851ec41d3fe9e5aff46e4bf9fb39b4d662e36d235a97bfec3827855b

Observation f80b0c76-ef3c-4d15-9b6a-9029202c8e7d · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.919074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:8e017021fd1e8ff6916701a24b7952c861db80ca8cb141c0b02d37070c26eb35

Observation 33b53e28-96ef-4d18-8419-7549eb4b4098 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.881624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:b62916e363644de8699e861133434b88e89f13727b720cff8daa1ed792132081

Observation 1653027f-19ef-47dd-a00b-ef4e2d8659da · outbound

This paper cites Pointer sentinel mixture models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pointer sentinel mixture models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.062287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:1aa17efa43a36c47b0498a975b5484ddd38397d8cb0d04f16f1977433d907c1c

Observation 162d2908-54e7-4457-b0be-4ee977d78c01 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement TACO: Topics in Algorithmic COde generation dataset

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.908258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:1d0a651a891cba7c700d90525f3e964cefcd1116493d039b2388ee3a31686aec

Observation 7ff6cb8c-caa1-41f1-add2-96631a92d043 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.992346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ba432771b72f45daafbfaeaecb5cebc2df5efe992bb70a733a18ae6cabb479db

Observation 27729f97-78dd-44f1-a1cf-f9b3cba67c3c · outbound

This paper cites Joint application placement and request routing optimization for dynamic edge computing service management.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Joint application placement and request routing optimization for dynamic edge computing service management

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.022246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:28d0ca892274f56609f4f0dbd2f22acdd0053d7bc0238605eecba53875a14acc

Observation 46ec477a-dcc3-4e53-aa45-6acc8f0b077d · outbound

This paper cites Task placement and resource allocation for edge machine learning: A gnn- based multi-agent reinforcement learning paradigm.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Task placement and resource allocation for edge machine learning: A gnn- based multi-agent reinforcement learning paradigm

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.973069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:95ca3936d4d4a65202fbedf886b6ac9087e97276ffb758d4954711e6f1f1f830

Observation 94421f3a-aee4-4696-bc92-8a53e4996ba2 · outbound

This paper cites Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.044212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:4348ba10380cf3de4f4e4f46eac356fb6bdcbae18b9b6d8b8545e9cd487546cf

Observation a4d1369b-13fd-4c1f-bf09-04ef30249892 · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.057367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:947a356938b318ed640a2236ba03bef53201fafdc48e57ebe057f731914157ce

Observation 295ed844-d8fd-44c3-abae-acb58ffc8c93 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.068135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:87930c54e9706dd9a95ad491878700661704d7cc3b3f090f97428e7c3c32262d

Observation a59f0663-8ab0-4e3e-9bba-eacf0386e657 · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large- scale moe models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Prophet: Fine-grained load balancing for parallel training of large- scale moe models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.038428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:016b72614a4d262f3cc3b2dec7ed083397bf5d587f349e3bb71e51a52a908c93

Observation 2bbbc822-bdac-496b-b4ae-c8a673107c8d · outbound

This paper cites Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.888643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:875cc6bca508a5611092c3f02ba54f8838a0ec5f7025dfb0ab31e244c30b062b

Observation 8223bee1-211b-4122-962f-abcb7fa044a7 · outbound

This paper cites Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.001256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:7c8270138561df9510bf344692a9efa0330403900c7bcc8dd9cf152f73821477

Observation 189a01be-0957-4ce8-acbf-ca8f9b7ad920 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Accelerating distributed {MoE} training and inference with lina

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.009910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:fc41352dcbbd1384815703725d57af7dd5942fe16d29d2abee329cef28004389

Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:52:51.896416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:7e8964fc8272ebade201d57a04f2d98617dbb5bc25f4675ab1d1c31cac3c70fc

Observation c8181ca5-3127-45e8-8eeb-d2c31f7b72f3 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.931759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:75e33a20fbb24c8b5932b9f1170f2b8f2671f8235312b248a36b8ade49235a31

Observation 91c32161-0484-4cd7-88be-11e5b78f1e3f · outbound

This paper cites AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.914141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ce242925828861fca1b7183b8eb2548723a3f6ab40ea1bb695d4f2950c0f24c4

Observation 63ab3982-2d7b-4a8d-aac7-51adcfe8dbe5 · outbound

This paper cites Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.988367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:5333a17a710233d73c8f7fac6135484cc33e0f993893b342ef7611ba2e8bc82b

Observation 0a824cf5-fbcf-4467-af21-b6c80fe85591 · outbound

This paper cites Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.977772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:f37fe474aba767e1ba8de732c5626d1050e7c57e665275e17c6a69006c57bec2

Observation 8926653f-8995-4258-9dae-2495de2d0545 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.902779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:3083805a5f07055fe6f5164193e1bd6183e7aab05519c4344c7a3d962cef942f

Observation df5f0a61-7f5b-4617-b094-43cf6ee7c39e · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pipemoe: Accelerating mixture- of-experts through adaptive pipelining

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.033806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:39118e6336463f328ab39e27c1f00abd8c0bb727e177c339564eb4e2eeb7f543

Observation 6b75f761-5b45-4cd4-bc95-07f70bce0efb · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.013559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:db603ec513f658a3a26069c5a4eb1bb6e5e587929892cc80386a1f0eb1c1c2e0

Observation 5a664ba4-c476-4d49-b517-72779e4d8c19 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tutel: Adaptive mixture-of-experts at scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.005928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:0d22b7d6a895254dd50d448aeddb6db2f695710a87ea6ac3ff13edfe8e574c2f

Pith citing papers

Observation b7337072-9053-4129-8709-020fe34ac8a7 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:44:54.930922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-08T13:43:17.950000Z digest=sha256:7b33759eabe4a55ed962609ce0b928c00f8cf8e9e145315ee094712f5f9f5ed7

Observation 67f9a1aa-d74f-4562-b519-63ae177f5500 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:44.161601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-07-11T01:06:46.426582Z digest=sha256:1d686adfdc82b44e4d4f1f2f02eed27e47c906299cefddadec313e7e59ee47ac

Observation f0ba5b1d-6ddb-4152-bc87-f0d5573af3fa · inbound

OrderMoE: An expert similarity driven distributed edge MoE inference cites this paper.

OrderMoE: An expert similarity driven distributed edge MoE inference Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:19.766584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:19.766584Z digest=sha256:e7b400e6151557503e27afe3dddb623bbdb1360f8386295bd910aa9808502125