Pith. sign in

Paper Citation Record · LEDGER

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing

As of 21 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2501.05313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05313 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:20:53.820333Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:44:35.764432Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:59:52.416593Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28a1cd29-1036-4468-a21d-7dd32bf0100d · outbound

This paper cites Ditto: Efficient serverless analytics with elastic parallelism,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Ditto: Efficient serverless analytics with elastic parallelism,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.206088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.688468Z digest=sha256:c793758f8988fa87fd1e30ea03be9737fae56475c58382d07f6648a4638b9a9c

Observation 3e45feef-c700-4a05-b636-adc00008ff91 · outbound

This paper cites Caerus:{NIMBLE} task scheduling for serverless analytics,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Caerus:{NIMBLE} task scheduling for serverless analytics,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.198006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.692333Z digest=sha256:aa891de0971f7a006c69273c74512b3d95e3f04f351110dc51b2cc927d351237

Observation fe5cba69-ecdd-47f5-8ce4-aa8f9ee0bfc8 · outbound

This paper cites Gillis: Serving large neural networks in serverless functions with automatic model partitioning,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Gillis: Serving large neural networks in serverless functions with automatic model partitioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.190524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.695981Z digest=sha256:655a3a8adb5395a033110a5512de69e46b097a2f9d54ec11dd8d97f4e231bc3a

Observation 26d66836-d16a-4308-b1d4-def24f09810e · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.180125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.698907Z digest=sha256:7b2a7e8b8bacac2744b6e17c70846ce099edf969b1772aad907c47acfbcc8970

Observation 99395268-1944-46b8-a81d-754c01baf32f · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.170509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.702762Z digest=sha256:38746520af0ecd1d75c5380c14b06a1f392d590fdfec96bfd20c7dec394fd023

Observation ba61c4e7-9be1-4e69-8154-0960f92433ca · outbound

This paper cites Tetris: Memory-efficient serverless inference through tensor sharing,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Tetris: Memory-efficient serverless inference through tensor sharing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.163119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.705949Z digest=sha256:dc224cde9e9851e131cf03ccf75e9d7e884de280501667238b18212483afa943

Observation c01bf16b-2b73-45d5-a259-a6d8a6f9d399 · outbound

This paper cites A holistic view on resource management in serverless computing environments: Taxonomy and future directions,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing A holistic view on resource management in serverless computing environments: Taxonomy and future directions,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.154377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.708796Z digest=sha256:03ea9108b02e76528f177b0986c503180b48a89416ce63ca6638fdf0d69a03c2

Observation 07159989-1728-4afa-9574-d80ce7352bb8 · outbound

This paper cites Serverless computing: What it is, and what it is not?.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Serverless computing: What it is, and what it is not?

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.146670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.711450Z digest=sha256:e9cd1b6cd81312370937509fcc65bca33e6acf0e0fdfee3dd8a820b2b983d6cd

Observation 096bfb59-57fe-4def-acb4-9f993500f1e5 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.139081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.713940Z digest=sha256:06988f3101f831ad3b41c5b75c30964dfe579ed9869b227df73efdd6a11513c0

Observation 778ab619-c3ee-4735-9725-ab3e69d6a83f · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.132700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.716540Z digest=sha256:7632cfb986eaa61d3d9be8ff945ae28ef2dac6d93fccae98b77d8adb1a02c0f9

Observation bf99ae0d-d527-4a27-9bab-e373c7f9edc3 · outbound

This paper cites Distributed Inference Performance Optimization for LLMs on CPUs.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Distributed Inference Performance Optimization for LLMs on CPUs

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:20:53.905278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.719084Z digest=sha256:30eb5ad2733165d85dfc56ff548e24595aa96e9a34a050b5c19a40f933117a08

Observation 09564e26-0e03-4d26-a745-253c77293041 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.722179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.722179Z digest=sha256:a90d58599d0183c6e2e8b08fccd9f0279d3b8a069e8fc7f85d0b498e4d03afd0

Observation 6e56372c-08fe-4be0-a6a5-36501605d03c · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Glam: Efficient scaling of language models with mixture-of-experts,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.121134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.724784Z digest=sha256:82dc851bd2780e539762a8281e49c47c8b56b605c4f731ceaa95b9b247873971

Observation e65c72d1-d112-4835-aa36-c38756206dc0 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.727400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.727400Z digest=sha256:7fa208d981333a88873205da6786e4151da1f29c3a01e42eb4630bb786c0cc56

Observation 760e01a8-e570-4ce1-b8fc-d330990656e5 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Accelerating distributed {MoE} training and inference with lina,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.729687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.729687Z digest=sha256:4edb017fb14f11af18cebb0b5bab5929a3c2831c6e71b0b1c48b5a8fb64fd67b

Observation 1871a911-8349-4889-b921-75d298cb4a76 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing FastMoE: A Fast Mixture-of-Expert Training System

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.732254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.732254Z digest=sha256:a6db1ae4eec57622526638f0e1f8367db2cb24e937380aa69f0c6a44b838a604

Observation de91a24c-ed2f-4448-96a2-3a774609a711 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.734548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.734548Z digest=sha256:34eeb4715690a389102cdfb7794fa826dd722a3c61b2d5c8cd9632e64f77e6d0

Observation 28e5439e-d17d-44b8-a7c0-db784630ec65 · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.736530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.736530Z digest=sha256:f6ab17394edcccd2662a47025628e14cf8dd74fd6ee93ba8ec2dfc1978001715

Observation 5e712609-ee81-4e3d-bbb8-1e9ca63a16a9 · outbound

This paper cites Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.738545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.738545Z digest=sha256:d8f0897680ed45560a133c430fbed0e3fcf0e23a07d50c89eb017f3adccbf9f7

Observation 653b872d-92bc-4d5d-9667-367355095c11 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.092651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.740738Z digest=sha256:2e6599261769413a81ad40749bd830418d5c6ec5f5504bd39385d506b6f800f3

Observation 99ff8aba-e580-4af3-b99c-aed1034c1c36 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.084261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.742664Z digest=sha256:b00d94236781762f66dbfc6d92dc8b6072ba0c273892f705bf3862428327d344

Observation 15594768-dd40-49ed-9643-3b46fb68bb18 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.074986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.744863Z digest=sha256:c67a4198c8c4c9ad4011509df76dca2da599bb3906a84a8f1e298f2311ef0507

Observation 48688f29-c12b-487c-a2e5-fca4e8c29671 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.746986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.746986Z digest=sha256:a9ba294e8ce813ea23237e58fc10703e35254d8e697c992d6a8fbec5f764efa3

Observation 594b7819-444c-4ac8-b55e-eb3a76d8db7f · outbound

This paper cites Mixture-of-experts with expert choice routing,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Mixture-of-experts with expert choice routing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.750283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.750283Z digest=sha256:544604946dc95fe2b534f5d7669c10fb96b68873eba7b80aa709ce30cb1b97c7

Observation fdc2c165-bbde-4e6c-bcd4-b8e8ec968d84 · outbound

This paper cites Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.753021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.753021Z digest=sha256:2c1ef4055eaecae2f3a9d8fd2b2442467462879f73bfd043ea54d2eb30fd2021

Observation 7f9b4138-b135-464f-a7e9-2c52aee0f78b · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.755719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.755719Z digest=sha256:36b228254b81346e4771ae8d51550f55f778312ffab9508d88341f54ea28c384

Observation 69bb6db3-0f94-485f-be87-68e5341f68f4 · outbound

This paper cites Amps-inf: Automatic model partitioning for serverless inference with cost efficiency,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Amps-inf: Automatic model partitioning for serverless inference with cost efficiency,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.056415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.758161Z digest=sha256:289646d6f9612d010339f9f8877133b46f116b251f60dd3f3f3042871a95c630

Observation 3c5ca793-feb2-4e62-898b-276f84cd90cb · outbound

This paper cites Batch: machine learning inference serving on serverless platforms with adaptive batching,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Batch: machine learning inference serving on serverless platforms with adaptive batching,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.047384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.761812Z digest=sha256:f3f97c2498a80cc45140c465eab4890b380b5167d3ed87d158ffe1dac35a8cb5

Observation f8c67685-34f3-46d6-9b76-5ac5dcfb4278 · outbound

This paper cites Infless: a native serverless system for low-latency, high-throughput inference,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Infless: a native serverless system for low-latency, high-throughput inference,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.038391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.765245Z digest=sha256:db0246283beea3e793297b5428e19405e744076ceea80332eb2c4975c94bd059

Observation b0610e7d-8f36-4f1a-815f-4e051742a95e · outbound

This paper cites {ServerlessLLM}:{Low-Latency} serverless inference for large language models,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing {ServerlessLLM}:{Low-Latency} serverless inference for large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.767668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.767668Z digest=sha256:dfe5b2bde2e1ee6d1751986acf3d3f0bca35000423e461509cab3a9767576500

Observation 5b4c526a-864f-4e33-9c8e-2945b4d8f9d8 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.024839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.770047Z digest=sha256:0b46ddbd662001a396da753b3938ae50e2f8e13428c38a47aa99957afececbc7

Observation 1579174c-405e-4444-8ef9-d2b4aedad689 · outbound

This paper cites A machine learning approach to measurement of text readability for efl learners using various linguistic features.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing A machine learning approach to measurement of text readability for efl learners using various linguistic features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:54.015584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.772660Z digest=sha256:5cfecff12f59598ebee45a25df7e0ddab74769b94aee9cf711b829bdb2a0e281

Observation 7084c89e-72ad-4a03-9345-e348ee5647be · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.775335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.775335Z digest=sha256:9f162455dacd1c9b79492b481cc15fed80ccc62b49c8fc03ba73e82c77c89757

Observation 7e9af875-223b-4052-850d-21d96502903d · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large- scale moe models,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Prophet: Fine-grained load balancing for parallel training of large- scale moe models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.777604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.777604Z digest=sha256:08633418ac182a24a8ae7da17daa5d3ccf8aac843b250098c07d12d43ff6dc8a

Observation fb91f1b8-60fe-49c7-9716-432843758ae5 · outbound

This paper cites Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.780065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.780065Z digest=sha256:a24fecdfe2017ecbdb41ebef1575fed39bb214fbf903216a07ae470ba6be10d0

Observation eb296c5d-cd3e-4424-a03f-fb5afbabead1 · outbound

This paper cites {INFaaS}: Automated model-less inference serving,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing {INFaaS}: Automated model-less inference serving,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.785064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.785064Z digest=sha256:319efaea18d3258af0f3f5808b758a33279cdac60dfc11556e792dfc43c70979

Observation 6dbe615f-bcc1-47c9-a3af-26ac54cbfe0e · outbound

This paper cites Global solution of non-convex quadrati- cally constrained quadratic programs,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Global solution of non-convex quadrati- cally constrained quadratic programs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.991613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.787505Z digest=sha256:c9af970cba3c41200d3567f12807bd8928f65dbab62fc3b50d498648a4fda594

Observation 303c83d7-be99-4ec7-8199-33ab62967382 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.983731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.790009Z digest=sha256:ebe95f0163c9fe5200574b6b7a824ce4784e02d2f86c5e9fbcfb8ee3acc42a80

Observation 9bb53892-2442-4d78-8c8f-cbcfb9f07d77 · outbound

This paper cites $\epsilon$-shotgun: $\epsilon$-greedy Batch Bayesian Optimisation.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing $\epsilon$-shotgun: $\epsilon$-greedy Batch Bayesian Optimisation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:20:53.863658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.792732Z digest=sha256:e9a8d3cc42848c25352a453c6bddabf753336b6b97d02d29e4a7642dc6af43a7

Observation 1bcbfc19-3070-4e08-aa3a-c51067f11c0b · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.975510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.796236Z digest=sha256:6a8acf9cd43cad4452cf29c1d6fb3067828f648890a11a19c5019c7fc9a28306

Observation 24529845-232c-4099-a163-618ee6806849 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.968575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.798604Z digest=sha256:738c013f6e9a773b277492a3ade3bf42cd6c4e09345f52f3bd8a50c2128345aa

Observation ead71402-e17e-48fa-a601-038ee0b7ae21 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.959714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.800919Z digest=sha256:9a3a0e3b604b06fb7a2e6ca9b32f97afa61a438fb066436a839765faa3b11b48

Observation b2f653a3-a508-451f-abd6-6aded08068fc · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.803064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.803064Z digest=sha256:c74e23ffc474745b18b4d492b48e5f9d23c846587ccfc20dd3b58d68aaeae40e

Observation e3b72791-f479-4dee-9829-36b27256d30e · outbound

This paper cites Language models are unsupervised multitask learners,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Language models are unsupervised multitask learners,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.805586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.805586Z digest=sha256:cd292cff4b83c4649bc0012ce47219159d6dfed26b9190b7632f4202b9d49e53

Observation ea25869d-652a-4594-931a-5b7ada24eace · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing bert2BERT: Towards Reusable Pretrained Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.807838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.807838Z digest=sha256:99f5f3795f85c2955459540ef1828fa38b2a6bbd9819de2462933e30ce23f661

Observation 106f0425-5f40-4800-904c-1f2a4c8fdbe9 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.945855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.810118Z digest=sha256:ad169b0b28f68270f6fdf6d9129f7846362981f79afcaf45c85e74de503a6381

Observation 7a620941-e904-4254-b74c-44e0e02d5fc1 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.936922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.812464Z digest=sha256:76c8b2e9bed90b4eec09918fb663af22c2e7ddc65edd62dfe5806c1be9943849

Observation 53f673ed-0667-40e5-8f70-5aa2302a08f6 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.927766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.815253Z digest=sha256:1cfedde204937168fd6d3a297df9cfd52b08784e468a482f20dd14435351f8d0

Observation a1909807-aa62-4f5d-9074-2ac06c8b3ae5 · outbound

This paper cites Algorithms for hyper- parameter optimization,.

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing Algorithms for hyper- parameter optimization,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:53.817806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:53.817806Z digest=sha256:bfd36508f996d232d81c4d290528f48f2008be1e531d13f4c00d64836004494b

Observation 7a494794-6cd5-4308-92c3-f85f86c72126 · outbound

This paper cites [Online].

Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing [Online]

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:20:53.914616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T21:20:53.820333Z digest=sha256:e77cc608b9fab47f27e6dc4f684559aafab377077e42e33e116954197447be0a

Pith citing papers

Observation 4f07f43a-5d99-407d-8ba2-9e134a82eb84 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.764432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.764432Z digest=sha256:c58e3aba92e565b05998932f6e28064967e202e7f9d1fb4fa5fd593431bd8924

Observation a55f8727-180e-46d9-b188-4c0d8ee21fee · inbound

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning cites this paper.

Hecto: Modular Sparse Experts for Adaptive and Interpretable Reasoning Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:59:52.473973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:59:50.912662Z digest=sha256:93fdd48e03f81cfff6e07d0177b7c19e0597c0f9321631b83418f8fe73b912ca