Pith. sign in

Paper Citation Record · LEDGER

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.11217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11217 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:53:58.740609Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:17.763102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T23:00:19.435484Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8da072c8-5949-4691-b6a2-fef7e5d32b8c · outbound

This paper cites Flashinfer: Kernel library for llm serving.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flashinfer: Kernel library for llm serving

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.210066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.569013Z digest=sha256:e33d07a1e6f8a83692ed1e4b592d5d141bb4c6a6aee55dea86ad7a495e7892d2

Observation fa5341d7-fbc9-4a3d-bbef-2b1e985a54e5 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.572830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.572830Z digest=sha256:0506d97b5916a37066c5189624de625535e136f49fc48da216c9f509a4afa985

Observation 009a2d51-8c25-4fdc-a176-d01a4b12d137 · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm in a flash: Efficient large language model inference with limited memory, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.201165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.576262Z digest=sha256:85a67d39739c5458fe8d0a0e981c2dce1a1ada026535fd71d0d87f9d037ec5af

Observation 9967325f-6af9-423a-b845-3c356f178664 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.579622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.579622Z digest=sha256:3b17e2ca8313e7e94511896736732961573f4c2c9eadf721f0f63c0261c140f5

Observation b0c8aec7-1d4f-4c0a-95cd-375f17aa571f · outbound

This paper cites Accelerating large language model decoding with speculative sampling, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating large language model decoding with speculative sampling, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.582688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.582688Z digest=sha256:680b27144bffddba3a54a0d42ddbc9e17c672134a550bc65030fec2eaae7a8f5

Observation 6f2e38fb-5ad5-4bbe-ab60-b9f11d7fcc7f · outbound

This paper cites Lifelong language pretraining with distribution-specialized experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Lifelong language pretraining with distribution-specialized experts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.181927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.585827Z digest=sha256:1114bc30278c332ad7914a3aca6704f759bd4df2a7aa0fd29a0055287ce05b3a

Observation 1a75a12b-1a28-4e22-ac01-7b08d93f3396 · outbound

This paper cites Spreadsheetcoder: Formula prediction from semi-structured context.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Spreadsheetcoder: Formula prediction from semi-structured context

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.172382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.589148Z digest=sha256:d74f3e22311876840691612435e5ab53c5857bb2f2884697647e183e76f26666

Observation cb96245e-f471-4cfe-8707-587d6323d85d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.592078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.592078Z digest=sha256:09996ca9f5f57fd042efb591ed3faf57b1506b2a04aa330ee8e331f80988e470

Observation 8ce5be6e-b065-41b1-88db-eda92ce41cc2 · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Generating long sequences with sparse transformers, 2019

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.595355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.595355Z digest=sha256:b679e2312e677036288d4b8a3c925dd754e5773ca699e1225634b9bcc23fe52f

Observation beb32e99-611c-4b2a-b42a-16685f72a4aa · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.598074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.598074Z digest=sha256:fcebe2d2ea08d6e38b77a7658c44e54f5720d28ac99dc3f368b568e631e76a99

Observation 7a11b14a-0f39-440e-86f8-c9814b8fa096 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.158305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.601425Z digest=sha256:ec13c43d10ec7a1ea7139635e3118e3541b0112f6a26bbfa7801879ba62a70a4

Observation 68aa2d8d-7e73-42d5-8009-6502458b876e · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.149069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.604536Z digest=sha256:975c25038145c9cd8795a06fcbb3f4895ab943857524c538e1d0b6b6b8f0a8ac

Observation 7cb30671-87f3-47e3-a01a-61078df5b0cb · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Glam: Efficient scaling of language models with mixture-of-experts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.140456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.607616Z digest=sha256:34e52e1650b8b279ef79c27f477fe5be8106ccd4e83aa018059004dbebc84c83

Observation 41d24f27-b2cf-4e63-88ba-f6b921fbf4ff · outbound

This paper cites The Llama 3 Herd of Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.610574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.610574Z digest=sha256:07d0c2dbe4fe3557edf9eefa9a0cda5661a0e2fe2dd7c6fc3de7489896dc720f

Observation 5eac09d1-0975-4c86-9970-0e15d2df344a · outbound

This paper cites Fast inference of mixture-of-experts language models with offloading, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference of mixture-of-experts language models with offloading, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.131718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.613730Z digest=sha256:7db3ed32c779083c6204f69b95b11c3f816e51101650a02e0a9b458eefee3c97

Observation 8fe4204f-144d-4044-8fdc-8b1f16d15c1d · outbound

This paper cites Dap- ple: A pipelined data parallel approach for training large models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Dap- ple: A pipelined data parallel approach for training large models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.122947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.616760Z digest=sha256:e7ce6c50605f5d58d8b104520c6548e0b3f4aed916f381131b3cc3b4815f5385

Observation 43ead05a-ff96-4587-b4b1-ce2d915667f8 · outbound

This paper cites Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.113719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.619628Z digest=sha256:14694ab4673e0f32a2dd8fc4d9a6da849f58d45f1509b2b510b76065fc7c69c7

Observation 4e0e2679-1602-4ac0-a4db-611b73ba8ab5 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.622530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.622530Z digest=sha256:3f33c35594eed99966a869e17e18e147b10ac74e58bbb507e72043d740635f16

Observation eb5a6805-341d-47e0-ac34-f47d35eca29d · outbound

This paper cites Hugging face accelerate.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hugging face accelerate

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.098032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.625275Z digest=sha256:bd98b6af8c4a228210843ef537884f4ab96f666fa32a2d558e6290ddb93215ed

Observation 42c246a0-b8cd-48f7-b8f5-c3d9d776cfe8 · outbound

This paper cites Intel(r) oneapi math kernel library (onemkl).

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Intel(r) oneapi math kernel library (onemkl)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.088910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.628410Z digest=sha256:6680c95e38f998a1c09415eb52996afea332bdb059fc5bece7be4e9840f64733

Observation ed49b8cd-f4f6-43d6-9ad9-1a9a7b7870d4 · outbound

This paper cites Jacobs, Michael I.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Jacobs, Michael I

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.079406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.631285Z digest=sha256:b18cd0ff985adc1fe913f908f8bc10105b829a019d7fcab28d4e1ba3ad40f3cf

Observation 2cddb5ee-09a3-4e40-a6ba-974e76d19948 · outbound

This paper cites Mixtral of Experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.634169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.634169Z digest=sha256:5ea2ca8a3b9d304a3a207c08ef7b1eb8387d0275ef5518d98a3b96e94373f002

Observation d529eedc-00ec-4315-ba86-75a4ad3b9ed9 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hierarchical mixtures of experts and the em algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.637627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.637627Z digest=sha256:40fe3a290314b9e6d5be07a3b1fc36450c38ab0e2773daaae5a7024a2394a94b

Observation d2e2f8a2-b873-41f7-ae23-52fc6bd3699d · outbound

This paper cites Fu, Christo- pher Ré, and Azalia Mirhoseini.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Christo- pher Ré, and Azalia Mirhoseini

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.064941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.640464Z digest=sha256:c6480d1f043d6459205c4ce0b523072a3df1ca221fe6bc3cb094f377a0e57aa8

Observation 27c18c25-64a7-4d95-96f2-a8546200e836 · outbound

This paper cites Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.055522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.643198Z digest=sha256:6eb9cba826982793fc42f6afabfbc7cd97d3642fd8a5166883320fc0950508a0

Observation bf8ed4c5-ecf2-4a2a-a662-4c00a31e70b5 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient memory management for large language model serving with pagedattention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.646145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.646145Z digest=sha256:a537a04850a88e959cb3625e6101c68773ecf1229ab00ee134584ee627862c28

Observation b60c6726-b44a-453b-8d85-dd8ad830cae8 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.649016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.649016Z digest=sha256:7ea9c59661ebdcd90f71d81f80cf1cec02fe6afaa6d9866f22d9ecb8e6d0fd7b

Observation 50cc8e15-4e9d-43c8-bd38-bf08e7a406a5 · outbound

This paper cites Fast inference from transformers via speculative decoding, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference from transformers via speculative decoding, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.652418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.652418Z digest=sha256:b39fcfc63f2f01b900c5d71a5487b5bab89357fcb81e458268422b8c41c8fca4

Observation 854a1e9f-aa4b-4011-9d9e-4a05206611dc · outbound

This paper cites Holistic Evaluation of Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Holistic Evaluation of Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.655265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.655265Z digest=sha256:f54b29bd52f8b1d1f0979ee065cf0cac971b999c1b1694c36fb59b84a5089339

Observation 0c52fa16-3192-49a1-959a-6a51a5ec991c · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.036481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.659136Z digest=sha256:6121e7b5516564eafede780289e8dae9b6cfbb32534095e38fabc1fc0ebbc96d

Observation ab0a877a-ac83-429c-9448-a8b5ceeeed51 · outbound

This paper cites Gonzalez, Ion Stoica, and Matei Zaharia.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Ion Stoica, and Matei Zaharia

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.661945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.661945Z digest=sha256:f2e17a00368463c36a57dcfa0b2b201ac86b16b8369207a59c1072c8989b5713

Observation 0ddcab18-85f7-4c14-b8f9-f58b2fed74fa · outbound

This paper cites https://mistral.ai/news/mixtral-8x22b/, April 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs https://mistral.ai/news/mixtral-8x22b/, April 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.022040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.664995Z digest=sha256:2de7f65695db6064ab5e6be0801babcc68ffc97af5f869d0650a9f80dd2690da

Observation 98ae3737-33c3-4c1c-acb9-a09cb2ce5d7a · outbound

This paper cites Can Foundation Models Wrangle Your Data?.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Can Foundation Models Wrangle Your Data?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.667936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.667936Z digest=sha256:06d0aeaa48d59c8362bbad805a0ce576c14a1a60cc27272fa59413c7847bd4ca

Observation caa36d69-df5c-4da8-ba84-fd06f4860082 · outbound

This paper cites Pipedream: generalized pipeline parallelism for dnn train- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pipedream: generalized pipeline parallelism for dnn train- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.013260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.671082Z digest=sha256:e01ab0f2d0748baf9d80776108d0911d95bdc87acf912a41dd80b16d63dc2c0f

Observation 0f061996-066e-48ae-b67f-4d82d6e4132a · outbound

This paper cites Efficient large- scale language model training on gpu clusters using megatron-lm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient large- scale language model training on gpu clusters using megatron-lm

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.673854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.673854Z digest=sha256:cef4d7e64e0f1411e5af44f6187eeb6d39f85f712ec4d148e2adb410f917b16f

Observation bbcac705-de64-4bce-89a2-325d5c149bab · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pytorch: An imperative style, high- performance deep learning library

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.676884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.676884Z digest=sha256:1d64b84876f62480a666385fbbed57dc4cf0dd6dcf9a7bf5ceafae8147386d46

Observation a48a4139-6f90-4c9a-b9df-fa69800813b2 · outbound

This paper cites Efficiently scaling transformer inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficiently scaling transformer inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.994340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.679584Z digest=sha256:29a353061dd8ced39b2e68af76f4f39c9c37741d646b3ca0cb6713d94b9e3d9b

Observation 3c40886d-2f0c-49a8-ba95-f17e086b9b8d · outbound

This paper cites Int4 decoding gqa cuda optimizations for llm inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Int4 decoding gqa cuda optimizations for llm inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.985437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.682282Z digest=sha256:a51956a9576d584dc34678a381240d0c48bdc93744013c935bda4b419cde5d9d

Observation 27ca07b3-ed1d-4218-815e-0adcd9ae4157 · outbound

This paper cites Accelerating transformer inference for translation via parallel decod- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating transformer inference for translation via parallel decod- ing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.976464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.685155Z digest=sha256:622d6925b2fa9e44093790f629ee734e62bb522b5b59bc3b5c2072177ed197ba

Observation 529a1180-301d-4b3f-b41a-d892698429ce · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.688100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.688100Z digest=sha256:0df151d9bf7d0c113d1a3491f0ca91eb3982c1df775ec27e0e2207ad2fa963fb

Observation c9d347d1-5518-461c-afb6-97737c26e6c1 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.691288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.691288Z digest=sha256:5689edddf8669258e2a7ad535edb9630fc1a67102fb05f8b861a7135a5c29de3

Observation d3d95062-058a-4ec2-9761-03a9e6ee0b21 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.694423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.694423Z digest=sha256:e4afe8d82220a2987b9ed243a31913574a46df8a88895e1fdcb10305d072bca8

Observation fe41c939-776a-4a03-ab5d-d5bde01a1396 · outbound

This paper cites Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.961970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.697301Z digest=sha256:962fd64fcf1120405ceb94b5156396a8cdf2b6bdf577227c9ee8e4939e33b2e1

Observation 01f3f65f-3a72-4fab-81e1-462b883fb1c3 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models, 2018.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Blockwise parallel decoding for deep autoregressive models, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.951762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.700423Z digest=sha256:69f182b91b6538a3185ac547994ee811deaf31dd45e2a7fe5d6ce4ebee6df427

Observation 65f40479-2123-4938-a987-30448f8f6236 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.703312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.703312Z digest=sha256:404da0e96f8738dd2c5272ddc09cb105c1d5206b30efedd4ac516cad68190e4c

Observation bc4e239d-4295-49b9-a2d4-0f7d79d70e05 · outbound

This paper cites Introducing dbrx: A new state-of-the-art open llm, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Introducing dbrx: A new state-of-the-art open llm, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.942656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.706692Z digest=sha256:47843e342b09558a62bef30ad45186b6384bb28abf1048038d413dfbbcf78339

Observation 8512b705-473b-4585-8a91-61a75c353df6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.709940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.709940Z digest=sha256:45f97aff5f5ff993ec5264a725634ccface1452a577977df468f50a30496b6b1

Observation 82f6bb67-35da-44ff-9b32-512a4dd9306a · outbound

This paper cites Patterson.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Patterson

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.933257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.712972Z digest=sha256:b2986c84e7b93efff21a0c9945eaafb48ffa30246a675410391528409bc37348

Observation 426a9ebc-8270-4434-89e8-8c451dcce4e3 · outbound

This paper cites Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.924325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.715786Z digest=sha256:a6a443cdc9a7226b28bc1a1f00d42cd12ad458cd0a02ea9e50d6cb8ddd3007d0

Observation 4d0a4246-cd10-47af-a22e-b2ef80cf4cbb · outbound

This paper cites Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.914493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.719047Z digest=sha256:43870ad718ba041328bb2e9070d9ce6de34cc3b068700f3c91858378df12e8d1

Observation 92fdc5f8-79f9-45cc-8fce-0ead5940fce9 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Orca: A distributed serving system for {Transformer-Based} generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.722241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.722241Z digest=sha256:6ccf825b67e8e7f77d2b59862383a5d490e7924383f1999942e62b6cbb4302c7

Observation 091b585a-d928-4c63-9bef-ad4582ae5f35 · outbound

This paper cites Llm inference unveiled: Survey and roofline model insights, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm inference unveiled: Survey and roofline model insights, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.725433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.725433Z digest=sha256:ac4f98e434ef147b2bbb1953ec31ba60fc926ef8acd9c25935b5fb54f0c897ef

Observation 96114ba4-97b0-4a2e-82ab-d502a5f31661 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs OPT: Open Pre-trained Transformer Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.728292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.728292Z digest=sha256:fac964f5d5199146d1c453d86833d0fab7c90e42062ccccf16c2dbb69f394a48

Observation edf02297-787e-4ff1-b456-6570380d7f81 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative in- ference of large language models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs H2o: Heavy-hitter oracle for efficient generative in- ference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.894605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.731529Z digest=sha256:7c6a4c4fb3ad087917e0e634a51bdc65d500610cff249217a22683b263fa3573

Observation 91bdd8b7-1edd-4155-b51d-1767b4b4eae9 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.885475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.734590Z digest=sha256:8595867b413be586f1c8464dc835d34f11266c6dd8f696ea1d3386d9f763c39c

Observation 07fb519d-789a-43b0-a989-ce9514a7e432 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Clark Barrett, and Ying Sheng

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.737467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.737467Z digest=sha256:e347dfb967ccbd89fe78b865de476618cf0503d2b2892734df529e373ac69da1

Observation 859ce6be-be1c-4784-b280-8697531af128 · outbound

This paper cites Mixture-of- experts with expert choice routing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixture-of- experts with expert choice routing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.870916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T18:53:58.740609Z digest=sha256:12df1113b4225098df7df3374e8bc717aa22458097b9f60d4909ca9dc2e2fd88

Pith citing papers

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · inbound

FloE: On-the-Fly MoE Inference on Memory-constrained GPU cites this paper.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:7f97a4a8857106c095504bd0084a26e54127180720e26462a9b02749bd9c429f