Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:53:58.740609Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.11217.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:53:58.740609Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:17.763102Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-15T23:00:19.435484Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8da072c8-5949-4691-b6a2-fef7e5d32b8c · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flashinfer: Kernel library for llm serving
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fa5341d7-fbc9-4a3d-bbef-2b1e985a54e5 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009a2d51-8c25-4fdc-a176-d01a4b12d137 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm in a flash: Efficient large language model inference with limited memory, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9967325f-6af9-423a-b845-3c356f178664 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c8aec7-1d4f-4c0a-95cd-375f17aa571f · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating large language model decoding with speculative sampling, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2e38fb-5ad5-4bbe-ab60-b9f11d7fcc7f · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Lifelong language pretraining with distribution-specialized experts
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1a75a12b-1a28-4e22-ac01-7b08d93f3396 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Spreadsheetcoder: Formula prediction from semi-structured context
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cb96245e-f471-4cfe-8707-587d6323d85d · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce5be6e-b065-41b1-88db-eda92ce41cc2 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Generating long sequences with sparse transformers, 2019
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb32e99-611c-4b2a-b42a-16685f72a4aa · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a11b14a-0f39-440e-86f8-c9814b8fa096 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 68aa2d8d-7e73-42d5-8009-6502458b876e · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7cb30671-87f3-47e3-a01a-61078df5b0cb · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Glam: Efficient scaling of language models with mixture-of-experts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41d24f27-b2cf-4e63-88ba-f6b921fbf4ff · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eac09d1-0975-4c86-9970-0e15d2df344a · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference of mixture-of-experts language models with offloading, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8fe4204f-144d-4044-8fdc-8b1f16d15c1d · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Dap- ple: A pipelined data parallel approach for training large models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 43ead05a-ff96-4587-b4b1-ce2d915667f8 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4e0e2679-1602-4ac0-a4db-611b73ba8ab5 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5a6805-341d-47e0-ac34-f47d35eca29d · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hugging face accelerate
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 42c246a0-b8cd-48f7-b8f5-c3d9d776cfe8 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Intel(r) oneapi math kernel library (onemkl)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ed49b8cd-f4f6-43d6-9ad9-1a9a7b7870d4 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Jacobs, Michael I
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2cddb5ee-09a3-4e40-a6ba-974e76d19948 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixtral of Experts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d529eedc-00ec-4315-ba86-75a4ad3b9ed9 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hierarchical mixtures of experts and the em algorithm
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e2f8a2-b873-41f7-ae23-52fc6bd3699d · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Christo- pher Ré, and Azalia Mirhoseini
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 27c18c25-64a7-4d95-96f2-a8546200e836 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bf8ed4c5-ecf2-4a2a-a662-4c00a31e70b5 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient memory management for large language model serving with pagedattention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b60c6726-b44a-453b-8d85-dd8ad830cae8 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50cc8e15-4e9d-43c8-bd38-bf08e7a406a5 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference from transformers via speculative decoding, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 854a1e9f-aa4b-4011-9d9e-4a05206611dc · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Holistic Evaluation of Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c52fa16-3192-49a1-959a-6a51a5ec991c · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ab0a877a-ac83-429c-9448-a8b5ceeeed51 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Ion Stoica, and Matei Zaharia
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ddcab18-85f7-4c14-b8f9-f58b2fed74fa · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs https://mistral.ai/news/mixtral-8x22b/, April 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 98ae3737-33c3-4c1c-acb9-a09cb2ce5d7a · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Can Foundation Models Wrangle Your Data?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa36d69-df5c-4da8-ba84-fd06f4860082 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pipedream: generalized pipeline parallelism for dnn train- ing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0f061996-066e-48ae-b67f-4d82d6e4132a · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient large- scale language model training on gpu clusters using megatron-lm
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcac705-de64-4bce-89a2-325d5c149bab · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pytorch: An imperative style, high- performance deep learning library
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48a4139-6f90-4c9a-b9df-fa69800813b2 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficiently scaling transformer inference
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3c40886d-2f0c-49a8-ba95-f17e086b9b8d · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Int4 decoding gqa cuda optimizations for llm inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 27ca07b3-ed1d-4218-815e-0adcd9ae4157 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating transformer inference for translation via parallel decod- ing
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 529a1180-301d-4b3f-b41a-d892698429ce · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d347d1-5518-461c-afb6-97737c26e6c1 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d95062-058a-4ec2-9761-03a9e6ee0b21 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe41c939-776a-4a03-ab5d-d5bde01a1396 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 01f3f65f-3a72-4fab-81e1-462b883fb1c3 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Blockwise parallel decoding for deep autoregressive models, 2018
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 65f40479-2123-4938-a987-30448f8f6236 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc4e239d-4295-49b9-a2d4-0f7d79d70e05 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Introducing dbrx: A new state-of-the-art open llm, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8512b705-473b-4585-8a91-61a75c353df6 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs LLaMA: Open and Efficient Foundation Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f6bb67-35da-44ff-9b32-512a4dd9306a · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Patterson
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 426a9ebc-8270-4434-89e8-8c451dcce4e3 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4d0a4246-cd10-47af-a22e-b2ef80cf4cbb · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92fdc5f8-79f9-45cc-8fce-0ead5940fce9 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Orca: A distributed serving system for {Transformer-Based} generative models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091b585a-d928-4c63-9bef-ad4582ae5f35 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm inference unveiled: Survey and roofline model insights, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96114ba4-97b0-4a2e-82ab-d502a5f31661 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs OPT: Open Pre-trained Transformer Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf02297-787e-4ff1-b456-6570380d7f81 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs H2o: Heavy-hitter oracle for efficient generative in- ference of large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 91bdd8b7-1edd-4155-b51d-1767b4b4eae9 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07fb519d-789a-43b0-a989-ce9514a7e432 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Clark Barrett, and Ying Sheng
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 859ce6be-be1c-4784-b280-8697531af128 · outbound
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixture-of- experts with expert choice routing
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · inbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.