Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:32.963734Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12417.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:32.963734Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T04:31:11.271729Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T04:33:03.804063Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1b83d114-1784-493f-9690-744ac60e2df6 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80582c9-8953-4a5a-bc28-185de4c14044 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59dddad4-1d4c-4ed6-b0c5-62a42d571780 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models A neural probabilistic language model
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71b4bcfb-9f00-4e92-afa8-ad95ab622190 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Datacenter power and energy management: past, present, and future
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8ead566-47c9-44d8-9279-225b98678ce9 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Language models are few-shot learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe6c22e-b416-4b89-bbbe-cbcd728c75d7 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Large scale distributed deep networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf21b206-e2b7-45a7-a376-5c26f5da9c52 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Quasar: Resource-efficient and qos-aware cluster management
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bad8eb6c-46f3-481c-80fb-f768b78f85bb · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1678eaa-a4de-4c35-b151-fd9c69584ff5 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46e54108-df6f-4716-a6b2-03918b500a69 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Megablocks: Efficient sparse training with mixture-of-experts
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 482cdcd6-4f65-4613-aab7-5327551c3ef7 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Character-based NMT with Transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4340967-aa22-4a69-aad8-fd6b8cf1e4c5 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models FastMoE: A Fast Mixture-of-Expert Training System
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e82041-bfbd-4914-b069-7b631afd185a · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5affdc3e-01fd-4096-a213-3589403847e8 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Tutel: Adaptive mixture-of- experts at scale
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb3db8f3-6b32-46db-b5a8-175c2bbc2252 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Adaptive mixtures of local experts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 880a320b-7746-4f15-a62a-0d8f087a5101 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Mixtral of Experts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a95c6e-66a4-40b7-ba63-0318e79fa7ca · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Scaling Laws for Neural Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4915bd77-dafe-4b2c-8794-96d864e8cee6 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Learning multiple layers of features from tiny images
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899c9652-4ce1-4421-91cf-c81dd3cd9781 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Subword regularization: Improving neural network translation models with multiple subword candidates
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9903387d-0a07-4e20-8ae6-13c96a2f4168 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a53ad049-1517-4573-a92b-deb66a80e2bb · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1b134cd-08ca-4b04-9b4a-9d45491791e4 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b49519fd-b6e4-4747-858f-eba378d20bbd · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Accelerating distributed MoE training and inference with lina
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d509d5f2-c70d-49d2-b1aa-6b4a4438f67b · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSeek-V3 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434cd7b8-ca7e-4e57-98d7-209d76973cb4 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Pointer sentinel mixture models, 2016
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee18e1d9-862c-4881-b67c-c18d9894008d · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95b2e0a3-21d7-4511-9f68-f2f545fc8063 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Pytorch: An imperative style, high-performance deep learning library
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3195126a-8790-45de-982a-aefa27955565 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 797a5ad8-591e-4927-a66e-8fffdc192611 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aef745bd-b596-4ae7-a9ee-7af36f6fdcb7 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Outrageously large neural networks: The sparsely- gated mixture-of-experts layer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 461ce789-0910-464d-b777-7519b71a159c · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1711014-cffb-4582-8751-962448885974 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Borg: the next generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3d7be92-1c27-4073-bfd5-2ddf0befb776 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Attention is all you need
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b515975f-470e-4213-941d-a8e93a5a0525 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Prophet: Fine-grained load balancing for parallel training of large-scale moe models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09193130-6c13-4302-a06f-cbb3dbe6ed3c · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Resource-efficient algorithms and systems of foundation models: A survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70c8a7b8-7217-4b16-a8ca-4bee9cf1fc76 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Qwen2.5 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13c2bcf-08ac-4428-969e-df9c5f0bc56f · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Harnessing the power of LLMs in practice: A survey on chatgpt and beyond
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33d07119-e7e6-4e3f-a417-77c4e826e6de · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25f8865e-cbd2-42ad-909a-9a5eeac9efb0 · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c367ef27-b1e2-4b77-9f1f-ef08e899122a · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e082bc-2146-42df-8ca6-c3ee6fb2de9b · outbound
HarMoEny: Efficient Multi-GPU Inference of MoE Models Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0685dec3-b41b-42bf-afb6-58de81036b73 · inbound
GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems HarMoEny: Efficient Multi-GPU Inference of MoE Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.