Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:24:00.465856Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2411.15419.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:24:00.465856Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0a280264-2777-44ee-86ec-191dfdd495f2 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Attention is all you need,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a04b83fb-5c21-464b-9275-38b39740cdbc · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7bdecca-5cc6-413b-8280-79330a94d8dd · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models are unsupervised multitask learners,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9301b913-d786-4804-a656-d5c6c6488c5d · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling Laws for Neural Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5496b08-8a18-4e12-88dc-df73a3efa4fe · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0bebad-8f71-4e95-9af0-e5d4172f6b1c · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f0ac79-2cb9-4f8b-991a-9287f134d4d2 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32743ec6-8d9a-4074-b121-3a21e783519c · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Tutel: Adaptive mixture-of-experts at scale,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47cd25e9-1ff7-4056-876b-015c3601e47c · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d67466c7-272d-4c1e-a165-0ac5b345c237 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Janus: A unified distributed training framework for sparse mixture-of-experts models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af105974-1e52-46b4-b291-0f139b6f2d39 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Accelerating distributed {MoE} training and inference with lina,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e63e08-7b37-4d80-a0b2-5e79652addb8 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c50de7e-42a4-428e-9e76-6fda5d12953d · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473b3e24-55b5-4b80-b7e8-0f6c585f1f46 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Dota: detect and omit weak attentions for scalable transformer acceleration,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494eb4b3-c4e4-4195-8db4-5ab5e9f58cd3 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c83d3174-2cf8-45c8-b86b-d079ce6cf3c1 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling vision with sparse mixture of experts,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda643f5-befb-4273-8c2b-16b4640fd7cd · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9ea7e2-6dc4-4fc8-ae7f-9b8ba5f3c86b · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d67b0c2-877f-446f-af4d-5dc95e558616 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Bagualu: targeting brain scale pretrained models with over 37 million cores,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718c5dc3-3874-4f11-91c0-c0f7d46d9915 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation FastMoE: A Fast Mixture-of-Expert Training System
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8456a3e5-5214-4706-a9f4-ec1966c179e8 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08a7bd0-15ac-456c-b82b-9fa41790159e · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e381844c-2fb7-4493-97fc-1a2769e9639b · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mixtral of Experts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5321b8d5-d25c-4dfa-a6ef-26b31711acf8 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Evaluating the stability of embedding- based word similarities,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4ae30d9-3e43-4c48-a267-ecaec1b2ff23 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Sentiment classification using docu- ment embeddings trained with cosine similarity,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe9f19ea-c65e-4219-8610-8065e01afaac · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Problems with cosine as a measure of embedding similarity for high frequency words,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 776d376f-a324-46e6-ba43-979288ca7ad6 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 362ef62c-6582-4a07-9ab1-e50b1462e7c6 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Turbotransformers: an efficient gpu serving system for transformer models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3744365b-3450-41ce-aece-6fa235c139d3 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc879ffb-3799-4dc5-be37-b8bde5b4553a · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pytorch: An imperative style, high-performance deep learning library,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5c1a8a-ccd1-4be1-b474-ff6e14732725 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deep graph library: Towards efficient and scalable deep learning on graphs,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 80472934-57cb-4614-aa99-b72ccb8aabf5 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pointer sentinel mix- ture models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23546902-2a1f-43e7-b732-5a6d90a56522 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea65bc08-baed-432c-8caf-74f28e9851c7 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7434d240-26bc-4341-9dce-c9dfaa2304a0 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models with image descriptors are strong few-shot video-language learners,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6816a34e-bba2-4804-90d8-13fc59af7298 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Multitask mixture of sequential experts for user activity streams,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b998076-e7ee-4d05-b0e6-200ffa91e923 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Taming Sparsely Activated Transformer with Stochastic Experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88639ef7-6f65-42fc-930c-4d361d9c077c · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Generalizable person re- identification with relevance-aware mixture of experts,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f692a302-b56d-4130-912b-043a87f3d6e0 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 326748dc-4b49-4e69-b1b5-d9002782496a · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Palm: Scal- ing language modeling with pathways,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e05fe1-472d-474a-b3cb-976103035ec4 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Glam: Efficient scaling of language models with mixture-of-experts,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e826fa77-eaf6-4493-a3c1-8246c9a1c15a · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Llama-moe: Building mixture-of-experts from llama with continual pre-training,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a9bce051-d954-492c-8611-0506c5d3fdf8 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3facce-caa7-4c67-8bb6-456edfd66297 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a663f9-4cbf-446d-81ee-c53354ab40bd · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GSPMD: General and Scalable Parallelization for ML Computation Graphs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b787e6d-8484-4896-9d4b-da3eeeffdc85 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Base layers: Simplifying training of large, sparse models,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3c2234-a201-4c3f-9b05-2121b62c75cc · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4298b2fd-0172-4e22-858c-05aaac22e60a · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ec9ec0-aa5f-4fda-ad27-8498ff592d33 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16bf56f-ec91-4130-9987-ba68b2356f8e · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6ba828-6fa6-4b12-b88d-7d5097de5b10 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a263021b-5801-415c-ab26-56f647fab923 · outbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.