Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2110.04260.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:01.551591Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T15:17:07.182288Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 202a855d-5903-44e1-bd12-8d1df84630b8 · inbound
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 52602f3f-2653-4d8f-b858-a60f96c09ede · inbound
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b998076-e7ee-4d05-b0e6-200ffa91e923 · inbound
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Taming Sparsely Activated Transformer with Stochastic Experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3528265-2cb6-43d4-8e75-46fc1d5bf4d8 · inbound
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing Taming Sparsely Activated Transformer with Stochastic Experts
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b59191aa-fddc-4ba4-b15e-aea3f8979729 · inbound
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning Taming Sparsely Activated Transformer with Stochastic Experts
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82837663-8add-443f-a6c0-3baf4e2fdac8 · inbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee6c30a-1aff-4e7b-9b8a-135de7eab909 · inbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbcd22f-5233-4ee4-ac32-627f4265ded2 · inbound
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Taming Sparsely Activated Transformer with Stochastic Experts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b4d82a-8da6-4d42-ba48-29c7ac74bfcf · inbound
Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation Taming Sparsely Activated Transformer with Stochastic Experts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543a92b1-cf1d-4faf-8ffe-6e152c431736 · inbound
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts Taming Sparsely Activated Transformer with Stochastic Experts
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.