Pith. sign in

Paper Citation Record · LEDGER

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation

As of 12 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2411.15419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15419 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:24:00.465856Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a280264-2777-44ee-86ec-191dfdd495f2 · outbound

This paper cites Attention is all you need,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.225425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.225425Z digest=sha256:8b360e0faa4ba1eb60675fe91c107d699ffdfc50001f34247ab7ce4a3ea2742c

Observation a04b83fb-5c21-464b-9275-38b39740cdbc · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.230885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.230885Z digest=sha256:98a2b834844d1e53afd6ec5413c3dcd7bec930e8c89d9ced3cb848dc72fb1bf7

Observation d7bdecca-5cc6-413b-8280-79330a94d8dd · outbound

This paper cites Language models are unsupervised multitask learners,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models are unsupervised multitask learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.236232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.236232Z digest=sha256:3c427fcd1a922493550cfc4f455ff3b3b5b5b37de2ca73a88e4075b1d949cd97

Observation 9301b913-d786-4804-a656-d5c6c6488c5d · outbound

This paper cites Scaling Laws for Neural Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling Laws for Neural Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.241305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.241305Z digest=sha256:d29c272af49f9090c0be2107c46458b4944e7a4a2c0dd922a3c1ab53b684ee85

Observation f5496b08-8a18-4e12-88dc-df73a3efa4fe · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.246516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.246516Z digest=sha256:758119a88447776c7a6efbb56336230d314e533321235e089ee4559f6554c63d

Observation 4e0bebad-8f71-4e95-9af0-e5d4172f6b1c · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.251780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.251780Z digest=sha256:aad076cb3f88ddfb33763f73b79efb63593eacba4948520c14f81275d48c0389

Observation 40f0ac79-2cb9-4f8b-991a-9287f134d4d2 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.257034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.257034Z digest=sha256:743be2f88b4b04f189eb3c7774a589bd2d7a17b8ae4bf96a930d8140bb227aaf

Observation 32743ec6-8d9a-4074-b121-3a21e783519c · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Tutel: Adaptive mixture-of-experts at scale,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.261755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.261755Z digest=sha256:d7f606d60b95b7e3e8d87b7b26759ce21ab594822d076a52ecf975827013b8ef

Observation 47cd25e9-1ff7-4056-876b-015c3601e47c · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:01.114174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.266533Z digest=sha256:12bb928fb1deeb6a2154cba8ca6911c02ddc9fec805f87155b77fb5fb858e894

Observation d67466c7-272d-4c1e-a165-0ac5b345c237 · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.271283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.271283Z digest=sha256:93bc34233111c02f6d7cc01c3340ff6916865271dd24deab6f33e9a52a84a257

Observation af105974-1e52-46b4-b291-0f139b6f2d39 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Accelerating distributed {MoE} training and inference with lina,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.276180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.276180Z digest=sha256:cebe802e5699fc32eff7b92c8706c466b18cb48a491b9f079b6540400ce41ac1

Observation 90e63e08-7b37-4d80-a0b2-5e79652addb8 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.280979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.280979Z digest=sha256:137593d8dc7f1e40eed12c549209bdf1b70409cedd62febc5253186e91397226

Observation 3c50de7e-42a4-428e-9e76-6fda5d12953d · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.285442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.285442Z digest=sha256:280955148eb53d712dd4c7486a2ef665c7d141804df721aab3c540a53d83861d

Observation 473b3e24-55b5-4b80-b7e8-0f6c585f1f46 · outbound

This paper cites Dota: detect and omit weak attentions for scalable transformer acceleration,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Dota: detect and omit weak attentions for scalable transformer acceleration,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.289979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.289979Z digest=sha256:053a7cdbee51f15263079fae4c532ba8b8b923cec58888c68634a8bdf5d7e679

Observation 494eb4b3-c4e4-4195-8db4-5ab5e9f58cd3 · outbound

This paper cites Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:01.042677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.294391Z digest=sha256:6d306c034bda57fb9e86295e806ca41a3049df7b2549ac5df7e531e22659a205

Observation c83d3174-2cf8-45c8-b86b-d079ce6cf3c1 · outbound

This paper cites Scaling vision with sparse mixture of experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling vision with sparse mixture of experts,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.298778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.298778Z digest=sha256:6f8fae25312ec209be59d3436e7d5abe7b532ea27b5cc157e596e2a63feb58b8

Observation cda643f5-befb-4273-8c2b-16b4640fd7cd · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.303294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.303294Z digest=sha256:45b5280ce2b0ecde460da6dddb1527be9cd9ecf56f4e0600a9f35976acaac4a3

Observation df9ea7e2-6dc4-4fc8-ae7f-9b8ba5f3c86b · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.308149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.308149Z digest=sha256:e01c43bdf2c69332ac89cd859d84020e9b50fbaaa5ab79ad64d80a8620d28727

Observation 9d67b0c2-877f-446f-af4d-5dc95e558616 · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.312424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.312424Z digest=sha256:e3f9f3a7986923586c8cb134d6fc2be5326ec6fcd07e42eabe6b7333852eb3b6

Observation 718c5dc3-3874-4f11-91c0-c0f7d46d9915 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation FastMoE: A Fast Mixture-of-Expert Training System

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.317008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.317008Z digest=sha256:726f644c1902b17430548a9e1f605f510696784969a385174de66acf103b896d

Observation 8456a3e5-5214-4706-a9f4-ec1966c179e8 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.322070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.322070Z digest=sha256:bb7fb8ddaaf2692793007432aa565cedbab81b686c2eddd2f2d7795fdbcf014b

Observation d08a7bd0-15ac-456c-b82b-9fa41790159e · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.326484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.326484Z digest=sha256:5291bc4266665c6520c0a95c86b38fb2bd0629dfae536fe0736e82ef7817dc5e

Observation e381844c-2fb7-4493-97fc-1a2769e9639b · outbound

This paper cites Mixtral of Experts.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mixtral of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.331127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.331127Z digest=sha256:f59b035301884f4941cc4bda7b8c73fddc25cb4957f9d80c8d7947930e999510

Observation 5321b8d5-d25c-4dfa-a6ef-26b31711acf8 · outbound

This paper cites Evaluating the stability of embedding- based word similarities,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Evaluating the stability of embedding- based word similarities,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.987117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.335822Z digest=sha256:b77973abef1eef9835dca21728c9a47d98aac4c52125eb4a854241baf58379c3

Observation f4ae30d9-3e43-4c48-a267-ecaec1b2ff23 · outbound

This paper cites Sentiment classification using docu- ment embeddings trained with cosine similarity,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Sentiment classification using docu- ment embeddings trained with cosine similarity,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.971076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.340229Z digest=sha256:5d4c89fee5404d3f39bad962d034722cb1919bda2f397a9ccf65349c8b505c5f

Observation fe9f19ea-c65e-4219-8610-8065e01afaac · outbound

This paper cites Problems with cosine as a measure of embedding similarity for high frequency words,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Problems with cosine as a measure of embedding similarity for high frequency words,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.955273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.344952Z digest=sha256:b2c280f76225300970054030297145420f662fe3b9337058dd991b177b2e8c20

Observation 776d376f-a324-46e6-ba43-979288ca7ad6 · outbound

This paper cites PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.349134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.349134Z digest=sha256:c9db8bb803e52a709e329ce9a30eec61c255a56ebff4cb393cf36c74c5f8f2b7

Observation 362ef62c-6582-4a07-9ab1-e50b1462e7c6 · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Turbotransformers: an efficient gpu serving system for transformer models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.354236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.354236Z digest=sha256:d990e24c9ca24988c699011a3e60624d6f763234dda1d14eac4e5d0cd6362fee

Observation 3744365b-3450-41ce-aece-6fa235c139d3 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.358622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.358622Z digest=sha256:da6f8e73bd55860b0f2e5c236d43e9b0c86b446929826dd67fa3e2c5463d193d

Observation dc879ffb-3799-4dc5-be37-b8bde5b4553a · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pytorch: An imperative style, high-performance deep learning library,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.362885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.362885Z digest=sha256:eef2b40cacca2141ca28700908c4ae8852ad9d794a969ec5505e1c2c441b3ab2

Observation dd5c1a8a-ccd1-4be1-b474-ff6e14732725 · outbound

This paper cites Deep graph library: Towards efficient and scalable deep learning on graphs,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deep graph library: Towards efficient and scalable deep learning on graphs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.910518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.367558Z digest=sha256:46787e177c5f5a3d443173f11606ab985b7ae988c90e9edfc6ec4dbad1a547b7

Observation 80472934-57cb-4614-aa99-b72ccb8aabf5 · outbound

This paper cites Pointer sentinel mix- ture models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pointer sentinel mix- ture models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.895148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.372344Z digest=sha256:011022d21643394deebc072c4dd7964d1c5c2753b499a44560c936493b385e93

Observation 23546902-2a1f-43e7-b732-5a6d90a56522 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.377195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.377195Z digest=sha256:f011b6a0a0689803e00c22e2d84aa28b5e5fbbc208504e30ebf9aebd399a9e78

Observation ea65bc08-baed-432c-8caf-74f28e9851c7 · outbound

This paper cites Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.878621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.382319Z digest=sha256:d38f9ff0fd242a80dcf1bb4546ff626801266e678e49e0134777df6ccedf24fe

Observation 7434d240-26bc-4341-9dce-c9dfaa2304a0 · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models with image descriptors are strong few-shot video-language learners,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.862671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.387010Z digest=sha256:445c644a93edc5f53bd6e7cfe3ef93f679b328360b6db0e2e3fc31c9c1f86bde

Observation 6816a34e-bba2-4804-90d8-13fc59af7298 · outbound

This paper cites Multitask mixture of sequential experts for user activity streams,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Multitask mixture of sequential experts for user activity streams,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.847940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.391978Z digest=sha256:80098c74db0818b90151d7d094646940b0e82de2a408f83dee89374e6dda37ad

Observation 1b998076-e7ee-4d05-b0e6-200ffa91e923 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Taming Sparsely Activated Transformer with Stochastic Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.396667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.396667Z digest=sha256:47d39759170edb3a8b9c44835816c2b97baf73346b92bbc1bc170734f44a5b2c

Observation 88639ef7-6f65-42fc-930c-4d361d9c077c · outbound

This paper cites Generalizable person re- identification with relevance-aware mixture of experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Generalizable person re- identification with relevance-aware mixture of experts,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.833098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.401465Z digest=sha256:4daa2eeefbc7a29a15c66e7a57894925240169cbf0c768ff5f09ba0f349f1c00

Observation f692a302-b56d-4130-912b-043a87f3d6e0 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.817749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.406077Z digest=sha256:755d52f4fd90fa7154a7839b18c5fe6d9dc5fbdcc4d40ba50178bd6c3e83a44d

Observation 326748dc-4b49-4e69-b1b5-d9002782496a · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Palm: Scal- ing language modeling with pathways,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.410376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.410376Z digest=sha256:9bc93cf169f97b9418c3b73503a01b7ac37f0763ad22973a46b1d75b8cb569ea

Observation 52e05fe1-472d-474a-b3cb-976103035ec4 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Glam: Efficient scaling of language models with mixture-of-experts,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.414670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.414670Z digest=sha256:8a9f160a25f3ed8929d0b4f1722b3715623fd56bc7437dc0ca40b6144c67eb8d

Observation e826fa77-eaf6-4493-a3c1-8246c9a1c15a · outbound

This paper cites Llama-moe: Building mixture-of-experts from llama with continual pre-training,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Llama-moe: Building mixture-of-experts from llama with continual pre-training,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.782311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.419324Z digest=sha256:fa8d0747589fed6abe4d8b47fd9e39d26c70281857a71b51bc09e404ce7db122

Observation a9bce051-d954-492c-8611-0506c5d3fdf8 · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.423688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.423688Z digest=sha256:6d8a0e94d08db2ecc59ec8e11cc810cce6d67b54063db460a13b5eec763f56df

Observation ef3facce-caa7-4c67-8bb6-456edfd66297 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.428858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.428858Z digest=sha256:eda48bc92e551cb34503691204467850db1357a9fea1e3b4eafc2b82255bd5de

Observation a6a663f9-4cbf-446d-81ee-c53354ab40bd · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.433734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.433734Z digest=sha256:fd3b98f4c7d4e7fbfa7ec11c08993d486ae13740f2900884535a3725a641689f

Observation 7b787e6d-8484-4896-9d4b-da3eeeffdc85 · outbound

This paper cites Base layers: Simplifying training of large, sparse models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Base layers: Simplifying training of large, sparse models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.438386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.438386Z digest=sha256:b0ebd89bd558d59c294778bb453ab31d3910a521a9c755c0cda13b0754707fc8

Observation 4c3c2234-a201-4c3f-9b05-2121b62c75cc · outbound

This paper cites fairseq: A Fast, Extensible Toolkit for Sequence Modeling.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation fairseq: A Fast, Extensible Toolkit for Sequence Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.442960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.442960Z digest=sha256:7f40632260f31905623edbbfa986f0b614eabcd214f30afdbf275f5c079c0257

Observation 4298b2fd-0172-4e22-858c-05aaac22e60a · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.447653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.447653Z digest=sha256:b603cd5ec7617fe098a5f50690880d02ef711e21eadf1b5d27ff510a26df136b

Observation c0ec9ec0-aa5f-4fda-ad27-8498ff592d33 · outbound

This paper cites Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.451989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.451989Z digest=sha256:ae49e6852d6558c84d147bf71b5490e27777edf23920fd3e26f6e9ea79d263c0

Observation c16bf56f-ec91-4130-9987-ba68b2356f8e · outbound

This paper cites MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.456236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.456236Z digest=sha256:7f5857109f8e8b94c49754c5fba609f78177c5f6438bbb1764586e4925b2045d

Observation cc6ba828-6fa6-4b12-b88d-7d5097de5b10 · outbound

This paper cites Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.737473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:24:00.460975Z digest=sha256:f1a9603657ba014dda763cb26959fe7caedbd4e900e18b2eebf4655fd99bd6ee

Observation a263021b-5801-415c-ab26-56f647fab923 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.465856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.465856Z digest=sha256:2907c6ba427646704601595dd9b0dfa26ca0ec23ad8785c4ac3f323ebeae2799

Pith citing papers

No inbound Pith citation observations are available.