Pith. sign in

Paper Citation Record · LEDGER

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2411.15419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15419 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:24:00.465856Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a280264-2777-44ee-86ec-191dfdd495f2 · outbound

This paper cites Attention is all you need,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.225425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.225425Z digest=sha256:dabbdb83ecd00ed9678a3833f19d9a9f8ce377656f1c0c2a040c4b0aca63a545

Observation a04b83fb-5c21-464b-9275-38b39740cdbc · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.230885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.230885Z digest=sha256:056ff365db11d179bc9aa35cfa3f277d2eff3ec1beb46c20250b534bf3f5a644

Observation d7bdecca-5cc6-413b-8280-79330a94d8dd · outbound

This paper cites Language models are unsupervised multitask learners,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models are unsupervised multitask learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.236232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.236232Z digest=sha256:513f9ad40c224123561627a03b63890cc755a485e25bd3786ca0de4d962d964e

Observation 9301b913-d786-4804-a656-d5c6c6488c5d · outbound

This paper cites Scaling Laws for Neural Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling Laws for Neural Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.241305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.241305Z digest=sha256:6717a9fb7bc1c52b4051e99d55dc40b17de16f6b2d798da770560171f6ddf1ac

Observation f5496b08-8a18-4e12-88dc-df73a3efa4fe · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.246516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.246516Z digest=sha256:be68b3547ee55a88898fe4a2841bd394f51e9d36da75de4d0ce64d5350ea80b9

Observation 4e0bebad-8f71-4e95-9af0-e5d4172f6b1c · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.251780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.251780Z digest=sha256:6dcaa859be5319daa8af1ac824eb0ef0ce9fe6474065be4796b17df3a9b8bcab

Observation 40f0ac79-2cb9-4f8b-991a-9287f134d4d2 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation ai scale,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.257034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.257034Z digest=sha256:457a705a0d236d07e0b63d6365d24a0b4a4c34a50cbbd6862217470dafe90e02

Observation 32743ec6-8d9a-4074-b121-3a21e783519c · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Tutel: Adaptive mixture-of-experts at scale,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.261755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.261755Z digest=sha256:145bc1d04e65cb68a734ef191a730e38d64fb5c465766578398cb85ef7b72bf0

Observation 47cd25e9-1ff7-4056-876b-015c3601e47c · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:01.114174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.266533Z digest=sha256:0b1505cd259b092c59b210c2fc9b0afdd674a183cbb6de1a806f1c0eaf8ce192

Observation d67466c7-272d-4c1e-a165-0ac5b345c237 · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of-experts models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Janus: A unified distributed training framework for sparse mixture-of-experts models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.271283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.271283Z digest=sha256:d28099431a65441fecefb8572fd0dab5b18f915f7518bb48d33943d60ed10143

Observation af105974-1e52-46b4-b291-0f139b6f2d39 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Accelerating distributed {MoE} training and inference with lina,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.276180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.276180Z digest=sha256:2d315385f351440a03c3d0b06c46951c402a1a678930372541e3969f318eb095

Observation 90e63e08-7b37-4d80-a0b2-5e79652addb8 · outbound

This paper cites Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Gating dropout: Communication-efficient regularization for sparsely activated transform- ers,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.280979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.280979Z digest=sha256:fad6d06ea4ba313578af8b706fa66e76fb9ae3e28d9f8b1d6e995a6667dd4f1a

Observation 3c50de7e-42a4-428e-9e76-6fda5d12953d · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.285442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.285442Z digest=sha256:c2e6bfee6c62e7289fe8ef5ce69090f273fa478c656eed9d86be05c8c44f14b9

Observation 473b3e24-55b5-4b80-b7e8-0f6c585f1f46 · outbound

This paper cites Dota: detect and omit weak attentions for scalable transformer acceleration,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Dota: detect and omit weak attentions for scalable transformer acceleration,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.289979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.289979Z digest=sha256:6d5896740d5121887d5e3dc6376a5fc05f510046769ab1f38f9e48dfb3e4732e

Observation 494eb4b3-c4e4-4195-8db4-5ab5e9f58cd3 · outbound

This paper cites Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:01.042677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.294391Z digest=sha256:3db31bc5fd8b035ff909a390c1d45582bcab5dec530ec49504f94f8752e1d3b1

Observation c83d3174-2cf8-45c8-b86b-d079ce6cf3c1 · outbound

This paper cites Scaling vision with sparse mixture of experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Scaling vision with sparse mixture of experts,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.298778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.298778Z digest=sha256:b9641470973cb8f1d6629f0a0df712056cb7fadb545f3456452241c00f3d82d4

Observation cda643f5-befb-4273-8c2b-16b4640fd7cd · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.303294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.303294Z digest=sha256:add7b91baecf7f42d0d61328ab546cfb0910384795f95d3a049cf86d1cbb1273

Observation df9ea7e2-6dc4-4fc8-ae7f-9b8ba5f3c86b · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.308149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.308149Z digest=sha256:0c8bb67b8de9d03ba45ab63a4054cdb4022d0db4cd59acc91064385d2a7d1dc9

Observation 9d67b0c2-877f-446f-af4d-5dc95e558616 · outbound

This paper cites Bagualu: targeting brain scale pretrained models with over 37 million cores,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Bagualu: targeting brain scale pretrained models with over 37 million cores,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.312424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.312424Z digest=sha256:36a205d43c694568cce5d2b1f5d59cbd955def60369382a13f660b871eb612ee

Observation 718c5dc3-3874-4f11-91c0-c0f7d46d9915 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation FastMoE: A Fast Mixture-of-Expert Training System

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.317008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.317008Z digest=sha256:90843ae5c47fbe4ef2ac485299144f789a10167ca8f433b884fe2358fea0b9d2

Observation 8456a3e5-5214-4706-a9f4-ec1966c179e8 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.322070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.322070Z digest=sha256:8dd9dbef5bf93a14f3a85b9c4e1390b7e86fe9f409964d16f5bdbe4fb71fbf87

Observation d08a7bd0-15ac-456c-b82b-9fa41790159e · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.326484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.326484Z digest=sha256:7f5dac2dbe57d2819383178e90d603abe291a333d3a6a274d398fd57877eb8dc

Observation e381844c-2fb7-4493-97fc-1a2769e9639b · outbound

This paper cites Mixtral of Experts.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mixtral of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.331127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.331127Z digest=sha256:66269dec518b000338c0f97374d81f0ff1c5653d201606553a1a33da8c463e8f

Observation 5321b8d5-d25c-4dfa-a6ef-26b31711acf8 · outbound

This paper cites Evaluating the stability of embedding- based word similarities,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Evaluating the stability of embedding- based word similarities,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.987117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.335822Z digest=sha256:f18228a0e5be6e7c1c519041f9df1138450b9df68d0fa5b7c54af8047358ef92

Observation f4ae30d9-3e43-4c48-a267-ecaec1b2ff23 · outbound

This paper cites Sentiment classification using docu- ment embeddings trained with cosine similarity,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Sentiment classification using docu- ment embeddings trained with cosine similarity,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.971076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.340229Z digest=sha256:60abf808923a86942bd8e2a367da5fd77f431139e7e0878afa7b7e3a32ab97f0

Observation fe9f19ea-c65e-4219-8610-8065e01afaac · outbound

This paper cites Problems with cosine as a measure of embedding similarity for high frequency words,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Problems with cosine as a measure of embedding similarity for high frequency words,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.955273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.344952Z digest=sha256:85dade19ad198b57e8c79108a9645f7e75d1c8dbaaa08f6246ee6481e6fd983d

Observation 776d376f-a324-46e6-ba43-979288ca7ad6 · outbound

This paper cites PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation PanGu-$\pi$: Enhancing Language Model Architectures via Nonlinearity Compensation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.349134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.349134Z digest=sha256:72c13411e2c4711e42cff5fe3bc83d95227cfcb50f7c858bfd7fa2ba2bc13c63

Observation 362ef62c-6582-4a07-9ab1-e50b1462e7c6 · outbound

This paper cites Turbotransformers: an efficient gpu serving system for transformer models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Turbotransformers: an efficient gpu serving system for transformer models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.354236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.354236Z digest=sha256:b904c5ca88ca5a056206d9b0959cafd5e916f846f0daa08bc7561a5f33a8ac7c

Observation 3744365b-3450-41ce-aece-6fa235c139d3 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.358622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.358622Z digest=sha256:e5a8905aea79b59373a85584daa6ee9387ab854b0defc8e9f0544f309dd5ab3d

Observation dc879ffb-3799-4dc5-be37-b8bde5b4553a · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pytorch: An imperative style, high-performance deep learning library,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.362885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.362885Z digest=sha256:671c206d7c069066336466c230a27376c9c25e5bd5bde5661dc0f0cd6fb3e8df

Observation dd5c1a8a-ccd1-4be1-b474-ff6e14732725 · outbound

This paper cites Deep graph library: Towards efficient and scalable deep learning on graphs,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Deep graph library: Towards efficient and scalable deep learning on graphs,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.910518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.367558Z digest=sha256:fa2b1761c3ac91a1fccd6920e81f1532295074eb113d077a3d1d871447994edb

Observation 80472934-57cb-4614-aa99-b72ccb8aabf5 · outbound

This paper cites Pointer sentinel mix- ture models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pointer sentinel mix- ture models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.895148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.372344Z digest=sha256:40b0af82b55529c515bec0df5b88cc35e241f6bef17d655251a0a6c23a21bdd9

Observation 23546902-2a1f-43e7-b732-5a6d90a56522 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.377195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.377195Z digest=sha256:41d5585622d14eb11148610b085ebb738df172dcc7d252b0eb13f847486097f9

Observation ea65bc08-baed-432c-8caf-74f28e9851c7 · outbound

This paper cites Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Samsum corpus: A human-annotated dialogue dataset for abstractive summarization,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.878621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.382319Z digest=sha256:02bc317180efeae9e9392a0bbc553077f47aed21d05002842906e032f0470744

Observation 7434d240-26bc-4341-9dce-c9dfaa2304a0 · outbound

This paper cites Language models with image descriptors are strong few-shot video-language learners,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Language models with image descriptors are strong few-shot video-language learners,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.862671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.387010Z digest=sha256:3ab3e760caf85732f65c32d8cf38b1ebe913d335b37a7b0773b1d6dcd42e1df1

Observation 6816a34e-bba2-4804-90d8-13fc59af7298 · outbound

This paper cites Multitask mixture of sequential experts for user activity streams,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Multitask mixture of sequential experts for user activity streams,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.847940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.391978Z digest=sha256:586d937894623a07cb009eb4f4309b6adb2534ad64ecdd790430acd62f6940f9

Observation 1b998076-e7ee-4d05-b0e6-200ffa91e923 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Taming Sparsely Activated Transformer with Stochastic Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.396667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.396667Z digest=sha256:554a6dc946e32246b52e21a1b8bcd1c8ea4ab136b4451313ae6f49a9230d4cc5

Observation 88639ef7-6f65-42fc-930c-4d361d9c077c · outbound

This paper cites Generalizable person re- identification with relevance-aware mixture of experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Generalizable person re- identification with relevance-aware mixture of experts,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.833098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.401465Z digest=sha256:bbfd1b5cbaf76c00f0e534103f37aa1a8bd6f6dd641dffc3bf94b828611ae609

Observation f692a302-b56d-4130-912b-043a87f3d6e0 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.817749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.406077Z digest=sha256:779ca84dfd060bf06bd070d40392369e2c7f3e1cac454c959c62cd24cb466fba

Observation 326748dc-4b49-4e69-b1b5-d9002782496a · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Palm: Scal- ing language modeling with pathways,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.410376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.410376Z digest=sha256:776517090399f77fb900aca038fc383a77dd00a41f7000f3f3dc05947260155f

Observation 52e05fe1-472d-474a-b3cb-976103035ec4 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Glam: Efficient scaling of language models with mixture-of-experts,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.414670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.414670Z digest=sha256:9c8de212009e7f90f2e720ae050aff85352b6087e0627b98d175046082d60355

Observation e826fa77-eaf6-4493-a3c1-8246c9a1c15a · outbound

This paper cites Llama-moe: Building mixture-of-experts from llama with continual pre-training,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Llama-moe: Building mixture-of-experts from llama with continual pre-training,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.782311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.419324Z digest=sha256:e67ffc49c5b5945835c187d0973733b22ff967b611631fbc17f8e739e0b98557

Observation a9bce051-d954-492c-8611-0506c5d3fdf8 · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.423688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.423688Z digest=sha256:d3365974e2dd6990db7debf2c2b28c43e0a213489825fe196cefa4c64da0b214

Observation ef3facce-caa7-4c67-8bb6-456edfd66297 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.428858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.428858Z digest=sha256:86057ae65dd0b65a83850d31f0e0c651abbf83f74cd84915fbab2a4eb2a73f4d

Observation a6a663f9-4cbf-446d-81ee-c53354ab40bd · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.433734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.433734Z digest=sha256:c830b61e1be9441dd8d41e34ad63636b8c48d6a012887d92edfb4a04e26874c1

Observation 7b787e6d-8484-4896-9d4b-da3eeeffdc85 · outbound

This paper cites Base layers: Simplifying training of large, sparse models,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Base layers: Simplifying training of large, sparse models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.438386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.438386Z digest=sha256:0bd559fc9a16f4697ec77345008dda0aa5cc55fc2182c1ce5a107a4ad28f17b7

Observation 4c3c2234-a201-4c3f-9b05-2121b62c75cc · outbound

This paper cites fairseq: A Fast, Extensible Toolkit for Sequence Modeling.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation fairseq: A Fast, Extensible Toolkit for Sequence Modeling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.442960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.442960Z digest=sha256:8a7e2ebcdb287f0f0c657659d336c680409f0205b94ba38f4186329a8cdda5d7

Observation 4298b2fd-0172-4e22-858c-05aaac22e60a · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Pipemoe: Accelerating mixture- of-experts through adaptive pipelining,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.447653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.447653Z digest=sha256:037a4d88a5d95a0e6d2cf2d482e9bbda19d92a61b58435ccac59f0de7a4e1bae

Observation c0ec9ec0-aa5f-4fda-ad27-8498ff592d33 · outbound

This paper cites Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.451989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.451989Z digest=sha256:f81cdf343e6e8c53b603896f50fae9f2239e2e2a453312e739712aa8d9eef7b3

Observation c16bf56f-ec91-4130-9987-ba68b2356f8e · outbound

This paper cites MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.456236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.456236Z digest=sha256:934d82adf3e429ee4d6351eb7d356c51876220e21e53f651add0fd454267e9ef

Observation cc6ba828-6fa6-4b12-b88d-7d5097de5b10 · outbound

This paper cites Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:24:00.737473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:24:00.460975Z digest=sha256:b184ac5d3ec63eaa74e741b7e971460053e6ab35b08c5004588315bce7e05df7

Observation a263021b-5801-415c-ab26-56f647fab923 · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,.

Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:00.465856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:00.465856Z digest=sha256:7a2426c44cdcb0189f3d7c7f931d577dad5a4e35098d237acb68dddf32ca964a

Pith citing papers

No inbound Pith citation observations are available.