Pith. sign in

Paper Citation Record · LEDGER

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.08650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08650 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:58.446447Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 155412e3-7d0f-44fa-acc7-9d39d244d27b · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.257784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.233177Z digest=sha256:8f01222c2a143fb93e76cd9c0d18cb5ecbda9b42f2e6cfd47655cbe823e6e3fc

Observation f7f081eb-159f-4567-ba8b-42f519014d53 · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.232328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.238758Z digest=sha256:62b65da6e859e79b510f90b721c02aaa7c00dc6dba0b28e4367557f9fe49b416

Observation 88128731-e303-4cc8-b293-7508c5abf31e · outbound

This paper cites 总 All-to-All 时间.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism 总 All-to-All 时间

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.213592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.245187Z digest=sha256:8ddefe5df3162be674322c39c39ad4b26591c79ae0ca3ddf4e080368ffd207cd

Observation 1b015d02-473e-477d-8c0c-00fff5e8c6e9 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.384791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.251293Z digest=sha256:67f36aaf73ea8d587cdc70bfb640ce34e363691048de82bec4378bd1541f9d4a

Observation 78903f03-9259-4bf6-9fd4-697cb33ce04b · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.257289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.257289Z digest=sha256:4eb2c90259e33567b0bb07905a19ca09b7f9b0c6cde0baee19b92256d87fdbaf

Observation 8f2b9ad3-4bc8-440b-b148-5692f207abeb · outbound

This paper cites A Survey on Inference Optimization Techniques for Mixture of Experts Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.262836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.262836Z digest=sha256:b250a1c1d9bca05bada00c5ca7c990a3989109e927ec9f01366b2495cad7ad66

Observation 61c66eaf-ae8f-4359-9f8e-fc8da2449351 · outbound

This paper cites Speed Always Wins: A Survey on Efficient Architectures for Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.267934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.267934Z digest=sha256:8920782ad212095bf40a88b20092fe3c0b5cd8601ce41782cc9913fceb8724e8

Observation 8984cb4b-f455-4def-9e56-094faf9ff28f · outbound

This paper cites Adaptive Mixtures of Local Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Adaptive Mixtures of Local Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.273267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.273267Z digest=sha256:19ed30db4df30f600a32e549fbfec6c8925df23e9cfddc1779dbc44f0e15e074

Observation 2a1ecfed-c799-4c62-858a-1ffa1c4724dc · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.279219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.279219Z digest=sha256:9d4497d00adbadbecf0cdaea987eba30e5fbe86070535b138184880af93eff7a

Observation cd8cec16-c0d2-493a-953e-44a51c39c2ad · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.284500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.284500Z digest=sha256:af0ae0d1ec4e444484d042641eeb750cb7c681f1079ee0d3f7936bd6f4008647

Observation c43a1dfd-0197-4799-aa11-1113006dd440 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.289591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.289591Z digest=sha256:842c23ae2585c7d7b506054aa5544a182e75188cfff4807dc77a2e0d10ea297d

Observation c45be21f-5096-447c-9f5f-f1fe7941a958 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.294305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.294305Z digest=sha256:4574d72098c97eedb5e04e8deece14ac52aa0895de9ae9d68e23b350fc71875b

Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · outbound

This paper cites BASE Layers: Simplifying Training of Large, Sparse Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.299078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.299078Z digest=sha256:4b38cd9d459fcf597b2e18b5236a7754aeb6cee2d9bae78992813017ea18cb54

Observation 32714d3f-e41a-46c4-90d8-7886c02c8b33 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture-of-Experts with Expert Choice Routing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.304689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.304689Z digest=sha256:b87fe93b094949e62496d957b2f093c066e2eda080c33a33fc4d25981d931568

Observation 61850748-812c-48c4-941f-dec0ffaa1b60 · outbound

This paper cites Mixtral of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixtral of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.311824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.311824Z digest=sha256:d83d30a9e9704e1001d8952365598b1eb45d82508f3208240ee42d8a59e3f9e4

Observation 31688b3b-bc2c-40b5-9c15-8cc59de059d1 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.316527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.316527Z digest=sha256:90cdd5bc5395720f06202f39817e6487b0fe042ab98f8496a11dfeee36d4310a

Observation 07ebc0aa-91b1-4fa2-a626-cbea7ee04639 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.321856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.321856Z digest=sha256:098d8e478148e948c2122a04b9416bb72c5a0a24094f34a18da663ad006dff7d

Observation 1b20fc12-e4e1-4461-94ac-0d948153fc82 · outbound

This paper cites Arctic: Snowflake’s Open-Source LLM.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Arctic: Snowflake’s Open-Source LLM

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.365891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.329082Z digest=sha256:fe4e2c19d84f7d2f8b9c292044dcbe079f965ede9f2301b6df78ce0f638b487e

Observation 192b8514-451f-4f73-8613-4e8da2c1d24a · outbound

This paper cites Qwen3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Qwen3 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.334915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.334915Z digest=sha256:556b17d5fd86d1f6f87fefcb859b137faa0399e6e8a1cf5f484c913c5b58f3d9

Observation 879cd6f2-c6d0-4b73-a76a-4efda08f2733 · outbound

This paper cites Mixture of A Million Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture of A Million Experts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.340157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.340157Z digest=sha256:ba9095fe77cbda363027ed5384a5cfd318ba4a4fcf7076d4019bce41847dcef3

Observation 1f73fa79-2c97-4189-945a-659621f98fed · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Kimi K2: Open Agentic Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.345266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.345266Z digest=sha256:28c9327bacc79398ff69e781363ae17d24fdfe974f5c68016a07ab1312d48289

Observation ab426b07-485c-4686-a67a-560098855d76 · outbound

This paper cites gpt-oss Model Card: Architecture.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism gpt-oss Model Card: Architecture

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.345701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.351257Z digest=sha256:033a7c0dcdaefdfa994e251f3aa69d11123100df737aeb05ac52ae43d92aff98

Observation d634a2cc-ec24-42c7-8c8e-5debaca623cb · outbound

This paper cites LongCat-Flash Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat-Flash Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.356618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.356618Z digest=sha256:7b73674f914aa6cc8ded2eea60026d8eed6dcb5cdf037c55304adf92255f8566

Observation 6b6f590d-e17a-4d7c-ab6a-4bd47c855fca · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.362126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.362126Z digest=sha256:63941cb9273d3f2cce79ea157913fef761bad233f367d0ccb7e07b4ecb6e667d

Observation 8f27f14c-fbd6-44e2-998d-1f1b0f54fc92 · outbound

This paper cites LongCat 2.0 Technical Blog.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat 2.0 Technical Blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.317922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.368767Z digest=sha256:7848ac504d62135c3f14c59a8760ea6f42193318c33ccd8ea38ccac64e2fb0bb

Observation 898a66da-b382-4426-9b70-927251efbefc · outbound

This paper cites MoHGE: Mixture of Heterogeneous Grouped Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MoHGE: Mixture of Heterogeneous Grouped Experts

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.297216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.374274Z digest=sha256:8d42f656804d59ce30d11de13711a0af65da50d20bdf16cfa5a4662482fa386c

Observation 6226d18c-7617-47dd-86bf-e488410062da · outbound

This paper cites GMoE: Global Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GMoE: Global Mixture-of-Experts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.277555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.380158Z digest=sha256:26a75ebb5afad450115c7e074f34ec7c87d7f1d6bf766c06a519a25668dfd84d

Observation 42e1fceb-5840-425e-a95c-c0374e5b5b72 · outbound

This paper cites Multi-Head LatentMoE and Head Parallel.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Multi-Head LatentMoE and Head Parallel

Reference 64

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:32:58.792734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.387367Z digest=sha256:638f10d300072e2c07f0069c54a781cb182f045708bc9ba94f3e3b99e64a6fbb

Observation 6ea333fd-a000-4440-a917-4a2dd234c26b · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.393621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.393621Z digest=sha256:15f592ecdcba8769a8e704a9661e26235c7fee0ccc9c8feb073035bea1588684

Observation 1a7fe1c8-47ff-462b-8b70-68986a7d781b · outbound

This paper cites Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:58.684185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:58.404138Z digest=sha256:4f73399ca2b30114ff39eb2227390e2e1590356002d0dd1e19c0273884d4cfd0

Observation 6bd2c283-d122-4b69-83c7-0fa779d27a17 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism From Sparse to Soft Mixtures of Experts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.409771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.409771Z digest=sha256:219f11017778a9f763087341a95ed6bc31922c103f76df9e58084cba3efe17cc

Observation fa229c4d-40aa-428a-80f9-51b3769444ac · outbound

This paper cites Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.414782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.414782Z digest=sha256:0aeee56561dee4785914f2e1a69ffe4fae7c939bf0816b379ccb9379d76b1bef

Observation d413417b-2018-4afc-b784-c10df13c43d0 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.419634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.419634Z digest=sha256:7282d878f1ab9e2780b46d04a52a0695f663334b9ec1732f0c29a66cf6acda8a

Observation 21f6a2cc-cb6e-41bf-a386-d66cf431e1e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V3 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.425384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.425384Z digest=sha256:46f2bdc357143cd08aa98651cfcf9c62b47cf5875eec19445ca3ae04098dff32

Observation 5e4dd52a-1fc9-48fe-891b-a29c91793666 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.430621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.430621Z digest=sha256:f5f81f9997b937417be60cc125afdc3b4b930545c420f0fcc50e748c869ada6b

Observation 47e0c6d9-4338-48fe-8537-2c9497c55301 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.435499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.435499Z digest=sha256:c1a17ea317028d2120ed7a32eb1bf6b191444e46fea3714bfa3f4f69df074520

Observation 96585586-e007-4b53-a86a-6917574ca7c4 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.441113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.441113Z digest=sha256:94fdb18315d1b961c768a94a306f60a259c70444dc5bb97d585b8396090b07b8

Observation a825bfe9-00c2-4a9d-aa07-a6cb2e9774a0 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Jamba: A Hybrid Transformer-Mamba Language Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.446447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.446447Z digest=sha256:beec8dadeccf15fee9778f75a22d04b8eacb08cd9fb493339bc3b5d7e9b7ab26

Pith citing papers

No inbound Pith citation observations are available.