Pith. sign in

Paper Citation Record · LEDGER

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.08650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08650 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:58.446447Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 155412e3-7d0f-44fa-acc7-9d39d244d27b · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.257784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.233177Z digest=sha256:9e6f9490c03224fcc142421ac0a93348682f17a60084cffd0baa14ddf43adca3

Observation f7f081eb-159f-4567-ba8b-42f519014d53 · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.232328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.238758Z digest=sha256:dfb334a77668e281b50cf94310a860840c16f2ba2ed2b1fc0ea1a9ab7cc7704e

Observation 88128731-e303-4cc8-b293-7508c5abf31e · outbound

This paper cites 总 All-to-All 时间.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism 总 All-to-All 时间

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.213592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.245187Z digest=sha256:16675480d275b006c5dbe0709d954e26a0fff385ca16594de2549574aae582bf

Observation 1b015d02-473e-477d-8c0c-00fff5e8c6e9 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.384791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.251293Z digest=sha256:59dc3556642c626d550738eeba7c1372c43568097a0ea503fc65c3c876ea8f86

Observation 78903f03-9259-4bf6-9fd4-697cb33ce04b · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.257289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.257289Z digest=sha256:ad488cb511e3d84699c7f6f057a9aca8d1f631a6163a37107855e5eea4fce3b3

Observation 8f2b9ad3-4bc8-440b-b148-5692f207abeb · outbound

This paper cites A Survey on Inference Optimization Techniques for Mixture of Experts Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.262836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.262836Z digest=sha256:1f70ac180464d464c8d8e666cb963dad5582e60b1f135df1f26aa95f87dab2dd

Observation 61c66eaf-ae8f-4359-9f8e-fc8da2449351 · outbound

This paper cites Speed Always Wins: A Survey on Efficient Architectures for Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.267934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.267934Z digest=sha256:b28afcf2c0c3abc6c817f38e4e4e80e507f509240e763b1171596c31eed5c04d

Observation 8984cb4b-f455-4def-9e56-094faf9ff28f · outbound

This paper cites Adaptive Mixtures of Local Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Adaptive Mixtures of Local Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.273267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.273267Z digest=sha256:19ed30db4df30f600a32e549fbfec6c8925df23e9cfddc1779dbc44f0e15e074

Observation 2a1ecfed-c799-4c62-858a-1ffa1c4724dc · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.279219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.279219Z digest=sha256:0c68bcee8d6d4f55bf9c552713568855cc6507c8e449570a36382834c118a679

Observation cd8cec16-c0d2-493a-953e-44a51c39c2ad · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.284500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.284500Z digest=sha256:af0ae0d1ec4e444484d042641eeb750cb7c681f1079ee0d3f7936bd6f4008647

Observation c43a1dfd-0197-4799-aa11-1113006dd440 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.289591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.289591Z digest=sha256:842c23ae2585c7d7b506054aa5544a182e75188cfff4807dc77a2e0d10ea297d

Observation c45be21f-5096-447c-9f5f-f1fe7941a958 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.294305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.294305Z digest=sha256:4574d72098c97eedb5e04e8deece14ac52aa0895de9ae9d68e23b350fc71875b

Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · outbound

This paper cites BASE Layers: Simplifying Training of Large, Sparse Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.299078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.299078Z digest=sha256:99aef885da5caa6f381702a53b30b7e8a9f42a072cfa441753d9d428d0ee359f

Observation 32714d3f-e41a-46c4-90d8-7886c02c8b33 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture-of-Experts with Expert Choice Routing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.304689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.304689Z digest=sha256:b8c2819c2cd0026e3a95a66a50eba6ebd134e6d040f861bd9800696e22ceb45d

Observation 61850748-812c-48c4-941f-dec0ffaa1b60 · outbound

This paper cites Mixtral of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixtral of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.311824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.311824Z digest=sha256:f9033c8e15dd4b28660374debb47f3d26cc90726c67587baced0055d5dd1d3e9

Observation 31688b3b-bc2c-40b5-9c15-8cc59de059d1 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.316527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.316527Z digest=sha256:5e676b51364188c97cc6bed260a7ea3c333aa35f1be7130a1d24d26fd1202145

Observation 07ebc0aa-91b1-4fa2-a626-cbea7ee04639 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.321856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.321856Z digest=sha256:098d8e478148e948c2122a04b9416bb72c5a0a24094f34a18da663ad006dff7d

Observation 1b20fc12-e4e1-4461-94ac-0d948153fc82 · outbound

This paper cites Arctic: Snowflake’s Open-Source LLM.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Arctic: Snowflake’s Open-Source LLM

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.365891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.329082Z digest=sha256:c342fc8d1960a368021c929c806af4af93170872638a203a231cb768293a2ba2

Observation 192b8514-451f-4f73-8613-4e8da2c1d24a · outbound

This paper cites Qwen3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Qwen3 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.334915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.334915Z digest=sha256:556b17d5fd86d1f6f87fefcb859b137faa0399e6e8a1cf5f484c913c5b58f3d9

Observation 879cd6f2-c6d0-4b73-a76a-4efda08f2733 · outbound

This paper cites Mixture of A Million Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture of A Million Experts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.340157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.340157Z digest=sha256:901a79364a718640b92df137412217c83e682e34f9d11aef25243dc789528d1f

Observation 1f73fa79-2c97-4189-945a-659621f98fed · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Kimi K2: Open Agentic Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.345266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.345266Z digest=sha256:f56ae9be60dd000d46cef81a99028986aa88af6e3d8c1d3ad299aaa6682ae107

Observation ab426b07-485c-4686-a67a-560098855d76 · outbound

This paper cites gpt-oss Model Card: Architecture.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism gpt-oss Model Card: Architecture

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.345701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.351257Z digest=sha256:9a4be94be42dba335cb98658f979e232c55ee692d2254542f289211938c1437b

Observation d634a2cc-ec24-42c7-8c8e-5debaca623cb · outbound

This paper cites LongCat-Flash Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat-Flash Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.356618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.356618Z digest=sha256:7b73674f914aa6cc8ded2eea60026d8eed6dcb5cdf037c55304adf92255f8566

Observation 6b6f590d-e17a-4d7c-ab6a-4bd47c855fca · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.362126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.362126Z digest=sha256:f25da9db4c37af734a5e897a57e4f63365fe2ea34df90c82eca9720013cd87ed

Observation 8f27f14c-fbd6-44e2-998d-1f1b0f54fc92 · outbound

This paper cites LongCat 2.0 Technical Blog.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat 2.0 Technical Blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.317922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.368767Z digest=sha256:3a0dce6bf59e9efa3ab60edeb868c96ae215baac7cd21cef6446b9081a70ff03

Observation 898a66da-b382-4426-9b70-927251efbefc · outbound

This paper cites MoHGE: Mixture of Heterogeneous Grouped Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MoHGE: Mixture of Heterogeneous Grouped Experts

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.297216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.374274Z digest=sha256:ca7c0250294281a4b05c86a1f45a016469317c264830704498590785756e699d

Observation 6226d18c-7617-47dd-86bf-e488410062da · outbound

This paper cites GMoE: Global Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GMoE: Global Mixture-of-Experts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.277555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.380158Z digest=sha256:e9511e103e2e321ecca4fc2e458331e26141b0b0978d801499e383ba9de8a629

Observation 42e1fceb-5840-425e-a95c-c0374e5b5b72 · outbound

This paper cites Multi-Head LatentMoE and Head Parallel.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Multi-Head LatentMoE and Head Parallel

Reference 64

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:32:58.792734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.387367Z digest=sha256:ada9f77c3ec2410076a4055119cc288c6e85577ffcea4b577413b2eae713a4cb

Observation 6ea333fd-a000-4440-a917-4a2dd234c26b · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.393621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.393621Z digest=sha256:471575aa68122a437f1fc89eed4a8404ba63520e5e9614ed48398c195e88e675

Observation 1a7fe1c8-47ff-462b-8b70-68986a7d781b · outbound

This paper cites Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:58.684185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.404138Z digest=sha256:ce35555361931817dbc797206a405645312ea61ad7e8216defdfc8a18e8e20da

Observation 6bd2c283-d122-4b69-83c7-0fa779d27a17 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism From Sparse to Soft Mixtures of Experts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.409771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.409771Z digest=sha256:a469e960a7dbebff31a9c8cc4ebcbc1ee2753ab85bfa4c2f1102484a7cbcfb0f

Observation fa229c4d-40aa-428a-80f9-51b3769444ac · outbound

This paper cites Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.414782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.414782Z digest=sha256:1ed5e62bc40874387ef36918b6820fe1834cd738dc6ff74a8973853a40c04ade

Observation d413417b-2018-4afc-b784-c10df13c43d0 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.419634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.419634Z digest=sha256:7282d878f1ab9e2780b46d04a52a0695f663334b9ec1732f0c29a66cf6acda8a

Observation 21f6a2cc-cb6e-41bf-a386-d66cf431e1e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V3 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.425384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.425384Z digest=sha256:58e54bad551aeb0aaf431db4a8419a063429ddbff161925b217ba923e4f67765

Observation 5e4dd52a-1fc9-48fe-891b-a29c91793666 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.430621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.430621Z digest=sha256:1ac33279df8229b7f665aa58fc1716023a0eb509f2c496d3dfc3027cf1d289bf

Observation 47e0c6d9-4338-48fe-8537-2c9497c55301 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.435499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.435499Z digest=sha256:50e9f2e2749e4f49641e51ec87c258a4db141a53d2a4b1bc89eedc86955752df

Observation 96585586-e007-4b53-a86a-6917574ca7c4 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.441113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.441113Z digest=sha256:91ce319fd154f2255ef551c4f6782824cf01691a6b329cb4e8309e3eebe921b6

Observation a825bfe9-00c2-4a9d-aa07-a6cb2e9774a0 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Jamba: A Hybrid Transformer-Mamba Language Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.446447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.446447Z digest=sha256:beec8dadeccf15fee9778f75a22d04b8eacb08cd9fb493339bc3b5d7e9b7ab26

Pith citing papers

No inbound Pith citation observations are available.