Pith. sign in

Paper Citation Record · LEDGER

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.08650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08650 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:58.446447Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 155412e3-7d0f-44fa-acc7-9d39d244d27b · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.257784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.233177Z digest=sha256:c55b8582f3a23c8511c51b9b07554e68b8433b007970dc2f41d72e1f23d06de3

Observation f7f081eb-159f-4567-ba8b-42f519014d53 · outbound

This paper cites an unresolved cited work.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:32:59.232328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.238758Z digest=sha256:ec535c268514b5d16a242ba6301378175ba71828352dc2ab174806b7b9159910

Observation 88128731-e303-4cc8-b293-7508c5abf31e · outbound

This paper cites 总 All-to-All 时间.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism 总 All-to-All 时间

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.213592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.245187Z digest=sha256:519b8e311085fbdbbf49dbb100832974fc2154bd82f5089f14610c2d2e5bbe15

Observation 1b015d02-473e-477d-8c0c-00fff5e8c6e9 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.384791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.251293Z digest=sha256:37fce87e4cabdeaa5fb76f600963d06a3413bd249cffd09f55c79b68fa0a93fc

Observation 78903f03-9259-4bf6-9fd4-697cb33ce04b · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.257289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.257289Z digest=sha256:6de45d37005eac6715e74078666890ce291ce15e253c070eb94df1d85fad8bf2

Observation 8f2b9ad3-4bc8-440b-b148-5692f207abeb · outbound

This paper cites A Survey on Inference Optimization Techniques for Mixture of Experts Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.262836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.262836Z digest=sha256:7961d3d77aedb861535f47227392f027268e3e0c69d6c358f45b53bacd49920e

Observation 61c66eaf-ae8f-4359-9f8e-fc8da2449351 · outbound

This paper cites Speed Always Wins: A Survey on Efficient Architectures for Large Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.267934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.267934Z digest=sha256:58e78b3581f5bb6d54b3bf390680255924185ee89858ecb056d700eb0be42abf

Observation 8984cb4b-f455-4def-9e56-094faf9ff28f · outbound

This paper cites Adaptive Mixtures of Local Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Adaptive Mixtures of Local Experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.273267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.273267Z digest=sha256:c2041c1b3c1ba25ff593c00753e01bc10679ee73f928853c0404cf132a9e5079

Observation 2a1ecfed-c799-4c62-858a-1ffa1c4724dc · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.279219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.279219Z digest=sha256:9cb769769750c368b83bec6d5d3c081153c7f8528113ab30ec17071ee1369ed0

Observation cd8cec16-c0d2-493a-953e-44a51c39c2ad · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.284500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.284500Z digest=sha256:ffc2ad155ce275c6da6a4a9c5dc88bb3a9a7e6ab75d7761bcd10d761321e20d4

Observation c43a1dfd-0197-4799-aa11-1113006dd440 · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.289591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.289591Z digest=sha256:91fb0036d99a672302f51020abdea502247b6777486b5264e6b93788c8ccf976

Observation c45be21f-5096-447c-9f5f-f1fe7941a958 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.294305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.294305Z digest=sha256:46cb4d8b4126fb8498bd51877e9b6048b91917d051c423c1aae42da5491dbf40

Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · outbound

This paper cites BASE Layers: Simplifying Training of Large, Sparse Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.299078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.299078Z digest=sha256:215531f0e5daab5eff8cdb47dbdfdc0840546ca0e933b15dc35a70885b2d6a91

Observation 32714d3f-e41a-46c4-90d8-7886c02c8b33 · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture-of-Experts with Expert Choice Routing

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.304689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.304689Z digest=sha256:31d6e5bd26fe300d6c9820b23ea3642a445a9d0874c59ff8d0dc7c530a818bc8

Observation 61850748-812c-48c4-941f-dec0ffaa1b60 · outbound

This paper cites Mixtral of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixtral of Experts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.311824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.311824Z digest=sha256:98367d9e1e1ec249ec9a617544d6c7051949aa9eb60894e866f86edd793f78b6

Observation 31688b3b-bc2c-40b5-9c15-8cc59de059d1 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.316527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.316527Z digest=sha256:06b9748120959b71a93bab9f517a0098b452b4b8e8976b6d6e4e5e2561cd854c

Observation 07ebc0aa-91b1-4fa2-a626-cbea7ee04639 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.321856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.321856Z digest=sha256:684c13664c734508046d0f7ccc7a373b26c9e73d3db0403d4eb87642f210d64f

Observation 1b20fc12-e4e1-4461-94ac-0d948153fc82 · outbound

This paper cites Arctic: Snowflake’s Open-Source LLM.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Arctic: Snowflake’s Open-Source LLM

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.365891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.329082Z digest=sha256:9b2f68bcc47d82aeb643a3c54cea7636f05db0ededc6bdf7334901f124911f97

Observation 192b8514-451f-4f73-8613-4e8da2c1d24a · outbound

This paper cites Qwen3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Qwen3 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.334915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.334915Z digest=sha256:f1672a34fec4dc73425cef49126ff299b50ea38b2458897c7abe2547c89758fd

Observation 879cd6f2-c6d0-4b73-a76a-4efda08f2733 · outbound

This paper cites Mixture of A Million Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture of A Million Experts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.340157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.340157Z digest=sha256:f356432d37c1e98b930eaf7e6eefac3e76c672e2be1ae67fb9d4782a9c78b327

Observation 1f73fa79-2c97-4189-945a-659621f98fed · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Kimi K2: Open Agentic Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.345266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.345266Z digest=sha256:bea7de6bd16bfe80d4a9368e5a1bbff00a7ddf175622de4c65fd332dafa3455e

Observation ab426b07-485c-4686-a67a-560098855d76 · outbound

This paper cites gpt-oss Model Card: Architecture.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism gpt-oss Model Card: Architecture

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.345701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.351257Z digest=sha256:40db25e9ee51dad22e2d5a6ac7faf519776d0ecc2bb788832d3060d025e0c93c

Observation d634a2cc-ec24-42c7-8c8e-5debaca623cb · outbound

This paper cites LongCat-Flash Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat-Flash Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.356618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.356618Z digest=sha256:930c305927966ded117bd4da7beb9cfbce28d26e03631dffa09ab1d69079b42d

Observation 6b6f590d-e17a-4d7c-ab6a-4bd47c855fca · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.362126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.362126Z digest=sha256:5889255ae63a560f3b5e314c775360273fe2b572c52d44d5bb8a5bf8333a72a4

Observation 8f27f14c-fbd6-44e2-998d-1f1b0f54fc92 · outbound

This paper cites LongCat 2.0 Technical Blog.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat 2.0 Technical Blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.317922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.368767Z digest=sha256:8691becf641b62ee246df0766a3b4b66840ce82094e1cb1a3b96fc6a8d6b0af8

Observation 898a66da-b382-4426-9b70-927251efbefc · outbound

This paper cites MoHGE: Mixture of Heterogeneous Grouped Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MoHGE: Mixture of Heterogeneous Grouped Experts

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.297216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.374274Z digest=sha256:57c108de6d9ab01bd35326c851b935f4cdc9fd380ca7d796efe2feec98c213e3

Observation 6226d18c-7617-47dd-86bf-e488410062da · outbound

This paper cites GMoE: Global Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GMoE: Global Mixture-of-Experts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:59.277555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.380158Z digest=sha256:9dca2a68e73b108389850094cb7a66f8fddf731b0c12ef4bdf4b901d09b85176

Observation 42e1fceb-5840-425e-a95c-c0374e5b5b72 · outbound

This paper cites Multi-Head LatentMoE and Head Parallel.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Multi-Head LatentMoE and Head Parallel

Reference 64

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:32:58.792734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.387367Z digest=sha256:420e9ecce06740163944adb3d3cfbfd52b94e3623f264030bfaafa4da5bdde1d

Observation 6ea333fd-a000-4440-a917-4a2dd234c26b · outbound

This paper cites Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.393621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.393621Z digest=sha256:1b9a98a19599250beb9b20a8ca22b19525e7891fc752c17732aa27da809678d4

Observation 1a7fe1c8-47ff-462b-8b70-68986a7d781b · outbound

This paper cites Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:58.684185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:32:58.404138Z digest=sha256:911997e19f42b712e5910db94125f784f04389579499d94e457d70297f7ef63a

Observation 6bd2c283-d122-4b69-83c7-0fa779d27a17 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism From Sparse to Soft Mixtures of Experts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.409771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.409771Z digest=sha256:db2f0631014bda29396c013630db9045a132aa8162642369b84bdf8948daa862

Observation fa229c4d-40aa-428a-80f9-51b3769444ac · outbound

This paper cites Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.414782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.414782Z digest=sha256:f863c7e90e9ac6f9fad0a6941d5da6aaa834fe3942125a4a8e751a7370f4ef88

Observation d413417b-2018-4afc-b784-c10df13c43d0 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.419634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.419634Z digest=sha256:2dfba3671aeed59813de49252d25cb2ffa0ba51049f86b3332039eabeb39f47d

Observation 21f6a2cc-cb6e-41bf-a386-d66cf431e1e2 · outbound

This paper cites DeepSeek-V3 Technical Report.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V3 Technical Report

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.425384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.425384Z digest=sha256:48ffb7b647e60787c58992b77937edd05c2510df0fb4347510c65d70eebf38d0

Observation 5e4dd52a-1fc9-48fe-891b-a29c91793666 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.430621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.430621Z digest=sha256:f17167470b6b394ec8756f045c11c0858bb9b6985cb92ef4ba8c5b99284abd2f

Observation 47e0c6d9-4338-48fe-8537-2c9497c55301 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.435499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.435499Z digest=sha256:c05f7b8c74d62b3b4aa517df559cdfac2c47174995e3c36e1d57a6a4e101b223

Observation 96585586-e007-4b53-a86a-6917574ca7c4 · outbound

This paper cites JetMoE: Reaching Llama2 Performance with 0.1M Dollars.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism JetMoE: Reaching Llama2 Performance with 0.1M Dollars

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.441113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.441113Z digest=sha256:6f23d9447ee481f13c0f83819adc3a0c673b422bee86f0f114c3a621038c9349

Observation a825bfe9-00c2-4a9d-aa07-a6cb2e9774a0 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Jamba: A Hybrid Transformer-Mamba Language Model

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.446447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.446447Z digest=sha256:9313f6bde498519cfa28824b66f36253be1667c736d03fffd535cbad0550ffc0

Pith citing papers

No inbound Pith citation observations are available.