Pith. sign in

Paper Citation Record · LEDGER

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

As of 16 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 0 inbound Pith citation observations for arXiv:2608.03457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03457 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:54:53.254253Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75965d8d-6ec2-4461-a75e-d3b7e7808433 · outbound

This paper cites Frontiers of Computer Science , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Frontiers of Computer Science , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.931183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.931183Z digest=sha256:c8aff61fb40f810401ee52fd0d5682033a4c1205472aed441540d1c74698b13d

Observation ba1ac7da-a0f8-40d7-a8e7-6c28e4cd711b · outbound

This paper cites The Llama 3 Herd of Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.935452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.935452Z digest=sha256:7dedb969d9c172b6c747423d6162a5a1bf6e4ea3d674730609527f483386c5a5

Observation 6206f049-45b6-4b47-a1f2-057f674e6eeb · outbound

This paper cites Qwen Technical Report.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.939622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.939622Z digest=sha256:eedb1f8b4adb6e37552a189e106ae11e25bea2bcfa5810037270bb893f5216e8

Observation cd1833d1-8ec0-444f-b627-d40ca7dd748f · outbound

This paper cites Qwen3 Technical Report.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.943691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.943691Z digest=sha256:399859d1301f9b27eebe5e1f5dacbe411e7fb2f14c54ae59e6deef6b63af6171

Observation 5120ac24-3eef-4f52-bd86-34eb8fa177e8 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.947620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.947620Z digest=sha256:049020f9162340842f6c4a9ccbffa6982fc033c13e9c943af4da569e9fdd4e52

Observation dd281ed2-ce47-4d9b-a082-15c36523ae7a · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.951619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.951619Z digest=sha256:491df514905a46a78bd93832245ca79a192fac7c89abf48201cfde8a10f09f1a

Observation 18685613-ac2e-4d71-98f0-9e35f00e16c3 · outbound

This paper cites DeepSeek-V3 Technical Report.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.955813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.955813Z digest=sha256:8e9dd1fb870d1740210601cb3ef38a19d8c1d1eb6aaf9654266d88a498807c02

Observation c8a59808-a0af-4ca8-a1f1-005b2ad354ce · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.959898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.959898Z digest=sha256:2e24f784f3ecfdbaa9f6ed602c63145e0de5500a10b7a8f3f70d0b0a9ca10aeb

Observation 22eb8403-d854-4463-95c4-ea3e3e228bac · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.963551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.963551Z digest=sha256:a572725fc84edf5194c401383d741b7b0d5513f5eb8bd92dd5766a5667ab6ddd

Observation e15a25f7-5d7c-4e94-969c-221164ee947f · outbound

This paper cites Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.967334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.967334Z digest=sha256:0c1aab364079a7c44da9e2a2966a08b96400f022ef351074fc7077ec126803fa

Observation dc40a1a4-1179-4ad7-a951-9077cd13a1ac · outbound

This paper cites GPT-4 Technical Report.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.970602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.970602Z digest=sha256:da71e9843df409b89257db8390d057bd8e535d8aa7b344bc2719b717a118afbb

Observation 1261dea4-a1ef-4ec4-a5d3-77413e3379a1 · outbound

This paper cites Deep Learning Scaling is Predictable, Empirically.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Deep Learning Scaling is Predictable, Empirically

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.974103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.974103Z digest=sha256:3af7980af3f30d46f48a320600be4be342ba74218f6df8c9a9bc87a57a9e189c

Observation fda5b109-84f8-41fb-ab85-fd56b6105446 · outbound

This paper cites Scaling Laws for Neural Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Scaling Laws for Neural Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.978034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.978034Z digest=sha256:0d4c2898ed8b9d8106a420edbd65a864b9e61cf8df58fae684a9fc18a223d99c

Observation 86b49540-92cb-4bb8-8f87-6361e045fad3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Training Compute-Optimal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.982080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.982080Z digest=sha256:1abfaa74306f3b2d6e0dc5528a9abb2da266b86e2a0a4c32d98eb7ceac219ab3

Observation 2efdf051-7fe8-40eb-9fe2-d19777b7cf01 · outbound

This paper cites International conference on machine learning , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International conference on machine learning , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.986024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.986024Z digest=sha256:48c44672a5c7fbcbe766fa7a4e7dc30027805b30a6e5f34b0654ae6b7a0e91cf

Observation 9cf074b6-ca93-445b-8e7b-e58043980931 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.989733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.989733Z digest=sha256:1583b26b95ee07e448d8f81aa25f88ccf445b9f1f3799292f13a100274db46c3

Observation 6084a2e6-0e5a-4548-bbba-024d20f45e6c · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.993025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.993025Z digest=sha256:b41ceeaf33f26199ffbd6b7455bb3f316aea882a8b9d398f8523b8deeba090a6

Observation bf369686-e9cc-4ac7-a6b0-1c2076095e8e · outbound

This paper cites Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:52.996839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:52.996839Z digest=sha256:e605ec03b8a55e9dfb70773f94ab6ba8959bc141f018f7700737e43af1ed8dd5

Observation 2c8905de-c029-4220-9843-7be93689f9a0 · outbound

This paper cites Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.000585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.000585Z digest=sha256:3c417385fcff96e0e680d2c938c799454fa4761199451c8c1d123119164c4a73

Observation be4cefbc-1fec-42f4-97d1-62342163d75a · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.004168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.004168Z digest=sha256:634751615d3fa1be36ed8a69f379c9082ca5b7b50f7e7c629af452c77899bdc1

Observation 3ba852a8-d13b-4cc7-882b-6c663d6732e0 · outbound

This paper cites 2018 , publisher=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models 2018 , publisher=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.007719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.007719Z digest=sha256:b352e426bb2cf019f559e4c1ef7c2c96caaf4a0276a2218178a27420dce0123d

Observation abe33383-4864-442b-938f-6f8911245f92 · outbound

This paper cites OpenAI blog , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models OpenAI blog , volume=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.011294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.011294Z digest=sha256:f9fdbd2f2ebc3a2a88745b68e43ad11c798007ea6d88394e15974d753d72c2af

Observation fc90df4e-66ed-479a-b946-09a2ff8ce8e2 · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.015095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.015095Z digest=sha256:322d0068fdd26d4c4564828e052d8e7ec5ebc353d2e6c8618796365d36368810

Observation d6e0ba70-a55b-4ca5-ab0c-c4ed9095b594 · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.018393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.018393Z digest=sha256:dae8baa7ae5953de479f3f7bca443bdac3c38a5cfe57976f671add156d56c113

Observation bbc2df56-26f8-47c8-a697-d0c3d954d205 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.021652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.021652Z digest=sha256:276dae1b5c03848e5400c3725c00944b89a150674e87f9b36b557e95e9631005

Observation b7b2eef6-8930-4cff-a021-6e94b1bfb31b · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.025027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.025027Z digest=sha256:977aaf0e83bdc3f25690af18b521b7dc442effad1a9e49adddbd5a8c41a7001c

Observation 94dde3b9-4d63-44c1-bbe4-9d8e1819f529 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.028466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.028466Z digest=sha256:7bf755cb75735c06170fa5c299552be07aa387049dcb42cca3b2671469cce4fb

Observation 64391c34-2502-4223-8574-6c02c1046dbb · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers) , pages=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.228695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.032576Z digest=sha256:111acf5522a50b4a66bc13efa163f1f5fa72ddfaff287b3de7d9119f55b85e64

Observation b5df5b26-3242-4dc1-b488-b105e9f40b5a · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.036796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.036796Z digest=sha256:a04156e1302d6bb2971136d29bcee6634358cd810e6197812b68a478deafa6e6

Observation 9ea58332-30cb-4e3a-bf35-893d7d6cdbcd · outbound

This paper cites Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.040691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.040691Z digest=sha256:913cf1811d1cd7fff89ccd86c4fa573d6e2aa98dc9807a80027e98f46708d338

Observation 64a39fda-29e4-44dd-9656-95222abf72b7 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.044551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.044551Z digest=sha256:8ca44ab0aa97424e8012a8cb29169b2bfaa327b82477abefec383fac64a01a5a

Observation e47d7054-432e-4b94-82dd-ff89b5a2b5f5 · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.048318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.048318Z digest=sha256:362fe8f3d5bfb9e7f74bb84003d1a100589685c175d1448d2c449335936cc273

Observation 2d85350a-7c88-4fa0-a872-c05b11f94d12 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in Neural Information Processing Systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.051918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.051918Z digest=sha256:d3d0dfb7c779858651a5f6356ebb5110082239effe520c00c017c224fe3d1b47

Observation 5df4ee44-4a60-43a5-8a27-60fbae1cd7ca · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.205519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.055822Z digest=sha256:17ba4281e1e1e8ed48aecd818e6d5268ae656270c8c5770dfc00db351ad6e5e6

Observation 6ec388bc-2f80-4282-a437-77eebb97e1f5 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.194608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.059491Z digest=sha256:9c1903c6b13783fdd37a24115659fc5672533c5baa19b418efe82965b24fb3e2

Observation 991f974d-388a-4041-ab5d-0d47ee751760 · outbound

This paper cites arXiv preprint arXiv:2510.03280 , year=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models arXiv preprint arXiv:2510.03280 , year=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.063186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.063186Z digest=sha256:02a04e46f733f8e8c8908a7462fa5c0146d525c5d43fe941270bf31cfe778f13

Observation ebd66ced-8da8-4cbb-882f-4d0d663ccc9a · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.184057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.066854Z digest=sha256:46f9252abf195702e451392261e3be72e25553e16d884b3024a6a8c38a4fe109

Observation caef4893-8236-4384-a843-1601999a7d3a · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.070324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.070324Z digest=sha256:5aa9f608d8b5a6f95f88958d924b174c4f4937c3fbaac9aaafcf753522e5c928

Observation b495b72f-0cbf-40d9-9922-322c895a6699 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in Neural Information Processing Systems , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.073982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.073982Z digest=sha256:d255c30515afbe80c110cc9e5f86b65524a688b171f9313329f58a3cb4683720

Observation b928897f-33da-4b3a-b410-5b6577d4c080 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.077088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.077088Z digest=sha256:0691ced00b526d0508ed8cb3f4d567b56261212b652fbf3b20312a9206c3bb52

Observation 8b4bf501-ab9c-4a01-b38e-4509805f46e5 · outbound

This paper cites Improved Large Language Diffusion Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Improved Large Language Diffusion Models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:54:53.715349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.080402Z digest=sha256:89190713b38f1d4d1c127452deb5ea79c3eff7e9767bd4b689f9f3e55e685c98

Observation e7b3a87d-d071-44f8-9d64-2122bdf209ca · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.153984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.083853Z digest=sha256:d79df589915c96e4accd2d3daadfb9c1cfdb08e63ba7d103762405fe9a73a236

Observation d5cb847b-dddf-4ede-9f40-ccb5db8a54e1 · outbound

This paper cites Dream 7B , url =.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Dream 7B , url =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.087165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.087165Z digest=sha256:833e7fe547c988c2c1c1ea1f4492af55ced59945aabb060f85704feb1f33b38c

Observation 6026ebe9-b289-448e-8c25-3cc802fd7b5a · outbound

This paper cites arXiv preprint arXiv:2509.24389 , year=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models arXiv preprint arXiv:2509.24389 , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.090534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.090534Z digest=sha256:5b44a49ac6fca50b9b8abab5be5e21a69d0232a546d2b6305125790aaa8c27bd

Observation e36e7b3b-9ea2-4c3b-9408-063d72d38c08 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.094173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.094173Z digest=sha256:24c4e5f699edbcc50b58997b5e660a6456d2c6763ee3145ecb7a61d034bbc53a

Observation 47124b80-aab9-46b0-8857-a98b7dbd21d8 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.097941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.097941Z digest=sha256:a3025aa9735262aed2f906e61f6b74162f4ed90d46171e16ed6766a1fdafcf10

Observation ec5f3e20-305d-4036-9425-4c39d8f7f448 · outbound

This paper cites A Survey on Diffusion Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models A Survey on Diffusion Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.102557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.102557Z digest=sha256:d259e52e5ba92d63a6835da290a058cdbed8132750e551b8cf477046fc0662e7

Observation 128f563b-0b1a-45cb-870f-a31c81a7745f · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.106073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.106073Z digest=sha256:4ef587cfcfab9e996ac002385ea4c72507cea9d935dd17f6702cb814a844160e

Observation 7947933a-bb29-4922-ac17-10bac9834349 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.129423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.109809Z digest=sha256:ff6609775b6ec296b08cf9d26ef6dbc045505bd175fc0d3de497a30a083ae625

Observation 50b3f40e-14c6-4748-9c2e-618650574541 · outbound

This paper cites arXiv e-prints , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models arXiv e-prints , pages=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.118248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.113077Z digest=sha256:ee0e62d20e87144ccbca8ca943577dbc945ff6c43f98c15b466d20384deef176

Observation bea49905-124e-454d-b4e4-13d014752b44 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.107425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.116586Z digest=sha256:447588b62d2d13cb959e3b62e7c8994fbbcfb662b312f53a5870519ae53cebef

Observation 97df7726-25bc-4991-8e62-b9f26c9b7749 · outbound

This paper cites arXiv preprint arXiv:2602.15014 , year=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models arXiv preprint arXiv:2602.15014 , year=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.119929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.119929Z digest=sha256:93a1e0cc1f3b7bef62406d749117fe0b0072ba9059ca7eeed857dd52a511db65

Observation 126e29bf-78bd-4fc2-a8a1-55a52d91e9e3 · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.123132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.123132Z digest=sha256:0906cd2b905b2b47d43a5864b59394495ef499288058272aa078d6f3deea3492

Observation d857c1f6-a465-4610-ab3f-14ece18875ea · outbound

This paper cites arXiv preprint arXiv:2511.03276 , year=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models arXiv preprint arXiv:2511.03276 , year=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.126756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.126756Z digest=sha256:ed2fe03561fa19f7859578dd196e7c2b16cd0dcbaa812df7506bca3dba3be6f5

Observation 8b974f91-7218-4993-86a9-99f6a18de55e · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.096808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.130377Z digest=sha256:152b202877f50dff55abaeb92c14e4a5ddb1bdfb92357d653a202689bdceca43

Observation 04a14416-6abf-4296-a821-17e29a4701d6 · outbound

This paper cites dMoE: dLLMs with Learnable Block Experts.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models dMoE: dLLMs with Learnable Block Experts

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:54:53.464595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.134040Z digest=sha256:f71413e5f127aa7103321bcc41c8029987d966b9671838634bed0accfb6f6e75

Observation b202179c-5108-443f-b91e-3c66846ee8b3 · outbound

This paper cites Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.137569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.137569Z digest=sha256:22fa9bd899f5d0387a7f6cea96ce2a5cc47d883af7f55767d18178f61c7b6779

Observation bf0ce0b0-b63b-4f20-b9b8-f61bc24469a3 · outbound

This paper cites DFlash: Block Diffusion for Flash Speculative Decoding.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DFlash: Block Diffusion for Flash Speculative Decoding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.141263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.141263Z digest=sha256:68c24c929bbaa3669dd89cfe97dcedba967d19ec485353e9b96e27c50d107b57

Observation 5ad64725-31d1-4a65-bc11-652075915d36 · outbound

This paper cites DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.144855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.144855Z digest=sha256:7b4f5944a3c499de026934527b7708bd705a5cb345ff8781c034809bac70cf98

Observation ea22e41c-ea76-4029-961d-782f7116a91f · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.148461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.148461Z digest=sha256:c68ae929f1e3ecca012051025ff4ee34d78275212274317247cb347fe897b61d

Observation 51d4c490-9b38-4c25-a012-5833966ad492 · outbound

This paper cites Decoupled Weight Decay Regularization.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.151806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.151806Z digest=sha256:15ebc1838af1e977f0d135423b6d7e5e2528a0d0f6ecafde4e5051c84ba50046

Observation 845867cb-d805-4ba5-b183-fb495e42741e · outbound

This paper cites GLU Variants Improve Transformer.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models GLU Variants Improve Transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.155681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.155681Z digest=sha256:ae2b9da899f77b6789c12b824268fc3a9afd07b29e36318d7106e82555508165

Observation ca6b9a35-6527-4bcc-b4cc-22d3c76976e7 · outbound

This paper cites Proceedings of the 2023 conference on empirical methods in natural language processing , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 2023 conference on empirical methods in natural language processing , pages=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.079835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.159783Z digest=sha256:35cc603ccb820c099d369c03c039da43c84cdbd0dfab2ce9320704240d38e578

Observation 83023f4b-91b7-4c9e-95ac-0af83c8640be · outbound

This paper cites Neurocomputing , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Neurocomputing , volume=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.163109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.163109Z digest=sha256:d966cb810a63c51b0f301884281c7b5b4fea75cdd25e0413c6ddb2301db9d2b0

Observation 95ee070a-6525-461d-8916-4814d13eb6f8 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.062887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.166414Z digest=sha256:11210b95893fa68ff7a5769f30c3debad7a13eef79e29c95c5a1047fea265013

Observation 02e049a8-2cec-46aa-ab57-c6bb771b34fa · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.051623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.169811Z digest=sha256:f23051958fa206a533b613883a01266becfcdfacef483307a4094a58278e2144

Observation 1bc886c4-2d67-4239-9e0a-ad01dd08c488 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Scaling Laws for Fine-Grained Mixture of Experts

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.172947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.172947Z digest=sha256:88220c3c20cc93a42b7d1f8201202727f7aeaf0e57e5989a18ffdd63c1f61451

Observation 22173e29-ed15-4a68-86f5-71243fd2025f · outbound

This paper cites Mixtral of Experts.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Mixtral of Experts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.176291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.176291Z digest=sha256:7074c7fa0095c99d5e7772f7ea212acc140f974b8ff875b11598cd879279def8

Observation d685a404-6ba5-4865-ad6a-a285e8e854f5 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.179929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.179929Z digest=sha256:458195b20a1a7b3505eb00bbbb6d84fbb62eaa44a4c2fc1d2e70312c531259de

Observation eae97c02-4547-4bfa-835f-624bdc6a14a2 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.183723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.183723Z digest=sha256:b91ea6f347ea29fbba5e971f2e9e328d1892238548d849d87120feaf1a44ac4c

Observation c13e2aa1-c830-4d9c-a9e9-2e35a3a3dd22 · outbound

This paper cites Journal of Machine Learning Research , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Journal of Machine Learning Research , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.187768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.187768Z digest=sha256:e02798baa9edf377c0dad20541500b4913c2b3bdd0f9433bda708eff91fb35d9

Observation 604ca863-9c4e-4f37-97f2-8fcb9b3ec316 · outbound

This paper cites Neural computation , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Neural computation , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.191060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.191060Z digest=sha256:fdf289d0cc377cd614011c8472e3c3e524a5fe353fa8deff670238d9cbb1f0d3

Observation 947ee811-a441-43d6-8425-c789a6daaf12 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.194432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.194432Z digest=sha256:d7fe28f7a8a0f5ac14df5a64b8e9f361ba377f77e336ee817d4e46d44a07618e

Observation fa441e14-d273-402f-bbd2-6bf854e1bf70 · outbound

This paper cites Muon is Scalable for LLM Training.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Muon is Scalable for LLM Training

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.197951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.197951Z digest=sha256:80ea0ed675fd38118a72b427f414411cdc05b7d1a47c0cec4a139c9ad6707c73

Observation a5a3322d-bf8b-4b3f-9a36-1a3258859e74 · outbound

This paper cites International conference on machine learning , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International conference on machine learning , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.201763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.201763Z digest=sha256:c3f3fbe719bc62ef1e3b2bdbe0ecef2a6dad85ce9b0ca20e75febd9a0e866b37

Observation 95d0a104-3ced-46eb-a038-32bfc6212eaf · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Measuring Massive Multitask Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.205217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.205217Z digest=sha256:afbe846ddefff563feb3e22d6dbf70a1d9aa28c5045b0d603aaa7d3888c19bd1

Observation 2422a981-6b46-49d5-b684-b18badab8041 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in Neural Information Processing Systems , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.208822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.208822Z digest=sha256:84ec4e9c6140fe32f8260c1530aa95687ded0e3513a757893961abf21807f386

Observation bc595300-3fab-4f3c-ad36-40fc9555a072 · outbound

This paper cites Advances in neural information processing systems , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Advances in neural information processing systems , volume=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.014803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.211938Z digest=sha256:f2d17504a17e7835ecaa7746350d77aa4d66a8ea15f1f50d32ee73b89a7edaec

Observation 0bc3ffe4-4c02-428f-951e-5f4f451b8726 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:54.003972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.216328Z digest=sha256:44d80cbee5e27cb8d3113f657b56888c650eaaa6789ffd1ccc8060f2c218c712

Observation b3c605a8-5ef1-4cc7-a74b-24060c9b396e · outbound

This paper cites Proceedings of the 57th annual meeting of the association for computational linguistics , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.219903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.219903Z digest=sha256:5f77bec456c55a3c371dfdf762313a8908926abd7f47b4618264e861037934fe

Observation 4d43f2dc-96b1-485a-80b2-636aafdc4cb3 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:54:53.986631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:54:53.223215Z digest=sha256:819daf26ce3fc7583f51ee3da2a30aeeb560bddce1cc6598ca970fb9bed73040

Observation 1d8e1eb1-a840-4bc0-9732-e8513ee7ae0c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Training Verifiers to Solve Math Word Problems

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.226271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.226271Z digest=sha256:d4539bad106d96c187d393f060afe4287a8fe1cbb66c63954887070585de6d44

Observation bd29b304-ae32-42cc-b23f-6dc2f5032224 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.229762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.229762Z digest=sha256:b36a9b7e28d8f60d79617dd35f636f03519de8cb857a2f6102df4e066e00937c

Observation a96fb5dc-0f73-45d4-8c81-af9c6890ee5f · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.233320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.233320Z digest=sha256:e4a77eec529a911963ea87f579c60efec5aa7ed5fee94e9256856f3091f1af7c

Observation bed8744b-f7a2-4afe-bb78-b81373ef9f4c · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.236855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.236855Z digest=sha256:f900f46a6bab3604e27cf7612c4bbef2db6517f152d8d767779ebbfbb76dc8c4

Observation 4ae4e16c-ad30-4a66-af79-3e4389f958de · outbound

This paper cites Program Synthesis with Large Language Models.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Program Synthesis with Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.240344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.240344Z digest=sha256:a99dc502b4e665bc3b1c3acee1b738bf12659a5ee2cf1eabff66fafbaa0b865b

Observation 878ca4bd-0806-4965-b514-8313bea22b41 · outbound

This paper cites MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.244102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.244102Z digest=sha256:db3bedfa996fb0d0ae508a95e42cd331d0ea6746c830324ec1fbad97b99035e5

Observation e3938718-344f-4cad-977e-a9f1ba294cc7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Evaluating Large Language Models Trained on Code

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.247943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.247943Z digest=sha256:e687cff72232519f8b1d8516b85c3ef58abe1d6bd1e5c12828a25ef8c25b1c9d

Observation 323cffd1-091f-470a-9e22-088ea29f9bd3 · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.251157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.251157Z digest=sha256:2a69a36a52ed673fae3cb71e1950f9815b6a113b820f3607e30d8ca0c02c1da4

Observation 1e354280-3689-4878-b3df-4cfb7fe5d97d · outbound

This paper cites International Conference on Learning Representations , volume=.

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models International Conference on Learning Representations , volume=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:54:53.254253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:54:53.254253Z digest=sha256:472534d4b02bc6f1c9ffbaa49dab9d80e07253499b4e62a7ea2998430ffba68f

Pith citing papers

No inbound Pith citation observations are available.