Pith. sign in

Paper Citation Record · LEDGER

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing

As of 19 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2509.09420.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09420 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:16:55.968542Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c023fad7-e42f-4af1-91e1-f30712da59c3 · outbound

This paper cites DeepSeek-V3 Technical Report.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:53.978741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:53.978741Z digest=sha256:cd7aff1c23b9127e96d1472e51979eb561143b998fca28ab60ee0846a9f9d2a5

Observation cc95a5cd-9a8a-4919-8671-0fd2e916564b · outbound

This paper cites Mixtral of Experts.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Mixtral of Experts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.034317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.034317Z digest=sha256:b66f851bb76ee80cb282cabff95c937c6f773b97ea29eb23154e150c7866c117

Observation c9223967-e2e5-404c-91f1-74b5f48ccdbc · outbound

This paper cites 1.1 computing’s energy problem (and what we can do about it),.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing 1.1 computing’s energy problem (and what we can do about it),

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.123728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.123728Z digest=sha256:1b15bf7c64f49d773be63997b9073b67017195577edbcb7c907a9dafdc259887

Observation 648d718b-9cb3-415e-8a45-de80bf678ca6 · outbound

This paper cites Near- memory computing on fpgas with 3d-stacked memories: Ap- plications, architectures, and optimizations,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Near- memory computing on fpgas with 3d-stacked memories: Ap- plications, architectures, and optimizations,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.207966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.207966Z digest=sha256:46d01aa2d0fe93f81a0c01fc84fc2a21e83d10f09eae9b7b70e15196d5561906

Observation 8c005d75-46d8-4b02-8e66-8267716ce092 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.346068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.346068Z digest=sha256:dc3ce21625ec31c1e4cdaccb8bc606b571fb629bae2767c503d2fb22af4a73a1

Observation 868f603e-6201-4cf9-b30d-af78ff293680 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.437114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.437114Z digest=sha256:dde5feb12d4d57c36351751b0d1bea8c17c620177d2c1fbb3d2cd8f502543808

Observation 35deddcf-3cdb-4a0d-97fc-b175cceb5950 · outbound

This paper cites Adapmoe: Adaptive sensitivity-based expert gating and man- agement for efficient moe inference,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Adapmoe: Adaptive sensitivity-based expert gating and man- agement for efficient moe inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.528051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.528051Z digest=sha256:6a06f88ae345276431b83cb311184ad89ca160551ffce60e7f6b941e5a32c45a

Observation ad2c06d2-4f6e-4a83-b3c2-cd3417e7323b · outbound

This paper cites Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.612324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.612324Z digest=sha256:bd46c9877003d9f068f92cce1fa8c3d3a075d63efa2434e297b5294a248d720f

Observation 21df77bd-7575-4408-817d-5eea93f53ce4 · outbound

This paper cites DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.707195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.707195Z digest=sha256:e392d89051f4c56c38c4e8bdddf43d6dc95763b9e3707b21921c0a9ae40edabd

Observation 50eddfa0-7cf9-4198-978a-26934580ec3e · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.812983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.812983Z digest=sha256:5adeb71627bb52d417abc2f6f96c67ccf6b7498446918c4c5c711db339b3eab4

Observation b750e429-49c5-437c-9696-c99f7cb15091 · outbound

This paper cites Qwen2.5 Technical Report.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:54.942933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:54.942933Z digest=sha256:54fe2e589702522f0f8d48476401c35d5da02084ae4ccfa4a96c1447d363952e

Observation cd6e5e2e-36bb-47d3-a2b7-56376da3465f · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.009024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.009024Z digest=sha256:272b52cbee9426b1cb7250678864bcb12c34317f8e6704fb9aaa773d8be81ef2

Observation 9964d58e-d9b0-44da-8b7a-ef29cd5b4439 · outbound

This paper cites HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.064846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.064846Z digest=sha256:8fd47c6662a2c52cd0c045a13fa4d8612c7415a56555f9b2bea7a4e0de1bb109

Observation 39938ced-e2cb-4a12-b64c-37d5f53e01ee · outbound

This paper cites EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.123761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.123761Z digest=sha256:d02f1532363b3073c8753a5b56fae45c2ad35c172427ed0443efc626975696d4

Observation 13b6405c-ff34-49f8-811e-c463f0119c21 · outbound

This paper cites MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.187513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.187513Z digest=sha256:c555d874ac2221aadfa8f8c931b12b319e5dcb3672950d9e335bf38afe44e56b

Observation b511cd5d-fc47-4c18-b843-9628e6c05bcf · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.265118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.265118Z digest=sha256:406df6ad118879343dc6c8fd036087acfff72fd446906ee7a4ab008c7162e1c3

Observation fdf83db3-38a7-4543-b181-a7275d6628fc · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.306591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.306591Z digest=sha256:b4441040472d9817ede3b655ba8a5eaf9fa335db26799cda749a8119bb7a341d

Observation b14e6847-d7fa-40b6-a19c-e66ec4f544c2 · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.337181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.337181Z digest=sha256:02bc86ba83ac5ef65f296e4197b9047e67261e541f3ff6f31b4569f310ca3ba0

Observation 5725ad46-9614-418c-a293-1b4b20c56505 · outbound

This paper cites Netmoe: Accelerating moe training through dynamic sample placement,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Netmoe: Accelerating moe training through dynamic sample placement,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.387250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.387250Z digest=sha256:040b26b2936d6f3876868a897814065b795ad937344ae3e75c2e411b4501faff

Observation 23b3792a-05cb-4a57-af79-7107bfd2c085 · outbound

This paper cites Samsung pim/pnm for transfmer based ai: Energy efficiency on pim/pnm cluster,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Samsung pim/pnm for transfmer based ai: Energy efficiency on pim/pnm cluster,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.469355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.469355Z digest=sha256:d1e81ad11d5717b29d1c63319093f5372f2cd7c732ea35a249ac8f7e4d45554f

Observation 1246a6e4-ffb2-4a65-8ab5-7feea4864155 · outbound

This paper cites A stacked embedded dram array for lpddr4/4x using hybrid bond- ing 3d integration with 34gb/s/1gb 0.88 pj/b logic-to-memory interface,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing A stacked embedded dram array for lpddr4/4x using hybrid bond- ing 3d integration with 34gb/s/1gb 0.88 pj/b logic-to-memory interface,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.511861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.511861Z digest=sha256:9d6a9cc5dc71e40032b0f5d71a3456ce9d1ee2bfec3fe4b5dc15adb8b9cc9f05

Observation ada569a1-975f-409f-b4ec-24430e250a57 · outbound

This paper cites 184qps/w 64mb/mm 2 3d logic-to-dram hybrid bonding with process-near-memory engine for recommendation system,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing 184qps/w 64mb/mm 2 3d logic-to-dram hybrid bonding with process-near-memory engine for recommendation system,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.584274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.584274Z digest=sha256:6359d51c2061b4a0cd51c13c14814ca16503105799cbe7161df8d6f7040a6650

Observation 7a647bae-3091-4ebf-b727-a4aa758f498f · outbound

This paper cites Exploiting similarity opportunities of emerging vision ai models on hybrid bonding architecture,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Exploiting similarity opportunities of emerging vision ai models on hybrid bonding architecture,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.636411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.636411Z digest=sha256:8fe9c09f1366d02bedd144edd45f0c1c0662b4b42cccf922bd7f8945de92c95c

Observation 625e7d78-ba1c-46a1-93b1-226d046f6478 · outbound

This paper cites H2-llm: Hardware-dataflow co- exploration for heterogeneous hybrid-bonding-based low-batch llm inference,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing H2-llm: Hardware-dataflow co- exploration for heterogeneous hybrid-bonding-based low-batch llm inference,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.703749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.703749Z digest=sha256:f5e3825ca6531aa8be8e58965ecd445ed460ede9839ab63efa91f900dfcd3316

Observation c2fdafc8-2394-4fb8-9d3f-6f5e16e651b8 · outbound

This paper cites Qwen2 Technical Report.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Qwen2 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.758949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.758949Z digest=sha256:1cf464e75b50fe3baa57e6b27ae3449bf69ceb41d2e2ce853ba0c5d119cc4558

Observation 6040e835-9af4-4422-a13c-ecb910596444 · outbound

This paper cites Collec- tive communication on architectures that support simultaneous communication over multiple links,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Collec- tive communication on architectures that support simultaneous communication over multiple links,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.802746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.802746Z digest=sha256:337b982355a96c3feca12fbca5c6a7c110d9733ef366d619e5f49d1b9a79e322

Observation 518b3eac-fd3e-4f57-8556-34896c311998 · outbound

This paper cites Astra- sim: Enabling sw/hw co-design exploration for distributed dl training platforms,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Astra- sim: Enabling sw/hw co-design exploration for distributed dl training platforms,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.913639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.913639Z digest=sha256:e1263f99de4d12de62d84519dd882333f0d7bdcefcbe243502b917ec4d1605ae

Observation 7b94c5cc-7628-465d-be71-6e634afd80d3 · outbound

This paper cites Judging llm-as- a-judge with mt-bench and chatbot arena,.

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing Judging llm-as- a-judge with mt-bench and chatbot arena,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:16:55.968542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:16:55.968542Z digest=sha256:9d0e66e6468dfaf0704d8d530cec68bd1a1a44982af7cf39f4bc34a762a159f9

Pith citing papers

No inbound Pith citation observations are available.