Pith. sign in

Paper Citation Record · LEDGER

MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2411.01016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.01016 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:53:43.594876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:14:57.406110Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 43ae9bfd-f2b9-4bd6-8214-c8482cc2ae76 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:36.127711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:36.127711Z digest=sha256:d2c229b8689027a3297bf798f9b6ea807e55a0d48653e669a994dec95f9f53d0

Observation 9199348c-1b5d-493f-9db0-fb834161f713 · inbound

Faster MoE LLM Inference for Extremely Large Models cites this paper.

Faster MoE LLM Inference for Extremely Large Models MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:53:43.594876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:53:43.594876Z digest=sha256:c98313eaff7b5116fabffe4ada36d35d5ad00625a47cb36254ed4e9e498987c5

Observation 4635aa8b-5752-4e4d-be90-438bdaba0c1d · inbound

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration cites this paper.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.243625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.243625Z digest=sha256:d5d167db60d09f6092e1781975191b5a051e77c07c73e85ac48e01519f4a2161

Observation 533e0c95-9a40-499f-9fc9-bb97515cf5ef · inbound

Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference cites this paper.

Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:20:58.277550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:20:58.277550Z digest=sha256:33976be6ef39b42bc3b694004a056f94b0b3ef541d137e9c84dbf7ac2d34cf3c

Observation 38e4c1b4-dff6-4853-919d-8e25c1d55aed · inbound

SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation cites this paper.

SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T18:58:34.901052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:58:34.901052Z digest=sha256:98840fd5ec11e3ba4531d06b9c49dd78c04957defd4f86776238dc5c0f06c81b

Observation 0272af74-5132-471d-816a-2bfd2ee463fa · inbound

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging cites this paper.

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:09.523179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:09.523179Z digest=sha256:2a34cf921bafb8833d21d87c5d112de622b3fda59d7546a610dd5d82bf0b212e

Observation adc35558-fd31-43f6-910d-11c348de5fea · inbound

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference cites this paper.

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:14.314029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:14.314029Z digest=sha256:a3ed38cad74af320b145e279a132e344424f1108852b7a1fea77f4272aa2e9f2

Observation c3df6d3c-33f2-43a4-89fc-75718c3ca28c · inbound

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs cites this paper.

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:25.224160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:25.224160Z digest=sha256:fdc6bed5a455e63ba9968a811fc22b924ccb55522728981a2a46b69c5de950ab

Observation c986254d-4e84-45a7-81e5-e58639686ea9 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.927468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.927468Z digest=sha256:b02e88b8ed1f002e7671636178ce71eb29fe4f04da1155f3ef0e61eb3b48f4f4

Observation ba5cfb40-d5cc-4cea-be5f-101809737376 · inbound

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits cites this paper.

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T09:32:55.511235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:32:55.511235Z digest=sha256:fd803ca467606e4f0b42aee6e3a64aed86fa7c27557d5ffc14597fdfb9609059

Observation cd1471e9-d9dd-431f-b701-928b761f5709 · inbound

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE cites this paper.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.646451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T14:34:48.524592Z digest=sha256:bf0ce54bfa4130764b773532ddb658ce07b28dd24e9f3c953e9f5f63a25e0dc3

Observation db7cfda7-5df8-4f90-82bf-f193ef37c6fb · inbound

Path-Constrained Mixture-of-Experts cites this paper.

Path-Constrained Mixture-of-Experts MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.276593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T09:16:06.226566Z digest=sha256:e2dd6ec62678c0484a4af4cc39aaa002cf7bb4abf630f4e761cb59257b312b11

Observation 95615c89-1ce1-4a93-948e-e8717e670abb · inbound

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving cites this paper.

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:13.280550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:16:16.466375Z digest=sha256:d733236cff00cc614284e7cb6903acc23e491cd78f149a83cfb3b98d0cd9db5f

Observation d58bc009-1d59-459c-9739-63e84a7a37cd · inbound

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization cites this paper.

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.466223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T15:38:18.616792Z digest=sha256:5a1420e3f61c09732dfaeeb98ad14f3d874a2b176a18dfd6019f00f1a9dc07e5

Observation 528da36e-b37d-4f86-af60-f3647c4a628a · inbound

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference cites this paper.

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.407554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T04:27:02.915854Z digest=sha256:6280c1fb322f44aa72fbc7b3b7dbcea19e77d2e3ba24c44ff75bd6fa55702ec2

Observation 49ffc646-cb1b-4427-b0e8-255ff2f9e063 · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:185fb741c68a6e9577bf164c22bb6a283eefde3c17dbbfcb164b337095453f76