Pith. sign in

Paper Citation Record · LEDGER

Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2407.00945.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00945 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:56:37.579860Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a4d1b1f-30b1-43db-aec9-d07997c616df · inbound

Mixture of Experts (MoE): A Big Data Perspective cites this paper.

Mixture of Experts (MoE): A Big Data Perspective Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-10T18:56:37.579860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:56:37.579860Z digest=sha256:33fd7893976b44af8572dccafa8367c7c853670adfbca7a729581dec4016f742

Observation 7660d774-09ea-415c-9f92-5a4c5e66991e · inbound

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging cites this paper.

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:06.051660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:06.051660Z digest=sha256:dbe397cf9dfe46f7e01a015d8d6fb041b917616c9f404b09a08b00e479b89a48

Observation b966d609-c9b1-400b-9d19-4d782f931642 · inbound

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs cites this paper.

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:25.147648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:25.147648Z digest=sha256:109d324b568d5690fd4a37831d79ece4ff5ac4b2fc62cec1b620fdc7e999bbb7

Observation 0b00ed5d-839b-4f7f-85aa-d9bdfebc41b8 · inbound

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE cites this paper.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.542805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T14:34:48.524592Z digest=sha256:5cd0ae09629edc0f31f7a304f5ca3169512af0a0d225782dd0577d9482c2cb4a

Observation 50b30048-2659-47a3-8189-2652e2a66b70 · inbound

Does a Global Perspective Help Prune Sparse MoEs Elegantly? cites this paper.

Does a Global Perspective Help Prune Sparse MoEs Elegantly? Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:50.688930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T19:06:25.626026Z digest=sha256:ed7bbfbef8946e963588aebabae25bdd0f39428a06a81739081b949cd9b7d041

Observation 0930bd18-5d90-4894-a874-b0fdf08145f4 · inbound

Temporally Extended Mixture-of-Experts Models cites this paper.

Temporally Extended Mixture-of-Experts Models Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:39:48.290062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:39:39.492135Z digest=sha256:15fcf662688648f5f377b50a535d7222f8c9af6a8699d5bc847dfbd72f53dc09

Observation 27d5642d-aa9e-4f3e-9fc4-2c17e568e113 · inbound

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts cites this paper.

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:04.829046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T05:54:32.496951Z digest=sha256:9531acd3a228e36f47c2be56e67516669e75fa8a2b1c1123666e4fb04a2340e1

Observation 29d0b9e4-6f75-4189-add2-25b448d22c8f · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.243574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:00:35.496822Z digest=sha256:47f62458669efb62d0520347e384034c0f6d0daa406fc1cbdb0d393806392234

Observation 489d6158-235b-4aa9-b203-60bb8436b6ef · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:25:00.060784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:22:45.702572Z digest=sha256:77bbc413bcdf912feda0b5d4cdaa670a66ac1b5e1af6d9e8b071cd8f776e362c

Observation e62a1f1a-b2f7-4f84-b6dc-bdce1d046c27 · inbound

dMoE: dLLMs with Learnable Block Experts cites this paper.

dMoE: dLLMs with Learnable Block Experts Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.115735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:50:51.900169Z digest=sha256:111753f61bc5c9cf3325e16d3d864c1af5e433fb3f1888d422866a684a4b3ce0

Observation 94868339-c023-4799-a1e7-cc7b5949fc0a · inbound

Less is MoE: Trimming Experts in Domain-Specialist Language Models cites this paper.

Less is MoE: Trimming Experts in Domain-Specialist Language Models Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:36:55.540848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T03:11:23.755739Z digest=sha256:525ef6c234367c8cc5cac24b48e3c2cce280c579d9375fb6038f2d0a35af4cdb

Observation e61384f3-e1c4-4f9c-ab82-797e14b04bb5 · inbound

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression cites this paper.

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T02:10:21.751821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T02:07:44.237002Z digest=sha256:b13b2d3c4a697095c7eff7836e3517890932368ef49102e211e49e578b160d0f

Observation 331452d3-39fa-4727-a2c2-478da7597b75 · inbound

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement cites this paper.

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T07:33:31.233659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:33:31.233659Z digest=sha256:29189eaee4bd648440c0ab7bae65d5778671ad1232c98de38229e502d59fd837

Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · inbound

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding cites this paper.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.515013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.515013Z digest=sha256:1012bfee62fc03781624b91ba672cb984b97515187e77c3f0d270e82d7140706