Pith. sign in

Paper Citation Record · LEDGER

MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.14361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.14361 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:37.745029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.550340Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · inbound

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement cites this paper.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:52:51.896416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:e6ceed114569414ccdf600eabcb87412ef5afc580f11e87eb901f832e54354f9

Observation 94c46255-a666-451d-b290-ce5b4aa13816 · inbound

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices cites this paper.

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:42.658889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:03:42.658889Z digest=sha256:ed1f370c9a974529819da2aacc1658cb483c40d7236e179473fd0aef06a80978

Observation ae163c1c-a6a1-43c3-b44a-7a64dd5ea6a2 · inbound

DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance cites this paper.

DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.231748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:33:55.795906Z digest=sha256:2650ca05ba7c2e9a1be53417c71de1ea15bbdc13c03ded2daba26e3436d31c70

Observation 7a2237d3-2723-437f-bcdc-a3fb5ea9870c · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:06.401918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:06.401918Z digest=sha256:0c0385bfb5331a7a4435ede337e3b6a1be895eceb6622291730366990ae5c564

Observation 16a1edf2-a8f0-4291-aec4-96a3df5d55e6 · inbound

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism cites this paper.

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:55:19.584295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:55:19.584295Z digest=sha256:3f943a2512f14dbc71fb5e0bb69fde941132e357b7300fdd5bf89e5939ad32cf

Observation 3696400c-f729-49d2-83b5-1c5d6cbeb036 · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.790153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:82b12066df659e8811eef489979755a2a8606f749b6e7a252c3d93efb60c911d

Observation 050f97cd-542a-4526-96d4-9f81eac3ebd9 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.207159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:29e99fe3836a234666d2e2bf5433b01bd5f938ee182d6277906f9342f3830e16

Observation da4e28d0-5ecc-43b7-a955-3f4bbfc5c878 · inbound

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving cites this paper.

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:18:13.312025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:16:16.466375Z digest=sha256:1f10c82d2b2d746654440bc17e8a00a8e6c405cce1070237f0f37deed0184008

Observation 0f51cc86-66b4-4384-adfd-2d1f08f02b43 · inbound

ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving cites this paper.

ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.627584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:28:26.169161Z digest=sha256:dc0b08149853b8ade2b2ad0da0f07be49f5aac4953cf58fb34d4b81fe96060e8

Observation 4c43c914-c0c0-44b0-afd7-2ed143816c41 · inbound

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy cites this paper.

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.447419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:46:52.811371Z digest=sha256:34c2099e624d5a4a6c3b53a043708c8dae205a7173182ef4d736fb2f034ae7db

Observation 5fbd1d3d-e3b1-4435-8059-2abfea64ee3b · inbound

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs cites this paper.

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.174701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:31:25.826205Z digest=sha256:a6e7df3ba99e71cc3cc71f46332acc3395662c3b354df2c31e6f85ff65e3b93d

Observation 794d1f12-71e5-4b83-9186-1f73e578f0ab · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.468196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:c9ea9f83a02bf9b88f8b12bb97ec61ce4d0952c6bad028f24949af05c03aaa5b

Observation 7a8d0425-079a-4154-afa6-63511dc03319 · inbound

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution cites this paper.

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.396632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:07:53.045082Z digest=sha256:ab89100bfe2492c71af111efab1dcc7faf1357a4682c0854acfb02ae5bf65539

Observation 81743ddd-ab30-4031-9717-3c676997fcbc · inbound

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG cites this paper.

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T02:12:58.326430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T02:10:57.582345Z digest=sha256:5310c1dac516011866c28023b71ce652fc53c3a31609e96962c17ae36b4a4075

Observation b4175d1d-845f-4f1c-b391-4bfa1a17e30f · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.693206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:baff577855c8c1d1ceff9e7ba3e1d00169d7a40732d352ab5d50d38d3d87fec8

Observation 7b1639c6-c879-4556-8598-1c975e26c877 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.552150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:12893c627fbffbf32b9b3e597559cd7a469acfc862f4b114f7a348d313f6e7c6

Observation 98e4b80d-de56-429b-b932-23f96c90d71d · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:707e8d9430260fdbadaf45682af04d07e203d0e816dc2b243cd0110bddfd9cfc

Observation 7feb9fdd-86cd-4459-b7ab-e676637d5b6d · inbound

BatchGen: An Architecture for Scalable and Efficient Batch Inference cites this paper.

BatchGen: An Architecture for Scalable and Efficient Batch Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:39.555548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:58:25.319255Z digest=sha256:ca10cc9cdd88dc71ddda90065770b66495a7ef73b0f9b175e211bfc50045af8d

Observation 0586afed-f7b3-4eb8-9d54-1716a183b7d3 · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.570088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:2260900aefd128d33f5ee33f4559391e57f95559b00fcaf30049f575ae6bee27

Observation 6863ad4a-8a91-4c67-b247-f838f9f292f3 · inbound

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices cites this paper.

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T13:40:50.742149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:40:50.742149Z digest=sha256:0f716f37ec30a7bbfa333d1f465fd03640e4d307b8282de75fd2cd83fb0ec661

Observation 43ad2203-067a-49c0-a69e-efe1f2f3fcfc · inbound

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference cites this paper.

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:37.574777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:37.574777Z digest=sha256:2a3846a2864b9b115b2e0c3601b43c3208fa16feb4ca581dc325cb2566503d95

Observation 7dbd1280-86be-4709-a542-8d32f1fd8085 · inbound

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference cites this paper.

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:29.052904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:37:29.052904Z digest=sha256:2a8d42e455fc8eb661defb453ef000e9e6a0d5863eba860bddf5907d1ee56025

Observation b5f181eb-5125-4db1-8aeb-b2973b4e5ba1 · inbound

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models cites this paper.

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:37.745029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:51:37.745029Z digest=sha256:e9b6304464e260b6ae393b51d3e77a1ca42b9d09ee79bf1bc3d7b39f7764d766