Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.14361.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:37.745029Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:39.550340Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · inbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94c46255-a666-451d-b290-ce5b4aa13816 · inbound
MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae163c1c-a6a1-43c3-b44a-7a64dd5ea6a2 · inbound
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a2237d3-2723-437f-bcdc-a3fb5ea9870c · inbound
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a1edf2-a8f0-4291-aec4-96a3df5d55e6 · inbound
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3696400c-f729-49d2-83b5-1c5d6cbeb036 · inbound
LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 050f97cd-542a-4526-96d4-9f81eac3ebd9 · inbound
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da4e28d0-5ecc-43b7-a955-3f4bbfc5c878 · inbound
FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f51cc86-66b4-4384-adfd-2d1f08f02b43 · inbound
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c43c914-c0c0-44b0-afd7-2ed143816c41 · inbound
Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fbd1d3d-e3b1-4435-8059-2abfea64ee3b · inbound
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 794d1f12-71e5-4b83-9186-1f73e578f0ab · inbound
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a8d0425-079a-4154-afa6-63511dc03319 · inbound
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81743ddd-ab30-4031-9717-3c676997fcbc · inbound
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4175d1d-845f-4f1c-b391-4bfa1a17e30f · inbound
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b1639c6-c879-4556-8598-1c975e26c877 · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98e4b80d-de56-429b-b932-23f96c90d71d · inbound
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7feb9fdd-86cd-4459-b7ab-e676637d5b6d · inbound
BatchGen: An Architecture for Scalable and Efficient Batch Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0586afed-f7b3-4eb8-9d54-1716a183b7d3 · inbound
WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6863ad4a-8a91-4c67-b247-f838f9f292f3 · inbound
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ad2203-067a-49c0-a69e-efe1f2f3fcfc · inbound
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbd1280-86be-4709-a542-8d32f1fd8085 · inbound
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f181eb-5125-4db1-8aeb-b2973b4e5ba1 · inbound
Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.