Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:12:49.662460Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2601.03992.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T12:12:49.662460Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fb908f54-7c62-445c-904c-a84efd3f76aa · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attention is all you need,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c069319-6e8e-4f6b-b780-899828ee8ad3 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Adaptive mixtures of local experts,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 040eb670-725b-49f8-aeab-6dc394e1cbd8 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Outrageously large neural networks: The sparsely- gated mixture-of-experts layer,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ba9011-0790-4f7f-900c-a4856e6ce203 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798e843b-507d-46db-9c8a-5795a5da64a2 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LLaMA: Open and Efficient Foundation Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f9bbad-12eb-4684-9c98-b3d3884a54c3 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems OPT: Open Pre-trained Transformer Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d96968-f5a7-424e-a377-1b2cd3680072 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2580717-12ba-485c-a366-3a679e19a3af · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A survey on mixture of experts in large language models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd75ece3-f377-482d-9eb4-1c96fe40361f · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) GeForce RTX 5080
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3bd0c9-0b32-4f35-bf22-b8e86318e18c · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Qwen3 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f595b27-c512-46d0-8bb1-d8aa7b4ae2b0 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d60b81f-dad6-4b85-90da-0421a5b5a144 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c79c1e-5af5-4994-ab70-2257e512a681 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fast Inference of Mixture-of-Experts Language Models with Offloading
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6305259-e3db-461a-8787-04ef004c4b67 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DAOP: Data-aware offloading and predictive pre- calculation for efficient moe inference,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78c45cc-9a64-4796-bc27-61a70a4db24a · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Moe-lightning: High-throughput moe inference on memory-constrained gpus,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb7a272-16ae-4c75-8266-941ea208f832 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d018a23c-8b61-4c64-8ffa-a5e838a69b61 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Monde: Mixture of near-data experts for large-scale sparse models,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5095e32c-e81e-4822-bcdc-f2f23f1d6f58 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b453858d-4086-434b-a15a-744210aede6d · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Co-designing binarized transformer and hardware accel- erator for efficient end-to-end edge deployment,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9156e491-9016-40bb-9e00-5ef4c67a2c51 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Language models at the edge: A survey on techniques, challenges, and applications,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdc3f5b-7058-474d-8dd4-8571e2cec33d · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9aab894-9168-4329-a9c2-3d5d1ab96678 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Spark: Scalable and precision-aware acceleration of neural networks via efficient encoding,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c14f80-ecbd-435b-aa5f-51eebba247bd · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Medusa: Simple llm inference acceleration framework with multiple decoding heads,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc085ba-503a-4bcf-bbbe-134d1b3aef5e · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Recnmp: Accelerating personalized recommendation with near-memory processing,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0bf0c4-6d8e-4874-af4e-e49eca8e287a · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23489549-cbba-4acd-bd44-1773d0a8c3ab · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attacc! unleashing the power of pim for batched transformer-based generative model inference,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8baebdbd-3461-4ad5-af9d-2e813169c9ca · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems PIMoE: Towards efficient moe transformer deployment on npu-pim system through throttle-aware task offloading,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a6ccdb-83b2-44d8-8705-3b301705dfe3 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36e300a-70f9-4da5-b29a-b73369e0cc3b · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems The true processing in memory accelerator,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8310dd-19bb-4e6b-ae5c-e2295098443e · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LP-Spec: Leveraging lpddr pim for efficient llm mobile speculative inference with architecture-dataflow co-optimization,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d59443-36f5-41f6-9dac-53b5672c11c8 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ndpage: Efficient address translation for near-data pro- cessing architectures via tailored page table,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43dcc1e0-9a50-4d7d-9a2b-346fdf6e1b1f · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Hyqa: Hybrid near-data processing platform for embed- ding based question answering system,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d22f975-30f8-4062-8ad2-fe8b650b36ae · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Near-memory parallel indexing and coalescing: Enabling highly efficient indirect access for spmv,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc330cc-7484-4e7e-bbb6-1184a963bc1a · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Um-pim: Dram-based pim with uniform & shared memory space,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9877c3c1-1c39-4302-b36c-c18530c54ea0 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Bramac: Compute-in-bram architectures for multiply- accumulate on fpgas,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320f2894-342f-4de9-af40-4b84f23cdf2e · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems An overview of processing-in-memory circuits for artificial intelligence and machine learning,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b97508-1b2b-421f-b195-90957737fd48 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Mixtral of Experts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58f14cc-db82-4710-a1c3-ea80e7dcac6d · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Aim: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc86f0e8-9fb7-4aff-9fbc-1eb2e95a0d7e · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) intel-core-i7-14700
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce93888-8b09-4c08-beed-03005f3e05ea · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ramulator 2.0: A modern, modular, and extensible dram simulator,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da4ecce-5822-4ad8-9ef6-7ccbcdd31636 · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c7e0cb-b925-4b9e-8dd4-6bfa3fa1c13d · outbound
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.