Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:18.175589Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2505.05950.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:18.175589Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-07T10:32:50.809236Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T09:36:25.890370Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 677af2a7-ab45-4cc1-a6cf-26fd2a0eb536 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab1eeb2-937e-47af-a381-8421766f1a6b · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a684dc7-b459-4971-9310-dbc1ee9f3a79 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e76686a2-ee19-4dfa-b3fa-d00fbaa16392 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Y., Rajbhandari, S., Awan, A
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76b3fe7-ac24-4bae-bad1-10d4219a1f68 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Shaji, A
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c5e29aa8-6157-4dfb-9fbe-580f2760c38a · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Active multi-task representation learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bfe644cd-5cd9-4aa9-abd9-01ff6d4eabea · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f07ce0-f174-4036-9d54-5bba9f6fc698 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327c293c-9f4a-49ae-95b0-424966f745af · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 727dac6a-b239-4e5c-b196-f08320653ecf · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47029fd0-e2e1-477d-88f6-0b58449e6d6d · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a525660-27d3-4001-9f46-2fccff125050 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Few-Shot Learning via Learning the Representation, Provably
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbdf6ffa-d052-4180-8de0-69afaa65e337 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725c3151-30e9-4283-8e9a-ce509f742b17 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU mixtral-offloading
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 24a75ecc-bda2-479a-a402-2e51d3870d32 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fast Inference of Mixture-of-Experts Language Models with Offloading
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ee202f-5eb4-4359-804f-d3be1ea19b8d · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c143d9d5-eff0-4231-8ad8-99bb8ed62d2c · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Alistarh, D
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bab7de65-c6fd-478b-9bbd-083636a3e1e6 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU A framework for few-shot language model evaluation, 07 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90db61db-82a7-4a3e-83e0-43e8603c7159 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Transformer feed-forward layers are key-value memories
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e977d2-584b-4820-88f5-c5ecb758d738 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5fd3d0-45d8-4e20-b9d1-6287cebae18b · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Accelerate: Training and inference at scale made simple, efficient and adaptable
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 41ee817b-2f93-4290-b901-4f84ed2fac7c · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b58a5f5-e855-4a40-b0c3-80347f028c66 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Measuring Massive Multitask Language Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0480f9b6-fb16-4211-b02d-6eb2c6daaf25 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 86d25de5-e46b-41be-8515-7729a727dab2 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Mixtral of Experts
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6483c91-a401-4d11-a347-f66fe7148c14 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ede6a3e-f848-411e-952d-fd84810af759 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9869fc6-dd19-4af2-89c0-e12f69f572a6 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f39fe42-3c72-43c4-b54d-70ead58ce9c4 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU CATS : Context-aware thresholding for sparsity in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ccd9857e-4fa1-42d1-835d-18a12ca7d368 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU InfiniGen : Efficient generative inference of large language models with dynamic KV cache management
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 721b679f-5e19-4d5b-8847-d32c9d840572 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Training-Free Activation Sparsity in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90acd159-9eae-44a2-be46-2df4ef83b4a5 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deja vu: Contextual sparsity for efficient llms at inference time
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12c6a7b1-854a-4a6f-9f98-cd25bd05f0a2 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU llama.cpp
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e1be1af-7e6a-4334-80b9-d51fcc0d4362 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Llm-pruner: On the structural pruning of large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fa720d6d-75bb-44eb-b3f7-6cdfe7b87cfb · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pointer sentinel mixture models, 2016
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8298ebcf-7fd1-4210-a33b-d5de6894be62 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepspeed-mii
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4996bea5-75f8-4329-9f3b-c6a878ea5b9c · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc1ba83-bfda-4c95-86b3-d3ca038e1f71 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU GPT-4 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d91d0e5e-58af-4a06-8580-15f3e665cb1e · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pytorch: An imperative style, high-performance deep learning library
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 02c19407-6e5e-4461-b6bc-a9ab78ee0377 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a578f77-24fc-484b-916a-488691c6535a · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Zero-infinity: breaking the gpu memory wall for extreme scale deep learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247c01b2-1b5e-4d2e-a3be-0c67cf207d01 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379bb1b7-2012-488b-85e8-dd56140309c9 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 743ac543-05d0-4955-94e1-eb0d5c5b87b2 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sharegpt
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3229b3bf-af95-4555-b7c4-e6dd604ab33e · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU GLU Variants Improve Transformer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7eeb3d-bca8-4529-b9cc-75833d975c4d · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79b653c8-7925-4814-85c9-3af318e79a75 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020fa894-8bee-4aa0-81da-03f28f859508 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f6bcbe-b074-4bbf-b084-00f63654d330 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProMoE: Fast MoE-based LLM Serving using Proactive Caching
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f9b36d-d5ba-4aa4-8804-3ff60ef17578 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc400941-896c-477b-a247-7d5c13026a17 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU A Simple and Effective Pruning Approach for Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020ce44c-6145-4969-b276-bd807a80f7a2 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e2268d-4170-42b0-8a56-6860916b885c · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b0b7a3-6ade-4a8f-bae6-82b815c30366 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sample Efficient Linear Meta-Learning by Alternating Minimization
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e445290-c782-4165-9ed4-658d4a891137 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU T., and Cox, D
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16e6808-ffae-4924-abbe-99bd90483fdc · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU On the theory of transfer learning: The importance of task diversity
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ab80d049-fdd9-4250-8266-c4ec26090e30 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Provable meta-learning of linear representations
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2300d282-1841-401c-aa6a-b3c072c8df32 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation adff6728-d4ce-4fc3-aca1-7179da0c7abd · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f415b45-5a5a-40e4-9337-3a562c093362 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Ananiadou, S
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc91911b-4e65-4197-9ca3-932ce986b3eb · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccd1920-6cc0-40bd-85c3-af1cdb5f9d27 · outbound
FloE: On-the-Fly MoE Inference on Memory-constrained GPU write newline
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513369c9-9c91-4e0b-ad07-1138bdc3b9cd · inbound
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving FloE: On-the-Fly MoE Inference on Memory-constrained GPU
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.