Pith. sign in

Paper Citation Record · LEDGER

FloE: On-the-Fly MoE Inference on Memory-constrained GPU

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2505.05950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05950 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:18.175589Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T10:32:50.809236Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:36:25.890370Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 677af2a7-ab45-4cc1-a6cf-26fd2a0eb536 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.731611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.731611Z digest=sha256:7625ab0b693a5bf54cdc942331fd7b6fe6f9400cbc6f6fb1c002033765accb3d

Observation bab1eeb2-937e-47af-a381-8421766f1a6b · outbound

This paper cites Phi-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.739273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.739273Z digest=sha256:018a3d43c7920469478de79f9089033665f9b63ad25f6d920fc17b728ccd87c8

Observation 3a684dc7-b459-4971-9310-dbc1ee9f3a79 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.745925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.745925Z digest=sha256:c12cadc88eec25b9bf7cc7bef1a10671c1562b692a03d6d2d2394fe812592888

Observation e76686a2-ee19-4dfa-b3fa-d00fbaa16392 · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Y., Rajbhandari, S., Awan, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.751415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.751415Z digest=sha256:bb265cc2a53bbdb63646e92ecbb702021978f92fcd23bbc11313114d574b89c1

Observation d76b3fe7-ac24-4bae-bad1-10d4219a1f68 · outbound

This paper cites and Shaji, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Shaji, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.224588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.757214Z digest=sha256:465c05678ceee1c9ebc8483b36ce9b4b63f28eb5463d4b8e48047e7d43799a56

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · outbound

This paper cites MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:7f97a4a8857106c095504bd0084a26e54127180720e26462a9b02749bd9c429f

Observation c5e29aa8-6157-4dfb-9fbe-580f2760c38a · outbound

This paper cites Active multi-task representation learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Active multi-task representation learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.199982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.769821Z digest=sha256:bf713074b8604f66946c593ae1968b667465c9a9aec47be20da9d77feaa16230

Observation bfe644cd-5cd9-4aa9-abd9-01ff6d4eabea · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.775307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.775307Z digest=sha256:eb64495fab615168d6f4c1c02b7545eec38928e74f2519f3b075d8196cc98d1e

Observation 52f07ce0-f174-4036-9d54-5bba9f6fc698 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.782233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.782233Z digest=sha256:00e4e4a5e4ef7e1d625ab5e796911285014d38133e036b3a65c120d8dabf2714

Observation 327c293c-9f4a-49ae-95b0-424966f745af · outbound

This paper cites D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.791097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.791097Z digest=sha256:0486e32123b7cddf5d2cbec74274b1d8990e131c68f6277a0139b0932c557987

Observation 727dac6a-b239-4e5c-b196-f08320653ecf · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.799618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.799618Z digest=sha256:d371e6442e384c4b09e7cf503be6b38da32466ac06c66867ba533b703f552933

Observation 47029fd0-e2e1-477d-88f6-0b58449e6d6d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.806520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.806520Z digest=sha256:678415262c5d5ed4b7a2c9f24aad181f5eccaf1db4ba2ab8e36f66fd065356f9

Observation 7a525660-27d3-4001-9f46-2fccff125050 · outbound

This paper cites Few-Shot Learning via Learning the Representation, Provably.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Few-Shot Learning via Learning the Representation, Provably

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.816788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.816788Z digest=sha256:5bd2cf5f49adc279da6e4ca62f67403bafad8b1ce38dc39270e22c97c54f6de7

Observation cbdf6ffa-d052-4180-8de0-69afaa65e337 · outbound

This paper cites The Llama 3 Herd of Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.823318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.823318Z digest=sha256:6ef4354cf222365556dcd73ac3ac0981ded0fc758bce3e0360a44db28df4ad93

Observation 725c3151-30e9-4283-8e9a-ce509f742b17 · outbound

This paper cites mixtral-offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU mixtral-offloading

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.148601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.828971Z digest=sha256:f447e25411f772e670b08dba6721d6f9e20711e1cf99ba8e391e87fd735b99ed

Observation 24a75ecc-bda2-479a-a402-2e51d3870d32 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.834440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.834440Z digest=sha256:c07f320f1f96c033d1240fba8883f718415e9593395fc6b996a4a7b0d0011038

Observation 84ee202f-5eb4-4359-804f-d3be1ea19b8d · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.840069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.840069Z digest=sha256:364333b74efa262a617ce2e7aa167336429eb1170eb4eafebfc0100efd3981ef

Observation c143d9d5-eff0-4231-8ad8-99bb8ed62d2c · outbound

This paper cites and Alistarh, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Alistarh, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.108972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.846718Z digest=sha256:590c3eb3d58a093e17a73aaa5b1f2a65dbf5a5d41a04f5e95ff3882688be4144

Observation bab7de65-c6fd-478b-9bbd-083636a3e1e6 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A framework for few-shot language model evaluation, 07 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.852078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.852078Z digest=sha256:e4037e3e2d8b30ea0250106a718a5b5dde71cb1ff9545b05c99b662bdafc2781

Observation 90db61db-82a7-4a3e-83e0-43e8603c7159 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.857019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.857019Z digest=sha256:b3aaca179fdf5ab8cc8d6b3083ace8b3ae81dd658d1743bd39ffea83a6c45e66

Observation 28e977d2-584b-4820-88f5-c5ecb758d738 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.863141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.863141Z digest=sha256:39f5a73a74626bbc0badf03e45b6ed0e7d1b4891ff99fd2024327243c2902690

Observation bb5fd3d0-45d8-4e20-b9d1-6287cebae18b · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.041897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.869011Z digest=sha256:19ef06cdc2d1f18e55e5a91d047e1d18f552931f7a4896fba430e03c4e8d6267

Observation 41ee817b-2f93-4290-b901-4f84ed2fac7c · outbound

This paper cites CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:00:19.156537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.875990Z digest=sha256:eb0b27de28ef39e09645dd9ae530de4e3816583a3279957396a1142d60050bd2

Observation 3b58a5f5-e855-4a40-b0c3-80347f028c66 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.881727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.881727Z digest=sha256:c2b9ab1661ab74926fe848017a43202519df72ea3227e93527785605ddefc27f

Observation 0480f9b6-fb16-4211-b02d-6eb2c6daaf25 · outbound

This paper cites Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.020055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.887270Z digest=sha256:68d55ef1319c807208a818f8961b41ae2dc54f707057fef16d969cb2a27251fc

Observation 86d25de5-e46b-41be-8515-7729a727dab2 · outbound

This paper cites Mixtral of Experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.896175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.896175Z digest=sha256:92ae5c6b3fd43e92b3e2e058cbf457c58bc09a013a2f533376b2e34e2e1aec1f

Observation c6483c91-a401-4d11-a347-f66fe7148c14 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.902621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.902621Z digest=sha256:a43d1e233b509da863df2406b0bbdc14a9ccc1dcdc45c756336c57282bc24682

Observation 1ede6a3e-f848-411e-952d-fd84810af759 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.909669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.909669Z digest=sha256:21aab786c80d2ae32df5984c72ca131fb9d7bf07f6674879fc62433f8402b7aa

Observation b9869fc6-dd19-4af2-89c0-e12f69f572a6 · outbound

This paper cites S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.964733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.918029Z digest=sha256:eac6c418972999501c7759d918ea64feefa328e26bbaece1ca85df244cb7c94d

Observation 2f39fe42-3c72-43c4-b54d-70ead58ce9c4 · outbound

This paper cites CATS : Context-aware thresholding for sparsity in large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CATS : Context-aware thresholding for sparsity in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.937766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.923383Z digest=sha256:ad30d7627e35d5dc17f20adddfa6cbef6284b375f1325b304ea42874e4876f48

Observation ccd9857e-4fa1-42d1-835d-18a12ca7d368 · outbound

This paper cites InfiniGen : Efficient generative inference of large language models with dynamic KV cache management.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU InfiniGen : Efficient generative inference of large language models with dynamic KV cache management

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.910507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.930729Z digest=sha256:9d7f4590408bfbbffa067a627d96524be399b565fb9f1eac9eb4cc3f926ddea9

Observation 721b679f-5e19-4d5b-8847-d32c9d840572 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Training-Free Activation Sparsity in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.937817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.937817Z digest=sha256:ca98a3959be44b4e13dc639de589c0972376f535058175e14e413be28856f128

Observation 90acd159-9eae-44a2-be46-2df4ef83b4a5 · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deja vu: Contextual sparsity for efficient llms at inference time

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.883330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.944379Z digest=sha256:51092985117973ab874c157075a0aaaf2aeb6ccfeec4c120dd6ccd84503e391a

Observation 12c6a7b1-854a-4a6f-9f98-cd25bd05f0a2 · outbound

This paper cites llama.cpp.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU llama.cpp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.857973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.952242Z digest=sha256:e2f2113fa20ca588492c81cc8d417712dba73f61f00e9aa6bcedff6637901fcc

Observation 2e1be1af-7e6a-4334-80b9-d51fcc0d4362 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Llm-pruner: On the structural pruning of large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.829297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.960800Z digest=sha256:8f0b37ddd498e47f96476b619a74493f6a44f2b9eac6368b20e808c891850062

Observation fa720d6d-75bb-44eb-b3f7-6cdfe7b87cfb · outbound

This paper cites Pointer sentinel mixture models, 2016.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pointer sentinel mixture models, 2016

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.966839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.966839Z digest=sha256:0d12fc0d716c72fb22336b8d3d45db3cbc3b25f38f656218ef4ab52c41c6ad51

Observation 8298ebcf-7fd1-4210-a33b-d5de6894be62 · outbound

This paper cites Deepspeed-mii.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepspeed-mii

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.787328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.973800Z digest=sha256:48faec5445ef4c00f83b86091faf0d73e070e37884987b6454801a36ca5e83ba

Observation 4996bea5-75f8-4329-9f3b-c6a878ea5b9c · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.983480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.983480Z digest=sha256:5aacfd46e0f746fb94f6c906a2f1f29e2ead4aa154379d84747ee98a47fa5d3b

Observation cdc1ba83-bfda-4c95-86b3-d3ca038e1f71 · outbound

This paper cites GPT-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.989790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.989790Z digest=sha256:84da9b782dd83fdd932be8dc47afc77b6869c9ec2d3ed663d06e45a2a887b74a

Observation d91d0e5e-58af-4a06-8580-15f3e665cb1e · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pytorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.759512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.995847Z digest=sha256:ffdd8e911e24480948696080cbc27947bcef34e8900ee3532dbf987f0ff6cbd3

Observation 02c19407-6e5e-4461-b6bc-a9ab78ee0377 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.005803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.005803Z digest=sha256:de560459a30809ec6e10292b4469c904218f44fe19386b9f6b3c3b59952b2632

Observation 9a578f77-24fc-484b-916a-488691c6535a · outbound

This paper cites Zero-infinity: breaking the gpu memory wall for extreme scale deep learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Zero-infinity: breaking the gpu memory wall for extreme scale deep learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.015066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.015066Z digest=sha256:e7a04f435cb41d5ed97aa3f2aec882219996e1e1ce839ce5f8a3ecad5e4c0c6f

Observation 247c01b2-1b5e-4d2e-a3be-0c67cf207d01 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.021867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.021867Z digest=sha256:4726e00bf2fbd94689567d70aeb760ae8a5f4155244852aa29442ffbe24d33d9

Observation 379bb1b7-2012-488b-85e8-dd56140309c9 · outbound

This paper cites Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.716571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.027790Z digest=sha256:63ad5483638996f6dab59345d0716d9977bb53a43f1471f32dd70fc17a1fefcc

Observation 743ac543-05d0-4955-94e1-eb0d5c5b87b2 · outbound

This paper cites Sharegpt.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sharegpt

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.694207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.033475Z digest=sha256:3b3b3709b70d5b8264795042d3c73626c8e5e498b986bd4d308b92f0482cfdfd

Observation 3229b3bf-af95-4555-b7c4-e6dd604ab33e · outbound

This paper cites GLU Variants Improve Transformer.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GLU Variants Improve Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.040770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.040770Z digest=sha256:797206736cc5a6c3af71f98cbc3e8bef708cd28e813eddc871bcfebff658979a

Observation 6c7eeb3d-bca8-4529-b9cc-75833d975c4d · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.670024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.046116Z digest=sha256:45f447b92405b826dc6da19597b526c7c76fefda8c9c1f733b0787dc5ef4fe77

Observation 79b653c8-7925-4814-85c9-3af318e79a75 · outbound

This paper cites SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.053935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.053935Z digest=sha256:8e78f726ea04041c3f197ebb1b8d0bef9a932e647abeb3512001162253d0dcb9

Observation 020fa894-8bee-4aa0-81da-03f28f859508 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.060634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.060634Z digest=sha256:1b2ff2218a36fb677006ce15c4588f2e1809e5b656ba65849ac757b8f628991b

Observation c8f6bcbe-b074-4bbf-b084-00f63654d330 · outbound

This paper cites ProMoE: Fast MoE-based LLM Serving using Proactive Caching.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProMoE: Fast MoE-based LLM Serving using Proactive Caching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.068102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.068102Z digest=sha256:37edc7364825c23713c6405e098ec2c3e4f3ec0a80f3ec5bb5c8a76181ff4cfe

Observation 84f9b36d-d5ba-4aa4-8804-3ff60ef17578 · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.080046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.080046Z digest=sha256:994f4c3daa8f7adf50ffcbd1f0d49688410e72f5a03bfcb64b3d9b5b0ce9a303

Observation fc400941-896c-477b-a247-7d5c13026a17 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A Simple and Effective Pruning Approach for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.088388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.088388Z digest=sha256:8f7efde4edf15775f33607e7e0f5db65cae2fb447caf82e7d4fc5b288f564199

Observation 020ce44c-6145-4969-b276-bd807a80f7a2 · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.099183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.099183Z digest=sha256:77c5d2867caf414b310b39b8874dbd393001ed33faba565b573d48977948d612

Observation c5e2268d-4170-42b0-8a56-6860916b885c · outbound

This paper cites Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.105373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.105373Z digest=sha256:1f0f0ba598e28e2b4e9602e21965ddf897ee930bda42d6d08c67e5adb0d97371

Observation 34b0b7a3-6ade-4a8f-bae6-82b815c30366 · outbound

This paper cites Sample Efficient Linear Meta-Learning by Alternating Minimization.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sample Efficient Linear Meta-Learning by Alternating Minimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.110331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.110331Z digest=sha256:069ca3eac8ef763ff656e408a4e7eb36e0865f79cfae44afb246d9162e16c36c

Observation 9e445290-c782-4165-9ed4-658d4a891137 · outbound

This paper cites T., and Cox, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU T., and Cox, D

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.117468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.117468Z digest=sha256:c3c8feabf909e3829c0a698c29a391726f2b0425ac94c200bdcd003087ddc06e

Observation e16e6808-ffae-4924-abbe-99bd90483fdc · outbound

This paper cites On the theory of transfer learning: The importance of task diversity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU On the theory of transfer learning: The importance of task diversity

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.626650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.125925Z digest=sha256:b32b3105c9da1e5f17356f844a1d1ed8b85726eb480a9c094ab46d77480cd71d

Observation ab80d049-fdd9-4250-8266-c4ec26090e30 · outbound

This paper cites Provable meta-learning of linear representations.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Provable meta-learning of linear representations

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.600418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.133442Z digest=sha256:e1aa84a8ee1d29af8879643deed4626a44434cb02283aed6a02c73d9c113dff2

Observation 2300d282-1841-401c-aa6a-b3c072c8df32 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:00:19.577922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.140080Z digest=sha256:98b8468c42a8aeccd1568c20482d8ea38add9dfe05543b5472d9ef4be13039bf

Observation adff6728-d4ce-4fc3-aca1-7179da0c7abd · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.147537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.147537Z digest=sha256:cc7c49e7f380660aa5af93bc326bb9a4937c7df81ae584dd3993a0d03f7afc0c

Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.153988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.153988Z digest=sha256:dc60e8c4cc41247252bcd718494b87448210ade8ea93743754da6b59392b7b0c

Observation 9f415b45-5a5a-40e4-9337-3a562c093362 · outbound

This paper cites and Ananiadou, S.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Ananiadou, S

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.161527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.161527Z digest=sha256:37e634f825c6fbaba5c7668e6b59e0e095e7e9c36569921a4567c07a30c0bebf

Observation cc91911b-4e65-4197-9ca3-932ce986b3eb · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.167395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.167395Z digest=sha256:ecdf003ab9fd0e39c751be31df2ef6d40e96b4688d2459d5a7c947701e076abe

Observation 9ccd1920-6cc0-40bd-85c3-af1cdb5f9d27 · outbound

This paper cites write newline.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU write newline

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.175589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.175589Z digest=sha256:0a6e080229a52772c9aca810d1e1eaff41835c434fde4e984e8efb328d5eb942

Pith citing papers

Observation 513369c9-9c91-4e0b-ad07-1138bdc3b9cd · inbound

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving cites this paper.

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving FloE: On-the-Fly MoE Inference on Memory-constrained GPU

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:25.893028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T10:32:50.809236Z digest=sha256:7a3569e8216b5607e35fb1d562d5b1791f3e63e8c9c40085f1e5e2dc9e9d4f86