Pith. sign in

Paper Citation Record · LEDGER

FloE: On-the-Fly MoE Inference on Memory-constrained GPU

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2505.05950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05950 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:18.175589Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T10:32:50.809236Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:36:25.890370Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 677af2a7-ab45-4cc1-a6cf-26fd2a0eb536 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.731611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.731611Z digest=sha256:ffa27e5ef7054000996f4873093ef772738238ff18bd0043d0b2964469129133

Observation bab1eeb2-937e-47af-a381-8421766f1a6b · outbound

This paper cites Phi-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.739273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.739273Z digest=sha256:5a2e587a8eb5536c41a883a6ac84b39ba2faf3229067e7987e7cf5da6e9e144d

Observation 3a684dc7-b459-4971-9310-dbc1ee9f3a79 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.745925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.745925Z digest=sha256:27c0e37343abfc924ce77af9fb7587b1c2b692f4f1db157a07411bc1d8688b9e

Observation e76686a2-ee19-4dfa-b3fa-d00fbaa16392 · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Y., Rajbhandari, S., Awan, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.751415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.751415Z digest=sha256:54f52461f7437d8900fa49cfc8470907587bfed90db090c3c169a81924c55eb2

Observation d76b3fe7-ac24-4bae-bad1-10d4219a1f68 · outbound

This paper cites and Shaji, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Shaji, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.224588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.757214Z digest=sha256:e98a2642259d80fdc1d94d269e0c8ad46f197a4ec45b659a3fa97eb0123b9bbe

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · outbound

This paper cites MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:c96e966f8e01392b0cf14260720660562bbf328ec13255502534911aa6750cc8

Observation c5e29aa8-6157-4dfb-9fbe-580f2760c38a · outbound

This paper cites Active multi-task representation learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Active multi-task representation learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.199982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.769821Z digest=sha256:9f9972dc3540564358c71bf7e552aad411afa93bc5a97948d3b6ccaccb4bfad2

Observation bfe644cd-5cd9-4aa9-abd9-01ff6d4eabea · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.775307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.775307Z digest=sha256:881101ef32ddf3ff72c295513d57b93d8340eb248642a864a933445dc0359ae5

Observation 52f07ce0-f174-4036-9d54-5bba9f6fc698 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.782233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.782233Z digest=sha256:02294e4088bee5c4ed99a994bd0fae517724293ccc1711bec41aff2e9adea7c1

Observation 327c293c-9f4a-49ae-95b0-424966f745af · outbound

This paper cites D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.791097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.791097Z digest=sha256:9f4e0bb2d410256fe4f48502a0c24990484ae82fb2d14ddac34946a05dfbf101

Observation 727dac6a-b239-4e5c-b196-f08320653ecf · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.799618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.799618Z digest=sha256:91e553566e4dc0a6da866c682e53ef0acf95a3e0bd3a3f3bf8a19c01c385ba0a

Observation 47029fd0-e2e1-477d-88f6-0b58449e6d6d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.806520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.806520Z digest=sha256:88f6a2c85d785d40c40bd97d7f550f5aa2108f9ec5f27e20612727305af4a7bd

Observation 7a525660-27d3-4001-9f46-2fccff125050 · outbound

This paper cites Few-Shot Learning via Learning the Representation, Provably.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Few-Shot Learning via Learning the Representation, Provably

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.816788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.816788Z digest=sha256:55871eba038268f6674c3e6a0d52a3d48fb9fed1c2db34c66c246adf733d67d2

Observation cbdf6ffa-d052-4180-8de0-69afaa65e337 · outbound

This paper cites The Llama 3 Herd of Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.823318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.823318Z digest=sha256:d23180df4246ca49f79aae3f833086dc4d1a29f03ddb88d93a9dabe25d44ab59

Observation 725c3151-30e9-4283-8e9a-ce509f742b17 · outbound

This paper cites mixtral-offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU mixtral-offloading

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.148601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.828971Z digest=sha256:83301709333087c1db75767ad8db561b64116a90fd032d7c3545b350ccc39e79

Observation 24a75ecc-bda2-479a-a402-2e51d3870d32 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.834440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.834440Z digest=sha256:0872707730b6d1df0a45ea320364f2e957fac58f6af82eed335284c36edb3307

Observation 84ee202f-5eb4-4359-804f-d3be1ea19b8d · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.840069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.840069Z digest=sha256:8476a289ca02542f50c8621c631d15c2ba7e4b3504be38235fdd9a14df498f31

Observation c143d9d5-eff0-4231-8ad8-99bb8ed62d2c · outbound

This paper cites and Alistarh, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Alistarh, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.108972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.846718Z digest=sha256:0ca01ec29f119da9383df0bc8d847aeb8484491731dcc51f30a66bb632d2a3f7

Observation bab7de65-c6fd-478b-9bbd-083636a3e1e6 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A framework for few-shot language model evaluation, 07 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.852078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.852078Z digest=sha256:c0616b6b75a87d8a3f58d1765fa95087568ca95cf7a594e955c6af4a6e11a2b6

Observation 90db61db-82a7-4a3e-83e0-43e8603c7159 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.857019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.857019Z digest=sha256:5e906a836b4ac5a6cbfc8f2c9669be3735fa54d53fdeca755ce764e17faa32fd

Observation 28e977d2-584b-4820-88f5-c5ecb758d738 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.863141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.863141Z digest=sha256:9ada038c80e75d13a996eb210c5f55b79d718904190a80df0ac71fbcc4fdfc34

Observation bb5fd3d0-45d8-4e20-b9d1-6287cebae18b · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.041897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.869011Z digest=sha256:6c53059e50522646d6f73b24f26e1049d0e5d721b29a6435a11b2d70c6d1bce1

Observation 41ee817b-2f93-4290-b901-4f84ed2fac7c · outbound

This paper cites CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:00:19.156537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.875990Z digest=sha256:916b11c372284cdb2792018360eb0a1fa51c40b444960a79ffe7c04156e5b59e

Observation 3b58a5f5-e855-4a40-b0c3-80347f028c66 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.881727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.881727Z digest=sha256:343bd13196cbad5a99a0a90b93fb3773c0bbb1f5ff9c9c9bc12ac6eadbd08e5b

Observation 0480f9b6-fb16-4211-b02d-6eb2c6daaf25 · outbound

This paper cites Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.020055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.887270Z digest=sha256:25170ef34aea51b264ea89fba9dbe2624ae7c2eceaf52c33d15d25884cb8c2d5

Observation 86d25de5-e46b-41be-8515-7729a727dab2 · outbound

This paper cites Mixtral of Experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.896175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.896175Z digest=sha256:1a7c47cae97ff739b6c2bf0ca507a3e3519aef56fdfdcb894d5a9e2c4ee2f216

Observation c6483c91-a401-4d11-a347-f66fe7148c14 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.902621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.902621Z digest=sha256:14d98374c6c05b4a15ae94b432fe5329448a47aebc0c99279355af71f9f05b80

Observation 1ede6a3e-f848-411e-952d-fd84810af759 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.909669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.909669Z digest=sha256:2e0e63d620317938c73b143a1bf3103ba24f9f173938c344080069066cc9e4b0

Observation b9869fc6-dd19-4af2-89c0-e12f69f572a6 · outbound

This paper cites S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.964733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.918029Z digest=sha256:669387651293fda199b28cbba4da6bada8445b5f02e0288df796df2d05f71901

Observation 2f39fe42-3c72-43c4-b54d-70ead58ce9c4 · outbound

This paper cites CATS : Context-aware thresholding for sparsity in large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CATS : Context-aware thresholding for sparsity in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.937766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.923383Z digest=sha256:6463e517cc23156dbc6628e06f478888064edf35aa5a962b94fc1253ae7b7093

Observation ccd9857e-4fa1-42d1-835d-18a12ca7d368 · outbound

This paper cites InfiniGen : Efficient generative inference of large language models with dynamic KV cache management.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU InfiniGen : Efficient generative inference of large language models with dynamic KV cache management

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.910507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.930729Z digest=sha256:6694f0adb0ae74cedc5a3da5d40a4fbbae269a89a54ed6328065b117bf271092

Observation 721b679f-5e19-4d5b-8847-d32c9d840572 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Training-Free Activation Sparsity in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.937817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.937817Z digest=sha256:bbb77a9b64a884ca4af418845c7ba0c2db1511fe227193f746de4aea979ca0d0

Observation 90acd159-9eae-44a2-be46-2df4ef83b4a5 · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deja vu: Contextual sparsity for efficient llms at inference time

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.883330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.944379Z digest=sha256:899bfcd174e5e7487f5e4a8edec8e07ac25053fad10078743c11ab5ea42bba8a

Observation 12c6a7b1-854a-4a6f-9f98-cd25bd05f0a2 · outbound

This paper cites llama.cpp.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU llama.cpp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.857973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.952242Z digest=sha256:7ba13da2c4b4dc3e6febfd0c73161f55a0509157fad1fd3d8e623e506cebddf8

Observation 2e1be1af-7e6a-4334-80b9-d51fcc0d4362 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Llm-pruner: On the structural pruning of large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.829297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.960800Z digest=sha256:d003ed6369349ab7fa9842bf026ca14939a006d2285020313898db0d4266eb36

Observation fa720d6d-75bb-44eb-b3f7-6cdfe7b87cfb · outbound

This paper cites Pointer sentinel mixture models, 2016.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pointer sentinel mixture models, 2016

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.966839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.966839Z digest=sha256:8db5aa8ac3d7ba8d368a3bfa83683b1bf572d2f0a96d99ac4732dcf526147c94

Observation 8298ebcf-7fd1-4210-a33b-d5de6894be62 · outbound

This paper cites Deepspeed-mii.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepspeed-mii

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.787328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.973800Z digest=sha256:39d14dc2526a6e4d11c53af2937c99ad363fabfefd2fcc1b7caed5c60fc10801

Observation 4996bea5-75f8-4329-9f3b-c6a878ea5b9c · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.983480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.983480Z digest=sha256:b6832e88d20225654cd81cfa417cae8841ac50fe40edc2791eb470ef5a9787c6

Observation cdc1ba83-bfda-4c95-86b3-d3ca038e1f71 · outbound

This paper cites GPT-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.989790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.989790Z digest=sha256:ef481ef94d9df867ea0eab758d85b87af3b40068d93e06de28bfbddc81e68500

Observation d91d0e5e-58af-4a06-8580-15f3e665cb1e · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pytorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.759512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.995847Z digest=sha256:90363005e53bf75470db7e267f18634c65ade9346ca4f9da5c936286af9e08d5

Observation 02c19407-6e5e-4461-b6bc-a9ab78ee0377 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.005803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.005803Z digest=sha256:43cd15801d3ff70c0ab55055e10527ef0b4698df31396ba5b1377b2c825453f9

Observation 9a578f77-24fc-484b-916a-488691c6535a · outbound

This paper cites Zero-infinity: breaking the gpu memory wall for extreme scale deep learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Zero-infinity: breaking the gpu memory wall for extreme scale deep learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.015066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.015066Z digest=sha256:b045b78ef879cb36128854a5a00821368e5d10659c9d2600c5a854e2d18960c0

Observation 247c01b2-1b5e-4d2e-a3be-0c67cf207d01 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.021867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.021867Z digest=sha256:903fb6eb1e02a91ca1a612f017aedae668340dd20f52d157eafc8809eff56812

Observation 379bb1b7-2012-488b-85e8-dd56140309c9 · outbound

This paper cites Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.716571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.027790Z digest=sha256:e66757d93d706dcf6afd57efc6271050c8099670be67ec30d38aa42e81d0e1bc

Observation 743ac543-05d0-4955-94e1-eb0d5c5b87b2 · outbound

This paper cites Sharegpt.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sharegpt

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.694207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.033475Z digest=sha256:8d2153aea6975c0ba5361b49e18067899a89a5e9dfd5ac700ca8a9894e877b93

Observation 3229b3bf-af95-4555-b7c4-e6dd604ab33e · outbound

This paper cites GLU Variants Improve Transformer.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GLU Variants Improve Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.040770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.040770Z digest=sha256:250b3eae509ac743a2ff9e1502c6c897b548e7c1b165f3fa15052d005f3e6050

Observation 6c7eeb3d-bca8-4529-b9cc-75833d975c4d · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.670024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.046116Z digest=sha256:0ad676d135ba222717574ca3485e2a94d4fe7263ba6b88dd72f5d1c09a31f0c6

Observation 79b653c8-7925-4814-85c9-3af318e79a75 · outbound

This paper cites SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.053935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.053935Z digest=sha256:ce0cbe6d9932579764b680e3c6880709d5fb8e6fb25bbc6561719b42c0666fab

Observation 020fa894-8bee-4aa0-81da-03f28f859508 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.060634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.060634Z digest=sha256:04d195c2adb28e5f0b27e36748af0c36d22a5ad6d2733a0ee7792fec1459386d

Observation c8f6bcbe-b074-4bbf-b084-00f63654d330 · outbound

This paper cites ProMoE: Fast MoE-based LLM Serving using Proactive Caching.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProMoE: Fast MoE-based LLM Serving using Proactive Caching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.068102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.068102Z digest=sha256:14bf9472a2c453b3afd152d82efd5848175a14c260878f07ea7275b231e353c4

Observation 84f9b36d-d5ba-4aa4-8804-3ff60ef17578 · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.080046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.080046Z digest=sha256:6a605d93cf38c1a3e39ae3736d267d47fb264566497ad52d0261a9cb8c2e361d

Observation fc400941-896c-477b-a247-7d5c13026a17 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A Simple and Effective Pruning Approach for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.088388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.088388Z digest=sha256:43089534b05246d94657e66608a4b0de8e47494abcafca123415fdc9f45f090a

Observation 020ce44c-6145-4969-b276-bd807a80f7a2 · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.099183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.099183Z digest=sha256:82d5be63c88e30074a5f1a7d4524e337891b986767f4f373f00e4499e2ea68bb

Observation c5e2268d-4170-42b0-8a56-6860916b885c · outbound

This paper cites Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.105373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.105373Z digest=sha256:a337fbf593b31c62e6041fe1a8071153ab41d019eb5e1503bd626e5e47eba493

Observation 34b0b7a3-6ade-4a8f-bae6-82b815c30366 · outbound

This paper cites Sample Efficient Linear Meta-Learning by Alternating Minimization.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sample Efficient Linear Meta-Learning by Alternating Minimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.110331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.110331Z digest=sha256:57beb831802f72e0630f8213566202d53fe179f8a49354535122f09e6fe43ff4

Observation 9e445290-c782-4165-9ed4-658d4a891137 · outbound

This paper cites T., and Cox, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU T., and Cox, D

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.117468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.117468Z digest=sha256:68b2e12d1f1a15d8b8d950d95bad6c8f8c6bd06ad9192ba26081775059525d3d

Observation e16e6808-ffae-4924-abbe-99bd90483fdc · outbound

This paper cites On the theory of transfer learning: The importance of task diversity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU On the theory of transfer learning: The importance of task diversity

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.626650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.125925Z digest=sha256:ae55d97f017171ba5811f0a48440912544d0918b3dde823110660f73afa8e7ff

Observation ab80d049-fdd9-4250-8266-c4ec26090e30 · outbound

This paper cites Provable meta-learning of linear representations.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Provable meta-learning of linear representations

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.600418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.133442Z digest=sha256:5fbd4acbdc1561eb455a3d7019f0381be809a4c529e277186461bece2e3c961a

Observation 2300d282-1841-401c-aa6a-b3c072c8df32 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:00:19.577922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.140080Z digest=sha256:c79633f8b0cd5b286333caa995048b7bb730779afca874310131bda43da1fa39

Observation adff6728-d4ce-4fc3-aca1-7179da0c7abd · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.147537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.147537Z digest=sha256:5eb2cbf532ecc8f538aee8587d7d0565cd946b56d7db88ff6b9c77f6524acc52

Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.153988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.153988Z digest=sha256:06b232770151e84a7e06a953706e6fbd9dc7426102817040c07bcce57b219c6c

Observation 9f415b45-5a5a-40e4-9337-3a562c093362 · outbound

This paper cites and Ananiadou, S.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Ananiadou, S

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.161527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.161527Z digest=sha256:a0830a41dc898154998ccbd1abd846b4e9ab1c2cf0322ab51667d60ff1c814f3

Observation cc91911b-4e65-4197-9ca3-932ce986b3eb · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.167395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.167395Z digest=sha256:8c171f590d1e8ef1549cd5d29197d11e6fabf9742623192e13bcb8412d33dec2

Observation 9ccd1920-6cc0-40bd-85c3-af1cdb5f9d27 · outbound

This paper cites write newline.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU write newline

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.175589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.175589Z digest=sha256:491d42480408bb4c7b8e267ddee69dc7c4ca39b25bf15bfaabd3180ff30e25d5

Pith citing papers

Observation 513369c9-9c91-4e0b-ad07-1138bdc3b9cd · inbound

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving cites this paper.

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving FloE: On-the-Fly MoE Inference on Memory-constrained GPU

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:25.893028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T10:32:50.809236Z digest=sha256:aabd599b52b7a03a6db9a037461d5e41da19c9cf1df5f094e915d020da105ea1