Pith. sign in

Paper Citation Record · LEDGER

HarMoEny: Efficient Multi-GPU Inference of MoE Models

As of 7 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12417 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:32.963734Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:31:11.271729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T04:33:03.804063Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved11
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b83d114-1784-493f-9690-744ac60e2df6 · outbound

This paper cites GPT-4 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.813852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.813852Z digest=sha256:7ff1d3a22af50222a748b35598f10e8e312301d3c4fb92d7c3b2e9256b5ab8ef

Observation f80582c9-8953-4a5a-bc28-185de4c14044 · outbound

This paper cites Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.504866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.818850Z digest=sha256:cdfb8d8402e1cd2718d3a0fea060a63d4581ba0c3a16f5ee38a7ae32a59bff64

Observation 59dddad4-1d4c-4ed6-b0c5-62a42d571780 · outbound

This paper cites A neural probabilistic language model.

HarMoEny: Efficient Multi-GPU Inference of MoE Models A neural probabilistic language model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.491982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.822866Z digest=sha256:c5cdb31e2e0ee476d7c8c18060ea929085cba16c683012f9824ce5378ecf23d4

Observation 71b4bcfb-9f00-4e92-afa8-ad95ab622190 · outbound

This paper cites Datacenter power and energy management: past, present, and future.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Datacenter power and energy management: past, present, and future

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.479177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.827146Z digest=sha256:d97c75cc889f14e03e69797d971ca18f890a3e0df0ebac01334603df44b9b34f

Observation d8ead566-47c9-44d8-9279-225b98678ce9 · outbound

This paper cites Language models are few-shot learners.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.831078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.831078Z digest=sha256:907b6f6a3c1ef5ab3e03b85abd89ee271f62d95fafdb4ae098fc89178146eb3a

Observation 2fe6c22e-b416-4b89-bbbe-cbcd728c75d7 · outbound

This paper cites Large scale distributed deep networks.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Large scale distributed deep networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.457504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.834897Z digest=sha256:1fb0104e346f8e8856397c204cb2fcaa46c89b0727794b6641704ab5644e1fa2

Observation bf21b206-e2b7-45a7-a376-5c26f5da9c52 · outbound

This paper cites Quasar: Resource-efficient and qos-aware cluster management.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Quasar: Resource-efficient and qos-aware cluster management

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.444890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.839116Z digest=sha256:989ecf314f6e553a33a89ff569b5ed62074d9ba3d478ba30f3b945a7c004d31f

Observation bad8eb6c-46f3-481c-80fb-f768b78f85bb · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.432026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.842730Z digest=sha256:9d3c5baa93525d19f768cd37152d8e44dfa0f9851e6c74bbe0239190af7d3309

Observation c1678eaa-a4de-4c35-b151-fd9c69584ff5 · outbound

This paper cites Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.417764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.846300Z digest=sha256:9c3cf31ec3f3f1a7553ee69971ca46c40248e3f3ac1e548a41006736ee887ab4

Observation 46e54108-df6f-4716-a6b2-03918b500a69 · outbound

This paper cites Megablocks: Efficient sparse training with mixture-of-experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megablocks: Efficient sparse training with mixture-of-experts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.404964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.850070Z digest=sha256:5e92c40ce770e1e3322b3b42ee3ded164cfff71c4a9eb8150742f28a1c12937d

Observation 482cdcd6-4f65-4613-aab7-5327551c3ef7 · outbound

This paper cites Character-based NMT with Transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Character-based NMT with Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:33.086400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.853759Z digest=sha256:2405b3daabce48b50a07b61185c549d40e749fc2b9c7e9505e6fde66ae29d282

Observation d4340967-aa22-4a69-aad8-fd6b8cf1e4c5 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

HarMoEny: Efficient Multi-GPU Inference of MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.858076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.858076Z digest=sha256:ff04f25537173f7155eef31517ac1a0803abd5a143d604102c0748c46472cee4

Observation 97e82041-bfbd-4914-b069-7b631afd185a · outbound

This paper cites Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.392572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.862163Z digest=sha256:de9b9f1771f83866033bcb043b41b23066544f1e1f9e268180423ef68c0f09dc

Observation 5affdc3e-01fd-4096-a213-3589403847e8 · outbound

This paper cites Tutel: Adaptive mixture-of- experts at scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Tutel: Adaptive mixture-of- experts at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.380243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.865806Z digest=sha256:62f94a60247b4ffcf182e0bae9420849dfe7f0e87b1da951f623e98a55b448a2

Observation eb3db8f3-6b32-46db-b5a8-175c2bbc2252 · outbound

This paper cites Adaptive mixtures of local experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Adaptive mixtures of local experts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.367056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.869364Z digest=sha256:5416da4592b7635b441d9415e86ffe52e9a60d70a1dfb18561a0ae2da53c37c5

Observation 880a320b-7746-4f15-a62a-0d8f087a5101 · outbound

This paper cites Mixtral of Experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Mixtral of Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.872804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.872804Z digest=sha256:f8697184b57b5388101a287585c1c112d338ef8471bf3782860a6dbf9e5f491c

Observation 60a95c6e-66a4-40b7-ba63-0318e79fa7ca · outbound

This paper cites Scaling Laws for Neural Language Models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.876451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.876451Z digest=sha256:21c47ba79ba1fd5a2d66910f3a7b752c2a1de5cdb8de6a119a3cbf9332fd62b4

Observation 4915bd77-dafe-4b2c-8794-96d864e8cee6 · outbound

This paper cites Learning multiple layers of features from tiny images.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Learning multiple layers of features from tiny images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.880325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.880325Z digest=sha256:93222d03fb038f5b8d5ce1feee5d82c5803ef5524b96887a7c8b801f8a0c77ad

Observation 899c9652-4ce1-4421-91cf-c81dd3cd9781 · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Subword regularization: Improving neural network translation models with multiple subword candidates

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.344951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.883909Z digest=sha256:38613756ae48b53430ddd27ffc9de7f42868e36b48c30b6e0c38c3f739ffdd3d

Observation 9903387d-0a07-4e20-8ae6-13c96a2f4168 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.332584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.887516Z digest=sha256:e46d85b7c46cf0d68fc73b284b1951cbdddf141510a2ac269ba6934d7aa380c2

Observation a53ad049-1517-4573-a92b-deb66a80e2bb · outbound

This paper cites AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.320398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.891089Z digest=sha256:81a5531cedb511595e5b5871fa91b4eb8739407cd2680786c442ba46929b4049

Observation a1b134cd-08ca-4b04-9b4a-9d45491791e4 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.307403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.894473Z digest=sha256:08380f55bb3614a8b79574805099f7c0c527ee2279fe9ece3566e43413ed40ce

Observation b49519fd-b6e4-4747-858f-eba378d20bbd · outbound

This paper cites Accelerating distributed MoE training and inference with lina.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Accelerating distributed MoE training and inference with lina

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.295114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.898045Z digest=sha256:ae8f6bb20c2022aa54a30a8427e20a99bd1a3007973dca75e2cec7fde32fc6cf

Observation d509d5f2-c70d-49d2-b1aa-6b4a4438f67b · outbound

This paper cites DeepSeek-V3 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.902330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.902330Z digest=sha256:8f2e2bc611db96d09d9e93290127a7441758b96637f23c3e3ae7c90ce82ac73d

Observation 434cd7b8-ca7e-4e57-98d7-209d76973cb4 · outbound

This paper cites Pointer sentinel mixture models, 2016.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pointer sentinel mixture models, 2016

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.905968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.905968Z digest=sha256:8f8353e0779e42e117f1761a85309234f0ece2aca9f3fb07ca56fa437fa840b9

Observation ee18e1d9-862c-4881-b67c-c18d9894008d · outbound

This paper cites Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.273719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.909635Z digest=sha256:5b40711c8e817b22ab545475cf13dd867cf8b1fac468b552dedf5c8c71734441

Observation 95b2e0a3-21d7-4511-9f68-f2f545fc8063 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pytorch: An imperative style, high-performance deep learning library

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.917296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.917296Z digest=sha256:a17947003fcf3c3b77268663d96d48dbf3c3358866887189e1f383ae2f5c098a

Observation 3195126a-8790-45de-982a-aefa27955565 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.239815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.920856Z digest=sha256:d660412fe440425a65b55d5a048b1e2618deb4033e336c09dde86b5cb0877a74

Observation 797a5ad8-591e-4927-a66e-8fffdc192611 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.227715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.924343Z digest=sha256:21b58753c6c1a339923dc01a73ab7880eaeeecaaed13c33adc6757276247cf6a

Observation aef745bd-b596-4ae7-a9ee-7af36f6fdcb7 · outbound

This paper cites Outrageously large neural networks: The sparsely- gated mixture-of-experts layer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Outrageously large neural networks: The sparsely- gated mixture-of-experts layer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.215236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.927724Z digest=sha256:92a441ce0cf5c13ee45c46a8431e78176040df2d42bbab304cc1342f96742fd0

Observation 461ce789-0910-464d-b777-7519b71a159c · outbound

This paper cites Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.202658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.931223Z digest=sha256:0adb0d47bcc803a7c855c340bf798011ea5b2106d0584644a9afa850f920d978

Observation a1711014-cffb-4582-8751-962448885974 · outbound

This paper cites Borg: the next generation.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Borg: the next generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.190565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.934674Z digest=sha256:763edcbc51ff06aec18069f59a995fbcb1fede9f7d72679632d1854107567d95

Observation e3d7be92-1c27-4073-bfd5-2ddf0befb776 · outbound

This paper cites Attention is all you need.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.177817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.938064Z digest=sha256:e22d604c3c93f8c9f6c1ab0b16244f0a34eb58e39751b4e49e50736c32d490c5

Observation b515975f-470e-4213-941d-a8e93a5a0525 · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large-scale moe models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Prophet: Fine-grained load balancing for parallel training of large-scale moe models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.164132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.941528Z digest=sha256:741328c5f145037be72878c9b8d6bd4cb06f9cd45fdeffdd841a9a63026dda08

Observation 09193130-6c13-4302-a06f-cbb3dbe6ed3c · outbound

This paper cites Resource-efficient algorithms and systems of foundation models: A survey.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Resource-efficient algorithms and systems of foundation models: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.151534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.945320Z digest=sha256:3ffa74ef6ab43aec872300f96a41a4c76e0a8684e9b511c757c1dddc64037d88

Observation 70c8a7b8-7217-4b16-a8ca-4bee9cf1fc76 · outbound

This paper cites Qwen2.5 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Qwen2.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.948934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.948934Z digest=sha256:b81b2b708ef8d159f8eb41ed3db5f7b1f9b6c5c64c2ff3a2e1fabc14043c991c

Observation f13c2bcf-08ac-4428-969e-df9c5f0bc56f · outbound

This paper cites Harnessing the power of LLMs in practice: A survey on chatgpt and beyond.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Harnessing the power of LLMs in practice: A survey on chatgpt and beyond

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.139056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.952538Z digest=sha256:a007bfef485658c4d43b4205b57e98104a5f8636c8d699483446b558721d761d

Observation 33d07119-e7e6-4e3f-a417-77c4e826e6de · outbound

This paper cites Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.126477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.956355Z digest=sha256:3a8ecea93e80b8dc0cba5e7d86e12d1680f4e33198ef7278a148680747cd6f49

Observation 25f8865e-cbd2-42ad-909a-9a5eeac9efb0 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization.

HarMoEny: Efficient Multi-GPU Inference of MoE Models SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.113747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.959823Z digest=sha256:07d80078eb45874d5326243c1e0fb9fddaa2c5c82c2d6586b26ff53a3241d8b8

Observation c367ef27-b1e2-4b77-9f1f-ef08e899122a · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.963734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.963734Z digest=sha256:4a992bcc59823fc5f464583efc2a6975d2639a83393997ebe1169f82d4ed39ab

Observation 32e082bc-2146-42df-8ca6-c3ee6fb2de9b · outbound

This paper cites an unresolved cited work.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T00:57:33.261045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:57:32.913583Z digest=sha256:cb114b222d359978020d977195f45d9a310148640b4e1173e00ad57a85a88483

Pith citing papers

Observation 0685dec3-b41b-42bf-afb6-58de81036b73 · inbound

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems cites this paper.

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems HarMoEny: Efficient Multi-GPU Inference of MoE Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:33:03.808270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T04:31:11.271729Z digest=sha256:f1106fea15bf4052b55dcb28954b431d415926a9414dd8b84a0f22209fcaf140