Pith. sign in

Paper Citation Record · LEDGER

A Survey on Inference Optimization Techniques for Mixture of Experts Models

As of 21 August 2026, this Paper Citation Record lists 100 of 233 outbound references and 13 inbound Pith citation observations for arXiv:2412.14219.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14219 v2

Coverage vector

measured 100 of 233 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:44:35.737640Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:09.991963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:14:57.408873Z

Reference resolution

100 of 233 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75556786-36ec-49de-8532-1b590561cf25 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.073722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.073722Z digest=sha256:91ab579c7fd43a812f2250c0666744795002f829ac66ecf684e756e1bea162f7

Observation 819cb0be-ee3f-4339-b017-737378d4e2a6 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

A Survey on Inference Optimization Techniques for Mixture of Experts Models On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.079132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.079132Z digest=sha256:09f4021fa7ca2fd9f6868342bc6e9eacebb60d15ca2619e0d5cce4d7cac06f20

Observation b5160ad1-ab34-4960-b716-80ed71cb2eab · outbound

This paper cites A., Jin, H., and Wu, Y.

A Survey on Inference Optimization Techniques for Mixture of Experts Models A., Jin, H., and Wu, Y

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.083638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.083638Z digest=sha256:7d06c8f35d2ee5d427c2fb1de3930a51a82e164fdb1165df06d47c4091b1f871

Observation 8e7a7fe5-b759-4f60-8dff-b0c56a332909 · outbound

This paper cites Pytorch, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Pytorch, 2024

Reference 4

Resolution
parse uncertain
no resolver link, observed 2026-08-11T12:44:35.088670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.088670Z digest=sha256:cb8ce4e511d310c4f923310f3d22b6cc39b61957c8abf96b2a624499fcea7671

Observation 4da7fc1f-b00d-47ef-bf27-42088972725b · outbound

This paper cites M., Neyshabur, B., and Zhai, X.

A Survey on Inference Optimization Techniques for Mixture of Experts Models M., Neyshabur, B., and Zhai, X

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.092783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.092783Z digest=sha256:f7c58b155b7a1d80e4f226ba981c9e8504b8dcc9e3e43c5c60bab1ab746d2399

Observation cc36abf9-4e7d-4445-ab1a-51e67f2edc21 · outbound

This paper cites Green ai: A preliminary empirical study on energy consumption in dl models across different runtime infrastructures.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Green ai: A preliminary empirical study on energy consumption in dl models across different runtime infrastructures

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.097360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.097360Z digest=sha256:ae5c7f77b3f889c8ec8160f1aa079c5d7d8bc4c4a68d26825cd2f254090c58b6

Observation ae11bb5a-b37c-4c47-b240-d9d4e73da5e6 · outbound

This paper cites Artificial analysis llm performance leaderboard, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Artificial analysis llm performance leaderboard, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.101848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.101848Z digest=sha256:6627517264dc9f448567e58eea389df7fb60066b3e37a1397bd1c658a37ab673

Observation 9963d321-045c-45cf-a97d-14c7a15f1a7f · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2023.

A Survey on Inference Optimization Techniques for Mixture of Experts Models The claude 3 model family: Opus, sonnet, haiku, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.105844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.105844Z digest=sha256:b8a81cea6c86a4410cbfb694c0367bfe6be4643d4f5d4627003ffbeece692202

Observation 8d2b4150-d179-4f3f-9f0b-af6a917fc177 · outbound

This paper cites Efficient Large Scale Language Modeling with Mixtures of Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.109794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.109794Z digest=sha256:dc3dc0020bb2e8fa6c2a49214cd648733ef0054989926014eabad1097253b7c0

Observation 7cd5cc8c-12a3-4e42-97fc-6d93e7143ee9 · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.114286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.114286Z digest=sha256:c81f6b24a7bddb5547c2a877505384efa23e80d034c61c5813bf9e6ad168c7c0

Observation 48aa5ea5-c10a-4959-903e-ba1ad3bf2ca5 · outbound

This paper cites E., Akella, A., and Wang, Z.Read-me: Refactorizing llms as router-decoupled mixture of experts with system co-design, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models E., Akella, A., and Wang, Z.Read-me: Refactorizing llms as router-decoupled mixture of experts with system co-design, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.119325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.119325Z digest=sha256:3ec78ba41ffade6bf231b366ebb67683075edf1556626eddc3ebd51b3401297d

Observation ee4a0f25-b9cf-4e47-a5ba-819afc92b3e6 · outbound

This paper cites Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.123801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.123801Z digest=sha256:33a89d82ca9ea1ae4f43c82819a34639b8fbb218cca2936d1fca1890e8dc3013

Observation 21065e8d-4f60-429c-892d-b9447403fd4b · outbound

This paper cites Authorea Preprints (2024).

A Survey on Inference Optimization Techniques for Mixture of Experts Models Authorea Preprints (2024)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.128321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.128321Z digest=sha256:4cf7d8edc78418777cdc910ca22059d47d92ac719d479a022d99f0b51ce9f663

Observation 704bd528-17da-4d98-8598-51a5d1566ea6 · outbound

This paper cites Moc-system: Efficient fault tolerance for sparse mixture-of-experts model training, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Moc-system: Efficient fault tolerance for sparse mixture-of-experts model training, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.132287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.132287Z digest=sha256:e668725ac0c9789d887df8eae67fe7e7d53604e1125022408850d0c42bdb273e

Observation f96f75df-b905-4a1c-b369-d7f6a433b973 · outbound

This paper cites E., Zaharia, M., and Stoica, I.Moe-lightning: High-throughput moe inference on memory-constrained gpus, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models E., Zaharia, M., and Stoica, I.Moe-lightning: High-throughput moe inference on memory-constrained gpus, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.135870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.135870Z digest=sha256:fcd746d4fe9b368cf283385d96f74a481c7bf42d5e07c4a320f012cb75c8e5bf

Observation 877d352a-259b-4c49-8629-2767b2fa037a · outbound

This paper cites Advances in Neural Information Processing Systems 35 (2022), 22173–22186.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Advances in Neural Information Processing Systems 35 (2022), 22173–22186

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.139850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.139850Z digest=sha256:88945fb69876784da1a4ca178427bd853800033d8dd2c1f68287f88ed9893c1f

Observation f9bbfdc1-2141-4094-852c-6aed723071c1 · outbound

This paper cites arXiv preprint arXiv:2410.08589 (2024).

A Survey on Inference Optimization Techniques for Mixture of Experts Models arXiv preprint arXiv:2410.08589 (2024)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.143261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.143261Z digest=sha256:ee68a333e7cedc21b0508cbf69f3e45d5bcea85a0ad9fb85dc2e6985f11914f3

Observation 0b99f528-462e-49d7-a05f-8956a01c7f87 · outbound

This paper cites Task-Specific Expert Pruning for Sparse Mixture-of-Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Task-Specific Expert Pruning for Sparse Mixture-of-Experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.146689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.146689Z digest=sha256:fddcbd8ffe1a51f2e6f9ac01ade8251d2cbca0f52ca4dd1ea673cd9e6ada1bfc

Observation cece13b7-bdb5-4c09-937c-f048a482fcd1 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Reproducible scaling laws for contrastive language-image learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.150268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.150268Z digest=sha256:41e398121d38f17dc6718f23a496c4cbf13f2d4c4eca5d9e35bb1fed3118e2cf

Observation 5623aee1-4c4a-42ae-b2c1-c7cb374a55bf · outbound

This paper cites W., Sutton, C., Gehrmann, S., et al.

A Survey on Inference Optimization Techniques for Mixture of Experts Models W., Sutton, C., Gehrmann, S., et al

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.154201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.154201Z digest=sha256:9d459ade163105c5353f8edba5ea82a92a143a0dfb6aad1e91277f198754e266

Observation 3f4913c2-b64e-48ee-ac41-693744a62138 · outbound

This paper cites A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.157741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.157741Z digest=sha256:0dcde71b8847b0de2efd6d447c3c18399606517868142616220e89ac56f479ac

Observation f6d64cac-06c8-4f90-8841-87a63ac5d28b · outbound

This paper cites an unresolved cited work.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.177680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.177680Z digest=sha256:9be9dd55d8cb4c302a6f0af490a61e6d7d4a75df7e991c423bd090d4e14c7025

Observation 5a5ecbdb-2329-480a-b43a-f438a6398fc7 · outbound

This paper cites Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.197818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.197818Z digest=sha256:9738dbd9d549e33e0b58b87aaf5bccf937e70c745e770fa39229d6838aa29a8b

Observation 3f3ce8e0-35e7-459a-aedd-c8bbdb309866 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

A Survey on Inference Optimization Techniques for Mixture of Experts Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.220112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.220112Z digest=sha256:3a08d4ed33734176de0fdf09f0bb0223ffe7690fdf541096055ffd36f839bbdb

Observation 9e132ec2-57b1-428c-aee1-b7347f94442f · outbound

This paper cites Innovating for Tomorrow: The Convergence of SE and Green AI.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Innovating for Tomorrow: The Convergence of SE and Green AI

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.240011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.240011Z digest=sha256:a065c26b8aff6b21e7fe8c6e8b6bdd272d233ce38a879ebc29ac25e661b43848

Observation f8d5e531-5c6e-4610-bf94-d5320e8c1a65 · outbound

This paper cites MoEUT: Mixture-of-Experts Universal Transformers.

A Survey on Inference Optimization Techniques for Mixture of Experts Models MoEUT: Mixture-of-Experts Universal Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.256899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.256899Z digest=sha256:0c554d2467eacec5df22b4bfcbb190d1a21e215f826262fae8dcef4edf7da135

Observation f7c9e1aa-fa4d-4b21-b550-578058234edc · outbound

This paper cites SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention.

A Survey on Inference Optimization Techniques for Mixture of Experts Models SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.261587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.261587Z digest=sha256:5015607a7426164c6f05acbe11a87af43ccc1605cbdcc8a7d6fde947fb77930a

Observation 8e3e9628-ed37-4f82-83a3-b272df83d33d · outbound

This paper cites In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) (Boston, MA, July 2023), USENIX Association, pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) (Boston, MA, July 2023), USENIX Association, pp

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.265805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.265805Z digest=sha256:e0939550f1042beb451a3cd356e6d97bb4b0a7a2ead90f0834fef8bb0792dbbf

Observation 07d6e292-912b-40ff-b8dc-9497e9e85446 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.269913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.269913Z digest=sha256:98d24cf7fa85864ec436c62c065fbb121966a3ae0d7aedfbf9b614ecc407df46

Observation f92c5aff-79a8-4c33-a1f6-63447148a713 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.273834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.273834Z digest=sha256:06c362daa2321a3095aaf37be1559d2386e46b06a893459ee953d14685e2bade

Observation f571f914-c158-4d77-bf54-3a425daa95a5 · outbound

This paper cites DeepSeek-V3 Technical Report.

A Survey on Inference Optimization Techniques for Mixture of Experts Models DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.277153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.277153Z digest=sha256:66ef021c3d776311409e304b5105fed5b1ed5505b12ea9fbfde970a64d963f8f

Observation 64aa3727-b185-4784-90da-1035a4137392 · outbound

This paper cites an unresolved cited work.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.280993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.280993Z digest=sha256:e6a3fe1d4a8e0ed774cc6d466e8419c16dd6e15ac34cde656a0a5ac4db3d3b50

Observation 2d8dade5-e848-4ab7-9560-8805a61075f9 · outbound

This paper cites XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.284328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.284328Z digest=sha256:124bb40b45c3456473f59e4e8d5bb023a5635f1d891a1fb20ac9bb0effec9a2a

Observation 4e9f4672-1f2a-4e46-8234-04387efb75f6 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Automatically constructing a corpus of sentential paraphrases

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.287945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.287945Z digest=sha256:caa42952d450d2a5a75edbd82dbb0cd234a140f97b42092d53bd1fd6d3a9cc29

Observation d0b204e6-0e98-410c-831c-8daaa84627e7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Survey on Inference Optimization Techniques for Mixture of Experts Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.291549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.291549Z digest=sha256:ad1d308f67e89acd8e1e4d6f592fd681a5fbfbeb022ae58d2cfc2cc77861d958

Observation 880ad4af-6edb-489f-a983-bea0e73d20ac · outbound

This paper cites In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2024), pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2024), pp

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.295373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.295373Z digest=sha256:357e2e66db9adff64e7df3ff8bfd7deadd024823feddad0bb0bc67d021cb3085

Observation 345194ac-4b9c-430e-9258-7cbb7fae0aec · outbound

This paper cites M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A.

A Survey on Inference Optimization Techniques for Mixture of Experts Models M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.298507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.298507Z digest=sha256:abda7eda25ef26d6693e20d2cfc2662a7b5132b679aaadaae39c251dd9d25ac0

Observation 31d55731-a4aa-4ede-be1f-31f70cc041cd · outbound

This paper cites Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.301772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.301772Z digest=sha256:b09e17171cb86b3394b81152465a8907215b73aed59706dbfb74edec3735ad23

Observation 47a8d856-94e5-490b-ae52-07e4d6c84349 · outbound

This paper cites Learning Factored Representations in a Deep Mixture of Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Learning Factored Representations in a Deep Mixture of Experts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.305018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.305018Z digest=sha256:039baa396973cae8d0d510fc25244d24c4349885ec631170faeae8d688bd3256

Observation d6be3993-e01f-4bed-890f-9fe9708749d7 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.308441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.308441Z digest=sha256:a7027292ed940e8ade590f82d98d2f3fc2d4f53433016155aa6ce1e65fd873ce

Observation a15d8e24-8217-467a-a609-7ddcea6ce950 · outbound

This paper cites LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.312020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.312020Z digest=sha256:e5618faada9a618119566099bc619a2738ecdf86e5788c88c4454c1501c5bd6c

Observation cb5e4eb4-18b4-497d-b3f5-1a9580d3f5b5 · outbound

This paper cites Advances in Neural Information Processing Systems 35 (2022), 28441–28457.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Advances in Neural Information Processing Systems 35 (2022), 28441–28457

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.315464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.315464Z digest=sha256:d8a6e8b2419eb07f10584b096a7e0c283c149b16eb06ca62929a85398ef8aa0a

Observation 0c6aa6d4-bf5e-4d24-b3f3-c390eefc1c35 · outbound

This paper cites A Review of Sparse Expert Models in Deep Learning.

A Survey on Inference Optimization Techniques for Mixture of Experts Models A Review of Sparse Expert Models in Deep Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.318591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.318591Z digest=sha256:92447ba1545c3ba90cab7f165ec7593b44c14d0fa2ee53c76986ee251b2c760a

Observation 077fe38b-bb39-4943-87f0-b24f7bb77761 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.326960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.326960Z digest=sha256:3b3aa690bace4681bd313ec551cdb6f53c60198fa2a696775d522ac06df722d3

Observation 9a04c506-e79c-40b4-8b5b-3dab7c107b41 · outbound

This paper cites QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.370331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.370331Z digest=sha256:26a913c4e1c8a73a21f8685126ff9529ec8b603cc178446ed707516e591359de

Observation 9ddd0bcc-00e8-4496-8648-877579f65487 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

A Survey on Inference Optimization Techniques for Mixture of Experts Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.412088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.412088Z digest=sha256:6a6d9380814b67845d1ae6b3a8c904aef783cb771be2464894c373149b186fce

Observation 47017505-543e-43c8-bbd3-000cbc5ac293 · outbound

This paper cites arXiv preprint arXiv:2412.07067 (2024).

A Survey on Inference Optimization Techniques for Mixture of Experts Models arXiv preprint arXiv:2412.07067 (2024)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.441234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.441234Z digest=sha256:131f894d2b588e8689f9abf683f96144f216c02f05238c7f4dfc8748237b0f9e

Observation 934038ce-097d-4c6b-a4dd-685fe481eb6b · outbound

This paper cites LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences.

A Survey on Inference Optimization Techniques for Mixture of Experts Models LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.444971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.444971Z digest=sha256:2bf6f78725a569361cfc7de9a67af49f6bc51bb4eff3900e887c7a0ea6088ab2

Observation 8d1e3641-3c36-4981-96da-d70770bcdfdb · outbound

This paper cites Higher Layers Need More LoRA Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Higher Layers Need More LoRA Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.448889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.448889Z digest=sha256:ff831b4fad5e09dabb4dd1441406e248bdde54759861f654e949b2f96dbd08e6

Observation 1707126b-88c3-4a63-ae03-3432f0526543 · outbound

This paper cites Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.452797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.452797Z digest=sha256:447de572c7964de33bee00c3f232dc6af257ae263c1f5acea906161c5d80baae

Observation f9dc8566-d88f-4f27-8b79-5f3a0d5e918b · outbound

This paper cites Llama.cpp, 2023.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Llama.cpp, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.457032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.457032Z digest=sha256:dad49a8eb027e9eb3887c0b5af27d3ca9fc5f6415982d3b6c08d32daef981ad6

Observation 6b57d0ed-92c0-43bb-b6e9-8e5a7681d2a5 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models MiniLLM: On-Policy Distillation of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.460526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.460526Z digest=sha256:4ffacb0390a34f06ef639c24282e15f42a98bb66e04bb1ac6348d0c1934faf4a

Observation b0867ddf-9f69-4902-9cce-0fc23997bd68 · outbound

This paper cites Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.464375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.464375Z digest=sha256:c81478bfadb89bacd907a2bb4275fbb66550640be4492ed09ed4d8f6679b1a78

Observation 4fd36e36-9173-4575-b923-fbe5f112f3c0 · outbound

This paper cites Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.467652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.467652Z digest=sha256:7f64c1be33073e16972867155cd503685801e50bde00d622e872c2a2ef4db273

Observation 27453547-5353-4a11-b56d-4a7916768262 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence 14 , 7 (1992), 751–769.

A Survey on Inference Optimization Techniques for Mixture of Experts Models IEEE Transactions on Pattern Analysis and Machine Intelligence 14 , 7 (1992), 751–769

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.471778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.471778Z digest=sha256:5c9111078e45895408ef7821ce6be5e5033116999d25f7444296197893187959

Observation daa71b01-c541-4ddf-8540-340cc1be5e8c · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

A Survey on Inference Optimization Techniques for Mixture of Experts Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.475551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.475551Z digest=sha256:e9e0013c258d2409956eb93f458284bc582e5bdf7cacdbb3183e4db062810d94

Observation b587bb20-9313-4cdf-a97c-33af2d58fc48 · outbound

This paper cites In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (New York, NY, USA, 2022), PPoPP ’22, Association for Computing Machinery, p.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (New York, NY, USA, 2022), PPoPP ’22, Association for Computing Machinery, p

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.479930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.479930Z digest=sha256:0046536a8e309ed95d1204e30c9eab22064ddf0d9aad38d359fa58b4a31b904f

Observation 8bf91338-3306-4131-922f-052ea349403a · outbound

This paper cites Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.484115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.484115Z digest=sha256:b58cc8563063b1cfea5fe1743574ed9999489fb0d48ceb094566959fc4bd23df

Observation c7612188-29eb-4c7c-8af2-be3e157e19b8 · outbound

This paper cites Merging Experts into One: Improving Computational Efficiency of Mixture of Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Merging Experts into One: Improving Computational Efficiency of Mixture of Experts

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.488134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.488134Z digest=sha256:9db7d91d088ee662e6449adec290c5aee0893e99015ca9cf94f5519dc64eb69d

Observation 5051b46b-430c-4e55-86fb-b5bb8acb8967 · outbound

This paper cites Y.Expertflow: Optimized expert activation and token allocation for efficient mixture-of-experts inference.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Y.Expertflow: Optimized expert activation and token allocation for efficient mixture-of-experts inference

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.492363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.492363Z digest=sha256:966f39a1bfbaca534806cc08f94899fbe9c3e2dfdcb8f34b40807803eb8d7fa0

Observation 00bd6203-44f4-4756-860b-c6716ed2d817 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Measuring Massive Multitask Language Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.496229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.496229Z digest=sha256:ca9177c26aa20060e3d266e23e4e3f6a3d9a4742e80a0c227345a6655c443fd3

Observation 99673fbc-4be1-41c3-8f37-c224edf4e022 · outbound

This paper cites Improving efficiency in multi-modal autonomous embedded systems through adaptive gating.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Improving efficiency in multi-modal autonomous embedded systems through adaptive gating

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.500350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.500350Z digest=sha256:0fd0d9f8e72d18267feb76e41234c634460e6289c548730d909334e9485b246a

Observation 25750da2-8fe7-44ab-98ae-952d45a1dfc3 · outbound

This paper cites Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.503688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.503688Z digest=sha256:b6bbda5c3e8c1a8b73f42729b94f7e15ae8fdec66261b4618b70122fbe085ae9

Observation de989d47-e93b-431f-bd0a-51b97962b5a8 · outbound

This paper cites Mixture Compressor for Mixture-of-Experts LLMs Gains More.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Mixture Compressor for Mixture-of-Experts LLMs Gains More

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.534899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.534899Z digest=sha256:6e2da2dc06cd530e57125297729df162054007a16edd6194d93b9d8eb8bb2947

Observation 110d3c23-bafd-4f95-be48-203e50c754b3 · outbound

This paper cites Experts Weights Averaging: A New General Training Scheme for Vision Transformers.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Experts Weights Averaging: A New General Training Scheme for Vision Transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.606784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.606784Z digest=sha256:a5b11d6f91f8dcf6fdc616a2ff95d34d8f63e4b3f2f0e7f87257974cf32f0599

Observation d3f91ee1-e91b-44e2-a5fe-75d0ea8cb9ab · outbound

This paper cites Llm-perf leaderboard, 2023.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Llm-perf leaderboard, 2023

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.610071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.610071Z digest=sha256:3a933f03fa69437621b4bbb475c30b21a9824bfb5c0c8e74a6a07c172aa5d08b

Observation 183aae9f-561b-41bc-a209-c93e3f7a4fc8 · outbound

This paper cites Transformers: State-of-the-art machine learning for jax, pytorch and tensorflow, 2023.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Transformers: State-of-the-art machine learning for jax, pytorch and tensorflow, 2023

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.612770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.612770Z digest=sha256:e6efe53173bc56fe0edf9010c4fdcd5cf71e17390672fc54d2e9dd77a57e5fc4

Observation 5d22193a-2a9a-46f7-a6a5-a7c39baf5f62 · outbound

This paper cites Proceedings of Machine Learning and Systems 5 (2023), 269–287.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Proceedings of Machine Learning and Systems 5 (2023), 269–287

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.616071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.616071Z digest=sha256:0dca97d6a974fd1a407e28250eb3685d9e95fe8b99502a20f2cd4e594eff8139

Observation 86d7c8c6-05bd-45d5-b224-a3b90c6493b9 · outbound

This paper cites In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) (2024), IEEE, pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) (2024), IEEE, pp

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.619326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.619326Z digest=sha256:ff4acee8434d388482893a022c97134e50d9084ceaaeddd3dd1f0b566ce290a2

Observation fa8d9af4-0683-4095-9b4a-9a0858086aea · outbound

This paper cites Comet: Learning cardinality constrained mixture of experts with trees and local search.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Comet: Learning cardinality constrained mixture of experts with trees and local search

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.622341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.622341Z digest=sha256:b91377c9211ef589c94c1db66b77d01606bcd7ef8af803da72bce264cee248ed

Observation 8848767c-4a59-4155-8d3b-1200fa7d5379 · outbound

This paper cites Mixture of Experts with Mixture of Precisions for Tuning Quality of Service.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Mixture of Experts with Mixture of Precisions for Tuning Quality of Service

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.625440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.625440Z digest=sha256:c2cbe2ab0a8903316c1e9d90c682e0793cb151013539f0bbafefbe5eb5565263

Observation b4a57f02-c501-493d-a320-d7bad3c1623e · outbound

This paper cites Adaptive mixture of local expert.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Adaptive mixture of local expert

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.629440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.629440Z digest=sha256:7de3ba0ccc3b47c8a6c9b620c6dc778d71547ef62d899087cfc66d4886c1597d

Observation 57ee8503-51c9-4916-bed0-8a172a4c5f4b · outbound

This paper cites Beyond data and model parallelism for deep neural networks.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Beyond data and model parallelism for deep neural networks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.633415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.633415Z digest=sha256:a4188b82aa8c6701053dc3151ad4ecdb6bb3c7acee2269f7e01eeff8124efdf1

Observation f964a7f0-13be-4585-b0f5-648267fc499c · outbound

This paper cites Mixtral of Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Mixtral of Experts

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.637206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.637206Z digest=sha256:bc403a9965ed1113069c9653b088e85970f9b634258a20d6c7bd9143e162bedc

Observation 4fe8223f-7d7c-4c93-90d5-0a1e009cae9c · outbound

This paper cites MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.640997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.640997Z digest=sha256:4e6eb20a87980aaa3dd397b2e7704d0d380f3a8947fda62fee1c41966c59d4d0

Observation 07e85e0e-df70-475d-b297-10f32600e46f · outbound

This paper cites arXiv preprint arXiv:2410.11842 (2024).

A Survey on Inference Optimization Techniques for Mixture of Experts Models arXiv preprint arXiv:2410.11842 (2024)

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.645375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.645375Z digest=sha256:c8915daff0b41bdcaaddc33dfa7de0adc25d7da35b5f74cd2a8cdbaa04619d4c

Observation dd2a1c99-3080-47b3-86bc-af93b0c267d7 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.649290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.649290Z digest=sha256:8f193cb22fe04e42bb1a61dbd3e78faa7f1ca7b29b542792c1099e892c934e75

Observation 30545522-c08b-4287-b349-ad973f3573fb · outbound

This paper cites Scaling Laws for Neural Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Scaling Laws for Neural Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.653280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.653280Z digest=sha256:0af99594cea54b6efbb7b0212b4612022c302168b43a05cb0753124d2b27d45f

Observation 86aa194b-7cce-4a77-b07b-fbeebff5134b · outbound

This paper cites In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (2018), pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (2018), pp

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.657280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.657280Z digest=sha256:ca10e9b379ed364021b99c3389c0859eb40373c85e4ea9865e051037bc04d009

Observation b1b81c50-da23-4aa4-af36-a7dc2f63ae41 · outbound

This paper cites LaDiMo: Layer-wise Distillation Inspired MoEfier.

A Survey on Inference Optimization Techniques for Mixture of Experts Models LaDiMo: Layer-wise Distillation Inspired MoEfier

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.661106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.661106Z digest=sha256:5422c9db83e5da6019af4f9fe56ffc99a91cad059e5df817445ecb95d1897776

Observation fee13037-c2bb-4c29-91bd-e2a9e166bcb5 · outbound

This paper cites Monde: Mixture of near-data experts for large-scale sparse models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Monde: Mixture of near-data experts for large-scale sparse models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.665220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.665220Z digest=sha256:9a2e5863c888bfaa50c25b9c4665bdeed4dce8a45361bfc8c61bbb10011b51db

Observation 8df0eb6f-749e-4d26-b793-51b5922b65fa · outbound

This paper cites Mixture of Quantized Experts (MoQE): Complementary Effect of Low-bit Quantization and Robustness.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Mixture of Quantized Experts (MoQE): Complementary Effect of Low-bit Quantization and Robustness

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.668872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.668872Z digest=sha256:efc1b224142b6698d0689a9b37e6047039e4eb1604b675b8fedbd16b9e4d603e

Observation e9d79fa4-3f3f-4126-b874-fc780814baf4 · outbound

This paper cites Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.672870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.672870Z digest=sha256:79431301af4f14955e6255a890754b174d4deed492e1e5ba4bba307d72228a7d

Observation a696b95e-0a83-4776-b663-616130dc9387 · outbound

This paper cites SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget.

A Survey on Inference Optimization Techniques for Mixture of Experts Models SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.676859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.676859Z digest=sha256:6660a533c1636f9b8b7ec306481899404698b291ae94af3a085a15da15de9c43

Observation 8f48895e-3490-4100-99b8-be762f3f40cb · outbound

This paper cites Learning multiple layers of features from tiny images.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Learning multiple layers of features from tiny images

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.680881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.680881Z digest=sha256:be6d6cfbc8fad89ae67c0ac1e0d8798a7800006ac8645c229889442c4e1ab39e

Observation a3626450-51f4-4da0-a006-bd44be5b1c17 · outbound

This paper cites H., Gonzalez, J.

A Survey on Inference Optimization Techniques for Mixture of Experts Models H., Gonzalez, J

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.684829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.684829Z digest=sha256:7c6b35da9e9d4e88a6d1ffeb84366b6933a38ee6eedb44263fc8d6bbf64778c5

Observation fc091daf-2eb9-4ebe-bdb8-99fdd26d9e63 · outbound

This paper cites STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning.

A Survey on Inference Optimization Techniques for Mixture of Experts Models STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.688612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.688612Z digest=sha256:caa91e997bb89a24b5073a95479b1ed4bda79f9fdc57900c3f81d2f4e2bf429d

Observation 94d95565-bdb4-4785-b4f2-7af1b80b0283 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

A Survey on Inference Optimization Techniques for Mixture of Experts Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.692765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.692765Z digest=sha256:21d5f2bfa5c2108da297fd79135ce46f957a1c87c70b8e2fd0e7d38a87146770

Observation 840caeb4-e9a7-4abc-a798-5c6a0bbc76c1 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Base layers: Simplifying training of large, sparse models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.696989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.696989Z digest=sha256:ad9a415248d0b3fdc82553d4bfc8bb81d1abcd98027ac22797443d9f6b25a900

Observation 32e2a295-61c1-43d3-90c6-58c39aa808e1 · outbound

This paper cites LLM Inference Serving: Survey of Recent Advances and Opportunities.

A Survey on Inference Optimization Techniques for Mixture of Experts Models LLM Inference Serving: Survey of Recent Advances and Opportunities

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.700698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.700698Z digest=sha256:eab20f6e703e6b30f0709f22abd731d9f7e90fbecf0a9356d49f8470fd95aca7

Observation e3c8b767-459a-4f21-b599-4c0bd8ac5f70 · outbound

This paper cites In 2023 USENIX Annual Technical Conference (USENIX ATC 23) (Boston, MA, July 2023), USENIX Association, pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In 2023 USENIX Annual Technical Conference (USENIX ATC 23) (Boston, MA, July 2023), USENIX Association, pp

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.704747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.704747Z digest=sha256:f7645fdb330fdf677bf1f583311eceedd6841ceca5647b573702b38f73efaa19

Observation e8686f65-3714-468d-b754-4f90b614ad53 · outbound

This paper cites an unresolved cited work.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.708699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.708699Z digest=sha256:cfa1bc0ca1ee1f50754048b1e4c890f41e898d3d144c3d9e7a393f6d86f22b5c

Observation 78e97e6c-b793-4704-a286-39ea059a9997 · outbound

This paper cites Adaptive Gating in Mixture-of-Experts based Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Adaptive Gating in Mixture-of-Experts based Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.711754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.711754Z digest=sha256:8deb6d81c650ccc2eb2125b42068505adeba988a50414a9b57967d8badcc07e6

Observation 5a23dcbb-c344-48ba-8baf-6940ca740e89 · outbound

This paper cites LocMoE: A Low-Overhead MoE for Large Language Model Training.

A Survey on Inference Optimization Techniques for Mixture of Experts Models LocMoE: A Low-Overhead MoE for Large Language Model Training

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.715001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.715001Z digest=sha256:54b6381d704c97171c938c2ee54052ceb26215cd9b8242aee9693a6ac9271863

Observation 99e56292-ea87-405a-a710-415b074b15ad · outbound

This paper cites Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.718615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.718615Z digest=sha256:3c9e552c5241dfed7a9432cf2a8dbfeac4e2429d7a2f4401e8228125c55dc115

Observation da4baa6e-d79f-48a0-a575-b395853e7b3c · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.721993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.721993Z digest=sha256:beeefcbbf438bcf12345f36c2c822b86a9c3df7e503dabed0c8613254fdcaa80

Observation fecbe6f1-cb96-47c5-99d6-5b94bb8b45f4 · outbound

This paper cites Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.725558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.725558Z digest=sha256:29e847d4a8e32a36b12d7562955c94d93cde5b43730b8eb6d1507f905fded420

Observation a78933dc-47ac-4271-b551-fc29047ec64d · outbound

This paper cites QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts.

A Survey on Inference Optimization Techniques for Mixture of Experts Models QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.729393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.729393Z digest=sha256:77fdcfd1065d1833727468f8245119ce52b18e63865d4b8d8e2e90d1ecffcd05

Observation a0a343e3-506a-4016-9847-3d808037e74d · outbound

This paper cites Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.733510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.733510Z digest=sha256:014588990bbea06659c4d5e5ce944ecf17298e77c7b3b57ee3fb12132182d963

Observation af71cf47-1ed2-435a-aea8-5bc6d4bca2e5 · outbound

This paper cites In International Conference on Machine Learning (2023), PMLR, pp.

A Survey on Inference Optimization Techniques for Mixture of Experts Models In International Conference on Machine Learning (2023), PMLR, pp

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.737640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.737640Z digest=sha256:9d60a400dae153269e3fb995bbd2fad1312b7687245b117c3eb2e1d85c561460

Pith citing papers

Observation e984fe05-a186-4a51-9352-57116c1f43a5 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:09.991963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:09.991963Z digest=sha256:b5333c319d7e4577af8ac801ea75053616111e9af5c9bf7bdf301526713d2b2d

Observation cad36c35-3160-41bb-8775-127ade344a8c · inbound

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models cites this paper.

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:32.448368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:32.448368Z digest=sha256:00a885bb732ff4b2718afd35bd47e6a62dc643da9d867e15f53895a88dbfc587

Observation 40b3f56d-e59a-4592-a9b2-cb357d70bb9d · inbound

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning cites this paper.

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:02.481239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:02.481239Z digest=sha256:b18bc63b91825c8a78b8bc1da5723f5916cc64f26bc1143d18b20d1aee5e4a54

Observation 0bc7d5a6-7732-4bbb-8062-79122b652db5 · inbound

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions cites this paper.

Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T16:20:59.357958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:20:59.357958Z digest=sha256:fbac6e973dc554f677c5ef4847b8af08696a93e33ffa30561e75f888b63c2e2a

Observation 7a15768f-cd61-47c5-aa42-83d95bffcd15 · inbound

MoPEQ: Mixture of Mixed Precision Quantized Experts cites this paper.

MoPEQ: Mixture of Mixed Precision Quantized Experts A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:41:32.103806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:41:32.103806Z digest=sha256:b6462f7be5c4042356d760efcec6d67f3b05722dd44f08df1da05d18ab2ed75b

Observation 1f47d478-6a7e-47f2-ac81-c96d1a6ac6a8 · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:05.392303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:05.392303Z digest=sha256:ea56a7edeec7dd1859d44d0a5298d37fed5fc094a5abdfa59967651a0c7899a4

Observation b5e59b19-0dac-4a67-b59a-07493690a00c · inbound

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE cites this paper.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T14:34:48.524592Z digest=sha256:8324e6151db027699447e1de0343070f61c1a459d772d0a41a2a6003a9bf7b05

Observation 92e6bf66-b973-4c95-a353-024bfdd4950d · inbound

Networking-Aware Energy Efficiency in Agentic AI Inference: A Survey cites this paper.

Networking-Aware Energy Efficiency in Agentic AI Inference: A Survey A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T17:39:02.634334Z digest=sha256:080d490a4fe93c79be209a9a8391386f8ea75637f3349a090f56812ddab56ae7

Observation abaf544c-2bfd-4e3a-ba8c-ba2a145de1a2 · inbound

GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking cites this paper.

GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T17:36:29.717792Z digest=sha256:2030988b493d589c6dcf8c1aee52058a291cf64431300597d29f00ca20fad191

Observation c25c1e7d-5234-41d2-8bdd-00651eb5c059 · inbound

Fast MoE Inference via Predictive Prefetching and Expert Replication cites this paper.

Fast MoE Inference via Predictive Prefetching and Expert Replication A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:25:59.268868Z digest=sha256:c90c326fbb4c1c4571fe9ab0e8100831bad39bf9d3d084475e8c46251cc53a81

Observation 626f439a-51d7-48d7-9930-d0c856a573a6 · inbound

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution cites this paper.

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:07:53.045082Z digest=sha256:352af7e9a1a5ad6188fc2ddbd8bb9b17da569610af6dd083d23f24e9ffe20581

Observation c389cbba-d360-4036-8a3b-84c8ab9d75bd · inbound

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference cites this paper.

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-08-11T02:23:21.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T04:27:02.915854Z digest=sha256:5ed8274b89e75a3509a02dcc84c037fe54917d5807b2458bdc28918784357361

Observation 8f2b9ad3-4bc8-440b-b148-5692f207abeb · inbound

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism cites this paper.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Inference Optimization Techniques for Mixture of Experts Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.262836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.262836Z digest=sha256:b250a1c1d9bca05bada00c5ca7c990a3989109e927ec9f01366b2495cad7ad66