Pith. sign in

Paper Citation Record · LEDGER

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

As of 20 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 8 inbound Pith citation observations for arXiv:2505.05799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05799 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:01:20.181224Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:23.206783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.791368Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fae7f16a-4076-4d7c-b3ac-4397f61496f8 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.057459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.057459Z digest=sha256:cc4e37345503855b25407c07050d947754e89c948b998ccbec7a617104475c15

Observation 442e46a2-6bc2-411c-84c1-41317f8a6c3e · outbound

This paper cites Low-bit quantization of neural networks for efficient inference.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Low-bit quantization of neural networks for efficient inference

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.523687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.062112Z digest=sha256:7def164d45f481fd77b60cbdd354eac151e583c45989c5d0ea7c9c0c125e8d68

Observation 2bbf663d-b615-4203-9c62-2ac844c7cb79 · outbound

This paper cites an unresolved cited work.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.065538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.065538Z digest=sha256:aaeb1fb7ab40dc8586f89237eaa58e2ab5359551e08522ee8f256b1a2a7d0b2b

Observation 13d58736-2609-45fd-a29c-c156249b2ea0 · outbound

This paper cites W., and Keutzer, K.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design W., and Keutzer, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.506413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.069871Z digest=sha256:1f46982a2109cec3e08b3354f89c57bc5ff622f0d9dd784d0ac6b367b692826c

Observation 3401228c-f079-463d-ada3-8a6ca46e1288 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.073536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.073536Z digest=sha256:e49d8ced6bebace7c2334c748c559a2de8d432dbbfbb70b0fe4133d8749cef11

Observation 6c448d4d-c0f8-492a-94d2-40e883942945 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.077587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.077587Z digest=sha256:ced4fe3976f7cd79a8f93201d58bea86622d9ea7c28fe62b5fadc092e159ecec

Observation 3d60be69-3fb9-49a9-b18a-05b46e5b8c58 · outbound

This paper cites MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.081983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.081983Z digest=sha256:271275ccb4a55179d46432d64c2c5b643ffc0c1fcc5e80dd2e9cac2ed338ead6

Observation 953946d5-49a6-46f0-b501-a78f5152ee61 · outbound

This paper cites Bounds on multiprocessing timing anomalies.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Bounds on multiprocessing timing anomalies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.494395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.085706Z digest=sha256:1c6f4adda87a4edb34b8b91ff33bab832f5567df001751df0f892971fbf17967

Observation e648adde-0231-4344-914a-9972409ac979 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.089202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.089202Z digest=sha256:1a5b925b09c00d2f87cb62e612bbbf266b40119a780f9884629bc65233732b6a

Observation 029dab77-d234-4476-b5fd-912743227866 · outbound

This paper cites Mixture Compressor for Mixture-of-Experts LLMs Gains More.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Mixture Compressor for Mixture-of-Experts LLMs Gains More

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.093251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.093251Z digest=sha256:18f38c84b66f7c72388d8f8a04e35cfc6b1569ae91083a71955d3348415b6bde

Observation 6b93a2e4-05c1-46cd-afb1-76b41fa05113 · outbound

This paper cites Mixtral of Experts.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Mixtral of Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.097145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.097145Z digest=sha256:d909d583e7cd681575c92b4518a281c575c568e3a839070766914ce8c5613fc5

Observation 41c61031-866a-4907-bf8c-50104bd6c46d · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.100645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.100645Z digest=sha256:1c32eb0efb724206b6c311a13448f9ec141c01bf1e0733100580800f4a51b799

Observation 7d6270ae-746d-425f-beab-4b82702672fd · outbound

This paper cites Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.108195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.108195Z digest=sha256:ed7b178ddd821a4e00811c6f11a5df93d81abc18dfe17402446a99ae1773f1da

Observation fcc5b887-e0e2-4a17-912a-fc68b7a8fbed · outbound

This paper cites H., Gonzalez, J.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design H., Gonzalez, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.111677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.111677Z digest=sha256:7c3f8397a34aaeff169d41cb9eeed062d3bf592fc8e80a063da676185b6541ed

Observation 9aa0f757-8182-4331-bd24-960600a3dd8e · outbound

This paper cites QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.115100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.115100Z digest=sha256:d0570e2749f3900c791ec75059b52719499e32dbc7204043e291911028732aa7

Observation 9b0cd6e3-4513-46c0-9e10-5dd0c9b3bcb1 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.476199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.118660Z digest=sha256:f38aa81d1cf4a87a76e7f195855c288f740468b1a45dc0c8955d05cb2c982d96

Observation 82e02a0e-00ff-4469-a085-791b4132f00a · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.121741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.121741Z digest=sha256:3e85c49dfcb434bdca32a97eaecb40a8c60a3e57224f8f5075ea88915b794e4b

Observation 8045fb33-47be-48ae-be19-9c8c422b18f5 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.125386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.125386Z digest=sha256:5818ee61bbb7b266488d81507afa62d52c968e42e39ad5b8c5fffb4b3fd90388

Observation c11c59ea-be45-44f3-ba3b-8b0c62cdea3c · outbound

This paper cites DeepSeek-V3 Technical Report.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.128688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.128688Z digest=sha256:1e7cd55305e3759c321b9e47d7758c8ccdaeb9682f39fca601d7c388aefc6596

Observation 9c216f39-8978-4db8-b4d6-5add35b02b26 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.131856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.131856Z digest=sha256:51dfeddb90e5461142b5508fd9f89aa1372e4ecdb29e76466fb9b83526f939ff

Observation 5546954f-c393-4fed-baf4-296e1ba9c294 · outbound

This paper cites Pointer sentinel mixture models, 2016.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Pointer sentinel mixture models, 2016

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.135454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.135454Z digest=sha256:03a5f24d9f57a3d4163ed1bec233196a08bd9f94aa8a606fd8b2b46f5d01ea98

Observation d406c120-e30f-4666-a268-cd6cdd663e32 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design OLMoE: Open Mixture-of-Experts Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.138698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.138698Z digest=sha256:a811ffd25fc2d6fa59a487effb9b0c8dbb3e336e3eddf6971dbce724e665a8ae

Observation bc90ad85-5426-4828-a92f-99e0c69fa85e · outbound

This paper cites Massive Activations in Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Massive Activations in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.142490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.142490Z digest=sha256:67c29f8ef9c9e4babc7f08141178a0e8398dee65e6f6399087b8bb7faa01c567

Observation 40010b15-ec52-4f49-b234-e096bfb73003 · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.145976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.145976Z digest=sha256:8a7c5a91327762b6c5107a81e08926a64252ec2ffedd7b8a05d24d051d8126c8

Observation 94be0a7a-2c94-49e8-bf56-a42d2a7e4936 · outbound

This paper cites CUTLASS , January 2023.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design CUTLASS , January 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.149211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.149211Z digest=sha256:419c10cc2c4fb12686f445e450afa00dec7e2ed802b571abc7d856db0caad8ae

Observation 8077026f-2022-4705-8605-6602c0fb99bd · outbound

This paper cites Introducing DBRX: A New State-of-the-Art Open LLM , 2024.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Introducing DBRX: A New State-of-the-Art Open LLM , 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.453101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.152779Z digest=sha256:e0a38841bf646649a6bbf1ce04809bed5c699bc1c5d01011ebb3e1c061adb202

Observation 766fb9b8-4356-41e6-a632-aadb94f155d4 · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Haq: Hardware-aware automated quantization with mixed precision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.442414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.156164Z digest=sha256:1d2504be5c90247dcdb5fff05cbc7e5a47b634a534620655d44fba030169275a

Observation 7a091d6f-6b83-4fd1-8f30-a549748a27f1 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Roofline: an insightful visual performance model for multicore architectures

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.431000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.159576Z digest=sha256:36a46d148ca739d87c3f6574fa20592fc426b4af8b0993a13253cbabd12981f1

Observation d898436a-7226-44cc-ae98-1d0efa43961d · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.163033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.163033Z digest=sha256:544f878c7fe645d5dbcc0f40b85105aec9ce113f696d8e8e40a11d707a540cde

Observation eddbd31d-f171-4b83-b690-17228831ae6f · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.166912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.166912Z digest=sha256:41b330160c95c3855430295be68a9cb81cd0cfa9cfddea67ec0741f55a0c6062

Observation bd7beb66-6f9b-40f3-a631-1496e62c40c7 · outbound

This paper cites Qwen2.5 Technical Report.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.170464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.170464Z digest=sha256:ac7f2533e1eed7804511c01b52aefe3ca42e6690387bcb930e481818dd998837

Observation 96252e40-ad26-42ad-b7a9-d50829ebd01a · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.173955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.173955Z digest=sha256:1c85c22a1e33e1a7efc523859195623e79d16bd39c2ef9b577a1cb96d478e924

Observation 1d7a95aa-ce8c-4afb-b10d-3c0656c2ece8 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Atom: Low-bit quantization for efficient and accurate llm serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.177528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.177528Z digest=sha256:f28f3df6eeba520a134b46583ab08b18b2f1aa238dbd16c6cd4be0a39b9eb58e

Observation 2d12f9fc-d808-40e2-a996-746172485f14 · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.181224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.181224Z digest=sha256:765cf3632b68b6a4a177311d68e3b0a33da2516f56e42d5542affbe8db580e45

Pith citing papers

Observation aac9bb44-f36d-488a-8da0-8cff6d5dd7b3 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.206783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.206783Z digest=sha256:60a3b0d412bb48bfabc0ffb324654c9f31848a8cd890504667a9aa43eb0cf2d6

Observation 4e76c524-1d36-4d4b-b0e7-1723665d2ccc · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:03.992193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:03.992193Z digest=sha256:7238d3165927f7f5292d7799c81f62570e1a01beb89638714d6d0d557125b6a8

Observation 3f50dad6-e6cb-4bef-a7a2-38beaa5ceeba · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.798871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:2793562a098b4edf35bc4f05b48d05c258c609246b2a8329a6bb27e8503022e2

Observation a1e71420-08dd-4415-bec4-d6d31c269da7 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.850528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.850528Z digest=sha256:092b56cc0b838a56b59ab453bd7530434901c9f5d3288bc3c3853c9a720c080e

Observation 1861e9f3-1d25-473e-916f-a8c08c0f001a · inbound

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering cites this paper.

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.205305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T07:09:48.239662Z digest=sha256:7d984cd689a85fb6e32dbe97445de3820117584aa52ba05a1a4e94bbc25ef66c

Observation abbf53c6-7b83-45a6-bc60-0514dfa14d96 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.704030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:9aaf5e747a2cf362767566a764809fd893a6271dcbd012c121fd1985ea0e5ef5

Observation 4315bd31-77f7-43ef-971f-904869774e49 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.963748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:7d4d68f28092007ca3e863f48e508f3c3f06349b5326068b1e87e971ea55d270

Observation f994a790-dfd9-4824-a686-9f7f09ec4ad9 · inbound

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization cites this paper.

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.793200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T07:06:32.220604Z digest=sha256:97d03696fc8b6211b092b5740513859a3a265eebae86671879d79ec770c7c487