Pith. sign in

Paper Citation Record · LEDGER

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration

As of 19 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2505.06481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06481 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:12.253030Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cb88a68-5a9c-499f-a41d-47a13801a8cd · outbound

This paper cites GPT-4 Technical Report.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.108866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.108866Z digest=sha256:aa337e1f0b22ed2b0d992ac8bc3d08c746368e52d4782f81d6f650ef109210d0

Observation 2977fc75-2541-4c08-9382-1b5a6ebbfae7 · outbound

This paper cites Instruction Tuning for Secure Code Generation.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Instruction Tuning for Secure Code Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.140853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.140853Z digest=sha256:c246f7a88821e937d3990cbaab5a75bc909d26665146e8776f2914d3ea58a4a4

Observation 926e7fdb-af0b-4f6d-a7fd-264c558c3acf · outbound

This paper cites Mixture of Experts with Mixture of Precisions for Tuning Quality of Service.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Mixture of Experts with Mixture of Precisions for Tuning Quality of Service

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.154979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.154979Z digest=sha256:651433998b23aac6eccf1ae057f07cf6db11fb15e38e7c85c19f4896e2b17884

Observation 0ad6bc08-7f39-478d-90c6-43f97f14df6f · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Averaging Weights Leads to Wider Optima and Better Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.159574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.159574Z digest=sha256:a3ce2f7c447ad6585649278642c5e8c08d4000aa91e379a6d504bd7d721ca3a6

Observation 67993dd1-3c23-42bc-81de-66d53d5ebf02 · outbound

This paper cites Dataless Knowledge Fusion by Merging Weights of Language Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Dataless Knowledge Fusion by Merging Weights of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.168712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.168712Z digest=sha256:db9c9a34639a25f03e5716dafa3beef653e2208132e353816404ac67a565aff7

Observation 180a2d0f-3b78-4c18-a98c-20115a7492bd · outbound

This paper cites Scalable and Efficient MoE Training for Multitask Multilingual Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Scalable and Efficient MoE Training for Multitask Multilingual Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.177984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.177984Z digest=sha256:c6aa858d02191056b81de44090bd6df1c05c86576a3b9f4ce263e059aed29db3

Observation bd693f6d-3a48-4236-8e8e-3dc39cbb3c5f · outbound

This paper cites K., El-Araby, E., and El-Ghazawi, T.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration K., El-Araby, E., and El-Ghazawi, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.977588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:45:12.182655Z digest=sha256:a359ed0c315a6256900fcb8c5d493ec0318cccdcb8730d0ec877651e771c9fef

Observation 1fc3f0ee-fe10-4aa9-a96a-fa192e03b7b9 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.191723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.191723Z digest=sha256:e8a346708bd7c2aff3c110b1824bdc71b9cc7994f41a0d77404669056b4b6e00

Observation e50477da-2adb-4129-8311-01b9b4a8f7cb · outbound

This paper cites A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.961418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:45:12.196486Z digest=sha256:0bf7cc01bdf5fa91f02bb25d7de81a54a04f7b982b057f1606a12bc19624b6d4

Observation e75e35e0-ff73-4501-ab77-ae34bc89b0ca · outbound

This paper cites Pointer Sentinel Mixture Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Pointer Sentinel Mixture Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.201264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.201264Z digest=sha256:8aa06a61e9c6a67ce15ba827bd61c9607f554724125ab0c64ce8c599dfa50a71

Observation f706776d-9051-4c0e-9b59-a2ef268d3098 · outbound

This paper cites Learning More Generalized Experts by Merging Experts in Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Learning More Generalized Experts by Merging Experts in Mixture-of-Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.210312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.210312Z digest=sha256:014a915bfbcf7138334fede009f67c004ac90b23feb19bacff1208f5dcbcac6a

Observation 28b23937-7fa0-437e-a296-b90d895a38f0 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.214771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.214771Z digest=sha256:9e000cc3220284fa3db9a2868f569a3744006994ee864d4a6b3be2fe3e8d5390

Observation 675643da-3237-4297-8020-47c5b86e76c0 · outbound

This paper cites Wortsman, M., Ilharco, G., Gadre, S.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Wortsman, M., Ilharco, G., Gadre, S

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-15T22:45:12.562004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:45:12.228731Z digest=sha256:05f6236b5ab7634838d0cffc9006988aaaccac47185e9cc9a897aad2f2a12bc8

Observation 97984be8-d51f-433c-b35f-b5b268f12e12 · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.233874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.233874Z digest=sha256:340b501a3dc7b2be3fab40684868590576d033b449e22aed891abc882c0cbb64

Observation 90bcabaa-a40b-4ee2-80a9-1e0eaeb73d1f · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration TIES-Merging: Resolving Interference When Merging Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.238873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.238873Z digest=sha256:6e088e7bf02f00242f01fd4fc49557b9ded926c60df44be0672b614619d70540

Observation 4635aa8b-5752-4e4d-be90-438bdaba0c1d · outbound

This paper cites MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.243625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.243625Z digest=sha256:586f63f0e7e1d0028ebf62f8a1db6e2542813d25d4b2b88dd62f8119a87d41df

Observation 51732976-2fb8-4f82-a106-ec094726bdf9 · outbound

This paper cites SurgeryV2: Bridging the Gap Between Model Merging and Multi-Task Learning with Deep Representation Surgery.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration SurgeryV2: Bridging the Gap Between Model Merging and Multi-Task Learning with Deep Representation Surgery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.248175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.248175Z digest=sha256:af0739b2b39207b6f3108cb578c6669ce077a8235581a3b348b9139ed4c2f7ee

Observation 4cbcd22f-5233-4ee4-ac32-627f4265ded2 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Taming Sparsely Activated Transformer with Stochastic Experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.253030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.253030Z digest=sha256:274695f153a56c9c5080c0de4500dfcdb7b9116a4ec2a76ac938476e84cde8c8

Observation d90afda7-6451-4fa6-a7cc-66e96165ae23 · outbound

This paper cites Mixtral of Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Mixtral of Experts

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.164099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.164099Z digest=sha256:f6dc04fef37efe3a2566542f26df2aa3d86cd7a2a4345937155465280422c613

Observation ebd18d54-e81c-405d-a247-9128159767ba · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.173343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.173343Z digest=sha256:f53a81b2deee55c5334b96f4bd93f84a050885c95babf246ee4691ab9846aba7

Observation 1e9d518d-105e-415c-bf48-de4c9ac13204 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.135322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.135322Z digest=sha256:6dace5404ab95ca5ccb0f18090eaa20ba0d313f786b8cca0f73d4d5363e3aaf8

Observation e560976b-51aa-4a73-8b09-283b64fb4c1c · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.150327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.150327Z digest=sha256:2624adb4c55590a9799bc9605c85ad8dc35ee4cc6ae0ed7590554092e9bf2fb9

Observation 1c64c3ed-d511-460e-a6db-c20dfd1dabe7 · outbound

This paper cites M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.187026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.187026Z digest=sha256:8025fee5e096da301cf3a36dbec3b61a06576a1abc3253122dccaf81e47db9c9

Observation 58335d6e-ac25-4847-819c-062afed78d62 · outbound

This paper cites SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.205697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.205697Z digest=sha256:c5dde942c8449a3d0e936ace3845bba51c663e274e23fce5e790c5ad89bc381a

Observation 2f83e62e-312f-41df-b7c3-fb75f1f40703 · outbound

This paper cites Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.219620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.219620Z digest=sha256:136fb192e3e207392ee15744483b3504f2eea49be9fe1350a1ab85b63a247a5f

Observation 7000e4ab-0d46-4197-a195-5babe867c87d · outbound

This paper cites Exploring in-memory accelerators and fp- gas for latency-sensitive dnn inference on edge servers.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Exploring in-memory accelerators and fp- gas for latency-sensitive dnn inference on edge servers

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.946045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:45:12.224258Z digest=sha256:d6a6d1bcc2f23806f12bbeab76180bbebe861c1523e44b4c6877bab74aec39e7

Observation a47a333e-07c6-4081-a2f7-b1d752c1d7cd · outbound

This paper cites A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.119688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.119688Z digest=sha256:35758c6af03ee8eb3db831a51baa5926f1a91c5b2f49853856932f930c317c40

Observation fcd730d3-e048-4e14-97a4-e59d216fa2a9 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.125138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.125138Z digest=sha256:f4a5ba986ec9dcbe5938f370a17bebd30ef7b041ec83e834275e491a149ea1e6

Observation b72950c3-4cfe-439e-af6d-2931c5b715e3 · outbound

This paper cites Parameter Competition Balancing for Model Merging.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Parameter Competition Balancing for Model Merging

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.130419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.130419Z digest=sha256:607e663cf857f57bbdef7de2b7167aa61b5da9414312b87df98c3b04dca74dba

Observation 5ac90b88-ef90-4698-8c4e-5138cb000122 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Measuring Massive Multitask Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.145498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.145498Z digest=sha256:23ea9ff61b347ebbc88d38d54c88a5652697a244f5323459a6311b27f9f8dd93

Observation 3b4c9814-32b5-4088-869a-dafd627ad67b · outbound

This paper cites Task-Specific Expert Pruning for Sparse Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Task-Specific Expert Pruning for Sparse Mixture-of-Experts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.114554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.114554Z digest=sha256:ed9e84b2cc65d4159bc61965f87d4fa69791df4a9084b65561e828f5bbd8f3ad

Pith citing papers

No inbound Pith citation observations are available.