Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:41:52.952675Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2511.04805.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:41:52.952675Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:34.294269Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T12:43:25.643835Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 01d8d632-0f50-416c-a227-1bb7d1a31f30 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference PIQA: Reasoning about Physical Commonsense in Natural Language
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b778f4-0e17-4462-a0bc-d74535b28087 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Retraining-free merging of sparse moe via hierarchical clustering, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991a0df8-9dd3-4937-bcc5-32c20e3974ef · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045acc1c-6900-406b-a9f8-8b148b8ae78c · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d95ac3-731f-4fa0-908f-8216daf7813a · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e71420-08dd-4415-bec4-d6d31c269da7 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d306b013-b76d-4eef-89b4-564c7cad30cc · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02cae48-d646-42a7-943f-7854944a495e · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Delta Decompression for MoE-based LLMs Compression
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00087d0a-f498-4b2b-8434-ae3b5b9bc33d · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f8a18a1-2ca8-4822-a646-55e5456bf810 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6576e63a-5b29-4b7a-b325-5a756d272590 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a063dfd-ee8a-45a6-869c-13db3188df59 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6e218f-ac0d-4a91-ad32-5c40fc7a4fff · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca93892-cceb-4536-a2cb-23743e9981af · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Mixtral of Experts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5345532e-07ac-4710-acc3-c0145d49c1b6 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference SqueezeLLM: Dense-and-Sparse Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f19fcee-eeb6-46a2-9a2f-59a5a0f17840 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Compressed sparse tiles for memory-efficient unstructured and semi-structured sparsity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8d09d90-56db-404f-952d-062cf6df0437 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a99e4a-4446-4a66-978b-cd849d8a7891 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa53dd31-ba35-43bd-a312-d9718e0365cf · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48c4271-7f56-4a8e-b8f5-47e105aec39a · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5df031a-ec70-4ec2-b6c7-fb95b160fc42 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd25960-a6da-476c-8de5-adbda4992407 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4f4a16-d4d7-4b78-845a-cd58e8b87437 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Pointer sentinel mixture models, 2016
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 025e935a-b34d-42c3-8a2b-1a75f9a5c104 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Ronald Miller
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe2b623-33f9-4e78-96a5-feb718932f01 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acdb41fb-8c1e-448f-beb2-a547391a88cb · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a676dbe-75d5-4327-b584-e9c86751f416 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b98c728-1004-4379-a87a-1534c1aad3cb · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference A Simple and Effective Pruning Approach for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b123788-4dea-4741-9699-a2a1fcd7b1fb · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011be3d4-145b-4233-99e7-dfedbfc066f5 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c986254d-4e84-45a7-81e5-e58639686ea9 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926af6f5-b454-492f-bdc7-4c4613c6064b · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021d6ffe-eb1a-4971-a46c-8aa30a6331d9 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c332b9a1-ca43-443a-bc6a-f37421567945 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference 70 URL https://arxiv.org/abs/2504.11651
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a97476-723d-43b8-bf54-07c85b7fceb1 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de22f6b-4fb0-432a-9842-f726d01d94dd · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference write newline
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cbb099-45b1-45f8-a50e-aff714f8f1d8 · outbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ece3df0-42ae-408b-bd84-f23684b18c69 · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c352232-09da-40ee-a84f-2162e5063b0b · outbound
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference However, their widespread deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of experts increases
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e6eba3-7f8f-4c47-9480-71d5faa9e20b · inbound
ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046135be-7fb7-4b0c-bec4-6263e30c6054 · inbound
Pruning and Distilling Mixture-of-Experts into Dense Language Models PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 98ebed38-9329-46d0-b19d-3d444bf03a0f · inbound
Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.