{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:GLLWBBEQIPRNTQHEMHK7VXNRGV","short_pith_number":"pith:GLLWBBEQ","schema_version":"1.0","canonical_sha256":"32d760849043e2d9c0e461d5faddb1354cb45192f449abeed97c0dc75f4e064d","source":{"kind":"arxiv","id":"2310.01334","version":2},"attestation_state":"computed","paper":{"title":"Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Mohit Bansal, Pingzhi Li, Prateek Yadav, Tianlong Chen, Yi-Lin Sung, Yu Cheng, Zhenyu Zhang","submitted_at":"2023-10-02T16:51:32Z","abstract_excerpt":"Sparsely activated Mixture-of-Experts (SMoE) has shown promise to scale up the learning capacity of neural networks, however, they have issues like (a) High Memory Usage, due to duplication of the network layers into multiple copies as experts; and (b) Redundancy in Experts, as common learning-based routing policies suffer from representational collapse. Therefore, vanilla SMoE models are memory inefficient and non-scalable, especially for resource-constrained downstream scenarios. In this paper, we ask: Can we craft a compact SMoE model by consolidating expert information? What is the best re"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.01334","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2023-10-02T16:51:32Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"f211e477a7538d4acdd825a577141151f7e4e8c993283c5a2133325ff61ef138","abstract_canon_sha256":"11f34890772a8f9146b9cd4f9df9231b682fd0bcb26f55b1304980d8898b6d49"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:55:59.843981Z","signature_b64":"yGqsSbLbYHukQGFncDPfs4euxsHvNjN7JyD79+UjR+ROppV3Swgxb4aWk5BzCbWtAHph3ucYpFLlwehT/yBXBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"32d760849043e2d9c0e461d5faddb1354cb45192f449abeed97c0dc75f4e064d","last_reissued_at":"2026-07-05T07:55:59.843474Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:55:59.843474Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Mohit Bansal, Pingzhi Li, Prateek Yadav, Tianlong Chen, Yi-Lin Sung, Yu Cheng, Zhenyu Zhang","submitted_at":"2023-10-02T16:51:32Z","abstract_excerpt":"Sparsely activated Mixture-of-Experts (SMoE) has shown promise to scale up the learning capacity of neural networks, however, they have issues like (a) High Memory Usage, due to duplication of the network layers into multiple copies as experts; and (b) Redundancy in Experts, as common learning-based routing policies suffer from representational collapse. Therefore, vanilla SMoE models are memory inefficient and non-scalable, especially for resource-constrained downstream scenarios. In this paper, we ask: Can we craft a compact SMoE model by consolidating expert information? What is the best re"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.01334","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.01334/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.01334","created_at":"2026-07-05T07:55:59.843538+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.01334v2","created_at":"2026-07-05T07:55:59.843538+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.01334","created_at":"2026-07-05T07:55:59.843538+00:00"},{"alias_kind":"pith_short_12","alias_value":"GLLWBBEQIPRN","created_at":"2026-07-05T07:55:59.843538+00:00"},{"alias_kind":"pith_short_16","alias_value":"GLLWBBEQIPRNTQHE","created_at":"2026-07-05T07:55:59.843538+00:00"},{"alias_kind":"pith_short_8","alias_value":"GLLWBBEQ","created_at":"2026-07-05T07:55:59.843538+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.05538","citing_title":"Less is MoE: Trimming Experts in Domain-Specialist Language Models","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18643","citing_title":"Post-Trained MoE Can Skip Half Experts via Self-Distillation","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29350","citing_title":"ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2411.08982","citing_title":"Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08738","citing_title":"SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18643","citing_title":"Post-Trained MoE Can Skip Half Experts via Self-Distillation","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2509.25041","citing_title":"GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2603.06003","citing_title":"EvoESAP: Non-Uniform Expert Pruning for Sparse MoE","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13997","citing_title":"HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08738","citing_title":"SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training","ref_index":42,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV","json":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV.json","graph_json":"https://pith.science/api/pith-number/GLLWBBEQIPRNTQHEMHK7VXNRGV/graph.json","events_json":"https://pith.science/api/pith-number/GLLWBBEQIPRNTQHEMHK7VXNRGV/events.json","paper":"https://pith.science/paper/GLLWBBEQ"},"agent_actions":{"view_html":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV","download_json":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV.json","view_paper":"https://pith.science/paper/GLLWBBEQ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.01334&json=true","fetch_graph":"https://pith.science/api/pith-number/GLLWBBEQIPRNTQHEMHK7VXNRGV/graph.json","fetch_events":"https://pith.science/api/pith-number/GLLWBBEQIPRNTQHEMHK7VXNRGV/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV/action/timestamp_anchor","attest_storage":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV/action/storage_attestation","attest_author":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV/action/author_attestation","sign_citation":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV/action/citation_signature","submit_replication":"https://pith.science/pith/GLLWBBEQIPRNTQHEMHK7VXNRGV/action/replication_record"}},"created_at":"2026-07-05T07:55:59.843538+00:00","updated_at":"2026-07-05T07:55:59.843538+00:00"}