{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:W32PNKH66ZHSO6PFCWZPJOHEWD","short_pith_number":"pith:W32PNKH6","schema_version":"1.0","canonical_sha256":"b6f4f6a8fef64f2779e515b2f4b8e4b0d6eb3708d0c6395667bb28281b6fa15d","source":{"kind":"arxiv","id":"2402.08562","version":1},"attestation_state":"computed","paper":{"title":"Higher Layers Need More LoRA Experts","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Baochen Sun, Chongyang Gao, Daiyi Peng, Jie Yang, Jinmeng Rao, Kezhen Chen, Ruibo Liu, VS Subrahmanian, Xiaoyuan Guo, Yawen Zhang","submitted_at":"2024-02-13T16:04:21Z","abstract_excerpt":"Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limited. Recent efforts integrate LoRA and Mixture-of-Experts (MoE) to improve the performance of PEFT methods. Despite promising results, research on improving the efficiency of LoRA with MoE is still in its early stages. Recent studies have shown that experts in the MoE architecture have different strengths and also exhibit some redundancy. Does this statement also apply to parameter-efficient MoE? In this paper, we int"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.08562","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-02-13T16:04:21Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"64a73de42be16890de94a5671f672a0de0130544e0f4e6fa58247bf7da0b2f65","abstract_canon_sha256":"59108d7df57260974db8d84e2cba5de7d1813688213384a95719f5f239e226f1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:44:48.311660Z","signature_b64":"cjlUew+FlERM1Qr4bCb2BMXx8HBe0uArcqj4vnaFd6ACxnxbGUy9X8XHcGkKcW4AfGfO8skdt1k38qUBwQT5BQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b6f4f6a8fef64f2779e515b2f4b8e4b0d6eb3708d0c6395667bb28281b6fa15d","last_reissued_at":"2026-07-05T07:44:48.311179Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:44:48.311179Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Higher Layers Need More LoRA Experts","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Baochen Sun, Chongyang Gao, Daiyi Peng, Jie Yang, Jinmeng Rao, Kezhen Chen, Ruibo Liu, VS Subrahmanian, Xiaoyuan Guo, Yawen Zhang","submitted_at":"2024-02-13T16:04:21Z","abstract_excerpt":"Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limited. Recent efforts integrate LoRA and Mixture-of-Experts (MoE) to improve the performance of PEFT methods. Despite promising results, research on improving the efficiency of LoRA with MoE is still in its early stages. Recent studies have shown that experts in the MoE architecture have different strengths and also exhibit some redundancy. Does this statement also apply to parameter-efficient MoE? In this paper, we int"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.08562","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.08562/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.08562","created_at":"2026-07-05T07:44:48.311247+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.08562v1","created_at":"2026-07-05T07:44:48.311247+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.08562","created_at":"2026-07-05T07:44:48.311247+00:00"},{"alias_kind":"pith_short_12","alias_value":"W32PNKH66ZHS","created_at":"2026-07-05T07:44:48.311247+00:00"},{"alias_kind":"pith_short_16","alias_value":"W32PNKH66ZHSO6PF","created_at":"2026-07-05T07:44:48.311247+00:00"},{"alias_kind":"pith_short_8","alias_value":"W32PNKH6","created_at":"2026-07-05T07:44:48.311247+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.21645","citing_title":"Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models","ref_index":129,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13161","citing_title":"A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18083","citing_title":"A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\\Delta$ Integration into Upcycled MoE","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2507.00029","citing_title":"LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13161","citing_title":"A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.26340","citing_title":"Adaptive and Fine-grained Module-wise Expert Pruning for Efficient LoRA-MoE Fine-Tuning","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07256","citing_title":"TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts","ref_index":15,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD","json":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD.json","graph_json":"https://pith.science/api/pith-number/W32PNKH66ZHSO6PFCWZPJOHEWD/graph.json","events_json":"https://pith.science/api/pith-number/W32PNKH66ZHSO6PFCWZPJOHEWD/events.json","paper":"https://pith.science/paper/W32PNKH6"},"agent_actions":{"view_html":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD","download_json":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD.json","view_paper":"https://pith.science/paper/W32PNKH6","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.08562&json=true","fetch_graph":"https://pith.science/api/pith-number/W32PNKH66ZHSO6PFCWZPJOHEWD/graph.json","fetch_events":"https://pith.science/api/pith-number/W32PNKH66ZHSO6PFCWZPJOHEWD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD/action/storage_attestation","attest_author":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD/action/author_attestation","sign_citation":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD/action/citation_signature","submit_replication":"https://pith.science/pith/W32PNKH66ZHSO6PFCWZPJOHEWD/action/replication_record"}},"created_at":"2026-07-05T07:44:48.311247+00:00","updated_at":"2026-07-05T07:44:48.311247+00:00"}