{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:63A4U2MUXMIGSAGHFWZHLQUQZT","short_pith_number":"pith:63A4U2MU","schema_version":"1.0","canonical_sha256":"f6c1ca6994bb106900c72db275c290ccf56de975154b8daba6b79e3c744fb233","source":{"kind":"arxiv","id":"2402.11187","version":2},"attestation_state":"computed","paper":{"title":"LaCo: Large Language Model Pruning via Layer Collapse","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Hai Zhao, Yifei Yang, Zouying Cao","submitted_at":"2024-02-17T04:16:30Z","abstract_excerpt":"Large language models (LLMs) based on transformer are witnessing a notable trend of size expansion, which brings considerable costs to both model training and inference. However, existing methods such as model quantization, knowledge distillation, and model pruning are constrained by various issues, including hardware support limitations, the need for extensive training, and alterations to the model internal structure. In this paper, we propose a concise layer-wise structured pruner called \\textit{Layer Collapse (LaCo)}, in which rear model layers collapse into a prior layer, enabling a rapid "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.11187","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-02-17T04:16:30Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"5bd59ca590728465bccf5982bb593ae9a30358e9d0213bfb4f3bd2ef85537a6b","abstract_canon_sha256":"9cc1587919000ae7062c5342044e0e1414da05d89c1f460af4fd7e5ed0c8434b"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:20:22.597163Z","signature_b64":"XTQcP5GG7zccA/+B+fUEgFxNiE/75x883scnvt/1lQt4AU+tMGXjNxHaX9WDfkDFfM/49k3MGjP7mIewYwdVDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"f6c1ca6994bb106900c72db275c290ccf56de975154b8daba6b79e3c744fb233","last_reissued_at":"2026-07-05T09:20:22.596650Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:20:22.596650Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LaCo: Large Language Model Pruning via Layer Collapse","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Hai Zhao, Yifei Yang, Zouying Cao","submitted_at":"2024-02-17T04:16:30Z","abstract_excerpt":"Large language models (LLMs) based on transformer are witnessing a notable trend of size expansion, which brings considerable costs to both model training and inference. However, existing methods such as model quantization, knowledge distillation, and model pruning are constrained by various issues, including hardware support limitations, the need for extensive training, and alterations to the model internal structure. In this paper, we propose a concise layer-wise structured pruner called \\textit{Layer Collapse (LaCo)}, in which rear model layers collapse into a prior layer, enabling a rapid "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.11187","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.11187/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.11187","created_at":"2026-07-05T09:20:22.596714+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.11187v2","created_at":"2026-07-05T09:20:22.596714+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.11187","created_at":"2026-07-05T09:20:22.596714+00:00"},{"alias_kind":"pith_short_12","alias_value":"63A4U2MUXMIG","created_at":"2026-07-05T09:20:22.596714+00:00"},{"alias_kind":"pith_short_16","alias_value":"63A4U2MUXMIGSAGH","created_at":"2026-07-05T09:20:22.596714+00:00"},{"alias_kind":"pith_short_8","alias_value":"63A4U2MU","created_at":"2026-07-05T09:20:22.596714+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.06574","citing_title":"Skip a Layer or Loop It? Learning Program-of-Layers in LLMs","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01544","citing_title":"CRePE: Convolution-aware Relative Importance in Post-training Pruning with Efficient Search","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26496","citing_title":"Dense2MoE: Pushing the Pareto Frontier of On-Device LLMs via Unified Pruning and Upcycling","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26538","citing_title":"CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08738","citing_title":"SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16234","citing_title":"No Free Swap: Protocol-Dependent Layer Redundancy in Transformers","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08738","citing_title":"SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training","ref_index":71,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12358","citing_title":"Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07271","citing_title":"Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions","ref_index":75,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT","json":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT.json","graph_json":"https://pith.science/api/pith-number/63A4U2MUXMIGSAGHFWZHLQUQZT/graph.json","events_json":"https://pith.science/api/pith-number/63A4U2MUXMIGSAGHFWZHLQUQZT/events.json","paper":"https://pith.science/paper/63A4U2MU"},"agent_actions":{"view_html":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT","download_json":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT.json","view_paper":"https://pith.science/paper/63A4U2MU","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.11187&json=true","fetch_graph":"https://pith.science/api/pith-number/63A4U2MUXMIGSAGHFWZHLQUQZT/graph.json","fetch_events":"https://pith.science/api/pith-number/63A4U2MUXMIGSAGHFWZHLQUQZT/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT/action/timestamp_anchor","attest_storage":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT/action/storage_attestation","attest_author":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT/action/author_attestation","sign_citation":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT/action/citation_signature","submit_replication":"https://pith.science/pith/63A4U2MUXMIGSAGHFWZHLQUQZT/action/replication_record"}},"created_at":"2026-07-05T09:20:22.596714+00:00","updated_at":"2026-07-05T09:20:22.596714+00:00"}