{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:7MWEPOGIC55GLQUL4FNTIWXRF5","short_pith_number":"pith:7MWEPOGI","schema_version":"1.0","canonical_sha256":"fb2c47b8c8177a65c28be15b345af12f58ab4ca1b643398073592a4ed824eb1f","source":{"kind":"arxiv","id":"2502.00922","version":1},"attestation_state":"computed","paper":{"title":"Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AR"],"primary_cat":"cs.LG","authors_text":"Arun George, Chinmay Hegde, Patrick Yubeaton, Pooria Taheri, Sai Qian Zhang, Shehab Naga, Siddharth Garg, Siddharth Joshi, Tareq Mahmoud, Tianhua Xia, Yasmein Khalil","submitted_at":"2025-02-02T21:23:42Z","abstract_excerpt":"As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \\emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \\emph{everywhere} -- cloud, disk, "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2502.00922","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2025-02-02T21:23:42Z","cross_cats_sorted":["cs.AR"],"title_canon_sha256":"08035ab0fcee0032f11fdb8fca9834ae0b502a834bb5d056932c8caeee63b4f9","abstract_canon_sha256":"02fd16dd4a2a7d810fb218f35c0295727cc75834adc6756587620a5be0ad8270"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:08:40.183238Z","signature_b64":"+0HpGM3T3iiOBTFnG2knbjr4yT0r8OvRhLkjiaC4DkmsjfOBHI5/qgXz+PjjqdGcqQ7/sCEvxT+1US7sMlqZAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fb2c47b8c8177a65c28be15b345af12f58ab4ca1b643398073592a4ed824eb1f","last_reissued_at":"2026-07-05T10:08:40.182766Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:08:40.182766Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AR"],"primary_cat":"cs.LG","authors_text":"Arun George, Chinmay Hegde, Patrick Yubeaton, Pooria Taheri, Sai Qian Zhang, Shehab Naga, Siddharth Garg, Siddharth Joshi, Tareq Mahmoud, Tianhua Xia, Yasmein Khalil","submitted_at":"2025-02-02T21:23:42Z","abstract_excerpt":"As they become more capable, large language models (LLMs) have continued to rapidly increase in size. This has exacerbated the difficulty in running state of the art LLMs on small, edge devices. Standard techniques advocate solving this problem through lossy compression techniques such as quantization or pruning. However, such compression techniques are lossy, and have been shown to change model behavior in unpredictable manners. We propose Huff-LLM, an \\emph{end-to-end, lossless} model compression method that lets users store LLM weights in compressed format \\emph{everywhere} -- cloud, disk, "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2502.00922","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2502.00922/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2502.00922","created_at":"2026-07-05T10:08:40.182825+00:00"},{"alias_kind":"arxiv_version","alias_value":"2502.00922v1","created_at":"2026-07-05T10:08:40.182825+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2502.00922","created_at":"2026-07-05T10:08:40.182825+00:00"},{"alias_kind":"pith_short_12","alias_value":"7MWEPOGIC55G","created_at":"2026-07-05T10:08:40.182825+00:00"},{"alias_kind":"pith_short_16","alias_value":"7MWEPOGIC55GLQUL","created_at":"2026-07-05T10:08:40.182825+00:00"},{"alias_kind":"pith_short_8","alias_value":"7MWEPOGI","created_at":"2026-07-05T10:08:40.182825+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.26558","citing_title":"Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding","ref_index":67,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2601.21198","citing_title":"ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03298","citing_title":"ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":28,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5","json":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5.json","graph_json":"https://pith.science/api/pith-number/7MWEPOGIC55GLQUL4FNTIWXRF5/graph.json","events_json":"https://pith.science/api/pith-number/7MWEPOGIC55GLQUL4FNTIWXRF5/events.json","paper":"https://pith.science/paper/7MWEPOGI"},"agent_actions":{"view_html":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5","download_json":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5.json","view_paper":"https://pith.science/paper/7MWEPOGI","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2502.00922&json=true","fetch_graph":"https://pith.science/api/pith-number/7MWEPOGIC55GLQUL4FNTIWXRF5/graph.json","fetch_events":"https://pith.science/api/pith-number/7MWEPOGIC55GLQUL4FNTIWXRF5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5/action/storage_attestation","attest_author":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5/action/author_attestation","sign_citation":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5/action/citation_signature","submit_replication":"https://pith.science/pith/7MWEPOGIC55GLQUL4FNTIWXRF5/action/replication_record"}},"created_at":"2026-07-05T10:08:40.182825+00:00","updated_at":"2026-07-05T10:08:40.182825+00:00"}