{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:ECWCXGQG5GZHZZCQMTHOU5L3ZI","short_pith_number":"pith:ECWCXGQG","schema_version":"1.0","canonical_sha256":"20ac2b9a06e9b27ce45064ceea757bca1bc9284a7034bf85d1441ef3378d12e9","source":{"kind":"arxiv","id":"2410.14731","version":2},"attestation_state":"computed","paper":{"title":"MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Bokai Lin, Hao Zhang, Siqi Kou, TianQi Hou, Xiaofeng Gao, Zhijie Deng, Zihao Zeng, Zipeng Xiao","submitted_at":"2024-10-16T08:34:51Z","abstract_excerpt":"KV cache has become a de facto technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical information for self-attention. As the size of the model and data grows, the KV cache can quickly become a bottleneck within the system in both storage and memory transfer. To address this, prior studies usually focus on the first three axes of the cache tensors for compression. This paper supplements them, focusing on the feature dimension axis, by utilizing low-rank projection matr"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.14731","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2024-10-16T08:34:51Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"8cda1c297bd53c7d36d906614c16e4c2d0a9bd07b2ebf67632304d4975fee1bf","abstract_canon_sha256":"699b9de8ada0133c8f58414dd02b0d1f97d68d53336bbfeb175fa62f3e8ff2e2"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:03:52.853375Z","signature_b64":"2kBEaIpG8U2qIGE04+mhv4Od8bSb5MEhK543o+98tNZXrzC3WJXd0DpjdpNNXLnjZmNiHz0E3clX54RmAfE8Cw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"20ac2b9a06e9b27ce45064ceea757bca1bc9284a7034bf85d1441ef3378d12e9","last_reissued_at":"2026-07-05T11:03:52.852898Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:03:52.852898Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Bokai Lin, Hao Zhang, Siqi Kou, TianQi Hou, Xiaofeng Gao, Zhijie Deng, Zihao Zeng, Zipeng Xiao","submitted_at":"2024-10-16T08:34:51Z","abstract_excerpt":"KV cache has become a de facto technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical information for self-attention. As the size of the model and data grows, the KV cache can quickly become a bottleneck within the system in both storage and memory transfer. To address this, prior studies usually focus on the first three axes of the cache tensors for compression. This paper supplements them, focusing on the feature dimension axis, by utilizing low-rank projection matr"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.14731","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.14731/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.14731","created_at":"2026-07-05T11:03:52.852955+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.14731v2","created_at":"2026-07-05T11:03:52.852955+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.14731","created_at":"2026-07-05T11:03:52.852955+00:00"},{"alias_kind":"pith_short_12","alias_value":"ECWCXGQG5GZH","created_at":"2026-07-05T11:03:52.852955+00:00"},{"alias_kind":"pith_short_16","alias_value":"ECWCXGQG5GZHZZCQ","created_at":"2026-07-05T11:03:52.852955+00:00"},{"alias_kind":"pith_short_8","alias_value":"ECWCXGQG","created_at":"2026-07-05T11:03:52.852955+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.18587","citing_title":"Dual Dimensionality for Local and Global Attention","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08382","citing_title":"STAR-KV: Low-Rank KV Cache Compression via Soft Thresholding for Adaptive Rank Control","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17757","citing_title":"OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19218","citing_title":"Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2509.21623","citing_title":"OjaKV: Context-Aware Online Low-Rank KV Cache Compression","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20682","citing_title":"Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11501","citing_title":"Quantization Dominates Rank Reduction for KV-Cache Compression","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI","json":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI.json","graph_json":"https://pith.science/api/pith-number/ECWCXGQG5GZHZZCQMTHOU5L3ZI/graph.json","events_json":"https://pith.science/api/pith-number/ECWCXGQG5GZHZZCQMTHOU5L3ZI/events.json","paper":"https://pith.science/paper/ECWCXGQG"},"agent_actions":{"view_html":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI","download_json":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI.json","view_paper":"https://pith.science/paper/ECWCXGQG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.14731&json=true","fetch_graph":"https://pith.science/api/pith-number/ECWCXGQG5GZHZZCQMTHOU5L3ZI/graph.json","fetch_events":"https://pith.science/api/pith-number/ECWCXGQG5GZHZZCQMTHOU5L3ZI/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI/action/timestamp_anchor","attest_storage":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI/action/storage_attestation","attest_author":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI/action/author_attestation","sign_citation":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI/action/citation_signature","submit_replication":"https://pith.science/pith/ECWCXGQG5GZHZZCQMTHOU5L3ZI/action/replication_record"}},"created_at":"2026-07-05T11:03:52.852955+00:00","updated_at":"2026-07-05T11:03:52.852955+00:00"}