{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:C5IW6MMPQBL7MWR3NZ3ZN2ZWOD","short_pith_number":"pith:C5IW6MMP","schema_version":"1.0","canonical_sha256":"17516f318f8057f65a3b6e7796eb3670c8351e91b49082219d035916970a6079","source":{"kind":"arxiv","id":"2305.17118","version":2},"attestation_state":"computed","paper":{"title":"Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Aditya Desai, Anastasios Kyrillidis, Anshumali Shrivastava, Fangshuo Liao, Victor Xie, Weitao Wang, Zhaozhuo Xu, Zichang Liu","submitted_at":"2023-05-26T17:39:58Z","abstract_excerpt":"Large language models(LLMs) have sparked a new wave of exciting AI applications. Hosting these models at scale requires significant memory resources. One crucial memory bottleneck for the deployment stems from the context window. It is commonly recognized that model weights are memory hungry; however, the size of key-value embedding stored during the generation process (KV cache) can easily surpass the model size. The enormous size of the KV cache puts constraints on the inference batch size, which is crucial for high throughput inference workload. Inspired by an interesting observation of the"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2305.17118","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2023-05-26T17:39:58Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"4cd24a5d8c92ddaa5cb7319041560e3f9ad302eff8b68df19df534dc3e0c89c8","abstract_canon_sha256":"998500a66f70fa8be549f9fad9fef35cfdccf3b4547373bc68b6206eb1d19bf6"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:45:36.384763Z","signature_b64":"08lo/kw4dwRibwyIRNrZFSr0SKim4dV84mO0UgBVFREE5/Nw1k9peKPbsKw9ZxQsc+PjSbry+sVXgW+WZdnhAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"17516f318f8057f65a3b6e7796eb3670c8351e91b49082219d035916970a6079","last_reissued_at":"2026-07-05T06:45:36.384165Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:45:36.384165Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Aditya Desai, Anastasios Kyrillidis, Anshumali Shrivastava, Fangshuo Liao, Victor Xie, Weitao Wang, Zhaozhuo Xu, Zichang Liu","submitted_at":"2023-05-26T17:39:58Z","abstract_excerpt":"Large language models(LLMs) have sparked a new wave of exciting AI applications. Hosting these models at scale requires significant memory resources. One crucial memory bottleneck for the deployment stems from the context window. It is commonly recognized that model weights are memory hungry; however, the size of key-value embedding stored during the generation process (KV cache) can easily surpass the model size. The enormous size of the KV cache puts constraints on the inference batch size, which is crucial for high throughput inference workload. Inspired by an interesting observation of the"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2305.17118","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2305.17118/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2305.17118","created_at":"2026-07-05T06:45:36.384231+00:00"},{"alias_kind":"arxiv_version","alias_value":"2305.17118v2","created_at":"2026-07-05T06:45:36.384231+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2305.17118","created_at":"2026-07-05T06:45:36.384231+00:00"},{"alias_kind":"pith_short_12","alias_value":"C5IW6MMPQBL7","created_at":"2026-07-05T06:45:36.384231+00:00"},{"alias_kind":"pith_short_16","alias_value":"C5IW6MMPQBL7MWR3","created_at":"2026-07-05T06:45:36.384231+00:00"},{"alias_kind":"pith_short_8","alias_value":"C5IW6MMP","created_at":"2026-07-05T06:45:36.384231+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08032","citing_title":"What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents","ref_index":75,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24033","citing_title":"RoPE-Aware Bit Allocation for KV-Cache Quantization","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20295","citing_title":"Token-Operations-Oriented Inference Optimization Techniques for Large Models","ref_index":118,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01065","citing_title":"GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09735","citing_title":"KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22337","citing_title":"Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22337","citing_title":"Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18053","citing_title":"Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2310.01801","citing_title":"Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs","ref_index":88,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14037","citing_title":"Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility","ref_index":70,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09735","citing_title":"KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06763","citing_title":"Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17935","citing_title":"How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05219","citing_title":"Sparse Prefix Caching for Hybrid and Recurrent LLM Serving","ref_index":17,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD","json":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD.json","graph_json":"https://pith.science/api/pith-number/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/graph.json","events_json":"https://pith.science/api/pith-number/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/events.json","paper":"https://pith.science/paper/C5IW6MMP"},"agent_actions":{"view_html":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD","download_json":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD.json","view_paper":"https://pith.science/paper/C5IW6MMP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2305.17118&json=true","fetch_graph":"https://pith.science/api/pith-number/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/graph.json","fetch_events":"https://pith.science/api/pith-number/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/action/storage_attestation","attest_author":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/action/author_attestation","sign_citation":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/action/citation_signature","submit_replication":"https://pith.science/pith/C5IW6MMPQBL7MWR3NZ3ZN2ZWOD/action/replication_record"}},"created_at":"2026-07-05T06:45:36.384231+00:00","updated_at":"2026-07-05T06:45:36.384231+00:00"}