{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:GHVQD25N6KYP2JAOBJ7FLXQKJW","short_pith_number":"pith:GHVQD25N","schema_version":"1.0","canonical_sha256":"31eb01ebadf2b0fd240e0a7e55de0a4dadd63ae3f1f7e77a1633b51015837eaa","source":{"kind":"arxiv","id":"2410.23317","version":1},"attestation_state":"computed","paper":{"title":"VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.DC","cs.PF"],"primary_cat":"cs.CV","authors_text":"Danylo Vashchilenko, Dezhan Tu, Panpan Xu, Yuzhe Lu","submitted_at":"2024-10-29T20:04:34Z","abstract_excerpt":"Vision-Language Models (VLMs) have demonstrated impressive performance across a versatile set of tasks. A key challenge in accelerating VLMs is storing and accessing the large Key-Value (KV) cache that encodes long visual contexts, such as images or videos. While existing KV cache compression methods are effective for Large Language Models (LLMs), directly migrating them to VLMs yields suboptimal accuracy and speedup. To bridge the gap, we propose VL-Cache, a novel KV cache compression recipe tailored for accelerating VLM inference. In this paper, we first investigate the unique sparsity patte"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.23317","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CV","submitted_at":"2024-10-29T20:04:34Z","cross_cats_sorted":["cs.AI","cs.CL","cs.DC","cs.PF"],"title_canon_sha256":"73513a1c441732fe545f096aec1a8a0a4932121e20f6d4f7e6433315fad4c88c","abstract_canon_sha256":"e57a5c4dd71f5597ee24f4196716eaa03224d5449b4da2fbc32e5d39e40045ed"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:28:58.258162Z","signature_b64":"rCaGM4U9V6USJjiWGYkQzOPf20djhamaDXVjaSO4cqd9H7k54nGiqpC9rrIrCeq0F7i+Wnv0uY22v7hSWMmTCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"31eb01ebadf2b0fd240e0a7e55de0a4dadd63ae3f1f7e77a1633b51015837eaa","last_reissued_at":"2026-07-05T09:28:58.257694Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:28:58.257694Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.DC","cs.PF"],"primary_cat":"cs.CV","authors_text":"Danylo Vashchilenko, Dezhan Tu, Panpan Xu, Yuzhe Lu","submitted_at":"2024-10-29T20:04:34Z","abstract_excerpt":"Vision-Language Models (VLMs) have demonstrated impressive performance across a versatile set of tasks. A key challenge in accelerating VLMs is storing and accessing the large Key-Value (KV) cache that encodes long visual contexts, such as images or videos. While existing KV cache compression methods are effective for Large Language Models (LLMs), directly migrating them to VLMs yields suboptimal accuracy and speedup. To bridge the gap, we propose VL-Cache, a novel KV cache compression recipe tailored for accelerating VLM inference. In this paper, we first investigate the unique sparsity patte"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.23317","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.23317/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.23317","created_at":"2026-07-05T09:28:58.257753+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.23317v1","created_at":"2026-07-05T09:28:58.257753+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.23317","created_at":"2026-07-05T09:28:58.257753+00:00"},{"alias_kind":"pith_short_12","alias_value":"GHVQD25N6KYP","created_at":"2026-07-05T09:28:58.257753+00:00"},{"alias_kind":"pith_short_16","alias_value":"GHVQD25N6KYP2JAO","created_at":"2026-07-05T09:28:58.257753+00:00"},{"alias_kind":"pith_short_8","alias_value":"GHVQD25N","created_at":"2026-07-05T09:28:58.257753+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08032","citing_title":"What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents","ref_index":115,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02775","citing_title":"AURA: Action-Gated Memory for Robot Policies at Constant VRAM","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2605.28115","citing_title":"CIVIC: End-to-End Sequence Compactness for Efficient Vision-Language Models","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29535","citing_title":"AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16439","citing_title":"KVCapsule: Efficient Sequential KV Cache Compression for Vision-Language Models with Asymmetric Redundancy","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19218","citing_title":"Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09649","citing_title":"Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2604.25642","citing_title":"Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2604.24391","citing_title":"FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching","ref_index":26,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW","json":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW.json","graph_json":"https://pith.science/api/pith-number/GHVQD25N6KYP2JAOBJ7FLXQKJW/graph.json","events_json":"https://pith.science/api/pith-number/GHVQD25N6KYP2JAOBJ7FLXQKJW/events.json","paper":"https://pith.science/paper/GHVQD25N"},"agent_actions":{"view_html":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW","download_json":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW.json","view_paper":"https://pith.science/paper/GHVQD25N","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.23317&json=true","fetch_graph":"https://pith.science/api/pith-number/GHVQD25N6KYP2JAOBJ7FLXQKJW/graph.json","fetch_events":"https://pith.science/api/pith-number/GHVQD25N6KYP2JAOBJ7FLXQKJW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW/action/storage_attestation","attest_author":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW/action/author_attestation","sign_citation":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW/action/citation_signature","submit_replication":"https://pith.science/pith/GHVQD25N6KYP2JAOBJ7FLXQKJW/action/replication_record"}},"created_at":"2026-07-05T09:28:58.257753+00:00","updated_at":"2026-07-05T09:28:58.257753+00:00"}