{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:JIC3GD76OA6LSDEOEI4ZBKYTTN","short_pith_number":"pith:JIC3GD76","schema_version":"1.0","canonical_sha256":"4a05b30ffe703cb90c8e223990ab139b61201b4b6b60333b4130bdb3ef5e7494","source":{"kind":"arxiv","id":"2311.18677","version":2},"attestation_state":"computed","paper":{"title":"Splitwise: Efficient generative LLM inference using phase splitting","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.DC"],"primary_cat":"cs.AR","authors_text":"Aashaka Shah, Chaojie Zhang, Esha Choukse, \\'I\\~nigo Goiri, Pratyush Patel, Ricardo Bianchini, Saeed Maleki","submitted_at":"2023-11-30T16:24:42Z","abstract_excerpt":"Recent innovations in generative large language models (LLMs) have made their applications and use-cases ubiquitous. This has led to large-scale deployments of these models, using complex, expensive, and power-hungry AI accelerators, most commonly GPUs. These developments make LLM inference efficiency an important challenge. Based on our extensive characterization, we find that there are two main phases during an LLM inference request: a compute-intensive prompt computation, and a memory-intensive token generation, each with distinct latency, throughput, memory, and power characteristics. Desp"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2311.18677","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.AR","submitted_at":"2023-11-30T16:24:42Z","cross_cats_sorted":["cs.DC"],"title_canon_sha256":"bb9c37d0b51537c3e5dc1c4fceb22328a90a838cbd2e635d5946284f5fdf5109","abstract_canon_sha256":"f0b8df3f60ccfeeb1cbd77623a16da4caad443ed5c43bc81fb4124028c64e9f2"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:20:51.858411Z","signature_b64":"z+DRrAA8lA7dAOoGPnJ2uCE3jnpv+B5W/bOx4ZGTSBGXJe633VWvCzWE5gRjRzglcSHae1a1S9jGiEkak2R8Dw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4a05b30ffe703cb90c8e223990ab139b61201b4b6b60333b4130bdb3ef5e7494","last_reissued_at":"2026-07-05T08:20:51.857860Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:20:51.857860Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Splitwise: Efficient generative LLM inference using phase splitting","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.DC"],"primary_cat":"cs.AR","authors_text":"Aashaka Shah, Chaojie Zhang, Esha Choukse, \\'I\\~nigo Goiri, Pratyush Patel, Ricardo Bianchini, Saeed Maleki","submitted_at":"2023-11-30T16:24:42Z","abstract_excerpt":"Recent innovations in generative large language models (LLMs) have made their applications and use-cases ubiquitous. This has led to large-scale deployments of these models, using complex, expensive, and power-hungry AI accelerators, most commonly GPUs. These developments make LLM inference efficiency an important challenge. Based on our extensive characterization, we find that there are two main phases during an LLM inference request: a compute-intensive prompt computation, and a memory-intensive token generation, each with distinct latency, throughput, memory, and power characteristics. Desp"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2311.18677","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2311.18677/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2311.18677","created_at":"2026-07-05T08:20:51.857921+00:00"},{"alias_kind":"arxiv_version","alias_value":"2311.18677v2","created_at":"2026-07-05T08:20:51.857921+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2311.18677","created_at":"2026-07-05T08:20:51.857921+00:00"},{"alias_kind":"pith_short_12","alias_value":"JIC3GD76OA6L","created_at":"2026-07-05T08:20:51.857921+00:00"},{"alias_kind":"pith_short_16","alias_value":"JIC3GD76OA6LSDEO","created_at":"2026-07-05T08:20:51.857921+00:00"},{"alias_kind":"pith_short_8","alias_value":"JIC3GD76","created_at":"2026-07-05T08:20:51.857921+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":29,"internal_anchor_count":3,"sample":[{"citing_arxiv_id":"2607.05876","citing_title":"Think Before You Grid-Search: Floor-First Triage for LLM Serving","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2607.08032","citing_title":"What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents","ref_index":93,"is_internal_anchor":true},{"citing_arxiv_id":"2607.05876","citing_title":"Think Before You Grid-Search: Floor-First Triage for LLM Serving","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18431","citing_title":"Beyond Prediction: Tail-Aware Scheduling for LLM Inference","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00466","citing_title":"ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17081","citing_title":"The Price of Anarchy in Disaggregated Inference","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01579","citing_title":"OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26666","citing_title":"PersistentKV: Page-Aware Decode Scheduling for Long-Context LLM Serving on Commodity GPUs","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00466","citing_title":"ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00735","citing_title":"ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00288","citing_title":"Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture","ref_index":121,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29708","citing_title":"Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29708","citing_title":"Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2505.09999","citing_title":"ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07985","citing_title":"Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.21847","citing_title":"CompPow: A Case for Component-level GPU Power Management","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20315","citing_title":"Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11333","citing_title":"MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2404.14294","citing_title":"A Survey on Efficient Inference for Large Language Models","ref_index":272,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11999","citing_title":"The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11232","citing_title":"Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11333","citing_title":"MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01708","citing_title":"SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving","ref_index":21,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN","json":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN.json","graph_json":"https://pith.science/api/pith-number/JIC3GD76OA6LSDEOEI4ZBKYTTN/graph.json","events_json":"https://pith.science/api/pith-number/JIC3GD76OA6LSDEOEI4ZBKYTTN/events.json","paper":"https://pith.science/paper/JIC3GD76"},"agent_actions":{"view_html":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN","download_json":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN.json","view_paper":"https://pith.science/paper/JIC3GD76","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2311.18677&json=true","fetch_graph":"https://pith.science/api/pith-number/JIC3GD76OA6LSDEOEI4ZBKYTTN/graph.json","fetch_events":"https://pith.science/api/pith-number/JIC3GD76OA6LSDEOEI4ZBKYTTN/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN/action/timestamp_anchor","attest_storage":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN/action/storage_attestation","attest_author":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN/action/author_attestation","sign_citation":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN/action/citation_signature","submit_replication":"https://pith.science/pith/JIC3GD76OA6LSDEOEI4ZBKYTTN/action/replication_record"}},"created_at":"2026-07-05T08:20:51.857921+00:00","updated_at":"2026-07-05T08:20:51.857921+00:00"}