{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:WN6AD3D37BBOB5UV4M5OAEPXL7","short_pith_number":"pith:WN6AD3D3","schema_version":"1.0","canonical_sha256":"b37c01ec7bf842e0f695e33ae011f75fdb17a13ea5df6e911998b06951939165","source":{"kind":"arxiv","id":"2503.06433","version":1},"attestation_state":"computed","paper":{"title":"Seesaw: High-throughput LLM Inference via Model Re-sharding","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.DC","authors_text":"Chenhao Jiang, Christina Giannoula, Gennady Pekhimenko, Kevin Song, Muralidhar Andoorveedu, Qidong Su, Wei Zhao, Xin Li, Zhanda Zhu","submitted_at":"2025-03-09T04:14:06Z","abstract_excerpt":"To improve the efficiency of distributed large language model (LLM) inference, various parallelization strategies, such as tensor and pipeline parallelism, have been proposed. However, the distinct computational characteristics inherent in the two stages of LLM inference-prefilling and decoding-render a single static parallelization strategy insufficient for the effective optimization of both stages. In this work, we present Seesaw, an LLM inference engine optimized for throughput-oriented tasks. The key idea behind Seesaw is dynamic model re-sharding, a technique that facilitates the dynamic "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2503.06433","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.DC","submitted_at":"2025-03-09T04:14:06Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"6f83b9db788e1fb994cedb5b584561deaa563d360c9dc606e7d5ef789ce837f0","abstract_canon_sha256":"b4275e4c565a16b981b38f75480af1399650aed199e668313f528a35ab991fd0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:27:33.624025Z","signature_b64":"KZ0orginYv9wrYyWWkir6NcoAotlgTnr2RCjcgHPsAH8OY3dG/UnqqYdojKqiHDDwHqJGB2SJf/WLmUCWoApDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b37c01ec7bf842e0f695e33ae011f75fdb17a13ea5df6e911998b06951939165","last_reissued_at":"2026-07-05T10:27:33.623463Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:27:33.623463Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Seesaw: High-throughput LLM Inference via Model Re-sharding","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.DC","authors_text":"Chenhao Jiang, Christina Giannoula, Gennady Pekhimenko, Kevin Song, Muralidhar Andoorveedu, Qidong Su, Wei Zhao, Xin Li, Zhanda Zhu","submitted_at":"2025-03-09T04:14:06Z","abstract_excerpt":"To improve the efficiency of distributed large language model (LLM) inference, various parallelization strategies, such as tensor and pipeline parallelism, have been proposed. However, the distinct computational characteristics inherent in the two stages of LLM inference-prefilling and decoding-render a single static parallelization strategy insufficient for the effective optimization of both stages. In this work, we present Seesaw, an LLM inference engine optimized for throughput-oriented tasks. The key idea behind Seesaw is dynamic model re-sharding, a technique that facilitates the dynamic "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2503.06433","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2503.06433/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2503.06433","created_at":"2026-07-05T10:27:33.623526+00:00"},{"alias_kind":"arxiv_version","alias_value":"2503.06433v1","created_at":"2026-07-05T10:27:33.623526+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2503.06433","created_at":"2026-07-05T10:27:33.623526+00:00"},{"alias_kind":"pith_short_12","alias_value":"WN6AD3D37BBO","created_at":"2026-07-05T10:27:33.623526+00:00"},{"alias_kind":"pith_short_16","alias_value":"WN6AD3D37BBOB5UV","created_at":"2026-07-05T10:27:33.623526+00:00"},{"alias_kind":"pith_short_8","alias_value":"WN6AD3D3","created_at":"2026-07-05T10:27:33.623526+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2511.09557","citing_title":"Understanding and Improving Communication Performance in Multi-node LLM Inference","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2509.19729","citing_title":"Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06046","citing_title":"Requests of a Feather Must Flock Together: Batch Size vs. Prefix Homogeneity in LLM Inference","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02189","citing_title":"PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers","ref_index":7,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7","json":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7.json","graph_json":"https://pith.science/api/pith-number/WN6AD3D37BBOB5UV4M5OAEPXL7/graph.json","events_json":"https://pith.science/api/pith-number/WN6AD3D37BBOB5UV4M5OAEPXL7/events.json","paper":"https://pith.science/paper/WN6AD3D3"},"agent_actions":{"view_html":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7","download_json":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7.json","view_paper":"https://pith.science/paper/WN6AD3D3","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2503.06433&json=true","fetch_graph":"https://pith.science/api/pith-number/WN6AD3D37BBOB5UV4M5OAEPXL7/graph.json","fetch_events":"https://pith.science/api/pith-number/WN6AD3D37BBOB5UV4M5OAEPXL7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7/action/storage_attestation","attest_author":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7/action/author_attestation","sign_citation":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7/action/citation_signature","submit_replication":"https://pith.science/pith/WN6AD3D37BBOB5UV4M5OAEPXL7/action/replication_record"}},"created_at":"2026-07-05T10:27:33.623526+00:00","updated_at":"2026-07-05T10:27:33.623526+00:00"}