{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:V5HYVXB7XWFE22SVXQ4RH6FKXA","short_pith_number":"pith:V5HYVXB7","schema_version":"1.0","canonical_sha256":"af4f8adc3fbd8a4d6a55bc3913f8aab80ed5607d544caf9c04b1a113ce3cbaac","source":{"kind":"arxiv","id":"2401.17644","version":5},"attestation_state":"computed","paper":{"title":"BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.PF"],"primary_cat":"cs.DC","authors_text":"Amelie Chi Zhou, Qiang Wang, Rui Guo, Xiaowen Chu, Xin He, Xin Wang, Xueze Kang, Yang Zheng, Yeju Zhou, Yuchu Fang, Yuhan Chen, Yuxin Wang, Zeyu Li, Zhenheng Tang","submitted_at":"2024-01-31T07:52:48Z","abstract_excerpt":"Serving systems for Large Language Models (LLMs) are often optimized to improve quality of service (QoS) and throughput. However, due to the lack of open-source LLM serving workloads, these systems are frequently evaluated under unrealistic workload assumptions. Consequently, performance may degrade when systems are deployed in real-world scenarios. This work presents BurstGPT, an LLM serving workload with 10.31 million traces from regional Azure OpenAI GPT services over 213 days. BurstGPT captures LLM serving characteristics from user, model and system perspectives: (1) User request concurren"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2401.17644","kind":"arxiv","version":5},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.DC","submitted_at":"2024-01-31T07:52:48Z","cross_cats_sorted":["cs.PF"],"title_canon_sha256":"afca5d25840deee4d6e639f5e8f19623faf94994010cd2d25d463a057e870851","abstract_canon_sha256":"5267b27f98f9cc2a03da03fe053840e55d251bb4ca0f38ff1216ab85aab9efb9"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:09:41.836436Z","signature_b64":"x/4JXAdx1xSwIw2khZhTklHxpkIqpZ/geSAxuY7iY0Ij0eLja95TPynInx5vrE8ro5BwSpvo0zfuw1bkcURfDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"af4f8adc3fbd8a4d6a55bc3913f8aab80ed5607d544caf9c04b1a113ce3cbaac","last_reissued_at":"2026-07-05T11:09:41.835909Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:09:41.835909Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.PF"],"primary_cat":"cs.DC","authors_text":"Amelie Chi Zhou, Qiang Wang, Rui Guo, Xiaowen Chu, Xin He, Xin Wang, Xueze Kang, Yang Zheng, Yeju Zhou, Yuchu Fang, Yuhan Chen, Yuxin Wang, Zeyu Li, Zhenheng Tang","submitted_at":"2024-01-31T07:52:48Z","abstract_excerpt":"Serving systems for Large Language Models (LLMs) are often optimized to improve quality of service (QoS) and throughput. However, due to the lack of open-source LLM serving workloads, these systems are frequently evaluated under unrealistic workload assumptions. Consequently, performance may degrade when systems are deployed in real-world scenarios. This work presents BurstGPT, an LLM serving workload with 10.31 million traces from regional Azure OpenAI GPT services over 213 days. BurstGPT captures LLM serving characteristics from user, model and system perspectives: (1) User request concurren"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2401.17644","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2401.17644/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2401.17644","created_at":"2026-07-05T11:09:41.835966+00:00"},{"alias_kind":"arxiv_version","alias_value":"2401.17644v5","created_at":"2026-07-05T11:09:41.835966+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2401.17644","created_at":"2026-07-05T11:09:41.835966+00:00"},{"alias_kind":"pith_short_12","alias_value":"V5HYVXB7XWFE","created_at":"2026-07-05T11:09:41.835966+00:00"},{"alias_kind":"pith_short_16","alias_value":"V5HYVXB7XWFE22SV","created_at":"2026-07-05T11:09:41.835966+00:00"},{"alias_kind":"pith_short_8","alias_value":"V5HYVXB7","created_at":"2026-07-05T11:09:41.835966+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.05272","citing_title":"Adaptive Inference Batching using Policy Gradients","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11690","citing_title":"Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2505.09999","citing_title":"ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2512.09472","citing_title":"WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06534","citing_title":"ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL","ref_index":72,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19945","citing_title":"GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15788","citing_title":"ADAPT: A Self-Calibrating Proactive Autoscaler for Container Orchestration","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2603.28680","citing_title":"A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2601.21351","citing_title":"Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2603.28680","citing_title":"A Techno-Economic Framework for Cost Modeling and Revenue Opportunities in Open and Programmable AI-RAN","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02821","citing_title":"When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06534","citing_title":"ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL","ref_index":73,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02821","citing_title":"When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08075","citing_title":"Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving","ref_index":26,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA","json":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA.json","graph_json":"https://pith.science/api/pith-number/V5HYVXB7XWFE22SVXQ4RH6FKXA/graph.json","events_json":"https://pith.science/api/pith-number/V5HYVXB7XWFE22SVXQ4RH6FKXA/events.json","paper":"https://pith.science/paper/V5HYVXB7"},"agent_actions":{"view_html":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA","download_json":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA.json","view_paper":"https://pith.science/paper/V5HYVXB7","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2401.17644&json=true","fetch_graph":"https://pith.science/api/pith-number/V5HYVXB7XWFE22SVXQ4RH6FKXA/graph.json","fetch_events":"https://pith.science/api/pith-number/V5HYVXB7XWFE22SVXQ4RH6FKXA/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA/action/timestamp_anchor","attest_storage":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA/action/storage_attestation","attest_author":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA/action/author_attestation","sign_citation":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA/action/citation_signature","submit_replication":"https://pith.science/pith/V5HYVXB7XWFE22SVXQ4RH6FKXA/action/replication_record"}},"created_at":"2026-07-05T11:09:41.835966+00:00","updated_at":"2026-07-05T11:09:41.835966+00:00"}