REVIEW 3 major objections 4 minor 70 references
Cloud abstractions for AI workloads
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read HarmonAIze argues that the root cause of AI workload inefficiency in multi-tenant clouds is the absence of abstractions for tenant–provider cooperation, and proposes a two-level abstraction design—micro-level, where providers execute…
desk verdict A clear, well-scoped vision paper for cooperative tenant-provider abstractions, but the load-bearing assumption that useful and non-sensitive infrastructure information can coexist is acknowledged, not solved. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a pair of abstraction levels. Micro-level abstractions keep API compatibility with existing AI libraries (such as collective communication interfaces) while letting tenants specify ranges and constraints, and they delegate data-path decisions to the provider, which can execute them in the hypervisor or on isolated infrastructure processing units. Macro-level abstractions give tenants a pub/sub interface to subscribe to infrastructure-level events (resource availability, failures) in aggregate or coarse-grained form, and allow multi-round negotiation between tenants and providers for strategic adaptations. The argument rests on the claim that AI workloads' predictability and tight synchronization make these cooperative control loops both feasible and high-value.
What would settle it
Run a controlled multi-tenant experiment comparing AI training jobs under HarmonAIze-style cooperative control (provider-side collectives and checkpointing, tenant-side scaling on provider events) against the same jobs on standard IaaS with static NCCL settings; if the cooperative cohort does not show measurably better job completion time, energy use, or failure recovery on realistic workloads with noisy neighbours, the core efficiency claim fails. A simpler disproof would be evidence that cloud providers cannot expose even aggregated infrastructure events at a granularity tenants can act on without leaking competitive information.
Extended reading notes
Core claim
HarmonAIze's central claim is that today's cloud fails AI workloads because tenants lack infrastructure insight and providers lack workload insight, and that the fix is a new layer of abstractions that splits control by timescale: providers execute micro-level control loops (algorithm selection, batch-size adjustment, checkpoint placement) using tenant-declared constraints, while tenants execute macro-level control loops (scaling, parallelization strategy, hyperparameters) using provider-supplied aggregated events. The paper argues this division matches who has the timely information and who has the strategic view, and surveys seven optimization opportunities (O1–O7) where cross-layer cooperation would amplify known gains. HarmonAIze deliberately focuses on AI workloads first because their iterative, accelerator-driven, and tightly synchronized nature makes them more predictable than general distributed systems, making the optimization problem tractable.
Load-bearing premise
Tenants and providers must both be willing to reveal enough—tenants their workload requirements, providers their infrastructure state—without leaking commercially sensitive details, and must trust each other enough to act on that shared information; if this mutual disclosure cannot happen, the cooperative control loops have nothing to feed on.
Editorial extensions
If this is right
- Providers take over fine-grained, data-path decisions such as collective algorithm selection, batch-size adjustment, and checkpoint placement, using tenant-specified ranges as guardrails.
- Tenants receive standardized, aggregate infrastructure events and can negotiate with providers, so scaling, parallelism, and hyperparameter choices can react to failures and congestion mid-run.
- Known optimization gains—up to 2.4× collective communication speedup, 22.61× mean time between failures, and 53% energy reduction—would become attainable in public multi-tenant clouds rather than only in tightly coupled single-operator systems.
- Adoption is incremental: tenants opt in by adding buy-in interfaces, while providers continue serving non-adopting tenants as today.
- The first proof-of-concept can be assembled from existing open tools: scheduler simulators, distributed-job emulators, collectives-as-a-service, and runtime adaptation systems.
Reading between the lines
- If the premise holds, the same micro/macro split is a candidate for other synchronous, iterative workloads such as HPC simulations, even though the paper deliberately restricts itself to AI.
- The buy-in deployment model implies a competitive dynamic the paper leaves implicit: early-adopting tenants should see measurable performance wins, which would push providers to standardize these interfaces to avoid churn.
- A testable prediction following from the micro-level design is that provider-side execution of collectives on SmartNICs or IPUs outperforms tenant-side algorithm selection, because only the provider observes real-time topology and load.
- The granularity of aggregate infrastructure events is the crux the paper flags but does not resolve: too coarse and tenants cannot act, too fine and providers leak sensitive information; finding the workable middle is an empirical question.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper proposes HarmonAIze, a set of micro- and macro-level cloud abstractions intended to enable cooperative optimization between AI workload tenants and cloud providers. The central idea is that tenant control loops should be informed by infrastructure-level insights while provider control loops are guided by workload-specific requirements, with micro-level abstractions supporting API-compatible, provider-executed fine-grained adaptations and macro-level abstractions providing tenant-executed strategic adaptations via interfaces such as pub/sub events. The paper identifies seven concrete optimization opportunities (O1--O7, in Section 4) and tabulates benefits from prior systems as representative of expected gains. It does not present an implementation or measurements; Section 5 outlines an incremental roadmap and briefly discusses adoption, timing, and open challenges including privacy.
Significance. If the HarmonAIze vision holds, it offers a useful organizing framework for research on tenant--provider cooperation in AI clouds, with a clear division of labor between micro-level provider control and macro-level tenant control. The paper's strengths are its concrete enumeration of seven optimization opportunities with up-to-date references, the explicit identification of building blocks for a prototype (e.g., Blox, MCCS, KungFu), and its honest treatment of adoption and standardization challenges. It is a genuine position paper: it frames testable hypotheses rather than claiming demonstrated results. The main significance is as a catalyst for discussion; the claimed benefits, however, are asserted rather than shown, and the privacy--utility tension is left unresolved.
major comments (3)
- [§3.1–3.2, §5] The central claim of Section 1 that HarmonAIze's two innovations are load-bearing requires an information channel that is simultaneously non-sensitive and actionable, but the paper never characterizes the granularity at which infrastructure insight remains actionable. Section 3.1 suggests aggregate or coarse-grained updates to safeguard sensitive details, and Section 3.2 reiterates that sensitive infrastructure details must not be exposed; however, for O2 (collective communication) and O3 (cluster scheduling), decisions such as algorithm selection and scaling require topology, link utilization, and other tenants' activity. Aggregating those signals to protect commercial confidentiality can remove precisely the information the tenant control loop needs. The benefits in Table 1 are from systems such as Pollux [43] and MCCS [56] that assume full visibility of both workload and infrastructure, so those numbers cannot be transferred without further argument to a regime with deliberately degraded information. This is a load-bearing gap; Section 5 lists privacy as open, but the authors should either provide a concrete privacy--utility analysis or moderate the claimed benefits accordingly.
- [§4, Table 1] Table 1 is introduced as 'select benefits that are representative of cooperative optimizations that we expect from HarmonAIze' (Section 5), but the cited systems are not instances of the HarmonAIze abstractions and in some cases require precisely the full visibility that HarmonAIze's coarse-grained interfaces deliberately withhold. For example, Pollux [43] co-adapts scheduling and batch size with direct control of cluster resources, and MCCS [56] offloads collectives as a service with infrastructure-level knowledge. As presented, the table conflates the headroom that cross-layer optimization could offer with the gains that HarmonAIze specifically would deliver. The paper should either relabel the table as evidence of headroom, or add a caveat that these numbers assume information quality that HarmonAIze's privacy constraints may not preserve.
- [§3.1, macro-level abstractions] The proposed pub/sub interface for macro-level updates mentions several possible mechanisms for resolving contention—priority-based allocation, reservation windows, or price-based auctioning—as though they were interchangeable (Section 3.1). These mechanisms impose very different incentive structures and information requirements: auctioning can elicit strategic behavior and reveal private valuations, reservation windows can reduce utilization, and priority rules require a notion of fairness across tenants. The paper's later claim of a 'natural alignment of incentives' (Section 5) therefore needs substantially more support; without it, the cooperative control-loop model is underspecified. At minimum, the authors should indicate which mechanism they envisage for the initial prototype described in Section 5.
minor comments (4)
- [O6] The text '22.61×%' in Section 4 (O6) contains a stray percentage sign; it should read '22.61×' or 'by a factor of 22.61'.
- [Figure 1] In Figure 1, 'MACRO-level' appears in all caps while 'micro-level' is lowercase; use consistent casing for the two levels.
- [Throughout] The abbreviation 'c.f.' (e.g., in Section 3) should be 'cf.' (confer), and 'check pointing' (Section 4, O6) should be 'checkpointing' for consistency.
- [Section 3.2] The phrase 'as much as a $30K [42]' in Section 4 (O6) is awkward; consider 'up to $30K [42]' or 'as much as $30K [42]'.
Circularity Check
No significant circularity: the paper is a vision/position paper whose central claims are not derived from fitted parameters or self-cited uniqueness results.
full rationale
HarmonAIze is a position paper proposing new cloud abstractions; it contains no equations, no fitted parameters, and no derivation chain that reduces a prediction to its inputs. The claimed benefits in Table 1 are explicitly attributed to external literature results (e.g., KungFu, Parcae, Pollux, MCCS), and the paper frames HarmonAIze as an enabler that may amplify those benefits, not as a system that reproduces them. Some cited works have overlapping authors with the present paper (e.g., KungFu, Workload Intelligence, and a few others), but these citations are used as motivation, prior context, or roadmap components rather than as load-bearing justification for the core claim. The paper explicitly identifies open challenges such as privacy and adoption, and it does not assert that its abstractions have been validated. Because no step in the paper's argument is forced by definition, by construction, or by a self-citation chain, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption AI workloads are iterative, accelerator-driven, and predictable at micro and macro timescales, making cooperative optimization tractable.
- domain assumption Tenants and providers have aligned incentives to adopt cooperative abstractions and share information.
- domain assumption Infrastructure information can be shared with tenants in a form that is useful yet does not leak sensitive commercial details.
- domain assumption API-compatible abstractions can be introduced with minimal code changes to existing AI frameworks and libraries.
invented entities (1)
-
HarmonAIze abstraction layer (micro and macro level interfaces)
Cite this review
Pith. "Pith review of Cloud abstractions for AI workloads." pith.science (2026). https://pith.science/paper/652DNDVS
@misc{pith2026250109562,
author = {Pith},
title = {Pith review of: Cloud abstractions for AI workloads},
year = {2026},
howpublished = {\url{https://pith.science/paper/652DNDVS}},
note = {Machine review of arXiv:2501.09562}
}
read the original abstract
AI workloads, often hosted in multi-tenant cloud environments, require vast computational resources but suffer inefficiencies due to limited tenant-provider coordination. Tenants lack infrastructure insights, while providers lack workload details to optimize tasks like partitioning, scheduling, and fault tolerance. We propose HarmonAIze to redefine cloud abstractions, enabling cooperative optimization for improved performance, efficiency, resiliency, and sustainability. We outline key opportunities and challenges this vision faces.
Figures
Reference graph
Works this paper leans on
-
[43]
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. InOSDI
work page 2021
-
[56]
Minjie Wang, Chien-chin Huang, and Jinyang Li. 2019. Supporting Very Large Models using Automatic Dataflow Graph Partitioning. In EuroSys
work page 2019
-
[1]
Ahmed M. Abdelmoniem and Marco Canini. 2021. DC2: Delay-aware Compression Control for Distributed Machine Learning. InINFOCOM
work page 2021
-
[2]
Saurabh Agarwal, Amar Phanishayee, and Shivaram Venkataraman
-
[3]
Amazon Web Services. 2024. AWS SageMaker. https://aws.amazon. com/sagemaker
work page 2024
-
[4]
Mohammadreza Bayatpour, Nick Sarkauskas, Hari Subramoni, Ja- hanzeb Maqbool Hashmi, and Dhabaleswar K. Panda. 2021. BluesMPI: Efficient MPI Non-blocking Alltoall Offloading Designs on Modern Cloud abstractions for AI workloads APSys ’25, October 12–13, 2025, Seoul, Republic of Korea BlueField Smart NICs. InHigh Performance Computing
work page 2021
-
[5]
Muhammad Bilal, Marco Canini, Rodrigo Fonseca, and Rodrigo Ro- drigues. 2023. With Great Freedom Comes Great Opportunity: Re- thinking Resource Allocation for Serverless Functions. InEuroSys
work page 2023
-
[6]
Weilin Cai, Le Qin, and Jiayi Huang. 2025. MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training. InASPLOS
work page 2025
Show all 70 references
-
[7]
Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu. 2022. TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism.IEEE Transactions on Parallel and Distributed Systems33, 8 (2022)
2022
-
[8]
Jiamin Cao, Yu Guan, Kun Qian, Jiaqi Gao, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, and Ennan Zhai. 2024. Crux: GPU- Efficient Communication Scheduling for Deep Learning Training. In SIGCOMM
2024
-
[9]
Shubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, and Srinidhi Viswanatha. 2020. Balancing Efficiency and Fairness in Heterogeneous GPU Clusters for Deep Learning. In EuroSys
2020
-
[10]
Jingrong Chen, Hong Zhang, Wei Zhang, Liang Luo, Jeffrey Chase, Ion Stoica, and Danyang Zhuo. 2022. NetHint: White-Box Networking for Multi-Tenant Data Centers. InNSDI
2022
-
[11]
Xiaoqi Chen, Shay Vargaftik, and Ran Ben Basat. 2024. When ML Train- ing Cuts Through Congestion: Just-in-Time Gradient Compression via Packet Trimming. InHotNets
2024
-
[12]
Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal, and Mosharaf Chowdhury. 2024. Reducing Energy Bloat in Large Model Training. InSOSP
2024
-
[13]
Bryce Cronkite-Ratcliff, Aran Bergman, Shay Vargaftik, Madhusud- han Ravi, Nick McKeown, Ittai Abraham, and Isaac Keslassy. 2016. Virtualized Congestion Control. InSIGCOMM
2016
-
[14]
Michael Dalton, David Schultz, Jacob Adriaens, Ahsan Arefin, Anshu- man Gupta, Brian Fahs, Dima Rubinstein, Enrique Cauich Zermeno, Erik Rubow, James Alexander Docauer, Jesse Alpert, Jing Ai, Jon Olson, Kevin DeCabooter, Marc de Kruijf, Nan Hua, Nathan Lewis, Nikhil Kasinadhun...
2018
-
[15]
Jiangfei Duan, Ziang Song, Xupeng Miao, Xiaoli Xi, Dahua Lin, Harry Xu, Minjia Zhang, and Zhihao Jia. 2024. Parcae: Proactive, Liveput- Optimized DNN Training on Preemptible Instances. InNSDI
2024
-
[16]
Benjamin Frank, Ingmar Poese, Yin Lin, Georgios Smaragdakis, Anja Feldmann, Bruce Maggs, Jannis Rake, Steve Uhlig, and Rick Weber
-
[17]
Google Cloud. 2024. Google Vertex AI. https://cloud.google.com/ vertex-ai
2024
-
[18]
Google Cloud. 2024. Optimize Training Performance with Reduction Server in Vertex AI. https://cloud.google.com/blog/topics/developers- practitioners/optimize-training-performance-reduction-server- vertex-ai
2024
-
[19]
Richard Graham, George Bosilca, Yong Qin, Bradley Settlemyer, Gi- lad Shainer, Craig Stunkel, Geoffroy Vallee, Brody Williams, Gerardo Cisneros-Stoianowski, Sebastian Ohlmann, and Markus Rampp. 2024. Optimizing Application Performance with BlueField: Accelerating Large-Message...
2024
-
[20]
Tongzhou Gu, Jiawei Fei, and Marco Canini. 2024. OmNICCL: Zero- cost Sparse AllReduce with Direct Cache Access and SmartNICs. In NAIC
2024
-
[21]
Tanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vi- jeev, Bhargav Gulavani, Nipun Kwatra, Ramachandran Ramjee, and Muthian Sivathanu. 2024. Just-In-Time Checkpointing: Low Cost Error Recovery from Deep Learning Training Failures. InEuroSys
2024
-
[22]
Keqiang He, Eric Rozner, Kanak Agarwal, Yu (Jason) Gu, Wes Felter, John Carter, and Aditya Akella. 2016. AC/DC TCP: Virtual Congestion Control Enforcement for Datacenter Networks. InSIGCOMM
2016
-
[23]
Tao He, Xue Li, Zhibin Wang, Kun Qian, Jingbo Xu, Wenyuan Yu, and Jingren Zhou. 2024. Unicron: Economizing Self-Healing LLM Training at Scale. arXiv:2401.00134 [cs.DC]
2024 arXiv
-
[24]
Joseph, Randy Katz, Scott Shenker, and Ion Stoica
Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, An- thony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. InNSDI
2011
-
[25]
Samuel Hsia, Alicia Golden, Bilge Acun, Newsha Ardalani, Zachary DeVito, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. 2024. MAD- Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems. InISCA
2024
-
[26]
Lexiang Huang, Anjaly Parayil, Jue Zhang, Xiaoting Qin, Chetan Bansal, Jovan Stojkovic, Pantea Zardoshti, Pulkit Misra, Eli Cortez, Raphael Ghelman, Íñigo Goiri, Saravan Rajmohan, Jim Kleewein, Ro- drigo Fonseca, Timothy Zhu, and Ricardo Bianchini. 2024. Work- load Intelligenc...
2024 arXiv
-
[27]
Intel. 2023. Intel Infrastructure Processing Unit (Intel IPU) ASIC E2000. https://www.intel.com/content/www/us/en/products/details/ network-io/ipu/e2000-asic.html
2023
-
[28]
Xianyan Jia, Le Jiang, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang, Xinyuan Li, Langshi Chen, Yong Li, Zhen Zheng, Xiaoyong Liu, and Wei Lin. 2022. Whale: Efficient Giant Model Training over Heteroge- neous GPUs. InUSENIX ATC
2022
-
[29]
Junchen Jiang, Xi Liu, Vyas Sekar, Ion Stoica, and Hui Zhang. 2014. EONA: Experience-Oriented Network Architecture. InHotNets
2014
-
[30]
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi...
2024
-
[31]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InSOSP
2023
-
[32]
Dacheng Li, Hongyi Wang, Eric Xing, and Hao Zhang. 2022. AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness. InNeurIPS
2022
-
[33]
Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase, and Minjia Zhang. 2024. Universal Check- pointing: Efficient and Flexible Checkpointing for Large Scale Dis- tributed Training. arXiv:2406.18820 [cs.DC]
2024 arXiv
-
[34]
Zhiqi Lin, Youshan Miao, Quanlu Zhang, Fan Yang, Yi Zhu, Cheng Li, Saeed Maleki, Xu Cao, Ning Shang, Yilei Yang, Weijiang Xu, Mao Yang, Lintao Zhang, and Lidong Zhou. 2024. nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training. InOSDI
2024
-
[35]
Banruo Liu, Mubarak Adetunji Ojewale, Yuhan Ding, and Marco Canini
-
[36]
Xuting Liu, Behnaz Arzani, Siva Kesava Reddy Kakarla, Liangyu Zhao, Vincent Liu, Miguel Castro, Srikanth Kandula, and Luke Marshall
-
[37]
Luo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis, Andrei- Octavian Brabete, and Peter Pietzuch. 2020. KungFu: Making Training APSys ’25, October 12–13, 2025, Seoul, Republic of Korea Canini et al. in Distributed Machine Learning Adaptive. InOSDI
2020
-
[38]
Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation. InAPSys
-
[39]
Mustafa Rafique, Franck Cap- pello, and Bogdan Nicolae
Avinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cap- pello, and Bogdan Nicolae. 2024. DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models. InHPDC
2024
-
[40]
InSIGCOMM
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem. InSIGCOMM
-
[41]
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravis- hankar Iyer. 2024. Queue Management for SLO-Oriented Large Lan- guage Model Serving. InSoCC
2024
-
[42]
Ilia Markov, Kaveh Alim, Elias Frantar, and Dan Alistarh. 2024. L- GreCo: Layerwise-adaptive Gradient Compression For Efficient Data- parallel Deep Learning. InMLSys
2024
-
[44]
NVIDIA. 2024. NVIDIA Collective Communication Library (NCCL). https://developer.nvidia.com/nccl
2024
-
[45]
Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Cre- spo, and Dan Dennison
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Cre- spo, and Dan Dennison. 2015. Hidden Technical Debt in Machine Learning Systems. InNeurIPS
2015
-
[46]
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, Chao Wang, Peng Wang, Pengcheng Zhang, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, and Dennis Cai. 2024. Alibaba HPN: A Data Center Network for Large Languag...
2024
-
[47]
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2025. DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency. InHPCA
2025
-
[48]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He
-
[49]
Yi Tay. 2024. Training Great LLMs Entirely from Ground Zero in the Wilderness. https://www.yitay.net/blog/training-great-llms-entirely- from-ground-zero-in-the-wilderness
2024
-
[50]
Guanhua Wang, Olatunji Ruwase, Bing Xie, and Yuxiong He. 2024. FastPersist: Accelerating Model Checkpointing in Deep Learning. arXiv:2406.13768 [cs.DC]
2024 arXiv
-
[51]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL]
2020 arXiv
-
[52]
Zhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang, Xinwei Fu, T. S. Eu- gene Ng, and Yida Wang. 2023. GEMINI: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints. InSOSP
2023
-
[53]
Tarnawski, Deepak Narayanan, and Amar Phanishayee
Jakub M. Tarnawski, Deepak Narayanan, and Amar Phanishayee. 2021. Piper: Multidimensional Planner for DNN Parallelization. InNeurIPS
2021
-
[54]
Carole-Jean Wu, Bilge Acun, Ramya Raghavendra, and Kim Hazelwood
-
[55]
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, New- sha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien-Hsin...
2022
-
[57]
Richard Yang, Arvind Krishnamurthy, Yanbin Grace Liu, and Abraham Silberschatz
Haiyong Xie, Y. Richard Yang, Arvind Krishnamurthy, Yanbin Grace Liu, and Abraham Silberschatz. 2008. P4P: Provider Portal for Applica- tions. InSIGCOMM
2008
-
[58]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Syl- vain Gugger, Ma...
-
[59]
In EMNLP: System Demonstrations
Transformers: State-of-the-Art Natural Language Processing. In EMNLP: System Demonstrations
-
[60]
Jie You, Jae-Won Chung, and Mosharaf Chowdhury. 2023. Zeus: Under- standing and Optimizing GPU Energy Consumption of DNN Training. InNSDI
2023
-
[61]
Beyond Efficiency: Scaling AI Sustainably.IEEE Micro44, 5 (2024)
2024
-
[62]
Xing, Joseph E
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Au- tomating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. InOSDI
2022
-
[63]
Yongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang, Ying Zhang, Matthew Lentz, and Danyang Zhuo. 2024. MCCS: A Service-based Approach to Collective Communication for Multi-Tenant Cloud. In SIGCOMM
2024
-
[65]
Jihao Xin, Ivan Ilin, Shunkang Zhang, Marco Canini, and Peter Richtárik. 2023. Kimad: Adaptive Gradient Compression with Band- width Awareness. InDistributedML
2023
-
[66]
Yifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu, Guoshuai Zhao, Shuguang Liu, Dong Zhong, Boris Pinzur, Jie Zhang, Yang Wang, Jithin Jose, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng, Yongqiang Xiong, and Lidong Zhou
-
[67]
InUSENIX ATC
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation. InUSENIX ATC
-
[69]
Yuxuan Zhao, Weikang Weng, Rob van Nieuwpoort, and Alexandru Uta. 2024. In Serverless, OS Scheduler Choice Costs Money: A Hybrid Scheduling Approach for Cheaper FaaS. InMiddleware
2024
-
[2013]
Pushing CDN-ISP Collaboration to the Limit.SIGCOMM Comput. Commun. Rev.43, 3 (2013)
2013
-
[2020]
ZeRO: Memory Optimizations Toward Training Trillion Param- eter Models. InSC
-
[2024]
InEu- roSys
Blox: A Modular Toolkit for Deep Learning Schedulers. InEu- roSys
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.