REVIEW 4 major objections 4 minor 97 references
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A cluster scheduler that learns from past workload traces cuts carbon emissions by about 57% without knowing job lengths ahead of time.
desk verdict A solid systems paper with a genuinely new provisioning-plus-scheduling mechanism, but the 'oracle' is a flawed heuristic—the near-optimal framing needs correction before the headline claims can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the marginal-throughput-per-carbon greedy oracle: for every job, time slot, and allowed scale, it computes $p_j(k)/CI_t$, the normalized throughput gained per unit of carbon, and allocates resources in descending order of that ratio subject to cluster capacity and per-queue delay. This gives a provably optimal schedule for homogeneous clusters with monotonically decreasing scaling profiles (Theorem 4.1), and it is what CarbonFlex replays over historical traces. The runtime counterpart is a case-based reasoning layer that maps a state vector (carbon intensity, its gradient and day-ahead rank, queue lengths, mean elasticity) to the oracle's capacity decision and scheduling threshold, using k-nearest-neighbor matching in a knowledge base.
What would settle it
Run CarbonFlex on a week whose job arrivals are drawn from a distribution outside the learning window (for example, doubling the arrival rate or inserting a new job class absent from the historical trace) and compare its emissions to the offline oracle for that same week; if the carbon savings fall to near the carbon-agnostic baseline or delay violations exceed the queue slack, the stable-distribution hypothesis fails.
Extended reading notes
Core claim
The central claim is that continuous learning over historical cluster data can drive near-optimal carbon-aware provisioning and scheduling of many parallel elastic jobs in a shared cluster, without a priori knowledge of job length or future arrivals. CarbonFlex simulates an offline oracle (Algorithm 1), a greedy scheduler that sorts possible job-time-scale assignments by marginal throughput per unit of carbon $p_j(k)/CI_t$, over a past window, and records for each observed system state the cluster capacity $m_t$ and a marginal-throughput threshold $\rho$. At runtime, it matches the current state to the nearest stored states and provisions the cluster accordingly, then schedules all jobs whose marginal throughput exceeds $\rho$. On AWS CPU and GPU clusters driven by Azure, Alibaba, and SURF traces across ten regions, the paper reports carbon savings of 51.4% and 57.5% over the carbon-agnostic baseline and within 6.6% and 2.1% of the offline oracle respectively.
Load-bearing premise
The method works only if the recent past resembles the near future: if the mix of jobs, arrival rates, or carbon conditions changes sharply between the learning window and the evaluation period, the learned decisions no longer match what the oracle would do.
Editorial extensions
If this is right
- Carbon savings of roughly half can be achieved without job-length knowledge, removing a key barrier to deploying carbon-aware scheduling in real batch clusters.
- Separating provisioning from scheduling means CarbonFlex's provisioning policy can be paired with other schedulers, and its scheduler with other provisioning curves such as Google's VCC.
- Because decisions are relearned continuously from a rolling window, the approach tracks gradual seasonal and workload changes rather than needing manual re-tuning.
- The benefit grows with workload elasticity and carbon-intensity variability: the more jobs can be scaled and the more volatile the grid, the larger the saving.
- Elastic scaling also lets schedulers push high-power jobs into low-carbon windows, which yields extra savings on GPU clusters where power profiles differ across jobs.
Reading between the lines
- The same imitate-the-oracle-on-history recipe could be applied to other time-varying costs, such as electricity prices or spot-instance prices, wherever batch jobs tolerate delay and elastic scaling.
- If the method transfers, cluster operators could use the oracle's learned mappings to quantify the carbon value of adding delay slack or elasticity, informing queue design and instance procurement.
- A natural stress test beyond the paper's ±20% shift experiments is to measure how quickly continuous learning recovers after an abrupt regime change, and whether a drift detector could trigger re-learning.
- For heterogeneous clusters, the state vector would need to include resource-type counts; the paper leaves that extension unexamined, but the same state-matching structure appears to generalize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CarbonFlex is a carbon-aware cluster resource manager that separates capacity provisioning from job scheduling. In a learning phase it replays historical workload and carbon-intensity traces through a greedy offline oracle (Algorithm 1), records mappings from system state to provisioned capacity and a marginal-throughput threshold, and stores these mappings in a knowledge base. At runtime it uses k-nearest-neighbor/case-based reasoning to retrieve similar states and applies the learned capacity and scheduling threshold. The paper implements the system on AWS ParallelCluster for CPU and GPU clusters and evaluates on Azure, Alibaba, and SURF traces across ten carbon regions, reporting roughly 57.5% carbon savings relative to a carbon-agnostic baseline and performance within 2.1% of the oracle.
Significance. The empirical contribution is substantial: a real deployment on AWS ParallelCluster, CPU and GPU prototypes, three publicly available workload traces, ten carbon-intensity regions, and comparisons against five baselines. The 51-57% savings over a carbon-agnostic baseline is the most independent and credible result in the paper. However, the theoretical optimality claim is not established, and the 'within 2.1% of oracle' metric is partly a self-consistency check. With corrected framing and a validated oracle, the system can be a useful practical contribution; as written, the near-optimality claims exceed what the evidence supports.
major comments (4)
- [Section 4.2, Theorem 4.1 and Algorithm 1] The optimality claim is unsupported and, under the paper's own definitions, false. If p_j(k) is the total normalized throughput at scale k (as Figure 2 suggests), the correct marginal carbon-efficiency of the k-th server is (p_j(k)-p_j(k-1))/CI_t, not p_j(k)/CI_t, since the carbon cost of k servers is k*CI_t under the fixed per-resource energy assumption of Section 5. If p_j(k) is instead the marginal throughput, as the text in Section 3 states, then Algorithm 1 is solving a time-indexed covering problem with per-job deadlines and a capacity constraint, not the single-resource separable concave allocation problem of reference [19]; the one-line reduction is not a proof. A concrete counterexample under the marginal-profile reading is: one job with l_j=2, k_min=1, k_max=2, p(1)=1, p(2)=0.5, M=3, and CI=[100,1]. Algorithm 1 assigns scale 2 in both slots, consuming 2*100 + 2*1 = 202 carbon units, while scale 1 in both slots completes the same required work with 100 + 1 = 101 units. Thus Algorithm 1 is not carbon-optimal, and the abstract's 'within 2.1% of an oracle' does not establish near-optimality. Footnote 2 also concedes that the stated optimality conditions do not hold in the evaluation, which further separates the theorem from the experimental setting.
- [Section 4.1 and Section 6.2] The 'within 2.1% of the oracle' metric is largely a self-consistency check. CarbonFlex's runtime is a KNN learner trained on the oracle's decisions, and under the stable-distribution hypothesis stated in Section 4.1 a learned mimic should approach its teacher. The independent, non-circular result is the 51-57% improvement over the carbon-agnostic baseline. The paper should present the oracle gap as learning error, not as a distance to optimality, and should validate the oracle separately against a true optimum, for example by exhaustive search on small instances, before using the gap as evidence of near-optimal performance.
- [Section 5, Eq. (1), versus Algorithm 1] The oracle ranks allocations by p_j(k)/CI_t, but Eq. (1) defines operational carbon as measured per-job energy times carbon intensity. For GPU workloads, Section 5 states that per-GPU energy is measured with nvidia-smi and varies across workloads, so a job with high normalized throughput per server but high absolute power is prioritized by Algorithm 1 even though it does not minimize the reported carbon metric. The theorem is stated only for homogeneous clusters, and the GPU evaluation is precisely the setting where the oracle's objective and the reported carbon objective diverge. Either Algorithm 1 should use p_j(k)/(E_js*CI_t) with job- and scale-specific energy, or the GPU results in Figure 7 should be described as relative to a throughput-based heuristic rather than a carbon-optimal oracle.
- [Section 6.3, Figure 9b] The evaluation reports that CarbonFlex and CarbonScaler violate the configured delay by 3.4 and 1.6 hours at small allowed delays, and it also notes that some oracle jobs exceed the deadline and are 'fixed' by extending the delay. This conflicts with the problem statement in Section 3, which requires jobs to complete within their queue-specific slack. Reported carbon savings at those settings are therefore partly achieved by relaxing the stated constraint. The paper should either enforce deadlines and separate feasible from infeasible schedules or explicitly report the SLO-violation trade-off for every configuration.
minor comments (4)
- [Section 1] The text says 'terrawatt-hours'; this should be 'terawatt-hours'.
- [Section 5, Simulation Environment paragraph] The sentence 'we utilize the first two weeks of the Azure for sampling the historical trace' is missing a noun and should be revised for clarity.
- [Algorithm 1 and Algorithm 3] The capacity checks in Algorithm 1 line 9 and Algorithm 3 line 7 use non-incremental overwriting assignments together with strict inequality conditions, so the invariant that total allocation never exceeds M or m_t is not clearly maintained; please state whether the condition is on the incremental or total allocation and correct the strict/equality handling.
- [Table 3] The footnote 'This application present our least scalable workload' is ungrammatical, and the asterisk placement should be checked against the intended row.
Circularity Check
No significant circularity: the 57% savings result is an independent held-out comparison, and the 'within 2.1% of oracle' metric is a transparent train/test imitation check rather than a derivation from its own inputs.
full rationale
The paper's main savings claim is not circular. CarbonFlex is evaluated on a held-out week: Section 6.1 states that the first two weeks of the Azure trace (or first seven weeks of Alibaba) are used for learning and a later, separate week is used for evaluation. The 57% reduction is measured against a carbon-agnostic FCFS baseline, which is external to CarbonFlex's fitted parameters and not forced by construction. The runtime policy is admittedly an imitation of CarbonFlex(Oracle): the learning phase stores (STATE -> m_t, rho) mappings produced by Algorithm 1, and the execution phase retrieves similar states (Algorithms 2-3). Therefore the 'within 2.1% of oracle' claim measures how well the student reproduces its teacher on a held-out week; it is a self-consistency metric, but the paper explicitly frames it as such and does not present it as an independent first-principles bound. Theorem 4.1's optimality claim cites an external greedy-optimality theorem [19] rather than relying on a self-citation; whether the cost model is correct (the skeptic's point that the ratio p_j(k)/CI_t omits the number of servers k) is a correctness/optimality objection, not circularity. Self-citations such as [27] are contextual and not load-bearing for the measured savings. Hence no step reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (7)
- network_energy_efficiency_eta_net =
0.1 W/Gbps
- cpu_resource_energy =
not specified (fixed per resource)
- KNN_neighbors_k =
5
- state_distance_threshold_delta =
not specified
- violation_tolerance_epsilon =
not specified
- historical_window_T =
2 weeks (7 for Alibaba)
- elasticity_profile_assignment =
random assignment from Table 3
assumptions (6)
- domain assumption The workload and carbon-intensity distributions during evaluation are representative of the historical learning window (stationarity).
- domain assumption Elastic scaling profiles p_j(k) of jobs are known in advance and accurate.
- domain assumption The greedy oracle in Algorithm 1 is optimal for the carbon-aware scheduling problem under the conditions in Theorem 4.1.
- domain assumption Switching cost (energy and emissions to scale cluster or jobs) is negligible.
- domain assumption Day-ahead carbon intensity forecasts are accurate.
- ad hoc to paper Monotonically decreasing marginal throughput profiles.
Cite this review
Pith. "Pith review of CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters." pith.science (2026). https://pith.science/paper/FVZ5GGGL
@misc{pith2026250518357,
author = {Pith},
title = {Pith review of: CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters},
year = {2026},
howpublished = {\url{https://pith.science/paper/FVZ5GGGL}},
note = {Machine review of arXiv:2505.18357}
}
abstract
Accelerating computing demand, largely from AI applications, has led to concerns about its carbon footprint. Fortunately, a significant fraction of computing demand comes from batch jobs that are often delay-tolerant and elastic, which enables schedulers to reduce carbon by suspending/resuming jobs and scaling their resources down/up when carbon is high/low. However, prior work on carbon-aware scheduling generally focuses on optimizing carbon for individual jobs in the cloud, and not provisioning and scheduling resources for many parallel jobs in cloud clusters. To address the problem, we present CarbonFlex, a carbon-aware resource provisioning and scheduling approach for cloud clusters. CarbonFlex leverages continuous learning over historical cluster-level data to drive near-optimal runtime resource provisioning and job scheduling. We implement CarbonFlex by extending AWS ParallelCluster to include our carbon-aware provisioning and scheduling algorithms. Our evaluation on publicly available industry workloads shows that CarbonFlex decreases carbon emissions by $\sim$57\% compared to a carbon-agnostic baseline and performs within 2.1\% of an oracle scheduler with perfect knowledge of future carbon intensity and job length.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[19]
Awi Federgruen and Henri Groenevelt. 1986. The Greedy Procedure for Resource Allocation Problems: Necessary and Sufficient Conditions for Optimality.Oper. Res.34, 6 (dec 1986), 909–918
1986
-
[1]
2023.https://github.com/PySlurm/pyslurm
2023
-
[2]
Sverre Aarseth
J. Sverre Aarseth. 1985. 12 - Direct Methods for N-Body Simulations. InMultiple Time Scales. Academic Press, 377–418.https://doi.org/10. 1016/B978-0-12-123420-1.50017-3
1985
-
[3]
Bilge Acun, Benjamin Lee, Fiodar Kazhamiaka, Kiwan Maeng, Udit Gupta, Manoj Chakkaravarthy, David Brooks, and Carole-Jean Wu
-
[4]
Amazon Web Services. 2024. ParallelCluster.https://docs.aws.amazon. com/parallelcluster/
2024
-
[5]
Pradeep Ambati, Noman Bashir, David Irwin, and Prashant Shenoy
-
[6]
Luiz Andre Barroso and Urs Hölzle. 2007. The Case for Energy- Proportional Computing.Computer(2007)
2007
-
[7]
Noman Bashir, Varun Gohil, Anagha Belavadi Subramanya, Moham- mad Shahrad, David Irwin, Elsa Olivetti, and Christina Delimitrou
Show all 97 references
-
[8]
Ermao Cai, Da-Cheng Juan, Dimitrios Stamoulis, and Diana Mar- culescu. 2017. Neuralpower: Predict and Deploy Energy-efficient Con- volutional Neural Networks. InAsian Conference on Machine Learning
2017
-
[9]
Carbon Offset Guide. 2024. Understanding Carbon Offsets. https://www.offsetguide.org/understanding-carbon-offsets/carbon- offset-programs/mandatory-voluntary-offset-markets/
2024
-
[10]
Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Tal- luri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, and Alexandru Iosup. 2024. Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis. In2024 IEEE 30...
2024
-
[11]
Wesley J Cole, Danny Greer, Paul Denholm, A Will Frazier, Scott Machen, Trieu Mai, Nina Vincent, and Samuel F Baldwin. 2021. Quan- tifying the challenge of reaching a 100% renewable energy power system for the United States.Joule5, 7 (2021), 1732–1748
2021
-
[12]
Isaías Comprés, Ao Mo-Hellenbrand, Michael Gerndt, and Hans- Joachim Bungartz. 2016. Infrastructure and API Extensions for Elastic Execution of MPI Applications. InProceedings of the 23rd European MPI Users’ Group Meeting(Edinburgh, United Kingdom)(EuroMPI ’16). 82–97.https://...
2016
-
[13]
Eli Cortez, Anand Bonde, Alexandre Muzio, Mark Russinovich, Mar- cus Fontoura, and Ricardo Bianchini. 2017. Resource Central: Un- derstanding and Predicting Workloads for Improved Resource Man- agement in Large Cloud Platforms. InProceedings of the 26th Sym- posium on Operatin...
2017
-
[14]
Marco D’Amico, Ana Jokanovic, and Julita Corbalan. 2019. Holistic Slowdown Driven Scheduling and Resource Management for Malleable Jobs. InProceedings of the 48th International Conference on Parallel Processing(Kyoto, Japan)(ICPP ’19). Article 31, 10 pages.https://doi. org/10....
2019
-
[15]
Alyssa Daniels. 2020. Environmental Leader, Google Signs PPA for 140MW from Solar Farm in Texas.https://www.environmentalleader. com/2020/09/google-candela-texas-solar-ppa/
2020
-
[16]
Smith, Nicole DeCario, and Will Buchanan
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan. 2022. Measuring the Carbon Intensity of AI in Cloud Instances. InProceedings of the 2022 ACM Conference o...
2022
-
[17]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis- senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Trans- formers for Image Reco...
2021 arXiv
-
[18]
Teads Engineering. 2021. Carbon Footprint Estimator for AWS Instances.https://engineering.teads.com/sustainability/carbon- footprint-estimator-for-aws-instances/Accessed: 2025-01-17
2021
-
[20]
Rodrigo, and Lavanya Ramakrishnan
William Fox, Devarshi Ghoshal, Abel Souza, Gonzalo P. Rodrigo, and Lavanya Ramakrishnan. 2017. E-HPC: a library for elastic resource management in HPC environments. InProceedings of the 12th Workshop on Workflows in Support of Large-Scale Science(Denver, Colorado) (WORKS ’17)....
2017 doi
-
[21]
Google. 2024. Enviromental Report 2024.https://www.gstatic.com/ gumdrop/sustainability/google-2024-environmental-report.pdf
2024
-
[22]
Viktor Urban Gsteiger, Pin Hong (Daniel) Long, Yiran (Jerry) Sun, Parshan Javanrood, and Mohammad Shahrad. 2024. Caribou: Fine- Grained Geospatial Shifting of Serverless Applications for Sustain- ability. InProceedings of the ACM SIGOPS 30th Symposium on Op- erating Systems Pr...
2024
-
[23]
Abhishek Gupta, Bilge Acun, Osman Sarood, and Laxmikant V. Kalé
-
[24]
Lee, David Brooks, and Carole-Jean Wu
Udit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei, Hsien-Hsin S. Lee, David Brooks, and Carole-Jean Wu. 2022. ACT: Designing Sus- tainable Computer Systems With An Architectural Carbon Model- ing Tool. InProceedings of the 49th Annual International Symposium on Computer Arch...
2022
-
[25]
Yuelin Han, Zhifeng Wu, Pengfei Li, Adam Wierman, and Shaolei Ren
-
[26]
Hanafy, Roozbeh Bostandoost, Noman Bashir, David Irwin, Mohammad Hajiesmaili, and Prashant Shenoy
Walid A. Hanafy, Roozbeh Bostandoost, Noman Bashir, David Irwin, Mohammad Hajiesmaili, and Prashant Shenoy. 2023. The War of the Efficiencies: Understanding the Tension between Carbon and Energy Optimization. InProceedings of the 2nd Workshop on Sustainable Com- puter Systems(...
2023
-
[27]
Hanafy, Qianlin Liang, Noman Bashir, David Irwin, and Prashant Shenoy
Walid A. Hanafy, Qianlin Liang, Noman Bashir, David Irwin, and Prashant Shenoy. 2023. CarbonScaler: Leveraging Cloud Workload Elasticity for Optimizing Carbon-Efficiency.Proc. ACM Meas. Anal. Comput. Syst.7, 3, Article 57 (dec 2023), 28 pages.https://doi.org/10. 1145/3626788
2023
-
[28]
Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin, and Prashant Shenoy
Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin, and Prashant Shenoy. 2024. Going Green for Less Green: Opti- mizing the Cost of Reducing Cloud Carbon Emissions. InProceedings of the 29th ACM International Conference on Architectural Support for Programmi...
2024
-
[29]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 770– 778.https://doi.org/10.1109/CVPR.2016.90
2016 doi
-
[30]
arXiv:2412.06288 [cs.CY]https://arxiv.org/abs/2412.06288
The Unpaid Toll: Quantifying the Public Health Impact of AI. arXiv:2412.06288 [cs.CY]https://arxiv.org/abs/2412.06288
-
[31]
2024.Electricity 2024
International Energy Agency. 2024.Electricity 2024. Technical Report. IEA, Paris.https://www.iea.org/reports/electricity-2024
2024
-
[32]
Romain Jacob and Laurent Vanbever. 2023. The Internet of Tomorrow Must Sleep More and Grow Old.SIGENERGY Energy Inform. Rev.3, 3 (Oct. 2023), 27–32.https://doi.org/10.1145/3630614.3630620
2023
-
[33]
Suhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao, Zhihao Jia, and Gregory R. Ganger. 2023. Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling. InProceedings of the 29th Symposium on Operating Systems Principles(Koblenz, Germany)(SOSP ’23). 642–6...
2023
-
[34]
Daniel Justus, John Brennan, Stephen Bonner, and Andrew Stephen McGough. 2018. Predicting the Computational Cost of Deep Learning Models. In2018 IEEE International Conference on Big Data (Big Data)
2018
-
[35]
Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, An- thony Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-grained Resource Sharing in the Data Center. In USENIX Symposium on Networked Systems Design and Implementation (NSDI)
2011
-
[36]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Ima- geNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, Vol. 25. Curran As- sociates, Inc.https://proceedings.neurips.cc/paper_files/paper/2012/ file/...
2012
-
[37]
Cranor, Elisabeth Moore, Nathan Debardeleben, and George Amvrosiadis
Michael Kuchnik, Jun Woo Park, Charles D. Cranor, Elisabeth Moore, Nathan Debardeleben, and George Amvrosiadis. 2019.This is why ML-driven cluster scheduling remains widely impractical. Technical Report. CMU-PDL-19-103
2019
-
[38]
Loïc Lannelongue, Jason Grealey, and Michael Inouye. 2021. Green Algorithms: Quantifying the Carbon Footprint of Computation.Ad- vanced Science8, 12 (2021), 2100707.https://doi.org/10.1002/advs. 202100707
2021 doi
-
[39]
Adam Lechowicz, Nicolas Christianson, Jinhang Zuo, Noman Bashir, Mohammad Hajiesmaili, Adam Wierman, and Prashant Shenoy. 2023. The Online Pause and Resume Problem: Optimal Algorithms and An Application to Carbon-Aware Load Shifting.Proceedings of the ACM on Measurement and An...
2023
-
[40]
Laxmikant V Kale and Sanjeev Krishnan. 1993. Charm++ a portable concurrent object oriented system based on c++. InProceedings of the eighth annual conference on Object-oriented programming systems, languages, and applications. 91–108
1993
-
[42]
Liuzixuan Lin, Rajini Wijayawardana, Varsha Rao, Hai Nguyen, Em- manuel Wedan GNIBGA, and Andrew A. Chien. 2024. Exploding AI Power Use: an Opportunity to Rethink Grid Planning and Man- agement. InProceedings of the 15th ACM International Conference on Future and Sustainable E...
2024
-
[43]
Sitaraman
Diptyaroop Maji, Prashant Shenoy, and Ramesh K. Sitaraman. 2022. CarbonCast: Multi-Day Forecasting of Grid Carbon Intensity. InPro- ceedings of the 9th ACM International Conference on Systems for Energy- Efficient Buildings, Cities, and Transportation(Boston, Massachusetts) (B...
2022
-
[44]
Jens Malmodin, Nina Lövehagen, Pernilla Bergmark, and Dag Lundén
-
[45]
Wenxue Li, Xiangzhou Liu, Yuxuan Li, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang, and Kai Chen. 2024. Understanding Communication Characteristics of Distributed Training. InProceedings of the 8th Asia-Pacific Workshop on Networking(Sydney, Australia) (APNet ’24)....
2024
-
[46]
Aliaga, Maribel Castillo, and Sergio Iserte
Iker Martín-Álvarez, José I. Aliaga, Maribel Castillo, and Sergio Iserte
-
[47]
1994.MPI: A Message-Passing Inter- face Standard
Message Passing Interface Forum. 1994.MPI: A Message-Passing Inter- face Standard. Technical Report. USA
1994
-
[48]
Rich Miller. 2022. Cloud Titans Were the Largest Buyers of Renewable Energy in 2021.https://www.datacenterfrontier.com/featured/article/ 11427604/cloud-titans-were-the-largest-buyers-of-renewable- energy-in-2021
2022
-
[49]
Anne-Cecile Orgerie, Marcos Dias de Assuncao, and Laurent Lefevre
-
[50]
ICT sector electricity consumption and greenhouse gas emissions – 2020 outcome.Telecommunications Policy48, 3 (2024), 102701.https: //doi.org/10.1016/j.telpol.2023.102701
2024
-
[51]
Electricity Maps. 2022. Electricity Map.https://www.electricitymap. org/map
2022
-
[52]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, 15 Conference’17, July 2017, Washington, DC, USA Hanafy et al. A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay
2017
-
[53]
Supercomput.80, 15 (July 2024), 23083–23119
Proteo: a framework for the generation and evaluation of mal- leable MPI applications.J. Supercomput.80, 15 (July 2024), 23083–23119. https://doi.org/10.1007/s11227-024-06277-5
2024 doi
-
[54]
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanx- iong Guo. 2018. Optimus: An Efficient Dynamic Resource Sched- uler for Deep Learning Clusters. InProceedings of the Thirteenth Eu- roSys Conference(Porto, Portugal)(EuroSys ’18). Article 3, 14 pages. https://doi.org/...
2018
-
[55]
Lucas Perotin, Chaojie Zhang, Rajini Wijayawardana, Anne Benoit, Yves Robert, and Andrew Chien. 2023. Risk-Aware Scheduling Al- gorithms for Variable Capacity Resources. InProceedings of the SC ’23 Workshops of The International Conference on High Performance Computing, Networ...
2023
-
[56]
Suraj Prabhakaran, Marcel Neumann, Sebastian Rinke, Felix Wolf, Ab- hishek Gupta, and Laxmikant V. Kale. 2015. A Batch System with Efficient Adaptive Scheduling for Malleable and Evolving Applica- tions. In2015 IEEE International Parallel and Distributed Processing Symposium. ...
2015 doi
-
[57]
Surv.46, 4, Article 47 (mar 2014), 31 pages.https://doi.org/10.1145/2532637
A Survey on Techniques for Improving the Energy Efficiency of Large-Scale Distributed Systems.ACM Comput. Surv.46, 4, Article 47 (mar 2014), 31 pages.https://doi.org/10.1145/2532637
2014 doi
-
[58]
Yosuke Oyama, Akihiro Nomura, Ikuro Sato, Hiroki Nishimura, Yuki- masa Tamatsu, and Satoshi Matsuoka. 2016. Predicting Statistics of Asynchronous SGD Parameters for a Large-scale Distributed Deep Learning System on GPU Supercomputers. In2016 IEEE International Conference on Bi...
2016
-
[59]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Brad- bury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...
2019
-
[60]
Lawrence M. Ruane. 1990. Process synchronization in the UTS kernel. Computing systems3, 3 (1990), 387–421
1990
-
[61]
Ian Schneider, Hui Xu, Stephan Benecke, David Patterson, Keguo Huang, Parthasarathy Ranganathan, and Cooper Elsworth. 2025. Life- Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends. arXiv:2502.01671 [cs.AR]https://arxiv.org/abs/ 2502.01671
2025 arXiv
-
[62]
Ziqian Pei, Chensheng Li, Xiaowei Qin, Xiaohui Chen, and Guo Wei
-
[63]
Smith, Alex Hubbard, Alex Newkirk, Nuoa Lei, Md Abu Bakar Siddik, Billie Holecek, Jonathan Koomey, Eric Masanet, and Dale Sartor
Arman Shehabi, Sarah J. Smith, Alex Hubbard, Alex Newkirk, Nuoa Lei, Md Abu Bakar Siddik, Billie Holecek, Jonathan Koomey, Eric Masanet, and Dale Sartor. 2024.2024 United States Data Center Energy Usage Report. Technical Report LBNL-2001637. Lawrence Berkeley National Lab (LBL...
2024
-
[64]
Sartor, Richard E
Arman Shehabi, Sarah Josephine Smith, Dale A. Sartor, Richard E. Brown, Magnus Herrlin, Jonathan G. Koomey, Eric R. Masanet, Natha- nial Horner, Ines Lima Azevedo, and William Linter. 2016.United States Data Center Energy Usage Report. Technical Report LBNL-1005775. Lawrence B...
2016
-
[65]
Shaohuai Shi, Qiang Wang, and Xiaowen Chu. 2018. Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs. In2018 IEEE 16th Intl Conf on Dependable, Autonomic and Secure Computing, 16th Intl Conf on Pervasive Intelligence and Computing, 4th Intl Conf...
2018 doi
-
[66]
Abel Souza, Noman Bashir, Jorge Murillo, Walid Hanafy, Qianlin Liang, David Irwin, and Prashant Shenoy. 2023. Ecovisor: A Virtual Energy System for Carbon-Efficient Applications. InProceedings of the 28th ACM International Conference on Architectural Support for Program- ming ...
2023 doi
-
[67]
Sparks, and Ameet S
Qi, Evan R. Sparks, and Ameet S. Talwalkar. 2017. Paleo: A Performance Model for Deep Neural Networks. InThe International Conference on Learning Representations (ICLR’17)
2017
-
[68]
Ganger, and Eric P
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. In15th USENIX Symposium on Operating Systems Design and Imple...
2021
-
[69]
Ana Radovanović, Ross Koningstein, Ian Schneider, Bokan Chen, Alexandre Duarte, Binz Roy, Diyue Xiao, Maya Haridasan, Patrick Hung, Nick Care, Saurav Talukdar, Eric Mullen, Kendal Smith, MariEllen Cottman, and Walfredo Cirne. 2023. Carbon-Aware Comput- ing for Datacenters.IEEE...
2023
-
[70]
Mingxing Tan and Quoc Le. 2021. EfficientNetV2: Smaller Models and Faster Training. InInternational conference on machine learning. PMLR, 10096–10106
2021
-
[71]
Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J
Ahmad Tarraf, Martin Schreiber, Alberto Cascajo, Jean-Baptiste Besnard, Marc-André Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Rui Machado, Marta Garcia-Gasulla, Martin Schulz, Paul Carpenter, Simon P...
2024
-
[72]
Prateek Sharma, Tian Guo, Xin He, David Irwin, and Prashant Shenoy
-
[73]
Energy Information Administration
U.S. Energy Information Administration. 2025. Virginia State Energy Profile.https://www.eia.gov/state/?sid=VAAccessed: 2025-01-21
2025
-
[74]
Ward Van Heddeghem, Filip Idzikowski, Willem Vereecken, Didier Colle, Mario Pickavet, and Piet Demeester. 2012. Power consumption modeling in optical multilayer networks.Photonic Network Communi- cations24, 2 (2012), 86–102.https://doi.org/10.1007/s11107-011-0370-7
2012 doi
-
[75]
Jan Verschelde. 2016. Parallel Iterative Methods for Linear Sys- tems.https://homepages.math.uic.edu/~jan/mcs572f16/mcs572notes/ lec19.html. Lecture notes for MCS 572, Numerical Methods for Partial Differential Equations, University of Illinois at Chicago
2016
-
[76]
Ian Watson and Farhi Marir. 1994. Case-based reasoning: A review. The Knowledge Engineering Review9, 4 (1994), 327–354.https://doi. org/10.1017/S0269888900007098
1994 doi
-
[77]
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Hetero- geneous GPU Clusters. In19th {USENIX} Symposium on Networked 16 CarbonFlex Confer...
2022
-
[78]
Thanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin, and Prashant Shenoy. 2024. On the Limitations of Carbon-Aware Tempo- ral and Spatial Workload Shifting in the Cloud. InProceedings of the Nineteenth European Conference on Computer Systems(Athens, Greece) (EuroSys ’...
2024
-
[79]
Jennifer Switzer, Gabriel Marcano, Ryan Kastner, and Pat Pannuto
-
[80]
InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023)
Junkyard Computing: Repurposing Discarded Smartphones to Minimize Carbon. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023). 400–412.https://doi.org/10.1...
2023
-
[81]
Seyedali Tabaeiaghdaei, Simon Scherrer, Jonghoon Kwon, and Adrian Perrig. 2023. Carbon-Aware Global Routing in Path-Aware Networks. InProceedings of the 14th ACM International Conference on Future Energy Systems(Orlando, FL, USA)(e-Energy ’23). Association for Computing Machin...
2023
-
[82]
Andy B Yoo, Morris A Jette, and Mark Grondona. 2003. Slurm: Simple Linux Utility for Resource Management. InWorkshop on Job Scheduling Strategies for Parallel Processing. Springer, New York, NY, USA, 44–60
2003
-
[83]
Jie You, Jae-Won Chung, and Mosharaf Chowdhury. 2023. Zeus: Under- standing and Optimizing GPU Energy Consumption of DNN Training. In20th USENIX Symposium on Networked Systems Design and Im- plementation (NSDI 23). USENIX Association, Boston, MA, 119–139. https://www.usenix.or...
2023
-
[84]
Haque, Zhi- jing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes
Muhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque, Zhi- jing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes
-
[85]
Chaojie Zhang and Andrew A. Chien. 2021. Scheduling Challenges for Variable Capacity Resources. InJob Scheduling Strategies for Parallel Processing, Dalibor Klusáček, Walfredo Cirne, and Gonzalo P. Rodrigo (Eds.). Springer International Publishing, Cham, 190–209
2021
-
[86]
Chien, and Sangwon Suh
Jiajia Zheng, Andrew A. Chien, and Sangwon Suh. 2020. Mitigating Curtailment and Carbon Emissions through Load Migration between Data Centers.Joule4, 10 (2020), 2208–2222.https://doi.org/10.1016/j. joule.2020.08.001 17
2020 doi
-
[91]
Philipp Wiesner, Ilja Behnke, Dominik Scheinert, Kordian Gontarska, and Lauritz Thamsen. 2021. Let’s Wait Awhile: How Temporal Work- load Shifting Can Reduce Carbon Emissions in the Cloud. InProceed- ings of the 22nd International Middleware Conference(Québec city, Canada)(Mid...
2021 doi
-
[92]
2023.Green Digital Transformation: How to Sustainably Close the Digital Divide and Harness Digital Tools for Climate Action
World Bank. 2023.Green Digital Transformation: How to Sustainably Close the Digital Divide and Harness Digital Tools for Climate Action. World Bank, Washington, DC.http://hdl.handle.net/10986/40653
2023
-
[93]
Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 533...
2020
-
[94]
Kaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang, and Kai Chen. 2025. GREEN: Carbon-efficient Resource Scheduling for Machine Learning Clusters. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 999– 1014.htt...
2025
-
[97]
Franklin, Scott Shenker, and Ion Stoica
Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient Distributed Datasets: A Fault-Tolerant Ab- straction for In-Memory Cluster Computing. In9th USENIX Symposium on Networke...
2012
-
[2011]
Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research12 (2011), 2825–2830
2011
-
[2014]
In2014 21st International Conference on High Performance Computing (HiPC)
Towards Realizing the Potential of Malleable Jobs. In2014 21st International Conference on High Performance Computing (HiPC). 1–10. https://doi.org/10.1109/HiPC.2014.7116905 14 CarbonFlex Conference’17, July 2017, Washington, DC, USA
2014
-
[2016]
InACM European Conference on Computer Systems (EuroSys) (EuroSys ’16)
Flint: Batch-Interactive Data-Intensive Processing for Transient Servers. InACM European Conference on Computer Systems (EuroSys) (EuroSys ’16). ACM, London, United Kingdom, Article 6, 15 pages
-
[2019]
Iteration Time Prediction for CNN in Multi-GPU Platform: Modeling and Analysis.IEEE Access(2019)
2019
-
[2020]
InProceedings of the Fifteenth Eu- ropean Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20)
Borg: The Next Generation. InProceedings of the Fifteenth Eu- ropean Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Article 30, 14 pages.https://doi.org/10.1145/3342195.3387517
-
[2021]
InProceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA)(SoCC ’21)
Good Things Come to Those Who Wait: Optimizing Job Waiting in the Cloud. InProceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA)(SoCC ’21). Association for Computing Machinery, New York, NY, USA, 229–242.https://doi.org/10.1145/3472883.3487007
-
[2024]
InProceedings of the 2024 ACM Symposium on Cloud Computing(Redmond, WA, USA)(SoCC ’24)
The Sunk Carbon Fallacy: Rethinking Carbon Footprint Met- rics for Effective Carbon-Aware Scheduling. InProceedings of the 2024 ACM Symposium on Cloud Computing(Redmond, WA, USA)(SoCC ’24). Association for Computing Machinery, New York, NY, USA, 542–551. https://doi.org/10.114...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.