Pith. sign in

REVIEW 4 major objections 4 minor 97 references

CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A cluster scheduler that learns from past workload traces cuts carbon emissions by about 57% without knowing job lengths ahead of time.

desk verdict A solid systems paper with a genuinely new provisioning-plus-scheduling mechanism, but the 'oracle' is a flawed heuristic—the near-optimal framing needs correction before the headline claims can be taken at face value. read the letter →

arxiv 2505.18357 v1 pith:FVZ5GGGL submitted 2025-05-23 cs.DC

classification cs.DC
keywords carbon-awareschedulingcloudclusterprovisioningelasticbatchjobshistoricallearningoperationalcarbonemissionsmarginalthroughputAWSParallelintensity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

CarbonFlex is a resource manager for cloud clusters that treats capacity provisioning and job scheduling as two separate decisions and applies elastic scaling to both. Its core proposal is to learn these decisions by replaying historical job and carbon-intensity traces through an offline greedy oracle, then mimic the oracle's choices at runtime using only the current system state. The paper reports that on real CPU and GPU clusters CarbonFlex reduces operational carbon emissions by roughly 57% compared to a carbon-agnostic scheduler and comes within 2.1% of an oracle that knows the future perfectly. The result matters because batch and machine-learning jobs are often delay-tolerant, so most of this saving is available without asking users for accurate job-length estimates.

What carries the argument

The load-bearing mechanism is the marginal-throughput-per-carbon greedy oracle: for every job, time slot, and allowed scale, it computes $p_j(k)/CI_t$, the normalized throughput gained per unit of carbon, and allocates resources in descending order of that ratio subject to cluster capacity and per-queue delay. This gives a provably optimal schedule for homogeneous clusters with monotonically decreasing scaling profiles (Theorem 4.1), and it is what CarbonFlex replays over historical traces. The runtime counterpart is a case-based reasoning layer that maps a state vector (carbon intensity, its gradient and day-ahead rank, queue lengths, mean elasticity) to the oracle's capacity decision and scheduling threshold, using k-nearest-neighbor matching in a knowledge base.

What would settle it

Run CarbonFlex on a week whose job arrivals are drawn from a distribution outside the learning window (for example, doubling the arrival rate or inserting a new job class absent from the historical trace) and compare its emissions to the offline oracle for that same week; if the carbon savings fall to near the carbon-agnostic baseline or delay violations exceed the queue slack, the stable-distribution hypothesis fails.

Watch

Extended reading notes

Core claim

The central claim is that continuous learning over historical cluster data can drive near-optimal carbon-aware provisioning and scheduling of many parallel elastic jobs in a shared cluster, without a priori knowledge of job length or future arrivals. CarbonFlex simulates an offline oracle (Algorithm 1), a greedy scheduler that sorts possible job-time-scale assignments by marginal throughput per unit of carbon $p_j(k)/CI_t$, over a past window, and records for each observed system state the cluster capacity $m_t$ and a marginal-throughput threshold $\rho$. At runtime, it matches the current state to the nearest stored states and provisions the cluster accordingly, then schedules all jobs whose marginal throughput exceeds $\rho$. On AWS CPU and GPU clusters driven by Azure, Alibaba, and SURF traces across ten regions, the paper reports carbon savings of 51.4% and 57.5% over the carbon-agnostic baseline and within 6.6% and 2.1% of the offline oracle respectively.

Load-bearing premise

The method works only if the recent past resembles the near future: if the mix of jobs, arrival rates, or carbon conditions changes sharply between the learning window and the evaluation period, the learned decisions no longer match what the oracle would do.

Editorial extensions

If this is right

  • Carbon savings of roughly half can be achieved without job-length knowledge, removing a key barrier to deploying carbon-aware scheduling in real batch clusters.
  • Separating provisioning from scheduling means CarbonFlex's provisioning policy can be paired with other schedulers, and its scheduler with other provisioning curves such as Google's VCC.
  • Because decisions are relearned continuously from a rolling window, the approach tracks gradual seasonal and workload changes rather than needing manual re-tuning.
  • The benefit grows with workload elasticity and carbon-intensity variability: the more jobs can be scaled and the more volatile the grid, the larger the saving.
  • Elastic scaling also lets schedulers push high-power jobs into low-carbon windows, which yields extra savings on GPU clusters where power profiles differ across jobs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same imitate-the-oracle-on-history recipe could be applied to other time-varying costs, such as electricity prices or spot-instance prices, wherever batch jobs tolerate delay and elastic scaling.
  • If the method transfers, cluster operators could use the oracle's learned mappings to quantify the carbon value of adding delay slack or elasticity, informing queue design and instance procurement.
  • A natural stress test beyond the paper's ±20% shift experiments is to measure how quickly continuous learning recovers after an abrupt regime change, and whether a drift detector could trigger re-learning.
  • For heterogeneous clusters, the state vector would need to include resource-type counts; the paper leaves that extension unexamined, but the same state-matching structure appears to generalize.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. CarbonFlex is a carbon-aware cluster resource manager that separates capacity provisioning from job scheduling. In a learning phase it replays historical workload and carbon-intensity traces through a greedy offline oracle (Algorithm 1), records mappings from system state to provisioned capacity and a marginal-throughput threshold, and stores these mappings in a knowledge base. At runtime it uses k-nearest-neighbor/case-based reasoning to retrieve similar states and applies the learned capacity and scheduling threshold. The paper implements the system on AWS ParallelCluster for CPU and GPU clusters and evaluates on Azure, Alibaba, and SURF traces across ten carbon regions, reporting roughly 57.5% carbon savings relative to a carbon-agnostic baseline and performance within 2.1% of the oracle.

Significance. The empirical contribution is substantial: a real deployment on AWS ParallelCluster, CPU and GPU prototypes, three publicly available workload traces, ten carbon-intensity regions, and comparisons against five baselines. The 51-57% savings over a carbon-agnostic baseline is the most independent and credible result in the paper. However, the theoretical optimality claim is not established, and the 'within 2.1% of oracle' metric is partly a self-consistency check. With corrected framing and a validated oracle, the system can be a useful practical contribution; as written, the near-optimality claims exceed what the evidence supports.

major comments (4)
  1. [Section 4.2, Theorem 4.1 and Algorithm 1] The optimality claim is unsupported and, under the paper's own definitions, false. If p_j(k) is the total normalized throughput at scale k (as Figure 2 suggests), the correct marginal carbon-efficiency of the k-th server is (p_j(k)-p_j(k-1))/CI_t, not p_j(k)/CI_t, since the carbon cost of k servers is k*CI_t under the fixed per-resource energy assumption of Section 5. If p_j(k) is instead the marginal throughput, as the text in Section 3 states, then Algorithm 1 is solving a time-indexed covering problem with per-job deadlines and a capacity constraint, not the single-resource separable concave allocation problem of reference [19]; the one-line reduction is not a proof. A concrete counterexample under the marginal-profile reading is: one job with l_j=2, k_min=1, k_max=2, p(1)=1, p(2)=0.5, M=3, and CI=[100,1]. Algorithm 1 assigns scale 2 in both slots, consuming 2*100 + 2*1 = 202 carbon units, while scale 1 in both slots completes the same required work with 100 + 1 = 101 units. Thus Algorithm 1 is not carbon-optimal, and the abstract's 'within 2.1% of an oracle' does not establish near-optimality. Footnote 2 also concedes that the stated optimality conditions do not hold in the evaluation, which further separates the theorem from the experimental setting.
  2. [Section 4.1 and Section 6.2] The 'within 2.1% of the oracle' metric is largely a self-consistency check. CarbonFlex's runtime is a KNN learner trained on the oracle's decisions, and under the stable-distribution hypothesis stated in Section 4.1 a learned mimic should approach its teacher. The independent, non-circular result is the 51-57% improvement over the carbon-agnostic baseline. The paper should present the oracle gap as learning error, not as a distance to optimality, and should validate the oracle separately against a true optimum, for example by exhaustive search on small instances, before using the gap as evidence of near-optimal performance.
  3. [Section 5, Eq. (1), versus Algorithm 1] The oracle ranks allocations by p_j(k)/CI_t, but Eq. (1) defines operational carbon as measured per-job energy times carbon intensity. For GPU workloads, Section 5 states that per-GPU energy is measured with nvidia-smi and varies across workloads, so a job with high normalized throughput per server but high absolute power is prioritized by Algorithm 1 even though it does not minimize the reported carbon metric. The theorem is stated only for homogeneous clusters, and the GPU evaluation is precisely the setting where the oracle's objective and the reported carbon objective diverge. Either Algorithm 1 should use p_j(k)/(E_js*CI_t) with job- and scale-specific energy, or the GPU results in Figure 7 should be described as relative to a throughput-based heuristic rather than a carbon-optimal oracle.
  4. [Section 6.3, Figure 9b] The evaluation reports that CarbonFlex and CarbonScaler violate the configured delay by 3.4 and 1.6 hours at small allowed delays, and it also notes that some oracle jobs exceed the deadline and are 'fixed' by extending the delay. This conflicts with the problem statement in Section 3, which requires jobs to complete within their queue-specific slack. Reported carbon savings at those settings are therefore partly achieved by relaxing the stated constraint. The paper should either enforce deadlines and separate feasible from infeasible schedules or explicitly report the SLO-violation trade-off for every configuration.
minor comments (4)
  1. [Section 1] The text says 'terrawatt-hours'; this should be 'terawatt-hours'.
  2. [Section 5, Simulation Environment paragraph] The sentence 'we utilize the first two weeks of the Azure for sampling the historical trace' is missing a noun and should be revised for clarity.
  3. [Algorithm 1 and Algorithm 3] The capacity checks in Algorithm 1 line 9 and Algorithm 3 line 7 use non-incremental overwriting assignments together with strict inequality conditions, so the invariant that total allocation never exceeds M or m_t is not clearly maintained; please state whether the condition is on the incremental or total allocation and correct the strict/equality handling.
  4. [Table 3] The footnote 'This application present our least scalable workload' is ungrammatical, and the asterisk placement should be checked against the intended row.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the 57% savings result is an independent held-out comparison, and the 'within 2.1% of oracle' metric is a transparent train/test imitation check rather than a derivation from its own inputs.

full rationale

The paper's main savings claim is not circular. CarbonFlex is evaluated on a held-out week: Section 6.1 states that the first two weeks of the Azure trace (or first seven weeks of Alibaba) are used for learning and a later, separate week is used for evaluation. The 57% reduction is measured against a carbon-agnostic FCFS baseline, which is external to CarbonFlex's fitted parameters and not forced by construction. The runtime policy is admittedly an imitation of CarbonFlex(Oracle): the learning phase stores (STATE -> m_t, rho) mappings produced by Algorithm 1, and the execution phase retrieves similar states (Algorithms 2-3). Therefore the 'within 2.1% of oracle' claim measures how well the student reproduces its teacher on a held-out week; it is a self-consistency metric, but the paper explicitly frames it as such and does not present it as an independent first-principles bound. Theorem 4.1's optimality claim cites an external greedy-optimality theorem [19] rather than relying on a self-citation; whether the cost model is correct (the skeptic's point that the ratio p_j(k)/CI_t omits the number of servers k) is a correctness/optimality objection, not circularity. Self-citations such as [27] are contextual and not load-bearing for the measured savings. Hence no step reduces by construction to its own inputs.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

CarbonFlex's runtime performance rests on two families of assumptions: (1) the workload and carbon environment is sufficiently stationary for historical learning to transfer, and (2) the offline oracle provides a useful near-optimal benchmark. Neither is independently verified beyond the paper's own experiments. The evaluation also depends on several hand-chosen constants (network energy efficiency, KNN hyperparameters) that are not systematically varied.

free parameters (7)
  • network_energy_efficiency_eta_net = 0.1 W/Gbps
    Used in Eq. (3) to estimate network energy; the paper notes prior values vary by three orders of magnitude, so this choice materially affects absolute carbon numbers though scheduling decisions use compute-only marginal throughput.
  • cpu_resource_energy = not specified (fixed per resource)
    CPU energy assumed constant per resource ('we assume a fixed value per resource'), a common simplification; affects absolute emissions but normalized savings are emphasized.
  • KNN_neighbors_k = 5
    Number of nearest state matches in the knowledge base; hardcoded, no sensitivity analysis reported.
  • state_distance_threshold_delta = not specified
    Used in Algorithm 2 to trigger maximum-capacity fallback; value not reported in the paper.
  • violation_tolerance_epsilon = not specified
    Delay-violation tolerance in Algorithm 2; value not reported.
  • historical_window_T = 2 weeks (7 for Alibaba)
    Length of trace replayed to the oracle; chosen per trace without a systematic study of its effect.
  • elasticity_profile_assignment = random assignment from Table 3
    Jobs in traces are randomly assigned scaling profiles from the profiled set; this injects an arbitrary modeling choice into the evaluation.
assumptions (6)
  • domain assumption The workload and carbon-intensity distributions during evaluation are representative of the historical learning window (stationarity).
    Stated as the paper's core hypothesis in Section 4.1: 'under the presence of a stable workload distribution, mimicking the decisions of an oracle provides similar carbon savings at runtime.' The evaluation uses temporally adjacent weeks from the same trace, and Section 6.6 shows savings drop when distributions shift by 20%.
  • domain assumption Elastic scaling profiles p_j(k) of jobs are known in advance and accurate.
    Section 3 states 'we assume that a job's elastic scaling profile is known, which can be learned from profiling or performance models.' The oracle and runtime algorithms both require these profiles.
  • domain assumption The greedy oracle in Algorithm 1 is optimal for the carbon-aware scheduling problem under the conditions in Theorem 4.1.
    Theorem 4.1 is used to claim 'near-optimal' performance; the proof maps to Federgruen and Groenevelt [19] without a formal derivation, and the paper's own footnote concedes the conditions often fail in practice.
  • domain assumption Switching cost (energy and emissions to scale cluster or jobs) is negligible.
    Assumption 3 in Theorem 4.1; scaling overheads are measured in Section 6.8 (checkpoint and restore seconds, EC2 provisioning minutes) but not included in the scheduler's objective.
  • domain assumption Day-ahead carbon intensity forecasts are accurate.
    Section 6.1: 'we assume knowledge of day-ahead carbon intensity, as prior work demonstrates that such forecasts are highly accurate.' This enables the CI_R state feature and the baselines.
  • ad hoc to paper Monotonically decreasing marginal throughput profiles.
    Assumption 1 in Theorem 4.1 for the oracle's optimality; real workloads in Figure 2 may exhibit non-monotonic behavior, and the paper does not check monotonicity for all profiles.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters." pith.science (2026). https://pith.science/paper/FVZ5GGGL

@misc{pith2026250518357,
  author       = {Pith},
  title        = {Pith review of: CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FVZ5GGGL}},
  note         = {Machine review of arXiv:2505.18357}
}
abstract

Accelerating computing demand, largely from AI applications, has led to concerns about its carbon footprint. Fortunately, a significant fraction of computing demand comes from batch jobs that are often delay-tolerant and elastic, which enables schedulers to reduce carbon by suspending/resuming jobs and scaling their resources down/up when carbon is high/low. However, prior work on carbon-aware scheduling generally focuses on optimizing carbon for individual jobs in the cloud, and not provisioning and scheduling resources for many parallel jobs in cloud clusters. To address the problem, we present CarbonFlex, a carbon-aware resource provisioning and scheduling approach for cloud clusters. CarbonFlex leverages continuous learning over historical cluster-level data to drive near-optimal runtime resource provisioning and job scheduling. We implement CarbonFlex by extending AWS ParallelCluster to include our carbon-aware provisioning and scheduling algorithms. Our evaluation on publicly available industry workloads shows that CarbonFlex decreases carbon emissions by $\sim$57\% compared to a carbon-agnostic baseline and performs within 2.1\% of an oracle scheduler with perfect knowledge of future carbon intensity and job length.

Figures

Figures reproduced from arXiv: 2505.18357 by the authors.

Figure 1
Figure 1. Carbon Intensity Variations in four locations in the first week of April 2022. impact of data centers consists of two main components: (i) operational emissions, which comprise the emissions gener￾ated from the energy consumed by the hardware and infras￾tructure during its operations, and (ii) embodied emissions, which consist of the emissions generated during the manu￾facturing and transporting of the computing har… view at source ↗
Figure 2
Figure 2. Elastic scaling profiles of different MPI and ma￾chine learning jobs that depict the marginal increase in throughput for each additional server. on varying the cluster capacity rather than job scheduling. For instance, these approaches do not utilize application elas￾ticity or explicitly address the demand bursts in low-carbon periods [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Overview of the learning and execution phases of CarbonFlex. Resource Allocation Provisioning Policy Scheduling Policy Carbon Intensity Time [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Representing the decisions made by Carbon￾Flex(Oracle) as a provisioning and scheduling policy. schedule over a past window 𝑇 . The algorithm takes histor￾ical job traces of 𝑇 -length (e.g., a week) containing 𝑁 jobs. Each job is characterized by an arrival time 𝑎𝑗 , j…
Figure 5
Figure 5. Figure 5: Diversity in selected Carbon Intensity traces [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Carbon emissions (a) and delay (b) across carbon-aware scheduling approaches for the CPU cluster. 150 C8 VMs, yielding a mean utilization of ∼50%, the com￾mon utilization across clusters [63]. In contrast, for the GPU cluster, our resource quota only allowed for 15 G6 …
Figure 7
Figure 7. Figure 7: Carbon emissions and savings (on-top) across carbon-aware scheduling approaches in a GPU cluster. The highest delays, however, are exhibited by scale-based approaches (e.g., CarbonFlex and CarbonScaler) as the use of provisioning in CarbonFlex may limit the cluster cap…
Figure 8
Figure 8. Figure 8: Impact of the maximum cluster capacity on the carbon savings. Key Takeaways: On CPU and GPU clusters, CarbonFlex yields carbon savings up to 57.5% and 20.8% compared to Carbon￾Agnostic and CarbonScaler, respectively. 6.3 Effect of Configurations This section demonstrat…
Figure 11
Figure 11. Figure 11: Carbon Savings across workload traces. Virginia US Sweden Texas US Germany Spain Ireland Den￾mark California US Ontario Canada South Australia 0 10 20 30 40 50 60 Carbon Savings (%) CarbonFlex(Oracle) CarbonFlex CarbonScaler [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 12
Figure 12. Figure 12: Carbon Savings (%) across locations under multi￾ple job queues. −20 −10 0 10 20 Workload Trace Shift (%) 0 15 30 45 60 Carbon Savings (%) [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Impact of distribution shifts. trace. The reason for these differences can be traced back to variations in job length, as Azure has a higher average job length compared to the other traces. This is also reflected in the disparities between elastic and non-elastic sche…
Figure 14
Figure 14. Figure 14: Comparing CarbonFlex with carbon-aware ca￾pacity provisioning. 6.5 Effect of Cloud Location As noted in Section 2.1, the supply mix significantly affects optimizing carbon emissions, where locations with a vari￾able carbon intensity typically result in higher carbon s…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

97 extracted references · 41 canonical work pages

  1. [19]

    Awi Federgruen and Henri Groenevelt. 1986. The Greedy Procedure for Resource Allocation Problems: Necessary and Sufficient Conditions for Optimality.Oper. Res.34, 6 (dec 1986), 909–918

  2. [1]

    2023.https://github.com/PySlurm/pyslurm

  3. [2]

    Sverre Aarseth

    J. Sverre Aarseth. 1985. 12 - Direct Methods for N-Body Simulations. InMultiple Time Scales. Academic Press, 377–418.https://doi.org/10. 1016/B978-0-12-123420-1.50017-3

  4. [3]

    Bilge Acun, Benjamin Lee, Fiodar Kazhamiaka, Kiwan Maeng, Udit Gupta, Manoj Chakkaravarthy, David Brooks, and Carole-Jean Wu

  5. [4]

    Amazon Web Services. 2024. ParallelCluster.https://docs.aws.amazon. com/parallelcluster/

  6. [5]

    Pradeep Ambati, Noman Bashir, David Irwin, and Prashant Shenoy

  7. [6]

    Luiz Andre Barroso and Urs Hölzle. 2007. The Case for Energy- Proportional Computing.Computer(2007)

  8. [7]

    Noman Bashir, Varun Gohil, Anagha Belavadi Subramanya, Moham- mad Shahrad, David Irwin, Elsa Olivetti, and Christina Delimitrou

Show all 97 references
  1. [8]

    Ermao Cai, Da-Cheng Juan, Dimitrios Stamoulis, and Diana Mar- culescu. 2017. Neuralpower: Predict and Deploy Energy-efficient Con- volutional Neural Networks. InAsian Conference on Machine Learning

  2. [9]

    Carbon Offset Guide. 2024. Understanding Carbon Offsets. https://www.offsetguide.org/understanding-carbon-offsets/carbon- offset-programs/mandatory-voluntary-offset-markets/

  3. [10]

    Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Tal- luri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, and Alexandru Iosup. 2024. Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis. In2024 IEEE 30...

  4. [11]

    Wesley J Cole, Danny Greer, Paul Denholm, A Will Frazier, Scott Machen, Trieu Mai, Nina Vincent, and Samuel F Baldwin. 2021. Quan- tifying the challenge of reaching a 100% renewable energy power system for the United States.Joule5, 7 (2021), 1732–1748

  5. [12]

    Isaías Comprés, Ao Mo-Hellenbrand, Michael Gerndt, and Hans- Joachim Bungartz. 2016. Infrastructure and API Extensions for Elastic Execution of MPI Applications. InProceedings of the 23rd European MPI Users’ Group Meeting(Edinburgh, United Kingdom)(EuroMPI ’16). 82–97.https://...

  6. [13]

    Eli Cortez, Anand Bonde, Alexandre Muzio, Mark Russinovich, Mar- cus Fontoura, and Ricardo Bianchini. 2017. Resource Central: Un- derstanding and Predicting Workloads for Improved Resource Man- agement in Large Cloud Platforms. InProceedings of the 26th Sym- posium on Operatin...

  7. [14]

    Marco D’Amico, Ana Jokanovic, and Julita Corbalan. 2019. Holistic Slowdown Driven Scheduling and Resource Management for Malleable Jobs. InProceedings of the 48th International Conference on Parallel Processing(Kyoto, Japan)(ICPP ’19). Article 31, 10 pages.https://doi. org/10....

  8. [15]

    Alyssa Daniels. 2020. Environmental Leader, Google Signs PPA for 140MW from Solar Farm in Texas.https://www.environmentalleader. com/2020/09/google-candela-texas-solar-ppa/

  9. [16]

    Smith, Nicole DeCario, and Will Buchanan

    Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan. 2022. Measuring the Carbon Intensity of AI in Cloud Instances. InProceedings of the 2022 ACM Conference o...

  10. [17]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis- senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Trans- formers for Image Reco...

  11. [18]

    Teads Engineering. 2021. Carbon Footprint Estimator for AWS Instances.https://engineering.teads.com/sustainability/carbon- footprint-estimator-for-aws-instances/Accessed: 2025-01-17

  12. [20]

    Rodrigo, and Lavanya Ramakrishnan

    William Fox, Devarshi Ghoshal, Abel Souza, Gonzalo P. Rodrigo, and Lavanya Ramakrishnan. 2017. E-HPC: a library for elastic resource management in HPC environments. InProceedings of the 12th Workshop on Workflows in Support of Large-Scale Science(Denver, Colorado) (WORKS ’17)....

  13. [21]

    Google. 2024. Enviromental Report 2024.https://www.gstatic.com/ gumdrop/sustainability/google-2024-environmental-report.pdf

  14. [22]

    Viktor Urban Gsteiger, Pin Hong (Daniel) Long, Yiran (Jerry) Sun, Parshan Javanrood, and Mohammad Shahrad. 2024. Caribou: Fine- Grained Geospatial Shifting of Serverless Applications for Sustain- ability. InProceedings of the ACM SIGOPS 30th Symposium on Op- erating Systems Pr...

  15. [23]

    Abhishek Gupta, Bilge Acun, Osman Sarood, and Laxmikant V. Kalé

  16. [24]

    Lee, David Brooks, and Carole-Jean Wu

    Udit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei, Hsien-Hsin S. Lee, David Brooks, and Carole-Jean Wu. 2022. ACT: Designing Sus- tainable Computer Systems With An Architectural Carbon Model- ing Tool. InProceedings of the 49th Annual International Symposium on Computer Arch...

  17. [25]

    Yuelin Han, Zhifeng Wu, Pengfei Li, Adam Wierman, and Shaolei Ren

  18. [26]

    Hanafy, Roozbeh Bostandoost, Noman Bashir, David Irwin, Mohammad Hajiesmaili, and Prashant Shenoy

    Walid A. Hanafy, Roozbeh Bostandoost, Noman Bashir, David Irwin, Mohammad Hajiesmaili, and Prashant Shenoy. 2023. The War of the Efficiencies: Understanding the Tension between Carbon and Energy Optimization. InProceedings of the 2nd Workshop on Sustainable Com- puter Systems(...

  19. [27]

    Hanafy, Qianlin Liang, Noman Bashir, David Irwin, and Prashant Shenoy

    Walid A. Hanafy, Qianlin Liang, Noman Bashir, David Irwin, and Prashant Shenoy. 2023. CarbonScaler: Leveraging Cloud Workload Elasticity for Optimizing Carbon-Efficiency.Proc. ACM Meas. Anal. Comput. Syst.7, 3, Article 57 (dec 2023), 28 pages.https://doi.org/10. 1145/3626788

  20. [28]

    Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin, and Prashant Shenoy

    Walid A. Hanafy, Qianlin Liang, Noman Bashir, Abel Souza, David Irwin, and Prashant Shenoy. 2024. Going Green for Less Green: Opti- mizing the Cost of Reducing Cloud Carbon Emissions. InProceedings of the 29th ACM International Conference on Architectural Support for Programmi...

  21. [29]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 770– 778.https://doi.org/10.1109/CVPR.2016.90

  22. [30]

    arXiv:2412.06288 [cs.CY]https://arxiv.org/abs/2412.06288

    The Unpaid Toll: Quantifying the Public Health Impact of AI. arXiv:2412.06288 [cs.CY]https://arxiv.org/abs/2412.06288

  23. [31]

    2024.Electricity 2024

    International Energy Agency. 2024.Electricity 2024. Technical Report. IEA, Paris.https://www.iea.org/reports/electricity-2024

  24. [32]

    Romain Jacob and Laurent Vanbever. 2023. The Internet of Tomorrow Must Sleep More and Grow Old.SIGENERGY Energy Inform. Rev.3, 3 (Oct. 2023), 27–32.https://doi.org/10.1145/3630614.3630620

  25. [33]

    Suhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao, Zhihao Jia, and Gregory R. Ganger. 2023. Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling. InProceedings of the 29th Symposium on Operating Systems Principles(Koblenz, Germany)(SOSP ’23). 642–6...

  26. [34]

    Daniel Justus, John Brennan, Stephen Bonner, and Andrew Stephen McGough. 2018. Predicting the Computational Cost of Deep Learning Models. In2018 IEEE International Conference on Big Data (Big Data)

  27. [35]

    Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, An- thony Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-grained Resource Sharing in the Data Center. In USENIX Symposium on Networked Systems Design and Implementation (NSDI)

  28. [36]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Ima- geNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems, Vol. 25. Curran As- sociates, Inc.https://proceedings.neurips.cc/paper_files/paper/2012/ file/...

  29. [37]

    Cranor, Elisabeth Moore, Nathan Debardeleben, and George Amvrosiadis

    Michael Kuchnik, Jun Woo Park, Charles D. Cranor, Elisabeth Moore, Nathan Debardeleben, and George Amvrosiadis. 2019.This is why ML-driven cluster scheduling remains widely impractical. Technical Report. CMU-PDL-19-103

  30. [38]

    Loïc Lannelongue, Jason Grealey, and Michael Inouye. 2021. Green Algorithms: Quantifying the Carbon Footprint of Computation.Ad- vanced Science8, 12 (2021), 2100707.https://doi.org/10.1002/advs. 202100707

  31. [39]

    Adam Lechowicz, Nicolas Christianson, Jinhang Zuo, Noman Bashir, Mohammad Hajiesmaili, Adam Wierman, and Prashant Shenoy. 2023. The Online Pause and Resume Problem: Optimal Algorithms and An Application to Carbon-Aware Load Shifting.Proceedings of the ACM on Measurement and An...

  32. [40]

    Laxmikant V Kale and Sanjeev Krishnan. 1993. Charm++ a portable concurrent object oriented system based on c++. InProceedings of the eighth annual conference on Object-oriented programming systems, languages, and applications. 91–108

  33. [42]

    Liuzixuan Lin, Rajini Wijayawardana, Varsha Rao, Hai Nguyen, Em- manuel Wedan GNIBGA, and Andrew A. Chien. 2024. Exploding AI Power Use: an Opportunity to Rethink Grid Planning and Man- agement. InProceedings of the 15th ACM International Conference on Future and Sustainable E...

  34. [43]

    Sitaraman

    Diptyaroop Maji, Prashant Shenoy, and Ramesh K. Sitaraman. 2022. CarbonCast: Multi-Day Forecasting of Grid Carbon Intensity. InPro- ceedings of the 9th ACM International Conference on Systems for Energy- Efficient Buildings, Cities, and Transportation(Boston, Massachusetts) (B...

  35. [44]

    Jens Malmodin, Nina Lövehagen, Pernilla Bergmark, and Dag Lundén

  36. [45]

    Wenxue Li, Xiangzhou Liu, Yuxuan Li, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang, and Kai Chen. 2024. Understanding Communication Characteristics of Distributed Training. InProceedings of the 8th Asia-Pacific Workshop on Networking(Sydney, Australia) (APNet ’24)....

  37. [46]

    Aliaga, Maribel Castillo, and Sergio Iserte

    Iker Martín-Álvarez, José I. Aliaga, Maribel Castillo, and Sergio Iserte

  38. [47]

    1994.MPI: A Message-Passing Inter- face Standard

    Message Passing Interface Forum. 1994.MPI: A Message-Passing Inter- face Standard. Technical Report. USA

  39. [48]

    Rich Miller. 2022. Cloud Titans Were the Largest Buyers of Renewable Energy in 2021.https://www.datacenterfrontier.com/featured/article/ 11427604/cloud-titans-were-the-largest-buyers-of-renewable- energy-in-2021

  40. [49]

    Anne-Cecile Orgerie, Marcos Dias de Assuncao, and Laurent Lefevre

  41. [50]

    ICT sector electricity consumption and greenhouse gas emissions – 2020 outcome.Telecommunications Policy48, 3 (2024), 102701.https: //doi.org/10.1016/j.telpol.2023.102701

  42. [51]

    Electricity Maps. 2022. Electricity Map.https://www.electricitymap. org/map

  43. [52]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, 15 Conference’17, July 2017, Washington, DC, USA Hanafy et al. A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay

  44. [53]

    Supercomput.80, 15 (July 2024), 23083–23119

    Proteo: a framework for the generation and evaluation of mal- leable MPI applications.J. Supercomput.80, 15 (July 2024), 23083–23119. https://doi.org/10.1007/s11227-024-06277-5

  45. [54]

    Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanx- iong Guo. 2018. Optimus: An Efficient Dynamic Resource Sched- uler for Deep Learning Clusters. InProceedings of the Thirteenth Eu- roSys Conference(Porto, Portugal)(EuroSys ’18). Article 3, 14 pages. https://doi.org/...

  46. [55]

    Lucas Perotin, Chaojie Zhang, Rajini Wijayawardana, Anne Benoit, Yves Robert, and Andrew Chien. 2023. Risk-Aware Scheduling Al- gorithms for Variable Capacity Resources. InProceedings of the SC ’23 Workshops of The International Conference on High Performance Computing, Networ...

  47. [56]

    Suraj Prabhakaran, Marcel Neumann, Sebastian Rinke, Felix Wolf, Ab- hishek Gupta, and Laxmikant V. Kale. 2015. A Batch System with Efficient Adaptive Scheduling for Malleable and Evolving Applica- tions. In2015 IEEE International Parallel and Distributed Processing Symposium. ...

  48. [57]

    Surv.46, 4, Article 47 (mar 2014), 31 pages.https://doi.org/10.1145/2532637

    A Survey on Techniques for Improving the Energy Efficiency of Large-Scale Distributed Systems.ACM Comput. Surv.46, 4, Article 47 (mar 2014), 31 pages.https://doi.org/10.1145/2532637

  49. [58]

    Yosuke Oyama, Akihiro Nomura, Ikuro Sato, Hiroki Nishimura, Yuki- masa Tamatsu, and Satoshi Matsuoka. 2016. Predicting Statistics of Asynchronous SGD Parameters for a Large-scale Distributed Deep Learning System on GPU Supercomputers. In2016 IEEE International Conference on Bi...

  50. [59]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Brad- bury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...

  51. [60]

    Lawrence M. Ruane. 1990. Process synchronization in the UTS kernel. Computing systems3, 3 (1990), 387–421

  52. [61]

    Ian Schneider, Hui Xu, Stephan Benecke, David Patterson, Keguo Huang, Parthasarathy Ranganathan, and Cooper Elsworth. 2025. Life- Cycle Emissions of AI Hardware: A Cradle-To-Grave Approach and Generational Trends. arXiv:2502.01671 [cs.AR]https://arxiv.org/abs/ 2502.01671

  53. [62]

    Ziqian Pei, Chensheng Li, Xiaowei Qin, Xiaohui Chen, and Guo Wei

  54. [63]

    Smith, Alex Hubbard, Alex Newkirk, Nuoa Lei, Md Abu Bakar Siddik, Billie Holecek, Jonathan Koomey, Eric Masanet, and Dale Sartor

    Arman Shehabi, Sarah J. Smith, Alex Hubbard, Alex Newkirk, Nuoa Lei, Md Abu Bakar Siddik, Billie Holecek, Jonathan Koomey, Eric Masanet, and Dale Sartor. 2024.2024 United States Data Center Energy Usage Report. Technical Report LBNL-2001637. Lawrence Berkeley National Lab (LBL...

  55. [64]

    Sartor, Richard E

    Arman Shehabi, Sarah Josephine Smith, Dale A. Sartor, Richard E. Brown, Magnus Herrlin, Jonathan G. Koomey, Eric R. Masanet, Natha- nial Horner, Ines Lima Azevedo, and William Linter. 2016.United States Data Center Energy Usage Report. Technical Report LBNL-1005775. Lawrence B...

  56. [65]

    Shaohuai Shi, Qiang Wang, and Xiaowen Chu. 2018. Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs. In2018 IEEE 16th Intl Conf on Dependable, Autonomic and Secure Computing, 16th Intl Conf on Pervasive Intelligence and Computing, 4th Intl Conf...

  57. [66]

    Abel Souza, Noman Bashir, Jorge Murillo, Walid Hanafy, Qianlin Liang, David Irwin, and Prashant Shenoy. 2023. Ecovisor: A Virtual Energy System for Carbon-Efficient Applications. InProceedings of the 28th ACM International Conference on Architectural Support for Program- ming ...

  58. [67]

    Sparks, and Ameet S

    Qi, Evan R. Sparks, and Ameet S. Talwalkar. 2017. Paleo: A Performance Model for Deep Neural Networks. InThe International Conference on Learning Representations (ICLR’17)

  59. [68]

    Ganger, and Eric P

    Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. In15th USENIX Symposium on Operating Systems Design and Imple...

  60. [69]

    Ana Radovanović, Ross Koningstein, Ian Schneider, Bokan Chen, Alexandre Duarte, Binz Roy, Diyue Xiao, Maya Haridasan, Patrick Hung, Nick Care, Saurav Talukdar, Eric Mullen, Kendal Smith, MariEllen Cottman, and Walfredo Cirne. 2023. Carbon-Aware Comput- ing for Datacenters.IEEE...

  61. [70]

    Mingxing Tan and Quoc Le. 2021. EfficientNetV2: Smaller Models and Faster Training. InInternational conference on machine learning. PMLR, 10096–10106

  62. [71]

    Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J

    Ahmad Tarraf, Martin Schreiber, Alberto Cascajo, Jean-Baptiste Besnard, Marc-André Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Rui Machado, Marta Garcia-Gasulla, Martin Schulz, Paul Carpenter, Simon P...

  63. [72]

    Prateek Sharma, Tian Guo, Xin He, David Irwin, and Prashant Shenoy

  64. [73]

    Energy Information Administration

    U.S. Energy Information Administration. 2025. Virginia State Energy Profile.https://www.eia.gov/state/?sid=VAAccessed: 2025-01-21

  65. [74]

    Ward Van Heddeghem, Filip Idzikowski, Willem Vereecken, Didier Colle, Mario Pickavet, and Piet Demeester. 2012. Power consumption modeling in optical multilayer networks.Photonic Network Communi- cations24, 2 (2012), 86–102.https://doi.org/10.1007/s11107-011-0370-7

  66. [75]

    Jan Verschelde. 2016. Parallel Iterative Methods for Linear Sys- tems.https://homepages.math.uic.edu/~jan/mcs572f16/mcs572notes/ lec19.html. Lecture notes for MCS 572, Numerical Methods for Partial Differential Equations, University of Illinois at Chicago

  67. [76]

    Ian Watson and Farhi Marir. 1994. Case-based reasoning: A review. The Knowledge Engineering Review9, 4 (1994), 327–354.https://doi. org/10.1017/S0269888900007098

  68. [77]

    Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Hetero- geneous GPU Clusters. In19th {USENIX} Symposium on Networked 16 CarbonFlex Confer...

  69. [78]

    Thanathorn Sukprasert, Abel Souza, Noman Bashir, David Irwin, and Prashant Shenoy. 2024. On the Limitations of Carbon-Aware Tempo- ral and Spatial Workload Shifting in the Cloud. InProceedings of the Nineteenth European Conference on Computer Systems(Athens, Greece) (EuroSys ’...

  70. [79]

    Jennifer Switzer, Gabriel Marcano, Ryan Kastner, and Pat Pannuto

  71. [80]

    InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023)

    Junkyard Computing: Repurposing Discarded Smartphones to Minimize Carbon. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2(Vancouver, BC, Canada)(ASPLOS 2023). 400–412.https://doi.org/10.1...

  72. [81]

    Seyedali Tabaeiaghdaei, Simon Scherrer, Jonghoon Kwon, and Adrian Perrig. 2023. Carbon-Aware Global Routing in Path-Aware Networks. InProceedings of the 14th ACM International Conference on Future Energy Systems(Orlando, FL, USA)(e-Energy ’23). Association for Computing Machin...

  73. [82]

    Andy B Yoo, Morris A Jette, and Mark Grondona. 2003. Slurm: Simple Linux Utility for Resource Management. InWorkshop on Job Scheduling Strategies for Parallel Processing. Springer, New York, NY, USA, 44–60

  74. [83]

    Jie You, Jae-Won Chung, and Mosharaf Chowdhury. 2023. Zeus: Under- standing and Optimizing GPU Energy Consumption of DNN Training. In20th USENIX Symposium on Networked Systems Design and Im- plementation (NSDI 23). USENIX Association, Boston, MA, 119–139. https://www.usenix.or...

  75. [84]

    Haque, Zhi- jing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes

    Muhammad Tirmazi, Adam Barker, Nan Deng, Md E. Haque, Zhi- jing Gene Qin, Steven Hand, Mor Harchol-Balter, and John Wilkes

  76. [85]

    Chaojie Zhang and Andrew A. Chien. 2021. Scheduling Challenges for Variable Capacity Resources. InJob Scheduling Strategies for Parallel Processing, Dalibor Klusáček, Walfredo Cirne, and Gonzalo P. Rodrigo (Eds.). Springer International Publishing, Cham, 190–209

  77. [86]

    Chien, and Sangwon Suh

    Jiajia Zheng, Andrew A. Chien, and Sangwon Suh. 2020. Mitigating Curtailment and Carbon Emissions through Load Migration between Data Centers.Joule4, 10 (2020), 2208–2222.https://doi.org/10.1016/j. joule.2020.08.001 17

  78. [91]

    Philipp Wiesner, Ilja Behnke, Dominik Scheinert, Kordian Gontarska, and Lauritz Thamsen. 2021. Let’s Wait Awhile: How Temporal Work- load Shifting Can Reduce Carbon Emissions in the Cloud. InProceed- ings of the 22nd International Middleware Conference(Québec city, Canada)(Mid...

  79. [92]

    2023.Green Digital Transformation: How to Sustainably Close the Digital Divide and Harness Digital Tools for Climate Action

    World Bank. 2023.Green Digital Transformation: How to Sustainably Close the Digital Divide and Harness Digital Tools for Climate Action. World Bank, Washington, DC.http://hdl.handle.net/10986/40653

  80. [93]

    Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 533...

  81. [94]

    Kaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang, and Kai Chen. 2025. GREEN: Carbon-efficient Resource Scheduling for Machine Learning Clusters. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, Philadelphia, PA, 999– 1014.htt...

  82. [97]

    Franklin, Scott Shenker, and Ion Stoica

    Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient Distributed Datasets: A Fault-Tolerant Ab- straction for In-Memory Cluster Computing. In9th USENIX Symposium on Networke...

  83. [2011]

    Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research12 (2011), 2825–2830

  84. [2014]

    In2014 21st International Conference on High Performance Computing (HiPC)

    Towards Realizing the Potential of Malleable Jobs. In2014 21st International Conference on High Performance Computing (HiPC). 1–10. https://doi.org/10.1109/HiPC.2014.7116905 14 CarbonFlex Conference’17, July 2017, Washington, DC, USA

  85. [2016]

    InACM European Conference on Computer Systems (EuroSys) (EuroSys ’16)

    Flint: Batch-Interactive Data-Intensive Processing for Transient Servers. InACM European Conference on Computer Systems (EuroSys) (EuroSys ’16). ACM, London, United Kingdom, Article 6, 15 pages

  86. [2019]

    Iteration Time Prediction for CNN in Multi-GPU Platform: Modeling and Analysis.IEEE Access(2019)

  87. [2020]

    InProceedings of the Fifteenth Eu- ropean Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20)

    Borg: The Next Generation. InProceedings of the Fifteenth Eu- ropean Conference on Computer Systems(Heraklion, Greece)(EuroSys ’20). Article 30, 14 pages.https://doi.org/10.1145/3342195.3387517

  88. [2021]

    InProceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA)(SoCC ’21)

    Good Things Come to Those Who Wait: Optimizing Job Waiting in the Cloud. InProceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA)(SoCC ’21). Association for Computing Machinery, New York, NY, USA, 229–242.https://doi.org/10.1145/3472883.3487007

  89. [2024]

    InProceedings of the 2024 ACM Symposium on Cloud Computing(Redmond, WA, USA)(SoCC ’24)

    The Sunk Carbon Fallacy: Rethinking Carbon Footprint Met- rics for Effective Carbon-Aware Scheduling. InProceedings of the 2024 ACM Symposium on Cloud Computing(Redmond, WA, USA)(SoCC ’24). Association for Computing Machinery, New York, NY, USA, 542–551. https://doi.org/10.114...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.