REVIEW 4 major objections 4 minor 113 references
Towards Decentralized and Sustainable Foundation Model Training with the Edge
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that training foundation models on idle smartphones and laptops can cut net carbon emissions 4 to 8 times versus cloud GPUs because the devices' large embodied carbon footprint is already paid for by ownership.
desk verdict Nice vision paper with genuinely new small-scale energy measurements, but the headline 4–8x carbon win rests on FLOPS-only equivalence and an idealized training model, so treat the number as an upper bound, not a prediction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a three-step carbon-accounting identity: total device carbon equals embodied carbon plus operational carbon; embodied carbon is counted at purchase and unchanged by use; and since edge devices are idle most of the time, their embodied carbon is already amortized by ownership. The paper operationalizes this with FLOPS-equivalence scaling—15 laptops or 69 smartphones equal one H100—and with an idealized distributed-training model that assumes each parameter is transmitted once and gradients are aggregated locally, isolating the energy comparison from the artifacts of specific parallelization methods.
What would settle it
A controlled end-to-end run of OPT-1.3B training on 15 M2 Pro laptops or 69 Snapdragon 888 phones over real Wi-Fi, with actual gradient exchanges, stragglers, and thermal throttling, measuring total energy and comparing with the H100's embodied-plus-operational carbon, would falsify the claim if the edge fleet's marginal energy exceeds about 25% of the cloud GPU's total footprint, the threshold at which the 4 times bound no longer holds.
Extended reading notes
Core claim
The paper's core discovery is a carbon-accounting opportunity: because edge devices already carry a large embodied carbon footprint from manufacturing and are mostly idle, the marginal cost of using them for foundation-model training is small compared with the full embodied-plus-operational footprint of a cloud GPU. Concretely, the authors measure that training OPT-125M takes 1.5 to 7.5 times less energy on a Snapdragon 888 smartphone or M2 Pro laptop than on an NVIDIA A5000 cloud GPU, and that distributed training of OPT-1.3B across such devices remains 1.5 to 5 times more energy-efficient even with WiFi communication. Using raw FLOPS as a compute-equivalence metric, they estimate that 15 M2 Pro laptops or 69 Snapdragon 888 smartphones match one H100 GPU's compute; when the cloud GPU's total carbon is replaced by the marginal operational carbon of these fleets, the net reduction is 4 to 8 times, or 3.5 to 6 times after including communication energy. The discovery is not a new training algorithm but an argument that the largest carbon cost of edge devices is already sunk, so offloading cloud training onto them is a net carbon win.
Load-bearing premise
The 4 to 8 times figure assumes that a laptop or smartphone matches cloud GPU capability on raw FLOPS and that the idealized training pattern—each parameter transmitted only once, no direct device-to-device traffic—can be realized with real edge devices.
Editorial extensions
If this is right
- If the 4 to 8 times reduction holds, foundation-model training could be moved in part from data centers onto user devices, shrinking both operational and embodied carbon on the cloud side.
- The case makes carbon-aware scheduling a first-class requirement: training must be routed to devices that are idle, charging, and on low-carbon grids, because those conditions keep the marginal carbon small.
- New training methods must blend data and pipeline parallelism with communication compression and fault tolerance tuned for heterogeneous, preemptible edge devices, rather than reusing cloud-style sharding unchanged.
- If accurate component-level energy monitoring can be built into edge devices, users and platforms could be rewarded for participating in low-carbon time windows, turning sustainability into an incentive rather than a cost.
- The amortization logic implies that keeping older devices in service longer strengthens the carbon case, making device longevity itself a climate lever.
Reading between the lines
- Editor's inference: The 4 to 8 times figure is an upper bound, because raw FLOPS overstates edge capability for large-model training; memory capacity, memory bandwidth, interconnect, straggler effects, and thermal throttling will reduce real per-device contribution, so the realistic reduction will likely be smaller for large pretraining runs.
- Editor's inference: The argument becomes stronger for fine-tuning and continual pretraining, where models are smaller and communication-lighter; a direct comparative study of edge versus cloud for fine-tuning would be a natural, low-cost extension.
- Editor's inference: Beyond carbon, the same offloading logic implies a decentralization benefit: if edge fleets can train usable models, the high capital barrier to foundation-model development drops, shifting control away from the few organizations that currently own large GPU clusters.
- Editor's inference: A testable extension would be to include the embodied carbon of network infrastructure and the marginal grid carbon of charging sessions in the ledger, then re-derive the 4 to 8 times factor with real device telemetry from a pilot deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that the collective spare capacity of edge devices (smartphones, laptops) can be used for decentralized foundation model training at lower net carbon cost than cloud GPUs. The argument has three parts: edge devices are more energy-efficient per operation; their embodied carbon is already incurred through ownership even when idle; and offloading cloud compute to idle edge devices therefore reduces total carbon. The quantitative case is made through single-device experiments (OPT-125M, Table 1), a distributed DT-FM experiment (OPT-1.3B, Table 2), an idealized distributed-training energy extrapolation (Fig. 3), and a lifecycle carbon comparison over three years (Figs. 4 and 5) that yields a claimed net reduction of 4–8× (smartphones vs. laptops). The paper also frames open challenges in distributed training, orchestration, energy monitoring, security, and incentives.
Significance. Strengths: the paper makes a concrete and falsifiable quantitative claim, reports real measurements rather than only speculation, uses publicly documented embodied-carbon data, and explicitly lists the assumptions of its idealized training model in footnote 1. The observation that idle personal devices carry amortizable embodied carbon is a legitimate and useful sustainability perspective, and the paper correctly notes in §5 that communication and thermal throttling can erode the benefits. Weaknesses: the headline number depends on FLOPS-based device-count equivalence and on a best-case communication model, both of which are unresolved. The paper is a vision/position paper, and the quantitative claim is presented as evidence for the vision rather than as a bounded result. If the authors add a sensitivity analysis and qualify the claim, the paper would make a solid contribution to the sustainability-for-AI discussion.
major comments (4)
- [§4.2, Fig. 5] The offloading calculation in Fig. 5 and the accompanying text ('15 (M2 Pro) laptops or 69 (Snapdragon 888) smartphones') equates a cloud GPU with edge devices on peak FLOPS. For transformer training, the binding constraints are memory capacity, memory bandwidth, and interconnect speed, not peak FLOPS. A roofline estimate using the H100's HBM bandwidth (~3 TB/s) versus the LPDDR5 bandwidth of a smartphone (~50 GB/s) gives a substantially larger device multiplier than 69 at realistic arithmetic intensities. If the required device count is 2–3× higher, the added operational and communication carbon in Fig. 5 scales proportionally and the headline 4–8× reduction drops below 2×. Please provide a sensitivity analysis over device count and show the claim as a function of the assumed FLOPS-to-bandwidth ratio.
- [Footnote 1, §4.2, Fig. 3] The idealized distributed training method used for Fig. 3 and Fig. 5 assumes no peer-to-peer transmission, local gradient aggregation, and exactly one transmission per parameter per batch. This is a strict lower bound on communication, not a system model. The paper's own §5 (Distributed training methods for the edge) states that 'communication overhead grows rapidly with scale and can potentially dominate energy usage.' The Table 2 DT-FM experiment includes communication energy, but only for OPT-1.3B under an assumed symmetric 10 MB/s bandwidth and a single network type; it does not support extrapolation to 15/69 devices or to larger models. Please either replace the idealized model with a realistic communication model for the headline claim, or explicitly qualify the 4–8× number as a best-case lower bound.
- [§4.2, Fig. 5] The accounting in Fig. 5 does not show how the 'additional 8 hour daily usage' interacts with the device count. Matching a 24/7 cloud GPU with devices available for only 8 h/day should require a duty-cycle factor; the text states 'assuming an additional 8 hour daily usage per device while charging' but never writes the formula that converts per-device FLOPS-hours into the 15 and 69 counts. In addition, the cloud baseline is set to zero communication energy ('we only consider a single cloud GPU in isolation'), whereas any real cloud training uses a cluster with data movement; this asymmetry favors the edge case and should be stated as an assumption or modeled.
- [Table 1, §4.2] The single-device energy measurements that feed the distributed estimates in Table 2 and Fig. 3 are reported without methodology details: no description of the power measurement tool, whether power is whole-device or SoC-only, how many runs were averaged, or error bars. The '1.5–7.5× lower energy' claim rests on this table, and the three numbers (10W/3510s, 15W/480s, 220W/250s) are presented as point estimates despite likely run-to-run and thermal-throttling variation. Please supply the measurement protocol and repeat-run statistics, or temper the quantitative claims accordingly.
minor comments (4)
- [§4.2, Table 1] Typographical errors: 'NVDIA A5000' in §4.2 and Table 1 should be 'NVIDIA A5000'.
- [Fig. 5] The labels '17 devices' and '81 devices' in Fig. 5 do not match the text's '15 laptops' and '69 smartphones' in §4.2; please reconcile.
- [§4.2] The footnote marker 'idealized1' in §4.2 should be spaced as 'idealized 1'.
- [Fig. 4, §4.2] Fig. 4's right y-axis is labeled 'Total tCO2e' but the absolute values are not clearly tied to the stated 3-year replacement cycle; please state whether the laptop/smartphone totals include the device plus charger or only the device.
Circularity Check
No significant circularity: the 4–8x carbon-reduction figure is an explicit arithmetic consequence of measured energy efficiencies, published carbon intensities, and stated device-count assumptions; no fitted parameter is renamed as a prediction and no load-bearing self-citation appears.
full rationale
The paper's central quantitative claim (4–8x net carbon reduction in §1 and §4.2) is an accounting identity: with edge devices assumed already owned (embodied carbon sunk), the additional footprint of 8h/day training is computed from experimentally measured device power (Table 1), published carbon intensities, and a FLOPS-based device-count equivalence (15 laptops/69 smartphones). None of these inputs is fitted to the target 4–8x number; the reduction is the quotient of the estimated cloud-GPU total (7 tCO2e over 3 years) and the computed edge operational addition. The energy-efficiency experiments on OPT-125m/1.3B are empirical inputs, not predictions. The idealized distributed method (footnote 1) is explicitly labeled as an assumption and is used to separate method effects; the paper's own §5 concedes communication and thermal-throttling caveats. There are no load-bearing self-citations: the SOTA DT-FM method [98], WiFi emission values [82], and carbon-intensity sources are external. No fitted parameter is relabeled as a prediction, and no result is defined in terms of the conclusion it supports. The FLOPS-equivalence and no-P2P assumptions are substantive validity risks, but they are not circularity.
Assumptions & free parameters
free parameters (5)
- Additional daily training hours per edge device =
8 hours
- Replacement cycle =
3 years
- Edge network bandwidth =
10 MB/s
- WiFi peak power =
0.5 W
- Grid carbon intensity =
North America and Europe average (2021-23)
assumptions (6)
- domain assumption Raw FLOPS is an adequate proxy for compute capability when comparing a cloud GPU with edge devices.
- ad hoc to paper An idealized distributed training method (no peer-to-peer transmission, each parameter sent once, local gradient aggregation) represents edge training.
- ad hoc to paper Edge devices can be used for an additional 8 hours per day while charging without damaging batteries or degrading user experience.
- domain assumption The carbon footprint of an edge device is incurred anyway by ownership, so only the marginal operational increase should be counted.
- domain assumption An H100 GPU in a data center, with one-eighth of a server's embodied carbon, represents cloud training.
- domain assumption Scaling laws imply exponential compute growth for linear accuracy gains, motivating the centralization problem.
Cite this review
Pith. "Pith review of Towards Decentralized and Sustainable Foundation Model Training with the Edge." pith.science (2026). https://pith.science/paper/RXSSHBEG
@misc{pith2026250701803,
author = {Pith},
title = {Pith review of: Towards Decentralized and Sustainable Foundation Model Training with the Edge},
year = {2026},
howpublished = {\url{https://pith.science/paper/RXSSHBEG}},
note = {Machine review of arXiv:2507.01803}
}
read the original abstract
Foundation models are at the forefront of AI research, appealing for their ability to learn from vast datasets and cater to diverse tasks. Yet, their significant computational demands raise issues of environmental impact and the risk of centralized control in their development. We put forward a vision towards decentralized and sustainable foundation model training that leverages the collective compute of sparingly used connected edge AI devices. We present the rationale behind our vision, particularly in support of its sustainability benefit. We further outline a set of challenges that need to be addressed to turn this vision into reality.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Bilge Acun, Benjamin Lee, Fiodar Kazhamiaka, Kiwan Maeng, Udit Gupta, Manoj Chakkaravarthy, David Brooks, and Carole-Jean Wu. 2023. Carbon Explorer: A Holistic Framework for Designing Carbon Aware Datacenters. In ASPLOS (2). ACM, 118–132
2023
-
[2]
Tersiteab Adem, Andrew McCrabb, Vidushi Goyal, and Valeria Bertacco. 2024. Evergreen: Comprehensive Carbon Model for Performance-Emission Tradeoffs. In IISWC. IEEE, 132–143
2024
-
[3]
Aioz. 2025. DePIN for Web3, Empowering a Fast, Secure and Decentralized Future. https://aioz.network/
2025
-
[4]
Anderson
David P. Anderson. 2020. BOINC: A Platform for Volunteer Computing. J. Grid Comput. 18, 1 (2020), 99–122
2020
-
[5]
Anderson, Carl Christensen, and Bruce Allen
David P. Anderson, Carl Christensen, and Bruce Allen. 2006. Grid resource management – Designing a runtime system for volunteer computing. In SC. ACM Press, 126
2006
-
[6]
Anderson, Jeff Cobb, Eric Korpela, Matt Lebofsky, and Dan Werthimer
David P. Anderson, Jeff Cobb, Eric Korpela, Matt Lebofsky, and Dan Werthimer
-
[7]
Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan
Lasse F. Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan. 2020. Carbontracker: Tracking and Predicting the Carbon Footprint of Training Deep Learning Models. arXiv:2007.03051
arXiv 2020
-
[8]
Apple. [n. d.]. Apple unveils M2 Pro and M2 Max: next-generation chips for next- level workflows. https://www.apple.com/uk/newsroom/2023/01/apple-unveils- m2-pro-and-m2-max-next-generation-chips-for-next-level-workflows/
2023
Show all 113 references
-
[9]
Apple. 2023. Product Environmental Report 16-inch MacBook Pro. https: //www.apple.com/environment/pdf/products/notebooks/16-inch_MacBook_P ro_PER_Oct2023.pdf
2023
-
[10]
Apple. 2023. Product Environmental Report iPhone 15 Pro and iPhone 15 Pro Max. https://www.apple.com/environment/pdf/products/iphone/iPhone_15_Pr o_and_iPhone_15_Pro_Max_Sept2023.pdf
2023
-
[11]
Arslan, Indrajeet Singh, Shailendra Singh, Harsha V
Mustafa Y. Arslan, Indrajeet Singh, Shailendra Singh, Harsha V. Madhyastha, Karthikeyan Sundaresan, and Srikanth V. Krishnamurthy. 2012. Computing while charging: building a distributed computing infrastructure using smartphones. In CoNEXT. ACM, 193–204
2012
-
[12]
Backlinko. [n. d.]. Smartphone Usage Statistics. https://backlinko.com/smartp hone-usage-statistics
-
[13]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Ming- Hsuan Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng,...
2025 arXiv
-
[14]
Bartoldson, Bhavya Kailkhura, and Davis W
Brian R. Bartoldson, Bhavya Kailkhura, and Davis W. Blalock. 2023. Compute- Efficient Deep Learning: Algorithmic Trends and Opportunities. J. Mach. Learn. Res. 24 (2023), 122:1–122:77
2023
-
[15]
Giovanni Bartolomeo, Mehdi Yosofie, Simon Bäurle, Oliver Haluszczynski, Nitin- der Mohan, and Jörg Ott. 2023. Oakestra: A Lightweight Hierarchical Orchestra- tion Framework for Edge Computing. In ATC. USENIX Association
2023
-
[16]
Hudson, Ehsan Adeli, Russ B
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Cr...
2021 arXiv
-
[17]
Nicholas J Borge. 2022. Deep pockets: The economics of deep learning and the emergence of new AI platforms . Ph. D. Dissertation. Massachusetts Institute of Technology
2022
-
[18]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[19]
Qingqing Cao, Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian, and Niranjan Balasubramanian. 2021. IrEne: Interpretable Energy Prediction for Transformers. In ACL/IJCNLP (1). Association for Computational Linguistics, 2145–2157
2021
-
[20]
Carbon Footprint. [n. d.]. COUNTRY SPECIFIC ELECTRICITY GRID GREEN- HOUSE GAS EMISSION FACTORS. https://www.carbonfootprint.com/
-
[21]
Cheng-Wei Ching, Xin Chen, Taehwan Kim, Bo Ji, Qingyang Wang, Dilma Da Silva, and Liting Hu. 2024. Totoro: A Scalable Federated Learning Engine for the Edge. In EuroSys. ACM, 182–199
2024
-
[22]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vi...
2023
-
[23]
Shaiful Alam Chowdhury, Stephanie Borle, Stephen Romansky, and Abram Hindle. 2019. GreenScaler: training software energy models with automatic test generation. Empir. Softw. Eng. 24, 4 (2019), 1649–1692
2019
-
[24]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. In NeurIPS
2023
-
[25]
Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, Denis Mazur, Ilia Kobelev, Yacine Jernite, Thomas Wolf, and Gennady Pekhimenko
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Quentin Lhoest, Anton Sinitsin, Dmitry Popov, Dmitry V. Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, Denis Mazur, Ilia Kobelev, Yacine Jernite, Thomas Wolf, and Gennady Pekhimenko. 20...
2021
-
[26]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdh- ery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussai...
2023
-
[27]
The Economist. [n. d.]. Artificial intelligence enters its industrial age. https: //www.economist.com/foundational-AI-pod
-
[28]
Bello-Maldonado, Bishwaranjan Bhattacharjee, Carlos H
Tamar Eilam, Pedro D. Bello-Maldonado, Bishwaranjan Bhattacharjee, Carlos H. A. Costa, Eun Kyung Lee, and Asser N. Tantawi. 2023. Towards a Methodology and Framework for AI Sustainability Metrics. In HotCarbon. ACM, 13:1–13:7
2023
-
[29]
Sannara Ek, François Portet, Philippe Lalanda, and Germán Vega. 2021. A Federated Learning Aggregation Algorithm for Pervasive Computing: Evaluation and Comparison. In PerCom. IEEE, 1–10
2021
-
[30]
Ahmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Osi, Prateek Sharma, Fan Chen, and Lei Jiang. 2024. LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models. In ICLR. OpenReview.net
2024
-
[31]
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, et al . 2022. Towards artificial general intelligence via a multimodal foundation model.Nature Communications 13, 1 (2022), 3094
2022
-
[32]
Firefox. 2025. Firefox Source Tree Documentation. https://firefox-source- docs.mozilla.org/performance/tools_power_rapl.html
2025
-
[33]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323
2022 arXiv
-
[34]
GitHub. [n. d.]. GitHub Copilot · Your AI pair programmer. https://github.com /features/copilot
-
[35]
Greenhouse Gas Protocol. [n. d.]. The GHG Protocol for Project Accounting. https://ghgprotocol.org/project-protocol
-
[36]
Viktor Urban Gsteiger, Pin Hong (Daniel) Long, Yiran (Jerry) Sun, Parshan Javanrood, and Mohammad Shahrad. 2024. Caribou: Fine-Grained Geospatial Shifting of Serverless Applications for Sustainability. In SOSP. ACM, 403–420
2024
-
[37]
Matt Hamblen. [n. d.]. Update: ChatGPT runs 10K Nvidia training GPUs with potential for thousands more. https://www.fierceelectronics.com/sensors/chatg pt-runs-10k-nvidia-training-gpus-potential-thousands-more
-
[39]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. In ICLR. OpenReview.net
2021
-
[40]
Rae, Oriol Vinyals, and Laurent Sifre
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
2022 arXiv
-
[41]
Intel. [n. d.]. Intel Confidential Computing Solutions. https://www.intel.com/co ntent/www/us/en/security/confidential-computing.html
-
[42]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lam- ple, Lélio Renard Lavaud, Lucile Saulnier, Marie-An...
-
[43]
Harry Jiang, Xiaoxi Zhang, and Carlee Joe-Wong. 2022. DOLL: Distributed OnLine Learning Using Preemptible Cloud Instances. SIGMETRICS Perform. Evaluation Rev. 50, 2 (2022), 21–23
2022
-
[44]
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna ACM SIGENERGY Energy Informatics Review Volume 1 Issue 1, November 2021 Potapenko, et al. 2021. Highly accurate protein struc...
2021
-
[45]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models. arXiv:2001.08361
2020 arXiv
-
[46]
Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, and Bing Liu
-
[47]
Bran Knowles. 2021. ACM TechBrief: Computing and climate change
2021
-
[48]
Eric Korpela, Dan Werthimer, David Anderson, Jeff Cobb, and Matt Leboisky
-
[49]
Manikanta Kotaru. 2023. Adapting Foundation Models for Operator Data Analytics. In Proceedings of the 22nd ACM Workshop on Hot Topics in Net- works, HotNets 2023, Cambridge, MA, USA, November 28-29, 2023 . ACM, 172–179. https://doi.org/10.1145/3626111.3628191
2023
-
[50]
Alper Goksoy, Sumit K
Gokul Krishnan, A. Alper Goksoy, Sumit K. Mandal, Zhenyu Wang, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, and Yu Cao. 2022. Big-Little Chiplets for In-Memory Acceleration of DNNs: A Scalable Heterogeneous Architecture. In ICCAD. ACM, 8:1–8:9
2022
-
[51]
Mandal, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y
Gokul Krishnan, Sumit K. Mandal, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, and Yu Cao. 2020. Interconnect-Aware Area and Energy Optimization for In-Memory Acceleration of DNNs. IEEE Des. Test 37, 6 (2020), 79–87
2020
-
[52]
KubeEdge. 2025. Kubernetes Native Edge Computing Framework. https://kube edge.io/
2025
-
[53]
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres
-
[54]
Ganti, and Vyas Sekar
Franck Le, Mudhakar Srivatsa, Raghu K. Ganti, and Vyas Sekar. 2022. Rethinking data-driven networking with foundation models: challenges and opportunities. In HotNets. ACM, 188–197
2022
-
[55]
Andersen, Jun Woo Park, Alexander J
Mu Li, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J. Shekita, and Bor-Yiing Su. 2014. Scaling Distributed Machine Learning with the Parameter Server. In OSDI. USENIX Association, 583–598
2014
-
[56]
Weijian Liu, Mingzhen Li, Guangming Tan, and Weile Jia. 2025. Mario: Near Zero-cost Activation Checkpointing in Pipeline Parallelism. In PPoPP. ACM, 197–211
2025
-
[57]
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2017. Communication-Efficient Learning of Deep Net- works from Decentralized Data. In AISTATS (Proceedings of Machine Learning Research, Vol. 54). PMLR, 1273–1282
2017
-
[58]
Sarah McQuate. 2023. Q&A: UW researcher discusses just how much energy ChatGPT uses. https://www.washington.edu/news/2023/07/27/how-much- energy-does-chatgpt-use/
2023
-
[59]
Mengistu and Dunren Che
Tessema M. Mengistu and Dunren Che. 2019. Survey and Taxonomy of Volunteer Computing. ACM Comput. Surv. 52, 3 (2019), 59:1–59:35
2019
-
[60]
MLCommons. [n. d.]. MLCube. https://github.com/mlcommons/mlcube
-
[61]
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Ef- ficient large-scale language model training on GPU cl...
2021
-
[62]
NetMind. 2025. AI powerhouse for the future. https://netmind.ai/home
2025
-
[63]
NVIDIA. [n. d.]. NVIDIA Confidential Computing Secure data and AI models in use. https://www.nvidia.com/en-gb/data-center/solutions/confidential- computing
-
[64]
NVIDIA. 2025. NVML API Reference. http://docs.nvidia.com/deploy/nvml- api/index.html
2025
-
[65]
OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774
2023 arXiv
-
[66]
Anderson
Pratyush Patel, Theo Gregersen, and Thomas E. Anderson. 2023. An Agile Pathway Towards Carbon-aware Clouds. In HotCarbon. ACM, 10:1–10:8
2023
-
[67]
Gilbert, Marco Gruteser, Efren Robles, Krishna Sekar, Yong Wei, and Tenghui Zhu
David Patterson, Jeffrey M. Gilbert, Marco Gruteser, Efren Robles, Krishna Sekar, Yong Wei, and Tenghui Zhu. 2024. Energy and Emissions of Machine Learning on Smartphones vs. the Cloud. Commun. ACM 67, 2 (2024), 86–97
2024
-
[68]
Patterson, Joseph Gonzalez, Urs Hölzle, Quoc V
David A. Patterson, Joseph Gonzalez, Urs Hölzle, Quoc V. Le, Chen Liang, Lluis- Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean
-
[69]
Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendric...
2021 arXiv
-
[70]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21 (2020), 140:1–140:67
2020
-
[71]
Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan. 2021. Adaptive Federated Optimization. In ICLR. OpenReview.net
2021
-
[72]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. In CVPR. IEEE, 10674–10685
2022
-
[73]
Max Ryabinin, Tim Dettmers, Michael Diskin, and Alexander Borzunov
-
[74]
Max Ryabinin and Anton Gusev. 2020. Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts. In NeurIPS
2020
-
[75]
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. 2022. Compute Trends Across Three Eras of Machine Learning. In IJCNN. IEEE, 1–8
2022
-
[76]
Mohammad Shahrad and David Wentzlaff. 2017. Towards Deploying Decommis- sioned Mobile Devices as Cheap Energy-Efficient Compute Nodes. In HotCloud. USENIX Association
2017
-
[77]
Tell, Yanqing Zhang, William J
Yakun Sophia Shao, Jason Clemons, Rangharajan Venkatesan, Brian Zimmer, Matthew Fojtik, Nan Jiang, Ben Keller, Alicia Klinefelter, Nathaniel Ross Pinck- ney, Priyanka Raina, Stephen G. Tell, Yanqing Zhang, William J. Dally, Joel S. Emer, C. Thomas Gray, Brucek Khailany, and St...
2019
-
[78]
Craig S. Smith. 2023. What Large Models Cost You – There Is No Free AI Lunch. https://www.forbes.com/sites/craigsmith/2023/09/08/what-large-models-cost- you--there-is-no-free-ai-lunch/?sh=2b6d10724af7
2023
-
[79]
Rachuri, Chenren Xu, and Emmanuel Munguia Tapia
Vijay Srinivasan, Saeed Moghaddam, Abhishek Mukherji, Kiran K. Rachuri, Chenren Xu, and Emmanuel Munguia Tapia. 2014. MobileMiner: mining your frequent patterns on your phone. In UbiComp. ACM, 389–400
2014
-
[80]
In ICML (Proceedings of Machine Learning Research, Vol
SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient. In ICML (Proceedings of Machine Learning Research, Vol. 202). PMLR, 29416–29440
-
[81]
Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020. ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding. In AAAI. AAAI Press, 8968–8975
2020
-
[82]
Jennifer Switzer, Gabriel Marcano, Ryan Kastner, and Pat Pannuto. 2023. Junk- yard Computing: Repurposing Discarded Smartphones to Minimize Carbon. In ASPLOS (2). ACM, 400–412
2023
-
[83]
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, and Guoqing Harry Xu. 2023. Bamboo: Making Pre- emptible Instances Resilient for Affordable Training of Large DNNs. In NSDI. USENIX Association, 497–513
2023
-
[84]
Hugo Touvron, Louis Martin, Kevin Stone, and et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288
2023 arXiv
-
[85]
UK Government. 2024. AI Foundation Models: initial review. https://www.gov. uk/cma-cases/ai-foundation-models-initial-review
2024
-
[86]
Pietzuch
Marcel Wagenländer, Guo Li, Bo Zhao, Luo Mai, and Peter R. Pietzuch. 2024. Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections. In SOSP. ACM, 195–210
2024
-
[87]
Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize from human feedback. arXiv:2009.01325 https://arxiv.org/abs/2009 .01325
2020 arXiv
-
[88]
Jaylen Wang, Udit Gupta, and Akshitha Sriraman. 2023. Peeling Back the Carbon Curtain: Carbon Optimization Challenges in Cloud Computing. In HotCarbon. ACM, 8:1–8:7
2023
-
[89]
Jue Wang, Binhang Yuan, Luka Rimanic, Yongjun He, Tri Dao, Beidi Chen, Christopher Ré, and Ce Zhang. 2022. Fine-tuning Language Models over Slow Networks using Activation Compression with Guarantees. arXiv:2206.01299 ACM SIGENERGY Energy Informatics Review Volume 1 Issue 1, No...
2022 arXiv
-
[90]
Xiaofei Wang, Yiwen Han, Victor C. M. Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. 2020. Convergence of Edge Computing and Deep Learning: A Comprehensive Survey. IEEE Commun. Surv. Tutorials 22, 2 (2020), 869–904
2020
-
[91]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned Language Models are Zero-Shot Learners. In ICLR. OpenReview.net
2022
-
[92]
Alexander Wood, Kayvan Najarian, and Delaram Kahrobaei. 2021. Homomorphic Encryption for Machine Learning in Medicine and Bioinformatics.ACM Comput. Surv. 53, 4 (2021), 70:1–70:35
2021
-
[93]
Lee, Bugra Akyildiz, Maximilian Balandat, Joe Spisak, Ravi Jain, Mike Rabbat, and Kim M
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga Behram, Jinshi Huang, Charles Bai, Michael Gschwind, Anurag Gupta, Myle Ott, Anastasia Melnikov, Salvatore Candido, David Brooks, Geeta Chauhan, Benjamin Lee, Hsien...
2022
-
[94]
Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Connor Holmes, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, and Yuxiong He. 2023. ZeRO++: Extremely Efficient Collective Communication for Giant Model Train- ing. arXiv:2306.10209
2023 arXiv
-
[95]
Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In ICML (Proceedings of Machine Learning Research, Vol. 202). PMLR, 38087–38099
2023
-
[96]
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, and Ion Stoica. 2023. SkyPilot: An Intercloud Broker for Sky Computing. In 20th USENIX Symposium on Networked Systems Design and...
2023
-
[97]
Shengyuan Ye, Liekang Zeng, Xiaowen Chu, Guoliang Xing, and Xu Chen. 2024. Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices. In MobiCom. ACM, 312–326
2024
-
[98]
Binhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang, Tri Dao, Beidi Chen, Percy Liang, Christopher Ré, and Ce Zhang. 2022. Decentralized Training of Foundation Models in Heterogeneous Environments. In NeurIPS
2022
-
[99]
Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, and Ruifeng Xu
-
[100]
Li Zhang, Zhe Fu, Boqing Shi, Xiang Li, Rujin Lai, Chenyang Yang, Ao Zhou, Xiao Ma, Shangguang Wang, and Mengwei Xu. 2024. More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCs. InUSENIX ATC. USENIX Association, 285–302
2024
-
[101]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang,...
2024 arXiv
-
[102]
Yancheng Zhang, Mengxin Zheng, Yuzhang Shang, Xun Chen, and Qian Lou
-
[103]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...
2023 arXiv
-
[104]
Xing, Joseph E
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yan- ping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Automating Inter- and Intra-Operator Par- allelism for Distributed Deep Learning. In OSDI. ...
2022
-
[105]
Barret Zoph and Quoc V. Le. 2017. Neural Architecture Search with Reinforce- ment Learning. In ICLR. OpenReview.net. ACM SIGENERGY Energy Informatics Review Volume 1 Issue 1, November 2021
2017
-
[107]
CPPO: Continual Learning for Reinforcement Learning with Human Feedback. In ICLR. https://openreview.net/forum?id=86zAUE80pP
-
[109]
Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuo- hui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoy...
2022 arXiv
-
[111]
In NeurIPS
HEPrune: Fast Private Training of Deep Neural Networks With Encrypted Data Pruning. In NeurIPS
-
[2001]
Computing in Science & Engineering 3, 1 (2001), 78–83
SETI@home – Massively Distributed Computing for SETI. Computing in Science & Engineering 3, 1 (2001), 78–83
2001
-
[2002]
SETI@home: An Experiment in Public-Resource Computing. Commun. ACM SIGENERGY Energy Informatics Review Volume 1 Issue 1, November 2021 ACM 45, 11 (2002), 56–61
2002
- [2019]
-
[2022]
Computer 55, 7 (2022), 18–28
The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. Computer 55, 7 (2022), 18–28
2022
-
[2023]
Continual Pre-training of Language Models. In ICLR. OpenReview.net
- [2024]
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.