REVIEW 4 major objections 5 minor 27 references
Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Mass adoption of diffusion-based AI art systems could add about 1.92 TWh per year for Stable Diffusion alone and 9.29 TWh per year across five major systems, comparable to the electricity use of small countries.
desk verdict A useful early call to study inference-phase energy in consumer AI art, but the headline TWh estimates are inflated by a usage assumption that contradicts the paper's own aggregate user data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a transparent bottom-up energy model: yearly energy equals peak graphics-processor power times hours of peak draw per user per day times the number of users times 365 days. Each factor is an explicit assumption rather than a hidden parameter. The per-image figure comes from an observed generation time of 20 seconds for a 1024x1024 image at 50 diffusion steps on a high-end consumer GPU, yielding 350 watt-hours divided by 180 images per hour, or about 1.94 Wh per image.
What would settle it
Measure real average GPU-hours per day among a representative sample of Stable Diffusion users, or obtain server-side telemetry from providers. If the observed mean is 0.1 hours per day rather than 1.5, the annual estimate for Stable Diffusion falls from about 1.92 TWh to about 0.13 TWh; a second check is to sample actual inference jobs across varied GPUs to test the 1.94 Wh-per-image figure.
Extended reading notes
Core claim
The central claim is a first-order estimate that moves attention from training cost to inference cost in consumer generative AI. Under the paper's explicit assumptions—350 W peak GPU draw, 1.5 hours of peak use per user per day, 10 million Stable Diffusion users—the arithmetic gives 1.92 TWh per year for Stable Diffusion alone; carrying the same per-user consumption to 48.5 million users across five diffusion-based art platforms gives 9.29 TWh per year. The authors present these as minimum-style estimates, not precise measurements, and note that open-source models like Stable Diffusion make the true user count hard to pin down because anyone can run the model locally.
Load-bearing premise
The calculation depends on the assumption that the average user runs a high-end GPU at peak power for 1.5 hours per day, a figure drawn from 42 self-selected survey respondents in six days; if the true average is only a few minutes per day, the yearly estimates would drop by roughly an order of magnitude.
Editorial extensions
If this is right
- If these numbers hold, everyday inference on diffusion art tools draws energy on the scale of small countries, broadening climate scrutiny from model training to mass consumer use.
- At roughly 1.94 Wh per image, the paper's observed power users—who generate hundreds or thousands of images per day, often by automated script—consume household-scale electricity through image generation alone.
- Providers' reported user counts and image totals become load-bearing climate data; without daily-active-user counts and hardware information, outside estimates of this technology's footprint remain rough.
- Because most surveyed users generate images for themselves and a large share never share them, much of the inference energy goes into outputs that are stored and discarded rather than actually consumed.
- The same trajectory that moves these tools from still images to video, animation, and 3D environments would multiply per-output compute, making the reported TWh figures a lower bound for the technology's growth path.
Reading between the lines
- The 1.5 hours-per-day usage assumption probably overstates the average user: the estimate rests on 42 self-selected respondents from AI-art Facebook communities, so if typical use is minutes per day rather than hours, the annual TWh figures fall by an order of magnitude or more.
- Climate harm depends on where the electricity is generated, so the same TWh consumed on a coal-heavy grid emits far more than on a renewable one; translating these energy numbers into emissions would require user-location data the paper does not have.
- The model is directly testable as a monitoring tool: if providers published aggregate GPU-hours per image, the per-image energy of about 1.94 Wh could be verified and tracked over time.
- The digital-waste framing suggests a testable behavioral lever: interfaces that discourage endless unshared iterations, such as previewing variations in a single batch rather than generating full images, could reduce inference energy without reducing user satisfaction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the rapid mass adoption of diffusion-based visual AI systems (Stable Diffusion, DALL-E 2, Midjourney, Dream, Lensa AI) has climate implications from inference-phase energy use. The authors provide a literature review, compile public data on users and daily image output in Table 1, and present calculations estimating 1.92 TWh/year for Stable Diffusion alone and 9.29 TWh/year for five systems. These estimates are based on a 350 W peak power draw for an RTX 3090, an assumed 1.5 h/day of peak-power use per user (derived from a 42-person Facebook survey in AI art communities), and reported user counts. The paper ends with recommendations for users, developers, and researchers, and discusses digital waste, overconsumption, and future research directions.
Significance. The topic is timely and the paper usefully draws attention to inference-phase energy use in consumer-facing generative AI, an area less studied than model training. The authors are transparent about their assumptions and about the lack of public data, and they compile useful system-level usage statistics. However, the quantitative centerpiece is not reliable: the per-user usage assumption is contradicted by the paper's own aggregate data in Table 1, and the headline TWh figures are therefore unsupported. The paper's main value at present is its qualitative call for transparency, measurement, and research on AI inference energy use, not its specific energy estimates.
major comments (4)
- [Energy Consumption Calculations (Assumption 2, Eqs. 1-3)] The assumed per-user usage of 1.5 h/day (about 285 images/day) is contradicted by the paper's own Table 1. DALL-E 2 has 3 million users producing 4 million images/day (about 1.33 images/user/day), and Dream has 10-100 million users producing 1.64 million images/day (about 0.016-0.164 images/user/day). At the paper's own energy per image of 1.94 Wh (Eq. 5), these aggregate ratios imply about 2.6 Wh/user/day for DALL-E 2, not 525 Wh/user/day (Eq. 1) - a discrepancy of roughly a factor of 200. Because Eqs. (1)-(3) and (6)-(8) all inherit this assumption, the headline estimates of 1.92 TWh and 9.29 TWh are not internally consistent with the data the paper itself reports.
- [Extrapolation to Other Systems (Eqs. 6-8)] The calculation treats 48.5 million 'total users' as if they were daily active users. The paper's Limitations section correctly notes that server members and app downloads differ from daily active users, but the arithmetic in Eqs. (6)-(8) does not account for this. For example, Midjourney's 12 million figure is based on Discord server members, and Dream's 10-100 million is based on downloads. Using a plausible daily-active-user fraction (e.g., 10-20%) would lower the five-system total by a corresponding factor. The estimate should be presented as a scenario with an explicit DAU assumption, or DAU data should be sought.
- [Assumption 1: Hardware and Eq. (1)] The calculation assumes each user runs the system at 350 W peak power draw for the full 1.5 h/day. The text acknowledges that 350 W is the peak Total Graphics Power and that it was observed for a specific 1024x1024, 50-step generation, but the arithmetic applies this peak value to the entire usage period. Real usage includes prompt entry, image viewing, and idle time, and the actual average draw would be lower. This assumption inflates every downstream figure, and the absence of a sensitivity analysis over power draw and usage duration leaves the robustness of the results unexamined.
- [Assumption 2: Duration of Use] The 1.5 h/day figure rests on a 6-day survey with 42 self-selected respondents recruited from AI-art Facebook communities. The paper itself characterizes these respondents as heavy users: most generate over 1000 images per week, and a substantial portion use automated scripts. Extrapolating 285 images/day to the entire user bases of services like Dream or Lensa - many of whose users are casual mobile users - is not statistically justified. The later claim that the estimates are a 'vast underestimation' refers to excluded cloud and embodied costs, so it does not address this likely overestimation of per-user usage.
minor comments (5)
- [Assumption 2 and Eq. (1)] The text derives approximately 95 minutes/day, but Eq. (1) uses 1.5 h; using 95 minutes would give about 0.554 kWh/day and change Eq. (3) to about 2.02 TWh/year. Please align the stated duration with the arithmetic.
- [Table 1] Table 1 would be much easier to read with explicit units and clearer footnotes; for example, the Stable Diffusion row's '10' is described in the text as daily users, but the footnote marker is ambiguous, and Dream's '10-100' million should be flagged as total downloads rather than active users.
- [Throughout] Spelling and naming are inconsistent: 'DallE-2' appears in Table 1 while 'DALL-E 2' is used in the text, and 'Midjournery' appears several times instead of 'Midjourney'. Please standardize.
- [Comparison to Other Digital Technologies] The value '91.31 TWh per year' for Bitcoin is time-sensitive; the sentence should include the date of the reading (e.g., February 2023) and ideally the Bitcoin Energy Consumption Index source with the exact retrieval date.
- [References] Several statistics are cited to blog posts and news articles without stable identifiers; adding DOIs or access dates would improve reproducibility, especially for the Statista and EIA data.
Circularity Check
No significant circularity: the energy estimates are explicit arithmetic from clearly stated assumptions, not derivations that embed their conclusions.
full rationale
The paper's central estimates (1.92 TWh/year for Stable Diffusion and 9.29 TWh/year for five systems) are transparent products of stated assumptions: 350 W peak power draw for an RTX 3090, a 20-second generation time per 1024x1024 image, an assumed 1.5 hours/day of peak-power use per user, and user counts from public reports. Each step is labeled as an assumption (e.g., Assumption 1, Assumption 2) and the paper explicitly calls for better data and lists limitations, including the difference between downloads and daily active users and uncertainty about cloud versus home hardware. No quantity is defined in terms of the target estimate, no parameter is fitted to the predicted outcome, and there is no load-bearing self-citation or imported uniqueness theorem. The concern raised by the skeptic — that the 285 images/day assumption conflicts with the paper's own Table 1 ratios, such as Dall-E2's ~1.3 images/user/day — is a substantive validity and data-quality criticism, but it is not circularity: the conclusion does not reduce to its inputs by construction; rather, one input may be poorly chosen. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- average daily usage time per user =
1.5 hours
assumptions (3)
- domain assumption GPU operates at peak power draw (350 W TGP for RTX 3090) for the entire usage duration
- ad hoc to paper All 48.5 million counted users are daily active users in the extrapolation
- ad hoc to paper Survey respondents are representative of the broader user population
Cite this review
Pith. "Pith review of Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption." pith.science (2026). https://pith.science/paper/HYL4QZLT
@misc{pith2026250518892,
author = {Pith},
title = {Pith review of: Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption},
year = {2026},
howpublished = {\url{https://pith.science/paper/HYL4QZLT}},
note = {Machine review of arXiv:2505.18892}
}
read the original abstract
Climate implications of rapidly developing digital technologies, such as blockchains and the associated crypto mining and NFT minting, have been well documented and their massive GPU energy use has been identified as a cause for concern. However, we postulate that due to their more mainstream consumer appeal, the GPU use of text-prompt based diffusion AI art systems also requires thoughtful considerations. Given the recent explosion in the number of highly sophisticated generative art systems and their rapid adoption by consumers and creative professionals, the impact of these systems on the climate needs to be carefully considered. In this work, we report on the growth of diffusion-based visual AI systems, their patterns of use, growth and the implications on the climate. Our estimates show that the mass adoption of these tools potentially contributes considerably to global energy consumption. We end this paper with our thoughts on solutions and future areas of inquiry as well as associated difficulties, including the lack of publicly available data.
Reference graph
Works this paper leans on
- [4]
-
[5]
Lensa AI, the app making ‘magic avatars,’ raises red flags for artists. TechCrunch. Proceedings of the 14th International Conference on Computational Creativity (ICCC’23) ISBN: 978-989-54160-5-9 271 https://techcrunch.com/2022/12/05/lensa-ai-app-store- magic-avatars-artists/ Huggingface. Stable Diffusion v1-4 Model Card. Hugging Face. https://huggingface....
work page 2022
- [10]
-
[12]
AI art apps are cluttering the app store’s top charts following Lensa AI’s success. TechCrunch. https://techcrunch.com/2022/12/12/ai-art- apps-are-cluttering-the-app-stores-top-charts-following- lensa-ais-success/ Perkins, D.N.; Drisse, M.B.; Nxele, T.; and Sly, P.D
work page 2022
-
[15]
Frontiers in Public Health9: https://doi.org/10.3389/fpubh.2021.641673 Mostaque, E
On the psychol- ogy of TikTok use: A first glimpse from empirical findings. Frontiers in Public Health9: https://doi.org/10.3389/fpubh.2021.641673 Mostaque, E. [@EMostaque]. 2022, Aug
-
[16]
https://twitter.com/Sta- bilityAI/status/1626164348220579840?cxt=HHwW- gIC9ofz2pZEtAAAA StabilityAI
We’re glad to be partnering with KrikeyApp [Tweet]. https://twitter.com/Sta- bilityAI/status/1626164348220579840?cxt=HHwW- gIC9ofz2pZEtAAAA StabilityAI. General FAQ. StabilityAI. https://stabil- ity.ai/faq StabilityAI. Stable Diffusion Launch Announcement. Sta- bilityAI. https://stability.ai/blog/stable-diffusion-announce- ment Strubell, E.; Ganesh, A.; a...
- [17]
- [18]
Show all 27 references
-
[19]
International Association of Online Engineering
Watch, share or create: The influence of personality traits and user motivation on Tik- Tok mobile video usage. International Association of Online Engineering. https://www.learntechlib.org/p/216454/ Pennington, A. 2022, Nov
2022
-
[22]
New Midjour- ney beta using Stable Diffusion looking absolutely awesome [Tweet]. Twitter. https://twitter.com/EMostaque/sta- tus/1561917541743841280?lang=en Moutaib, M.; Fattah, M.; and Farhaoui, Y
- [23]
-
[24]
Circular Economy1:100011 StabilityAI
Disentangling the worldwide web of e-waste and climate change co-benefits. Circular Economy1:100011 StabilityAI. 2022, Oct
2022
-
[25]
StabilityAI
Stability AI Announces $101 Million in Funding for Open-Source Artificial Intelligence. StabilityAI. https://stability.ai/blog/stability-ai-announces- 101-million-in-funding-for-open-source-artificial-intelli- gence StabilityAI. [@StabilityAI]. 2023, Feb
2023
-
[26]
arXiv:1906.02243v1 Walton, J
Energy and policy considerations for deep learning in NLP. arXiv:1906.02243v1 Walton, J. 2023, Jan
1906 arXiv
-
[27]
Statista
Lensa AI app downloads worldwide [Infographic]. Statista. https://www.statista.com/statis- tics/1350961/lensa-ai-app-downloads-worldwide/ Ceci, L. 2023b, Jan
-
[29]
NABAmplify
You might actually find a muse in the machine (or Midjourney). NABAmplify. https://amplify.nabshow.com/articles/ic-the-muse-in-the- machine/ Perez, S. 2022, Dec
2022
-
[30]
https://blog.google/prod- ucts/google-play/google-plays-best-apps-and-games-of- 2022/ Luccioni, A.S
Google Play’s best apps and games of 2022.Google - The Keyword. https://blog.google/prod- ucts/google-play/google-plays-best-apps-and-games-of- 2022/ Luccioni, A.S. and Hernandez-Garcia, A
2022
-
[1973]
2022, Nov
Uses and gratification theory.Public Opinion Quarterly37:509-523 Kelly, K. 2022, Nov
2022
-
[2000]
2022, Nov
What can be done to reduce overconsumption?Ecological Economics32(1):27- 41 Burnes, A. 2022, Nov
2022
-
[2005]
2021, Nov
Global perspectives on e-waste.Environmental Impact Assessment Review 25(5):436-458 Wombo.AI [@wombo.ai]. 2021, Nov
2021
-
[2011]
Devel- opmental neurotoxicants in e-waste: An emerging health concern.Environmental Health Perspectives119(4): https://doi.org/10.1289/ehp.1002452 deVries, A
-
[2014]
2023, Feb
E-waste: A global hazard.Annals of Global Health 80(4):286-295 PRNewswire. 2023, Feb
2023
-
[2019]
arXiv:1910.09700v2 Lewis, N
Quantifying the carbon emissions of machine learn- ing. arXiv:1910.09700v2 Lewis, N. 2022, Dec
1910 arXiv
-
[2020]
arXiv:2102.02622 Hatmaker, T
The climate impact of ICT: A review of estimates, trends and regulations. arXiv:2102.02622 Hatmaker, T. 2022, Dec
2022 arXiv
-
[2021]
and Vatanparast, R
Artificial Intelligence to improve the food and agriculture sector.Journal of Food Quality.https://doi.org/10.1155/2021/5584754 Bietti, E. and Vatanparast, R
2021 doi
-
[2022]
International Journal of Information Management 63:102456 Energy Information Administration
Climate change and the COP26: Are digital technologies and information management part of the prob- lem or the solution? An editorial reflection and call to action. International Journal of Information Management 63:102456 Energy Information Administration. International – Ele...
2022
-
[2023]
arXiv:2302.08476v1 Montag, C.; Yang, H.; and Elhai, J.D
Counting carbon: A survey of factors influencing the emissions of ma- chine learning. arXiv:2302.08476v1 Montag, C.; Yang, H.; and Elhai, J.D
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.