Pith. sign in

REVIEW 4 major objections 5 minor 27 references

Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Mass adoption of diffusion-based AI art systems could add about 1.92 TWh per year for Stable Diffusion alone and 9.29 TWh per year across five major systems, comparable to the electricity use of small countries.

desk verdict A useful early call to study inference-phase energy in consumer AI art, but the headline TWh estimates are inflated by a usage assumption that contradicts the paper's own aggregate user data. read the letter →

arxiv 2505.18892 v1 pith:HYL4QZLT submitted 2025-05-24 cs.CY cs.AI

classification cs.CYcs.AI
keywords diffusionmodelsgenerativeAIartinferenceenergyclimateimpactconsumptionGPUusagedigitalwastemassadoption
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the energy used while people run diffusion-based AI image generators, not only the energy spent training them, needs to be part of climate accounting for AI. Combining the peak power draw of a high-end GPU, a survey-derived estimate of 1.5 hours of daily use per user, and reported user counts, it calculates that Stable Diffusion alone could consume about 1.92 TWh per year and five popular diffusion art systems together about 9.29 TWh per year. Those amounts are comparable to the annual electricity consumption of entire small countries, which is the paper's way of showing that consumer-scale image generation can have country-scale energy demand. The paper also finds that most surveyed users create images mainly for themselves and often never share them, and it frames this as digital overconsumption and waste. It closes by calling for provider transparency on daily active users and hardware, and for more research on the psychological pull that keeps people generating.

What carries the argument

The carrying mechanism is a transparent bottom-up energy model: yearly energy equals peak graphics-processor power times hours of peak draw per user per day times the number of users times 365 days. Each factor is an explicit assumption rather than a hidden parameter. The per-image figure comes from an observed generation time of 20 seconds for a 1024x1024 image at 50 diffusion steps on a high-end consumer GPU, yielding 350 watt-hours divided by 180 images per hour, or about 1.94 Wh per image.

What would settle it

Measure real average GPU-hours per day among a representative sample of Stable Diffusion users, or obtain server-side telemetry from providers. If the observed mean is 0.1 hours per day rather than 1.5, the annual estimate for Stable Diffusion falls from about 1.92 TWh to about 0.13 TWh; a second check is to sample actual inference jobs across varied GPUs to test the 1.94 Wh-per-image figure.

Watch

Extended reading notes

Core claim

The central claim is a first-order estimate that moves attention from training cost to inference cost in consumer generative AI. Under the paper's explicit assumptions—350 W peak GPU draw, 1.5 hours of peak use per user per day, 10 million Stable Diffusion users—the arithmetic gives 1.92 TWh per year for Stable Diffusion alone; carrying the same per-user consumption to 48.5 million users across five diffusion-based art platforms gives 9.29 TWh per year. The authors present these as minimum-style estimates, not precise measurements, and note that open-source models like Stable Diffusion make the true user count hard to pin down because anyone can run the model locally.

Load-bearing premise

The calculation depends on the assumption that the average user runs a high-end GPU at peak power for 1.5 hours per day, a figure drawn from 42 self-selected survey respondents in six days; if the true average is only a few minutes per day, the yearly estimates would drop by roughly an order of magnitude.

Editorial extensions

If this is right

  • If these numbers hold, everyday inference on diffusion art tools draws energy on the scale of small countries, broadening climate scrutiny from model training to mass consumer use.
  • At roughly 1.94 Wh per image, the paper's observed power users—who generate hundreds or thousands of images per day, often by automated script—consume household-scale electricity through image generation alone.
  • Providers' reported user counts and image totals become load-bearing climate data; without daily-active-user counts and hardware information, outside estimates of this technology's footprint remain rough.
  • Because most surveyed users generate images for themselves and a large share never share them, much of the inference energy goes into outputs that are stored and discarded rather than actually consumed.
  • The same trajectory that moves these tools from still images to video, animation, and 3D environments would multiply per-output compute, making the reported TWh figures a lower bound for the technology's growth path.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 1.5 hours-per-day usage assumption probably overstates the average user: the estimate rests on 42 self-selected respondents from AI-art Facebook communities, so if typical use is minutes per day rather than hours, the annual TWh figures fall by an order of magnitude or more.
  • Climate harm depends on where the electricity is generated, so the same TWh consumed on a coal-heavy grid emits far more than on a renewable one; translating these energy numbers into emissions would require user-location data the paper does not have.
  • The model is directly testable as a monitoring tool: if providers published aggregate GPU-hours per image, the per-image energy of about 1.94 Wh could be verified and tracked over time.
  • The digital-waste framing suggests a testable behavioral lever: interfaces that discourage endless unshared iterations, such as previewing variations in a single batch rather than generating full images, could reduce inference energy without reducing user satisfaction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This position paper argues that the rapid mass adoption of diffusion-based visual AI systems (Stable Diffusion, DALL-E 2, Midjourney, Dream, Lensa AI) has climate implications from inference-phase energy use. The authors provide a literature review, compile public data on users and daily image output in Table 1, and present calculations estimating 1.92 TWh/year for Stable Diffusion alone and 9.29 TWh/year for five systems. These estimates are based on a 350 W peak power draw for an RTX 3090, an assumed 1.5 h/day of peak-power use per user (derived from a 42-person Facebook survey in AI art communities), and reported user counts. The paper ends with recommendations for users, developers, and researchers, and discusses digital waste, overconsumption, and future research directions.

Significance. The topic is timely and the paper usefully draws attention to inference-phase energy use in consumer-facing generative AI, an area less studied than model training. The authors are transparent about their assumptions and about the lack of public data, and they compile useful system-level usage statistics. However, the quantitative centerpiece is not reliable: the per-user usage assumption is contradicted by the paper's own aggregate data in Table 1, and the headline TWh figures are therefore unsupported. The paper's main value at present is its qualitative call for transparency, measurement, and research on AI inference energy use, not its specific energy estimates.

major comments (4)
  1. [Energy Consumption Calculations (Assumption 2, Eqs. 1-3)] The assumed per-user usage of 1.5 h/day (about 285 images/day) is contradicted by the paper's own Table 1. DALL-E 2 has 3 million users producing 4 million images/day (about 1.33 images/user/day), and Dream has 10-100 million users producing 1.64 million images/day (about 0.016-0.164 images/user/day). At the paper's own energy per image of 1.94 Wh (Eq. 5), these aggregate ratios imply about 2.6 Wh/user/day for DALL-E 2, not 525 Wh/user/day (Eq. 1) - a discrepancy of roughly a factor of 200. Because Eqs. (1)-(3) and (6)-(8) all inherit this assumption, the headline estimates of 1.92 TWh and 9.29 TWh are not internally consistent with the data the paper itself reports.
  2. [Extrapolation to Other Systems (Eqs. 6-8)] The calculation treats 48.5 million 'total users' as if they were daily active users. The paper's Limitations section correctly notes that server members and app downloads differ from daily active users, but the arithmetic in Eqs. (6)-(8) does not account for this. For example, Midjourney's 12 million figure is based on Discord server members, and Dream's 10-100 million is based on downloads. Using a plausible daily-active-user fraction (e.g., 10-20%) would lower the five-system total by a corresponding factor. The estimate should be presented as a scenario with an explicit DAU assumption, or DAU data should be sought.
  3. [Assumption 1: Hardware and Eq. (1)] The calculation assumes each user runs the system at 350 W peak power draw for the full 1.5 h/day. The text acknowledges that 350 W is the peak Total Graphics Power and that it was observed for a specific 1024x1024, 50-step generation, but the arithmetic applies this peak value to the entire usage period. Real usage includes prompt entry, image viewing, and idle time, and the actual average draw would be lower. This assumption inflates every downstream figure, and the absence of a sensitivity analysis over power draw and usage duration leaves the robustness of the results unexamined.
  4. [Assumption 2: Duration of Use] The 1.5 h/day figure rests on a 6-day survey with 42 self-selected respondents recruited from AI-art Facebook communities. The paper itself characterizes these respondents as heavy users: most generate over 1000 images per week, and a substantial portion use automated scripts. Extrapolating 285 images/day to the entire user bases of services like Dream or Lensa - many of whose users are casual mobile users - is not statistically justified. The later claim that the estimates are a 'vast underestimation' refers to excluded cloud and embodied costs, so it does not address this likely overestimation of per-user usage.
minor comments (5)
  1. [Assumption 2 and Eq. (1)] The text derives approximately 95 minutes/day, but Eq. (1) uses 1.5 h; using 95 minutes would give about 0.554 kWh/day and change Eq. (3) to about 2.02 TWh/year. Please align the stated duration with the arithmetic.
  2. [Table 1] Table 1 would be much easier to read with explicit units and clearer footnotes; for example, the Stable Diffusion row's '10' is described in the text as daily users, but the footnote marker is ambiguous, and Dream's '10-100' million should be flagged as total downloads rather than active users.
  3. [Throughout] Spelling and naming are inconsistent: 'DallE-2' appears in Table 1 while 'DALL-E 2' is used in the text, and 'Midjournery' appears several times instead of 'Midjourney'. Please standardize.
  4. [Comparison to Other Digital Technologies] The value '91.31 TWh per year' for Bitcoin is time-sensitive; the sentence should include the date of the reading (e.g., February 2023) and ideally the Bitcoin Energy Consumption Index source with the exact retrieval date.
  5. [References] Several statistics are cited to blog posts and news articles without stable identifiers; adding DOIs or access dates would improve reproducibility, especially for the Statista and EIA data.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the energy estimates are explicit arithmetic from clearly stated assumptions, not derivations that embed their conclusions.

full rationale

The paper's central estimates (1.92 TWh/year for Stable Diffusion and 9.29 TWh/year for five systems) are transparent products of stated assumptions: 350 W peak power draw for an RTX 3090, a 20-second generation time per 1024x1024 image, an assumed 1.5 hours/day of peak-power use per user, and user counts from public reports. Each step is labeled as an assumption (e.g., Assumption 1, Assumption 2) and the paper explicitly calls for better data and lists limitations, including the difference between downloads and daily active users and uncertainty about cloud versus home hardware. No quantity is defined in terms of the target estimate, no parameter is fitted to the predicted outcome, and there is no load-bearing self-citation or imported uniqueness theorem. The concern raised by the skeptic — that the 285 images/day assumption conflicts with the paper's own Table 1 ratios, such as Dall-E2's ~1.3 images/user/day — is a substantive validity and data-quality criticism, but it is not circularity: the conclusion does not reduce to its inputs by construction; rather, one input may be poorly chosen. Accordingly, the circularity score is 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The estimates rest on two load-bearing assumptions: a 1.5h/day average usage time fitted from a small, biased survey, and total user counts treated as daily active users. No invented entities are introduced.

free parameters (1)
  • average daily usage time per user = 1.5 hours
    Derived from a 6-day Facebook survey with 42 respondents; the paper assumes 2000 images/week, translating to about 95 minutes/day, rounded to 1.5 hours. This is the key multiplier for all energy estimates.
assumptions (3)
  • domain assumption GPU operates at peak power draw (350 W TGP for RTX 3090) for the entire usage duration
    Assumption 1: Hardware, the paper states the RTX 3090 draws 350W peak and assumes users run it at that level for 1.5h/day.
  • ad hoc to paper All 48.5 million counted users are daily active users in the extrapolation
    The extrapolation multiplies per-user daily energy by total user counts from downloads/server members, implicitly treating them as daily active users despite the paper acknowledging this differs in a later limitation.
  • ad hoc to paper Survey respondents are representative of the broader user population
    The 42 respondents from AI art Facebook groups are self-selected power users; the paper itself later notes a difference with Midjourney's official polls.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption." pith.science (2026). https://pith.science/paper/HYL4QZLT

@misc{pith2026250518892,
  author       = {Pith},
  title        = {Pith review of: Climate Implications of Diffusion-based Generative Visual AI Systems and their Mass Adoption},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HYL4QZLT}},
  note         = {Machine review of arXiv:2505.18892}
}
read the original abstract

Climate implications of rapidly developing digital technologies, such as blockchains and the associated crypto mining and NFT minting, have been well documented and their massive GPU energy use has been identified as a cause for concern. However, we postulate that due to their more mainstream consumer appeal, the GPU use of text-prompt based diffusion AI art systems also requires thoughtful considerations. Given the recent explosion in the number of highly sophisticated generative art systems and their rapid adoption by consumers and creative professionals, the impact of these systems on the climate needs to be carefully considered. In this work, we report on the growth of diffusion-based visual AI systems, their patterns of use, growth and the implications on the climate. Our estimates show that the mass adoption of these tools potentially contributes considerably to global energy consumption. We end this paper with our thoughts on solutions and future areas of inquiry as well as associated difficulties, including the lack of publicly available data.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 23 canonical work pages

  1. [4]

    Statista

    Lensa AI in-app purchase revenue 2021-2022 [Infographic]. Statista. https://www.sta- tista.com/statistics/1350980/lensa-ai-in-app-revenue- worldwide/ Cervenka, E. 2022, Oct

  2. [5]

    TechCrunch

    Lensa AI, the app making ‘magic avatars,’ raises red flags for artists. TechCrunch. Proceedings of the 14th International Conference on Computational Creativity (ICCC’23) ISBN: 978-989-54160-5-9 271 https://techcrunch.com/2022/12/05/lensa-ai-app-store- magic-avatars-artists/ Huggingface. Stable Diffusion v1-4 Model Card. Hugging Face. https://huggingface....

  3. [10]

    HowToGeek

    Here’s the PC hardware you should buy for Stable Diffusion. HowToGeek. https://www.how- togeek.com/853529/hardware-for-stable-diffusion/ Lim, T. 2022, Nov

  4. [12]

    TechCrunch

    AI art apps are cluttering the app store’s top charts following Lensa AI’s success. TechCrunch. https://techcrunch.com/2022/12/12/ai-art- apps-are-cluttering-the-app-stores-top-charts-following- lensa-ais-success/ Perkins, D.N.; Drisse, M.B.; Nxele, T.; and Sly, P.D

  5. [15]

    Frontiers in Public Health9: https://doi.org/10.3389/fpubh.2021.641673 Mostaque, E

    On the psychol- ogy of TikTok use: A first glimpse from empirical findings. Frontiers in Public Health9: https://doi.org/10.3389/fpubh.2021.641673 Mostaque, E. [@EMostaque]. 2022, Aug

  6. [16]

    https://twitter.com/Sta- bilityAI/status/1626164348220579840?cxt=HHwW- gIC9ofz2pZEtAAAA StabilityAI

    We’re glad to be partnering with KrikeyApp [Tweet]. https://twitter.com/Sta- bilityAI/status/1626164348220579840?cxt=HHwW- gIC9ofz2pZEtAAAA StabilityAI. General FAQ. StabilityAI. https://stabil- ity.ai/faq StabilityAI. Stable Diffusion Launch Announcement. Sta- bilityAI. https://stability.ai/blog/stable-diffusion-announce- ment Strubell, E.; Ganesh, A.; a...

  7. [17]

    Bloomberg

    StabilityAI raises seed round at $1 billion value. Bloomberg. https://www.bloomberg.com/news/articles/2022-10-17/dig- ital-media-firm-stability-ai-raises-funds-at-1-billion-value Freitag, C.; Berners-Lee, M.; Widdicks, K.; Knowles, B.; Blair, G.; and Friday, A

  8. [18]

    2022, Nov

    Internet of things: Energy consumption and data storage.Procedia Computer Science175:609-614 OpenAI. 2022, Nov

Show all 27 references
  1. [19]

    International Association of Online Engineering

    Watch, share or create: The influence of personality traits and user motivation on Tik- Tok mobile video usage. International Association of Online Engineering. https://www.learntechlib.org/p/216454/ Pennington, A. 2022, Nov

  2. [22]

    New Midjour- ney beta using Stable Diffusion looking absolutely awesome [Tweet]. Twitter. https://twitter.com/EMostaque/sta- tus/1561917541743841280?lang=en Moutaib, M.; Fattah, M.; and Farhaoui, Y

  3. [23]

    arXiv:1907.10597v3 Singh, N

    Green AI. arXiv:1907.10597v3 Singh, N. and Ogunseitan, O.A

  4. [24]

    Circular Economy1:100011 StabilityAI

    Disentangling the worldwide web of e-waste and climate change co-benefits. Circular Economy1:100011 StabilityAI. 2022, Oct

  5. [25]

    StabilityAI

    Stability AI Announces $101 Million in Funding for Open-Source Artificial Intelligence. StabilityAI. https://stability.ai/blog/stability-ai-announces- 101-million-in-funding-for-open-source-artificial-intelli- gence StabilityAI. [@StabilityAI]. 2023, Feb

  6. [26]

    arXiv:1906.02243v1 Walton, J

    Energy and policy considerations for deep learning in NLP. arXiv:1906.02243v1 Walton, J. 2023, Jan

  7. [27]

    Statista

    Lensa AI app downloads worldwide [Infographic]. Statista. https://www.statista.com/statis- tics/1350961/lensa-ai-app-downloads-worldwide/ Ceci, L. 2023b, Jan

  8. [29]

    NABAmplify

    You might actually find a muse in the machine (or Midjourney). NABAmplify. https://amplify.nabshow.com/articles/ic-the-muse-in-the- machine/ Perez, S. 2022, Dec

  9. [30]

    https://blog.google/prod- ucts/google-play/google-plays-best-apps-and-games-of- 2022/ Luccioni, A.S

    Google Play’s best apps and games of 2022.Google - The Keyword. https://blog.google/prod- ucts/google-play/google-plays-best-apps-and-games-of- 2022/ Luccioni, A.S. and Hernandez-Garcia, A

  10. [1973]

    2022, Nov

    Uses and gratification theory.Public Opinion Quarterly37:509-523 Kelly, K. 2022, Nov

  11. [2000]

    2022, Nov

    What can be done to reduce overconsumption?Ecological Economics32(1):27- 41 Burnes, A. 2022, Nov

  12. [2005]

    2021, Nov

    Global perspectives on e-waste.Environmental Impact Assessment Review 25(5):436-458 Wombo.AI [@wombo.ai]. 2021, Nov

  13. [2011]

    Devel- opmental neurotoxicants in e-waste: An emerging health concern.Environmental Health Perspectives119(4): https://doi.org/10.1289/ehp.1002452 deVries, A

  14. [2014]

    2023, Feb

    E-waste: A global hazard.Annals of Global Health 80(4):286-295 PRNewswire. 2023, Feb

  15. [2019]

    arXiv:1910.09700v2 Lewis, N

    Quantifying the carbon emissions of machine learn- ing. arXiv:1910.09700v2 Lewis, N. 2022, Dec

  16. [2020]

    arXiv:2102.02622 Hatmaker, T

    The climate impact of ICT: A review of estimates, trends and regulations. arXiv:2102.02622 Hatmaker, T. 2022, Dec

  17. [2021]

    and Vatanparast, R

    Artificial Intelligence to improve the food and agriculture sector.Journal of Food Quality.https://doi.org/10.1155/2021/5584754 Bietti, E. and Vatanparast, R

  18. [2022]

    International Journal of Information Management 63:102456 Energy Information Administration

    Climate change and the COP26: Are digital technologies and information management part of the prob- lem or the solution? An editorial reflection and call to action. International Journal of Information Management 63:102456 Energy Information Administration. International – Ele...

  19. [2023]

    arXiv:2302.08476v1 Montag, C.; Yang, H.; and Elhai, J.D

    Counting carbon: A survey of factors influencing the emissions of ma- chine learning. arXiv:2302.08476v1 Montag, C.; Yang, H.; and Elhai, J.D

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.