Pith. sign in

REVIEW 6 cited by

Holistically Evaluating the Environmental Impact of Creating Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05804 v1 pith:4TLB775F submitted 2025-03-03 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords impactmodelenvironmentaldevelopmentmodelspowertrainingdevelopers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released 493 metric tons of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed 2.769 million liters of water, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to ~50% of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between ~15% and ~85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

    cs.MM 2026-07 conditional novelty 6.5 of 10

    Energy of text-to-video diffusion models is predicted from architectural first principles and observable generation parameters with under 3% MAPE, without needing weights or model size.

  2. How Low Can We Go? Minimum Spectroscopic Requirements For Supernova Subtype Classification

    astro-ph.IM 2026-07 accept novelty 6.0 of 10

    ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.

  3. Misinformation by Omission: The Need for More Environmental Transparency in AI

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Environmental disclosure for notable AI models peaked in 2022 and then declined, and out-of-context energy and emissions estimates now dominate media coverage.

  4. HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation

    cs.SD 2026-04 unverdicted novelty 5.0 of 10

    HAFM uses a hierarchical autoregressive model with dual-rate HuBERT and EnCodec tokens to generate coherent instrumental music from vocals, achieving FAD 2.08 on MUSDB18 while matching prior systems with fewer parameters.

  5. Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming

    cs.SE 2025-08 reject novelty 4.0 of 10

    On 150 LeetCode problems, GPT-4.0 and DeepSeek-Reasoner beat three 3B-parameter models on correctness and speed; the 52% energy-efficiency claim counts any of three SLMs on correct outputs, not a per-model advantage.

  6. The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations

    cs.AI 2025-07 reject novelty 4.0 of 10

    In a public LLM comparison arena, showing users that the larger model consumes more energy caused about 46% of users who preferred it to say they would switch to the smaller model.

Pith tools