Pith. sign in

REVIEW 2 major objections 2 minor 1 cited by

Training the most popular open-source AI models has already emitted about 60,000 metric tons of carbon.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-04 01:53 UTC pith:EZFLDGAM

load-bearing objection The paper supplies a usable first aggregate for popular HF models at 60k tCO2e and releases code, but the tiered regressions on incomplete metadata are the part that needs the most scrutiny. the 2 major comments →

arxiv 2605.01549 v2 pith:EZFLDGAM submitted 2026-05-02 cs.CY

Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale

classification cs.CY
keywords carbon emissionsAI trainingHugging Faceopen-source modelssustainabilityemissions estimationcarbon intensity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper develops a method to estimate aggregate carbon emissions from training open-source AI models on the Hugging Face platform using whatever metadata is available. It applies a tiered estimation process backed by regressions to fill gaps in energy, compute, and emissions data, then introduces AI training carbon intensity as a per-compute efficiency measure. The resulting figure for models exceeding 5,000 downloads is 6.0×10^4 metric tons. A sympathetic reader would care because AI training is expanding quickly yet lacks consistent carbon accounting, so a scalable way to quantify the current footprint sets a concrete baseline for tracking growth and potential policy responses.

Core claim

Using available emissions, energy, compute, and model metadata from Hugging Face, a tiered estimation approach supported by empirical regressions produces an aggregate training emissions total of approximately 6.0×10^4 metric tons for the most popular open-source models, while AI training carbon intensity provides a comparable efficiency metric across trainings.

What carries the argument

The tiered approach to emissions estimation, which uses empirical regressions to derive reliable values from incomplete model metadata on energy use and compute.

Load-bearing premise

The tiered approach supported by empirical regressions accurately estimates emissions from incomplete metadata without introducing large systematic bias.

What would settle it

Direct measurement of actual training emissions for a sample of models whose full metadata later becomes available, then comparison against the tiered estimates for the same models.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Aggregate emissions from the most-downloaded open-source models already stand at roughly 60,000 metric tons.
  • AI training carbon intensity allows direct comparison of sustainability efficiency between different model trainings.
  • The method supplies a practical template for carbon reporting standards that the AI industry could adopt without requiring full reproduction of every training run.
  • Future disclosures with higher completeness would reduce reliance on the regression tier and tighten the overall estimate.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Widespread use of this estimation style could create pressure on developers to publish more complete training metadata as a norm.
  • The 60,000-ton figure provides a reference point for comparing AI training emissions against other large-scale computing activities such as cryptocurrency mining or data-center operations.
  • If model downloads continue to concentrate on a small set of popular checkpoints, total emissions could scale roughly with the number of new high-download releases rather than with every uploaded model.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper introduces a tiered estimation framework to quantify aggregate training carbon emissions of open-source AI models on the Hugging Face platform from unevenly disclosed metadata on energy, compute, and emissions. It uses empirical regressions to impute missing values, defines an AI training carbon intensity (ATCI) metric, and reports that training the most popular models (those with over 5,000 downloads) has produced approximately 6.0×10^4 metric tons of CO2e. Data and code are released publicly.

Significance. If the tiered imputation and regressions are shown to be unbiased, the work supplies a reproducible, scalable method for carbon accounting at the scale of an entire model repository without requiring re-execution of training runs. The open release of data and code is a clear strength that enables independent verification and extension.

major comments (2)
  1. [Methods / Results on aggregate estimate] The aggregate result of 6.0×10^4 tCO2e rests on regressions that impute missing metadata; the manuscript must demonstrate that these regressions were validated on held-out data independent of the fitted parameters, or provide explicit error propagation and sensitivity analysis showing that any circularity does not materially affect the reported total (see abstract and the section describing the tiered approach).
  2. [Abstract and empirical regressions subsection] No validation statistics, cross-validation scores, or uncertainty intervals for the regression-based imputations are visible in the abstract or summary description; without these, the reliability assessment claimed for the tiered method cannot be evaluated.
minor comments (2)
  1. Clarify the exact definition and units of ATCI early in the text and ensure it is used consistently in all figures and tables.
  2. Add a table or appendix listing the number of models per tier and the fraction of metadata imputed at each level.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback emphasizing the need for rigorous validation of our regression-based imputations. We address each major comment below and will revise the manuscript accordingly to strengthen the presentation of reliability metrics.

read point-by-point responses
  1. Referee: [Methods / Results on aggregate estimate] The aggregate result of 6.0×10^4 tCO2e rests on regressions that impute missing metadata; the manuscript must demonstrate that these regressions were validated on held-out data independent of the fitted parameters, or provide explicit error propagation and sensitivity analysis showing that any circularity does not materially affect the reported total (see abstract and the section describing the tiered approach).

    Authors: We agree that the current manuscript does not include explicit held-out validation, cross-validation, or sensitivity analysis for the regressions used in the tiered approach. Although the text describes empirical regressions to assess estimation reliability, these do not encompass the requested validation steps or error propagation. In the revised version we will add a dedicated validation subsection that applies the released code and data to perform held-out testing, reports cross-validation scores, and conducts sensitivity analysis on imputation parameters to quantify impact on the aggregate 6.0×10^4 tCO2e estimate. revision: yes

  2. Referee: [Abstract and empirical regressions subsection] No validation statistics, cross-validation scores, or uncertainty intervals for the regression-based imputations are visible in the abstract or summary description; without these, the reliability assessment claimed for the tiered method cannot be evaluated.

    Authors: We concur that the abstract and summary description omit these specific statistics. We will revise the abstract to include key validation metrics (e.g., cross-validation R² and uncertainty ranges) and expand the empirical regressions subsection with the corresponding scores, intervals, and reliability assessment details. revision: yes

Circularity Check

0 steps flagged

No significant circularity identified

full rationale

The paper's central aggregate estimate of ~6.0×10^4 tCO2e relies on a tiered procedure for incomplete metadata, with empirical regressions used only to assess estimation reliability rather than to generate the reported total by construction. No equations or steps reduce the final numerical result to a fit on the same data being estimated, no self-citations are load-bearing for uniqueness or ansatz, and the derivation supplies code and data for external verification. The method remains self-contained against external benchmarks with no quoted reduction of prediction to input.

Axiom & Free-Parameter Ledger

1 free parameters · 0 axioms · 0 invented entities

The estimate rests on the assumption that regressions fitted to partial disclosures can reliably impute missing values; no independent external benchmarks for the regressions are mentioned in the abstract.

free parameters (1)
  • regression coefficients for tiered metadata imputation
    Empirical regressions used to handle incomplete emissions, energy, and compute disclosures.

pith-pipeline@v0.9.1-grok · 5789 in / 1090 out tokens · 24201 ms · 2026-07-04T01:53:30.947712+00:00 · methodology

0 comments
read the original abstract

The scaling-law era has transformed artificial intelligence (AI) from research into a global industry, but its rapid growth also raises concerns over energy usage, carbon emissions, and environmental sustainability. Unlike traditional sectors, the AI industry still lacks systematic carbon accounting methods that support large-scale estimates without reproducing the original training process. This leaves open questions about how large the problem is today and how large it might be in the near future. Given its central role in hosting open-source AI models, the Hugging Face (HF) platform provides a large-scale and publicly accessible corpus for carbon accounting. We estimate aggregate training emissions of HF open-source models using available emissions, energy, compute, and model metadata. To address uneven disclosure quality, we introduce a tiered approach to handle incomplete metadata, supported by empirical regressions that assess estimation reliability. We further introduce AI training carbon intensity (ATCI, emissions per compute), a metric to assess the sustainability efficiency of model training. Our results show that training the most popular open-source models (with over 5,000 downloads) has already resulted in approximately $6.0\times10^4$ metric tons of carbon emissions. Overall, this paper provides a scalable, empirically grounded framework for estimating training emissions from incomplete disclosures and informing future carbon reporting standards in the AI industry. Data and code are available at https://github.com/insait-institute/HuggingCarbon.

Figures

Figures reproduced from arXiv: 2605.01549 by Jing Qiu, Jinjin Gu, Junhua Zhao, Ruibo Ming, Xinlei Wang.

Figure 1
Figure 1. Figure 1: Estimated training emissions from 5,234 Hugging Face models, compared with equivalent real-world scales (cars (Tiegte et al., 2021), homes (Eurostat, 2025), cement plants (IEA, 2025), and trees (Franklin Jr & Pindyck, 2024); tree absorption = 18 kg CO2/tree/year). efforts, making them a valuable proxy for estimating emis￾sions in practice. By accounting for the training emissions of these models, we seek t… view at source ↗
Figure 2
Figure 2. Figure 2: Scatter plot of estimated training emissions ver￾sus expected FLOPs, with regression fits for different ac￾celerator families (A: NVIDIA A100/A800, H: H100/H800, Others). Both axes are log-scaled. The fitted model is log(Etrain) = −39.25 + 0.85 log(EFregion) + 0.83 log(F total train ) − 0.83 I{H-family} + 0.63 I{Others}. Results indicate ∼0.83 FLOPs elasticity. Relative to the A-family (baseline), the H-fa… view at source ↗
Figure 3
Figure 3. Figure 3: Global accumulative training emissions of AI models (downloads 5,000+). abstracting away from model size or absolute compute cost, and therefore provides a normalized metric to compare across modalities and training paradigms. Overall, the re￾sults highlight that (a) modality matters: vision/multimodal training is more carbon-intensive per compute; (b) lifecy￾cle practice matters: finetuned variants exhibi… view at source ↗
Figure 5
Figure 5. Figure 5: Projected training emissions of HF models at scale from 2024 to 2035. sharply over time. From 2020–2021 to 2024–2025, annual emissions increased from only ∼4.3 × 102 tCO2e to more than 4.1 × 104 tCO2e, reflecting nearly two orders of mag￾nitude growth within five years. The composition of these emissions also shifted substantially. Early periods were dominated by CV and multi-modal models, but NLP ac￾tivit… view at source ↗
Figure 6
Figure 6. Figure 6: Total Amount of Models with Self-Disclosed Emission in Hugging Face. Out of more than 2.1 million repositories, only 2,422 include a structured carbon emissions field (co2 eq emissions) and just 126 mention energy use or emissions in their README files, highlighting a disclosure rate below 0.2%. Curating Models with Reliable Disclosed Emissions: • Total number of repositories in the accounting scope: In Hu… view at source ↗
Figure 7
Figure 7. Figure 7 [PITH_FULL_IMAGE:figures/full_fig_p019_7.png] view at source ↗
Figure 8
Figure 8. Figure 8 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Distribution of Estimated Emissions of Hugging Face Models (5,000+ Downloads) [PITH_FULL_IMAGE:figures/full_fig_p025_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How Low Can We Go? Minimum Spectroscopic Requirements For Supernova Subtype Classification

    astro-ph.IM 2026-07 accept novelty 6.0

    ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.

Reference graph

Works this paper leans on

19 extracted references · 19 canonical work pages · cited by 1 Pith paper

  1. [1]

    Official Hugging Face model cards, including structured fields (e.g., compute used, hardware, carbon emissions), author-provided notes, and embedded configuration snippets

  2. [2]

    These files provide parameter counts, layer depths, hidden sizes, patch sizes, and other FLOPs-relevant attributes

    Repository configuration files, such asconfig.json, tokenizer/vision encoder configs, and architecture descriptors. These files provide parameter counts, layer depths, hidden sizes, patch sizes, and other FLOPs-relevant attributes

  3. [3]

    2k”, “1.2M

    External authoritative sources, including official technical reports, GitHub repositories, and arXiv papers referenced in the model cards. When multiple sources were available, we applied a deterministic priority order (direct disclosure →config-derived→paper-derived→regression-estimated). Table 4.Metadata disclosure sparsity across the Hugging Face model...

  4. [4]

    When available, we directly record the number of trainable parameters (Nparams)

    Model classification and parameter extraction.Each model is classified as encoder-only (e.g., BERT), decoder-only (e.g., GPT, LLaMA), or encoder–decoder (e.g., T5). When available, we directly record the number of trainable parameters (Nparams). If parameters are missing, we infer them from architecture descriptors such as hidden size, number of layers, a...

  5. [5]

    For others, we distinguish: •Full-parameter Fine Tuning (FT):N params is the full count

    Effective parameter count adjustments.For pretraining we set Nparams to the full parameter count. For others, we distinguish: •Full-parameter Fine Tuning (FT):N params is the full count. • Parameter-efficient Fine Tuning (PEFT)(e.g., LoRA/adapters): we substitute Nparams by the number ofactive trainableparameters Ntrainable and include a forward-pass reus...

  6. [6]

    16 Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale

    Token accounting.When Ntokens is not directly reported, we infer it from dataset size and epochs, or reconstruct it from step geometry: Ntokens ≈S×G,(4) whereG=W×A×L×B.(5) S denotes the total number of training steps, W denotes the world size (number of devices), A denotes the gradient accumulation steps,Ldenotes the average sequence length, andBdenotes t...

  7. [7]

    Baseline FLOPs estimate.Let Nparams denote the number of (active) trainable parameters and Ntokens the number of training tokens effectively processed. The baseline lower-bound follows (Kaplan et al., 2020): FLOPsbase ≈c arch ×N params ×N tokens,(6) where carch ∈[5,12] accounts for architectural differences in the ratio of attention and feed-forward compu...

  8. [8]

    Let H, W, P, C be the input image height, width, patch size, and channels, respectively, and letd, L, r be the model’s hidden dimension, number of layers, and MLP expansion ratio

    ViT and CLIP models.For Vision Transformer (ViT) based models, we first calculate the FLOPs for a single forward step by summing the contributions from the patch embedding layer and the subsequent Transformer blocks. Let H, W, P, C be the input image height, width, patch size, and channels, respectively, and letd, L, r be the model’s hidden dimension, num...

  9. [9]

    This includes contributions from 2D convolutions (Mconv), self-attention (MSA), and cross-attention (MCA) layers

    Diffusion models.For U-Net-based models (e.g., Stable Diffusion), the MACs for a single denoising step are calculated by summing the compute across all layers in the U-Net’s down-sampling, middle, and up-sampling blocks. This includes contributions from 2D convolutions (Mconv), self-attention (MSA), and cross-attention (MCA) layers. The total FLOPs are th...

  10. [10]

    v4-128”, “v3-8

    Transformers.For Transformer-based models such as large vision-language models, where the architecture is predomi- nantly a large language model processing multimodal tokens, the total training FLOPs are approximated as: FT ransf ormers = 6×N·D(12) where N represents the number of model parameters and D is the total number of tokens in the training data. ...

  11. [11]

    Direct runtime:Iftraining hoursare disclosed, emissions are computed directly from reported wall-clock or GPU- hours multiplied by hardware power draw. 2.Imputed runtime:If training duration isnotdisclosed but total FLOPs are available, we back-compute runtime as T= Ftrain PeakTFLOPs×R eff ×N acc ,(13) where Ftrain is expected FLOPs, Reff is throughput ef...

  12. [12]

    Imputed hardware family (A100/A800/H100/TPU/AMD),

  13. [13]

    Throughput/efficiency variance inθ GPU across implementations and parallelism setups,

  14. [14]

    Datacenter amplification uncertainty (A time),

  15. [15]

    Because FLOPs is known whileK eff andEF region are imputed, Tier-2 inherits moderate uncertainty

    Regional EF uncertainty due to missing or ambiguous geography. Because FLOPs is known whileK eff andEF region are imputed, Tier-2 inherits moderate uncertainty. Tier-3: Neither FLOPs Nor Runtime DisclosedTier-3 models require the heaviest imputation. Total FLOPs must be estimated from model parameters via a scaling-law style approximation: F total train ≈...

  16. [16]

    Scaling-law coefficient variance (the proportionality constantc is architecture- and corpus-specific and absorbs variation in effective token counts and training stages)

  17. [17]

    Hardware inference as in Tier-2 (accelerator family, utilization, and datacenter amplification folded intoK eff)

  18. [18]

    Regional EF uncertainty when geography is missing or coarse

  19. [19]

    Since both F total train and Keff must be imputed, and each term enters multiplicatively, Tier-3 accumulates the largest theoretical error

    Compounded multiplicative propagation acrossF total train ,K eff, andEF region. Since both F total train and Keff must be imputed, and each term enters multiplicatively, Tier-3 accumulates the largest theoretical error. Plugging representative relative uncertainties as shown in Table.9 into Eq. 7 yields ∆E E ≈ s ∆F F 2 + ∆Keff Keff 2 + ∆EF EF 2 ∼0.9–1.5,(...