REVIEW 2 major objections 2 minor 1 cited by
Training the most popular open-source AI models has already emitted about 60,000 metric tons of carbon.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-07-04 01:53 UTC pith:EZFLDGAM
load-bearing objection The paper supplies a usable first aggregate for popular HF models at 60k tCO2e and releases code, but the tiered regressions on incomplete metadata are the part that needs the most scrutiny. the 2 major comments →
Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Using available emissions, energy, compute, and model metadata from Hugging Face, a tiered estimation approach supported by empirical regressions produces an aggregate training emissions total of approximately 6.0×10^4 metric tons for the most popular open-source models, while AI training carbon intensity provides a comparable efficiency metric across trainings.
What carries the argument
The tiered approach to emissions estimation, which uses empirical regressions to derive reliable values from incomplete model metadata on energy use and compute.
Load-bearing premise
The tiered approach supported by empirical regressions accurately estimates emissions from incomplete metadata without introducing large systematic bias.
What would settle it
Direct measurement of actual training emissions for a sample of models whose full metadata later becomes available, then comparison against the tiered estimates for the same models.
If this is right
- Aggregate emissions from the most-downloaded open-source models already stand at roughly 60,000 metric tons.
- AI training carbon intensity allows direct comparison of sustainability efficiency between different model trainings.
- The method supplies a practical template for carbon reporting standards that the AI industry could adopt without requiring full reproduction of every training run.
- Future disclosures with higher completeness would reduce reliance on the regression tier and tighten the overall estimate.
Where Pith is reading between the lines
- Widespread use of this estimation style could create pressure on developers to publish more complete training metadata as a norm.
- The 60,000-ton figure provides a reference point for comparing AI training emissions against other large-scale computing activities such as cryptocurrency mining or data-center operations.
- If model downloads continue to concentrate on a small set of popular checkpoints, total emissions could scale roughly with the number of new high-download releases rather than with every uploaded model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a tiered estimation framework to quantify aggregate training carbon emissions of open-source AI models on the Hugging Face platform from unevenly disclosed metadata on energy, compute, and emissions. It uses empirical regressions to impute missing values, defines an AI training carbon intensity (ATCI) metric, and reports that training the most popular models (those with over 5,000 downloads) has produced approximately 6.0×10^4 metric tons of CO2e. Data and code are released publicly.
Significance. If the tiered imputation and regressions are shown to be unbiased, the work supplies a reproducible, scalable method for carbon accounting at the scale of an entire model repository without requiring re-execution of training runs. The open release of data and code is a clear strength that enables independent verification and extension.
major comments (2)
- [Methods / Results on aggregate estimate] The aggregate result of 6.0×10^4 tCO2e rests on regressions that impute missing metadata; the manuscript must demonstrate that these regressions were validated on held-out data independent of the fitted parameters, or provide explicit error propagation and sensitivity analysis showing that any circularity does not materially affect the reported total (see abstract and the section describing the tiered approach).
- [Abstract and empirical regressions subsection] No validation statistics, cross-validation scores, or uncertainty intervals for the regression-based imputations are visible in the abstract or summary description; without these, the reliability assessment claimed for the tiered method cannot be evaluated.
minor comments (2)
- Clarify the exact definition and units of ATCI early in the text and ensure it is used consistently in all figures and tables.
- Add a table or appendix listing the number of models per tier and the fraction of metadata imputed at each level.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback emphasizing the need for rigorous validation of our regression-based imputations. We address each major comment below and will revise the manuscript accordingly to strengthen the presentation of reliability metrics.
read point-by-point responses
-
Referee: [Methods / Results on aggregate estimate] The aggregate result of 6.0×10^4 tCO2e rests on regressions that impute missing metadata; the manuscript must demonstrate that these regressions were validated on held-out data independent of the fitted parameters, or provide explicit error propagation and sensitivity analysis showing that any circularity does not materially affect the reported total (see abstract and the section describing the tiered approach).
Authors: We agree that the current manuscript does not include explicit held-out validation, cross-validation, or sensitivity analysis for the regressions used in the tiered approach. Although the text describes empirical regressions to assess estimation reliability, these do not encompass the requested validation steps or error propagation. In the revised version we will add a dedicated validation subsection that applies the released code and data to perform held-out testing, reports cross-validation scores, and conducts sensitivity analysis on imputation parameters to quantify impact on the aggregate 6.0×10^4 tCO2e estimate. revision: yes
-
Referee: [Abstract and empirical regressions subsection] No validation statistics, cross-validation scores, or uncertainty intervals for the regression-based imputations are visible in the abstract or summary description; without these, the reliability assessment claimed for the tiered method cannot be evaluated.
Authors: We concur that the abstract and summary description omit these specific statistics. We will revise the abstract to include key validation metrics (e.g., cross-validation R² and uncertainty ranges) and expand the empirical regressions subsection with the corresponding scores, intervals, and reliability assessment details. revision: yes
Circularity Check
No significant circularity identified
full rationale
The paper's central aggregate estimate of ~6.0×10^4 tCO2e relies on a tiered procedure for incomplete metadata, with empirical regressions used only to assess estimation reliability rather than to generate the reported total by construction. No equations or steps reduce the final numerical result to a fit on the same data being estimated, no self-citations are load-bearing for uniqueness or ansatz, and the derivation supplies code and data for external verification. The method remains self-contained against external benchmarks with no quoted reduction of prediction to input.
Axiom & Free-Parameter Ledger
free parameters (1)
- regression coefficients for tiered metadata imputation
read the original abstract
The scaling-law era has transformed artificial intelligence (AI) from research into a global industry, but its rapid growth also raises concerns over energy usage, carbon emissions, and environmental sustainability. Unlike traditional sectors, the AI industry still lacks systematic carbon accounting methods that support large-scale estimates without reproducing the original training process. This leaves open questions about how large the problem is today and how large it might be in the near future. Given its central role in hosting open-source AI models, the Hugging Face (HF) platform provides a large-scale and publicly accessible corpus for carbon accounting. We estimate aggregate training emissions of HF open-source models using available emissions, energy, compute, and model metadata. To address uneven disclosure quality, we introduce a tiered approach to handle incomplete metadata, supported by empirical regressions that assess estimation reliability. We further introduce AI training carbon intensity (ATCI, emissions per compute), a metric to assess the sustainability efficiency of model training. Our results show that training the most popular open-source models (with over 5,000 downloads) has already resulted in approximately $6.0\times10^4$ metric tons of carbon emissions. Overall, this paper provides a scalable, empirically grounded framework for estimating training emissions from incomplete disclosures and informing future carbon reporting standards in the AI industry. Data and code are available at https://github.com/insait-institute/HuggingCarbon.
Figures
Forward citations
Cited by 1 Pith paper
-
How Low Can We Go? Minimum Spectroscopic Requirements For Supernova Subtype Classification
ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.
Reference graph
Works this paper leans on
-
[1]
Official Hugging Face model cards, including structured fields (e.g., compute used, hardware, carbon emissions), author-provided notes, and embedded configuration snippets
-
[2]
Repository configuration files, such asconfig.json, tokenizer/vision encoder configs, and architecture descriptors. These files provide parameter counts, layer depths, hidden sizes, patch sizes, and other FLOPs-relevant attributes
-
[3]
External authoritative sources, including official technical reports, GitHub repositories, and arXiv papers referenced in the model cards. When multiple sources were available, we applied a deterministic priority order (direct disclosure →config-derived→paper-derived→regression-estimated). Table 4.Metadata disclosure sparsity across the Hugging Face model...
work page 2025
-
[4]
When available, we directly record the number of trainable parameters (Nparams)
Model classification and parameter extraction.Each model is classified as encoder-only (e.g., BERT), decoder-only (e.g., GPT, LLaMA), or encoder–decoder (e.g., T5). When available, we directly record the number of trainable parameters (Nparams). If parameters are missing, we infer them from architecture descriptors such as hidden size, number of layers, a...
-
[5]
For others, we distinguish: •Full-parameter Fine Tuning (FT):N params is the full count
Effective parameter count adjustments.For pretraining we set Nparams to the full parameter count. For others, we distinguish: •Full-parameter Fine Tuning (FT):N params is the full count. • Parameter-efficient Fine Tuning (PEFT)(e.g., LoRA/adapters): we substitute Nparams by the number ofactive trainableparameters Ntrainable and include a forward-pass reus...
work page 2023
-
[6]
16 Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale
Token accounting.When Ntokens is not directly reported, we infer it from dataset size and epochs, or reconstruct it from step geometry: Ntokens ≈S×G,(4) whereG=W×A×L×B.(5) S denotes the total number of training steps, W denotes the world size (number of devices), A denotes the gradient accumulation steps,Ldenotes the average sequence length, andBdenotes t...
-
[7]
Baseline FLOPs estimate.Let Nparams denote the number of (active) trainable parameters and Ntokens the number of training tokens effectively processed. The baseline lower-bound follows (Kaplan et al., 2020): FLOPsbase ≈c arch ×N params ×N tokens,(6) where carch ∈[5,12] accounts for architectural differences in the ratio of attention and feed-forward compu...
work page 2020
-
[8]
ViT and CLIP models.For Vision Transformer (ViT) based models, we first calculate the FLOPs for a single forward step by summing the contributions from the patch embedding layer and the subsequent Transformer blocks. Let H, W, P, C be the input image height, width, patch size, and channels, respectively, and letd, L, r be the model’s hidden dimension, num...
-
[9]
Diffusion models.For U-Net-based models (e.g., Stable Diffusion), the MACs for a single denoising step are calculated by summing the compute across all layers in the U-Net’s down-sampling, middle, and up-sampling blocks. This includes contributions from 2D convolutions (Mconv), self-attention (MSA), and cross-attention (MCA) layers. The total FLOPs are th...
-
[10]
Transformers.For Transformer-based models such as large vision-language models, where the architecture is predomi- nantly a large language model processing multimodal tokens, the total training FLOPs are approximated as: FT ransf ormers = 6×N·D(12) where N represents the number of model parameters and D is the total number of tokens in the training data. ...
work page 2000
-
[11]
Direct runtime:Iftraining hoursare disclosed, emissions are computed directly from reported wall-clock or GPU- hours multiplied by hardware power draw. 2.Imputed runtime:If training duration isnotdisclosed but total FLOPs are available, we back-compute runtime as T= Ftrain PeakTFLOPs×R eff ×N acc ,(13) where Ftrain is expected FLOPs, Reff is throughput ef...
-
[12]
Imputed hardware family (A100/A800/H100/TPU/AMD),
-
[13]
Throughput/efficiency variance inθ GPU across implementations and parallelism setups,
-
[14]
Datacenter amplification uncertainty (A time),
-
[15]
Because FLOPs is known whileK eff andEF region are imputed, Tier-2 inherits moderate uncertainty
Regional EF uncertainty due to missing or ambiguous geography. Because FLOPs is known whileK eff andEF region are imputed, Tier-2 inherits moderate uncertainty. Tier-3: Neither FLOPs Nor Runtime DisclosedTier-3 models require the heaviest imputation. Total FLOPs must be estimated from model parameters via a scaling-law style approximation: F total train ≈...
-
[16]
Scaling-law coefficient variance (the proportionality constantc is architecture- and corpus-specific and absorbs variation in effective token counts and training stages)
-
[17]
Hardware inference as in Tier-2 (accelerator family, utilization, and datacenter amplification folded intoK eff)
-
[18]
Regional EF uncertainty when geography is missing or coarse
-
[19]
Compounded multiplicative propagation acrossF total train ,K eff, andEF region. Since both F total train and Keff must be imputed, and each term enters multiplicatively, Tier-3 accumulates the largest theoretical error. Plugging representative relative uncertainties as shown in Table.9 into Eq. 7 yields ∆E E ≈ s ∆F F 2 + ∆Keff Keff 2 + ∆EF EF 2 ∼0.9–1.5,(...
work page 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.