{"id":"ff95a8cf-2c34-4970-88f1-b832a2cfafcb","arxiv_id":"2506.15572","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Environmental disclosure for notable AI models peaked in 2022 and then declined, and out-of-context energy and emissions estimates now dominate media coverage.","lead":"This paper documents that environmental transparency of notable AI models has declined since 2022, and shows how out-of-context estimates of AI energy use and emissions spread as misinformation. It argues for standardized measurement and disclosure so that users, regulators, and buyers can compare AI systems by environmental impact.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'Indirect Disclosure' category counts open-weights releases as environmental transparency, so the post-2022 decline may measure the shift to closed APIs rather than reduced environmental reporting. A stricter recoding and stratification test is required.","rationale":"The paper argues that environmental transparency in AI has declined since 2022 and that decontextualized estimates have become misinformation. The empirical trend in Section 'Environmental Transparency Trends' is load-bearing for the first claim. The reader's weakest assumption identified the Epoch dataset's curated nature and the authors' manual coding as the key risk. I agree those are real limitations, but the more basic problem is definitional: the Indirect Disclosure category treats open-weight releases as environmental transparency even when no environmental data are provided. Because the post-2022 era is exactly when notable models shifted from open-weights research releases to closed commercial APIs, the measured decline may be a restatement of the open/closed shift rather than a change in environmental reporting behavior. This does not require accusing the authors of bias; the category definition itself creates the conflation. The concrete re-analysis with a stricter definition and within-strata control would settle whether the decline is genuine. The paper's myth tracing and policy recommendations are valuable and do not depend on the exact slope of the trend, so I would not reject the paper. A conditional verdict is appropriate, conditioned on releasing the coded data and re-running the trend under a disclosure-only definition that does not count open weights as indirect environmental disclosure. If the re-analysis confirms the decline within both open and closed models, the empirical core is robust; if not, the paper should be reframed as a commentary on closed-model opacity rather than a documented transparency decline. The reader's original conditional verdict is consistent with this, so I set verdict_should_be to CONDITIONAL rather than REJECT or ACCEPT.","tokens_in":10336,"tokens_out":9397,"duration_ms":115142,"concrete_test":"Obtain the authors' coded dataset or independently recode all 754 Epoch AI Notable AI Models into three stricter categories: Direct = explicit energy or GHG figures reported; Indirect = explicit training compute (FLOPs, hardware, duration) reported without environmental numbers; No Disclosure = neither, ignoring whether weights are open. Compute yearly percentages and absolute counts under this definition, then stratify by open-weights vs closed-weights status. If the post-2022 increase in No Disclosure persists within both strata, the transparency decline is robust; if it appears only in the closed-weights stratum, the paper's headline trend is an artifact of the open/closed divide. Report how the trend changes when the original Indirect Disclosure criterion is replaced by this stricter definition.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In the section 'Environmental Transparency Trends,' Indirect Disclosure is defined to include any model whose developers 'released their model weights, allowing external estimates of training or inference impacts.' This operationally equates open-weight licensing with environmental transparency: a model with open weights and zero environmental information is classified as more transparent than a closed model that reports training compute or emissions but does not release weights. The headline claim—that transparency improved through 2022 and then reversed—therefore tracks the post-ChatGPT shift from open-weights research releases to closed commercial APIs, which is not the same as a decline in environmental disclosure. The authors themselves attribute the reversal to 'the introduction of increasingly commercial and proprietary models after 2022,' confirming that the measure is entangled with the proprietary/open divide. This is a construct-validity problem: the metric is called 'environmental impact transparency' but partly measures model accessibility. Even with perfect inter-rater reliability and a published codebook, this definitional choice could manufacture the observed trend. The reader's concern about coding reliability and Epoch's curated sample is related but secondary; the more fundamental issue is whether the category definitions support the central inference that users have less environmental information over time.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper examines environmental transparency in AI model development and deployment. Using Epoch AI's Notable AI Models dataset (754 models from 2010 to Q1 2025), the authors classify each model into three transparency categories: Direct Disclosure (energy/GHG emissions reported), Indirect Disclosure (training compute data or open weights enabling external estimates), and No Disclosure. They report that transparency improved through 2022, then reversed, with the majority of notable models in early 2025 falling into the no-disclosure category. An analysis of OpenRouter traffic for May 2025 finds that 84% of LLM token usage flows through no-disclosure models. The paper then investigates two widely circulated claims—that training a model emits as much CO2 as five cars over their lifetime, and that a ChatGPT query uses ten times more energy than a Google search—and traces their origins and misrepresentation. The authors also present a media analysis of 100 articles covering ChatGPT energy consumption, finding that 75% relayed energy estimates without uncertainty or sourcing. The paper concludes with recommendations for measurement, standards, verification, and policy.","tokens_in":10553,"tokens_out":6674,"duration_ms":72021,"significance":"If the documented trend is robust, the paper makes an important contribution by showing that environmental opacity in AI is worsening as impacts grow, and it identifies concrete misinformation pathways. The paper's careful treatment of the 'five cars' estimate—distinguishing the upper-bound neural architecture search scenario from typical training workloads—is a genuine strength, as is the provenance tracing of the 'ten times a Google search' claim. The analysis produces falsifiable claims: the transparency trend can be re-tested with alternative coding schemes, and the media analysis percentages can be audited. The recommendation section provides actionable, if high-level, pathways. However, the paper's own empirical contributions rest on a category definition that conflates open-weights licensing with environmental disclosure, and on manual coding without released data or reliability checks; these need to be addressed before the central claim can be considered established.","major_comments":[{"comment":"The definition of Indirect Disclosure—'Developers provided training compute data or released their model weights, allowing external estimates of training or inference impacts'—conflates model accessibility with environmental transparency. Under this coding, a model with open weights and no environmental information whatsoever is classified as more transparent than a closed model that reports training compute or emissions. Because the post-2022 reversal is attributed to 'the introduction of increasingly commercial and proprietary models,' the headline trend may simply track the shift from open-weights research releases to closed APIs rather than a decline in environmental reporting. Please re-analyze the data with a stricter coding scheme that separates release of weights from release of compute/emissions data (e.g., treating open weights alone as no disclosure when no compute or energy data are provided), and report the trend under both codings. If the decline attenuates or disappears, the central claim must be revised.","section":"Environmental Transparency Trends (definition of Indirect Disclosure)"},{"comment":"The classification of all 754 Epoch models and the coding of the 100 news articles are performed by the authors without a published codebook, released raw data, or inter-rater reliability checks. The quantitative headline claims—such as 84% of OpenRouter usage through no-disclosure models, and the period-by-period percentages in Figure 1—cannot be independently verified. Please make the coded dataset and codebook available, and provide reliability statistics (e.g., dual coding of a random sample) or, at minimum, a sensitivity analysis showing that the conclusions are robust to plausible coding disagreements.","section":"Environmental Transparency Trends / Methods (manual coding)"},{"comment":"The media analysis described in the section 'Investigating the Urban Legends' reports that 75% of 100 articles relayed energy estimates without uncertainty or sourcing, but the methodology is under-specified: the sample was taken from a Google News search for 'ChatGPT energy consumption' as of April 11, 2025, but the text does not state how the 100 articles were selected from the results, what inclusion/exclusion criteria were used, how the coding categories were defined, or whether coding was performed by more than one person. These details are needed to assess whether the 75% figure is reliable and generalizable. Please add a methods paragraph describing article selection, coding rules, and coder agreement.","section":"Investigating the Urban Legends (media analysis)"},{"comment":"The paper states that '84% of LLM usage is through models with no disclosure' based on OpenRouter's top 20 models in May 2025, but the text does not clarify whether this 84% is computed only over the top 20 or over total tokens on the platform. If OpenRouter data are limited to the top 20, the claim should be qualified as 'the share among the top 20 most-used models' and the paper should discuss whether the top 20 dominate overall token volume. Without this caveat, the statistic overstates the representativeness of the snapshot.","section":"OpenRouter analysis (Figure 2)"}],"minor_comments":[{"comment":"The text contains a typo: 'This period includes the the work of Strubell et al.' — the duplicate 'the' should be removed.","section":"Environmental Transparency Trends"},{"comment":"The sentence 'This remark was used was the basis of an estimate published in October 2023' contains an extra 'was'; it should read 'This remark was used as the basis...'.","section":"Investigating the Urban Legends"},{"comment":"Several entries in Table 1 are listed as '?' (e.g., Gemma 2B+9B energy and GHG, Llama 3 70B energy), and the 'Max/Min Variance' row is not defined; please add a note explaining the table's conventions and what the variance row represents.","section":"Appendix Table 1"},{"comment":"The authors recommend the AI Energy Score project without a conflict-of-interest statement; given that two authors are directly involved in that project, a formal disclosure should be included.","section":"Introduction / How to improve environmental impact disclosures"},{"comment":"The caption of Figure 1 does not specify what the y-axis represents (percentage of models, count, or other); please state this explicitly in the caption or text.","section":"Figure 1"},{"comment":"Some references (e.g., Han et al. [24], Schneider et al. [25], Morrison et al. [26]) are given as bare arXiv numbers without full citation details; please standardize these entries to journal or arXiv format.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for cs.CY and addresses a timely topic, but the empirical claims are likely to receive close scrutiny. The central construct-validity issue in the transparency categories is fixable with a re-analysis, and the lack of reproducibility should be addressed before publication. I would also advise the editor to require an explicit conflict-of-interest statement regarding the authors' roles in the AI Energy Score project and any sponsored funding, as the paper recommends that tool without qualification. The paper should not be rejected on these grounds, but the current version is not yet publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper does more than argue: it traces the origins of three AI-environment myths, codes 754 Epoch AI models for environmental disclosure, looks at May 2025 OpenRouter usage, and content-codes 100 news articles on ChatGPT energy claims. The myth tracing is the strongest part. The \"five cars\" discussion is careful—it correctly separates the neural architecture search upper bound from average training runs, and notes that recent Llama/Gemma reports exceed the estimate. The \"ten times a Google search\" chain is well documented: a secondhand John Hennessy remark in 2023, a 3 Wh estimate based on a 2009 Google blog post, and 75% of the 100 articles relaying numbers without caveats. That media analysis alone justifies reading the paper.\n\nThe transparency trend is plausible but not fully auditable. The coding is manual, no codebook or inter-rater reliability is reported, and Epoch's \"Notable AI Models\" is a convenience sample of state-of-the-art/cited models, not a random sample. More fundamentally, the Indirect Disclosure category counts \"released their model weights\" as environmental transparency. That conflates open-weight licensing with disclosure of environmental information. A model with zero environmental reporting can be classified as Indirect simply because its weights are public, while a closed model that reports training compute might also be Indirect but is arguably more transparent on the environmental dimension. So the post-2022 decline could partially measure the shift from open research releases to closed APIs. The authors acknowledge the shift but don't test whether the trend survives a stricter recoding.\n\nThis doesn't sink the paper. The OpenRouter finding—84% of token usage through no-disclosure models—is more robust to the objection, and the recommendations are reasonable even if some (AI Energy Score) come with a conflict of interest the authors should state. What's needed is a released dataset, a codebook, inter-rater checks, and a sensitivity analysis that treats open weights alone as non-disclosure.\n\nWho is this for? Researchers and policymakers working on AI sustainability disclosure and procurement standards, plus anyone tracking how scientific estimates become public misinformation. It deserves a serious referee, but the referee should push for the data release and the recoding before acceptance.\n\nRecommendation: send to peer review with requests for the codebook and sensitivity analysis.","headline":"Useful myth-tracing and a plausible but under-audited transparency trend; the Indirect Disclosure category needs a sensitivity check before I'd trust the headline.","tokens_in":11051,"tokens_out":2912,"would_cite":true,"duration_ms":34395,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI models have grown less transparent about their environmental costs since 2022, and the figures that fill the information gap are misinformation grown from decontextualized estimates.","keywords":["environmental transparency","AI energy consumption","greenhouse gas emissions","model disclosure","misinformation","Epoch AI notable models","LLM inference"],"falsifier":"Independently re-code a random sample of 100 models released in 2022 and 100 released in the first quarter of 2025 for direct, indirect, or no environmental disclosure using a pre-registered codebook; if the share of no-disclosure models is not higher in the later period, the claimed decline would be an artifact of the dataset curation or the manual coding. A complementary check would count how many of the top-20 OpenRouter models' publishers release energy or GHG figures in their model cards and compare the token-weighted disclosure share with the paper's figures.","tokens_in":10165,"feed_emoji":"🌍","tokens_out":6731,"duration_ms":70873,"temperature":0.7,"pith_summary":"This paper argues that the AI industry is becoming less transparent about the environmental costs of its models just as those costs are growing. Combining Epoch AI's list of 754 notable models from 2010 to early 2025, the authors find that direct disclosure of energy use or greenhouse-gas emissions peaked in 2022 and then declined, so that by the first quarter of 2025 most notable models disclosed nothing usable. The authors then trace three widely repeated numbers, the \"five cars\" estimate for training emissions, the claim that a ChatGPT query uses ten times the energy of a Google search, and the claim that AI can cut global emissions by 5 to 10 percent, to show how best-effort, decontextualized estimates become misinformation. If they are right, the paper establishes that the environmental transparency problem is not just missing data but active misunderstanding, and that users and policymakers are making decisions from figures that do not represent actual impacts.","feed_headline":"AI models are getting less transparent about their energy use","feed_subtitle":"By early 2025 most notable models disclose nothing, and 84% of LLM use is opaque.","key_machinery":"The central machinery is a three-category transparency classification applied to a curated dataset: Direct Disclosure (the developer reports energy or GHG numbers), Indirect Disclosure (training compute or model weights are released, allowing outsiders to estimate impacts), and No Disclosure (neither is available). The same classification applied to OpenRouter's top-20 monthly LLM usage list produces the headline usage split of 84 percent no disclosure, 14 percent indirect, and 2 percent direct. This coding is what turns the claim of declining transparency into a measurable trend rather than an anecdote, and the provenance tracing of the \"five cars\", \"3 Wh\", and \"5 to 10 percent\" figures is the second mechanism showing how isolated estimates become public misinformation.","core_discovery":"The paper's central discovery is a documented inversion: methodological tools for measuring AI's environmental impact have improved since 2019, yet industry disclosure has gone backwards. Classifying every model in Epoch AI's Notable AI Models dataset, 754 models from 2010 through the first quarter of 2025, into Direct Disclosure, Indirect Disclosure, or No Disclosure, the authors report that direct disclosure of energy or GHG data peaked in 2022 at 10 percent of notable models, and that by Q1 2025 the majority of notable models again fell into No Disclosure. A complementary snapshot of OpenRouter traffic in May 2025 shows that 84 percent of LLM token usage flowed through no-disclosure models, 14 percent through indirect-disclosure models, and only 2 percent through direct-disclosure models. The authors also show that the three most widespread quantitative claims about AI's climate footprint each trace to a single qualified or speculative estimate that was stripped of its context in repeated media coverage, with 75 percent of 100 sampled news articles relaying energy figures for ChatGPT queries without citing sources or expressing uncertainty.","pith_inferences":["An extension the authors leave implicit is that the same three-category coding could be applied to fine-tuning and API-serving layers rather than only base models, which would likely push the no-disclosure share even higher because most fine-tuning services do not report energy use.","A testable extension is a longitudinal audit of model cards on major hosting platforms, which could verify whether the post-2022 decline is driven by proprietary labs or by open-weight releases increasingly choosing not to disclose.","If mandatory reporting under the EU's sustainability disclosure rules is enforced, the paper's framework predicts that transparency will improve among labs that serve EU customers, providing a natural experiment for the claim that omission is a choice rather than a technical necessity.","The paper implies that the \"misinformation by omission\" mechanism generalizes beyond these three myths, so fact-checking infrastructure for AI-energy claims would be a low-cost mitigation even before formal standards arrive."],"forward_implications":["Users of the most-used LLM APIs cannot estimate the energy or carbon cost of the queries they make, so procurement decisions based on efficiency are effectively impossible.","Policymakers relying on \"five cars\" or \"ten times a Google search\" as if they were measured facts are likely to misallocate effort in climate regulations.","Because direct disclosures peaked in 2022 and then fell, attempts to regulate AI environmental impacts cannot rely on industry self-reporting in its current form.","The evidence that actual pretraining emissions for Gemma and Llama 3 exceed the \"five cars\" figure by roughly 4 times and 40 times shows that even the widely criticized estimate understates today's largest training runs.","Standardized, verifiable reporting frameworks and procurement requirements are the direct corollary of the finding that omission creates misinformation."],"supporting_citations":[{"why":"Supplies the full list of 754 notable AI models and their metadata on which the three-category transparency trend is built.","marker":"[27]"},{"why":"Supplies May 2025 OpenRouter traffic and token counts that yield the 84/14/2 percent usage split.","marker":"[28]"},{"why":"Origin of the \"five cars\" estimate and the baseline against which later pretraining emissions figures are compared.","marker":"[12]"},{"why":"Source of John Hennessy's remark that an LLM exchange costs roughly ten times a keyword search.","marker":"[39]"},{"why":"Published the \"approximately 3 Wh per LLM interaction\" estimate that the media repeat as fact.","marker":"[40]"},{"why":"Source of the 0.0003 kWh per Google search figure used in the ten-times comparison.","marker":"[41]"},{"why":"Reports 1247.61 tCO2e for Gemma pretraining, over four times the \"five cars\" estimate.","marker":"[34]"},{"why":"Reports 11,390 tCO2e for the Llama 3 family, over forty times the \"five cars\" estimate.","marker":"[35]"},{"why":"Origin of the claim that AI can mitigate 5 to 10 percent of global emissions, repeated by media and research.","marker":"[48]"}],"fun_headline_variants":["AI's environmental transparency is going backwards","As AI grows, its environmental transparency shrinks","AI's energy use is more opaque than ever","Why AI's climate impact is becoming less visible","The AI transparency gap: more models, less disclosure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' classification of 754 curated models into three transparency buckets faithfully represents the AI industry's environmental disclosure, rather than an artifact of which models Epoch AI selected or how the authors coded them.","fun_headline_variants_meta":{"raw":{"variants":["AI's environmental transparency is going backwards","As AI grows, its environmental transparency shrinks","AI's energy use is more opaque than ever","Why AI's climate impact is becoming less visible","The AI transparency gap: more models, less disclosure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1730,"prompt_tokens":918,"completion_tokens":812,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":742}},"tokens_in":534,"tokens_out":812,"duration_ms":8953,"temperature":1.0,"reasoning_tokens":742,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:52:44.967670+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-code a random sample of 100 models released in 2022 and 100 released in the first quarter of 2025 for direct, indirect, or no environmental disclosure using a pre-registered codebook; if the share of no-disclosure models is not higher in the later period, the claimed decline would be an artifact of the dataset curation or the manual coding. A complementary check would count how many of the top-20 OpenRouter models' publishers release energy or GHG figures in their model cards and compare the token-weighted disclosure share with the paper's figures.","supporting_citations":[{"cited_title":"Openrouter leaderboard rankings","cited_arxiv_id":null,"evidence_quote":"Supplies May 2025 OpenRouter traffic and token counts that yield the 84/14/2 percent usage split."},{"cited_title":"& Hutchinson, R","cited_arxiv_id":null,"evidence_quote":"Origin of the claim that AI can mitigate 5 to 10 percent of global emissions, repeated by media and research."}],"review_version":1}