REVIEW 21 cited by
LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The carbon footprint associated with large language models (LLMs) is a significant concern, encompassing emissions from their training, inference, experimentation, and storage processes, including operational and embodied carbon emissions. An essential aspect is accurately estimating the carbon impact of emerging LLMs even before their training, which heavily relies on GPU usage. Existing studies have reported the carbon footprint of LLM training, but only one tool, mlco2, can predict the carbon footprint of new neural networks prior to physical training. However, mlco2 has several serious limitations. It cannot extend its estimation to dense or mixture-of-experts (MoE) LLMs, disregards critical architectural parameters, focuses solely on GPUs, and cannot model embodied carbon footprints. Addressing these gaps, we introduce \textit{\carb}, an end-to-end carbon footprint projection model designed for both dense and MoE LLMs. Compared to mlco2, \carb~significantly enhances the accuracy of carbon footprint estimations for various LLMs. The source code is released at \url{https://github.com/SotaroKaneda/MLCarbon}.
Forward citations
Cited by 21 Pith papers
-
Throttling Web Agents Using Reasoning Gates
Rebus-based reasoning gates, puzzles built from random word/domain clue sets, impose token costs on LM web agents that are up to 9.2x the generator's cost.
-
CEO-DC: Driving Decarbonization in HPC Data Centers with Actionable Insights
A decision framework using new carbon and price efficiency metrics shows most AI platform improvements cannot keep pace with demand growth, and short upgrade cycles need carbon prices far above current levels.
-
Analysis of Propaganda in Tweets From Politically Biased Sources
Journalists at politically extreme news outlets tweet propaganda-like language more often than those at mild outlets, and large language models outperform a fine-tuned BERT in detecting it.
-
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
SITAlign is an inference-time constrained decoder that maximizes a primary reward while enforcing thresholds on secondary rewards, and it reports better primary-reward win-tie rates than weighted-objective decoding.
-
Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
Small open-weight language models show factual knowledge of ionic liquids but fail on reasoning-focused entailment tests in a new 5,920-example benchmark for carbon capture.
-
EcoServe: Designing Carbon-Aware AI Inference Systems
EcoServe combines four strategies (reuse, rightsize, reduce, recycle) in an ILP optimizer to cut modeled carbon emissions for LLM serving by up to 47% while keeping SLOs.
-
Developing LLM-based Multi-Agent Systems in Software Engineering: A Mixed-Method Experience Report
A mixed-method experience report on LLM-based multi-agent frameworks for software engineering: broad feature coverage, weak monitoring support, and no clear quality winner on a README-summarization task, with incomple...
-
Calculating Software's Energy Use and Carbon Emissions: A Survey of the State of Art, Challenges, and the Way Ahead
A structured survey of 21 software energy and carbon calculation tools, organized as Monitoring, Estimation, or Black-Box approaches, with a component-wise comparison and a list of open challenges.
-
Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights
At 56B total parameters, fine-grained MoE with smaller, more numerous experts beats standard Switch and Mixtral-style MoE on validation loss and average downstream accuracy at matched FLOPs.
-
Exploring Anthropomorphism in Conversational Agents for Environmental Sustainability
A 26-person lab study found an LLM-based laundry-scheduling chatbot increased users' self-reported energy self-efficacy, while the personified version increased rapport but not self-efficacy.
-
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
Measuring LLM inference across workloads, frameworks, and GPUs shows that efficiency optimizations like vLLM and CUDA graphs can cut energy use by up to 73% versus an unoptimized PyTorch baseline.
-
Beyond Vision: How Large Language Models Interpret Facial Expressions from Valence-Arousal Values
LLMs given only valence and arousal values classify facial expressions poorly but can generate free-text emotion descriptions that align with human annotations under Word2Vec and BERT similarity, though not under a ge...
-
Exploring the sustainable scaling of AI dilemma: A projective study of corporations' AI environmental impacts
A corporate AI portfolio LCA model projects that high generative-AI adoption could increase AI electricity use about 24-fold by 2030.
-
Kryptonite-N: Machine Learning Strikes Back
The authors demonstrate that the Kryptonite-N challenge datasets are solvable by logistic regression with polynomial expansion and L1 regularization, and identify their construction as a high-dimensional XOR problem w...
-
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair
On a filtered Persian-Hindi parallel corpus, phrase-based SMT yields a BLEU of 66.3, ahead of the transformer NMT's 53.7.
-
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
Across 47 therapeutic LLM configurations, the safest models had estimated energy use up to 60 times higher than slightly less safe efficient models, and extra reasoning did not reliably improve safety.
-
Performance is not All You Need: Sustainability Considerations for Algorithms
The paper introduces FMS and ASC, composite sustainability scores that fuse accuracy and energy consumption, and evaluates them on multiple vision tasks.
-
Generating HomeAssistant Automations Using an LLM-based Chatbot
LLM-based chatbots, especially GPT models, generate mostly valid HomeAssistant automation routines and are perceived as more engaging than rule-based chatbots, but green prompts show only qualitative, not quantitative...
-
Addressing the sustainable AI trilemma: a case study on LLM agents and RAG
LLM-dependent memory operations in agents and RAG consume orders of magnitude more energy than vector methods, and resource-constrained hardware pays higher energy for lower quality.
-
CarbonChat: Large Language Model-Based Corporate Carbon Emission Analysis and Climate Knowledge Q&A System
CarbonChat combines self-prompting RAG and text-to-SQL for carbon-emission report analysis, reporting internal ablation gains but no external baselines or released artifacts.
-
A Survey on Inference Optimization Techniques for Mixture of Experts Models
A structured survey of MoE inference optimization that categorizes existing techniques into model, system, and hardware levels and summarizes reported speedups and memory savings.
Discussion (0). Continue with ORCID to comment.