REVIEW 7 cited by
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requirements of modern LLMs, more and more top-of-the-line GPUs are being deployed to serve these models. Energy availability has come to the forefront as the biggest challenge for data center expansion to serve these models. In this paper, we present the trade-offs brought up by making energy efficiency the primary goal of LLM serving under performance SLOs. We show that depending on the inputs, the model, and the service-level agreements, there are several knobs available to the LLM inference provider to use for being energy efficient. We characterize the impact of these knobs on the latency, throughput, as well as the energy. By exploring these trade-offs, we offer valuable insights into optimizing energy usage without compromising on performance, thereby paving the way for sustainable and cost-effective LLM deployment in data center environments.
Forward citations
Cited by 7 Pith papers
-
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
AFlex combines operator-level disaggregation with per-operator DVFS to cut LLM serving energy per token by up to 49% without violating P90 TTFT/TPOT SLOs.
-
Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
vLLM attention kernel and prefix caching drive model- and task-dependent energy and latency effects, with no universal best config, and can unexpectedly shift measured accuracy.
-
ELASTIC: Event-Tracking Data Synchronization in Soccer Without Annotated Event Locations
The abstract claims ELASTIC outperforms prior soccer event-tracking synchronizers on 2,134 annotated events, but the submitted full text is a different manuscript, leaving the central claim unverifiable.
-
GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs
An eco-feedback browser extension for ChatGPT raises user awareness of energy and water use, but a nine-participant study finds limited effect on query frequency.
-
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
Retraining LLM tokenizers on chatbot conversation data reduces token counts by 5-10% on conversational text with minimal impact on general text.
-
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
A survey of LLM routing and hierarchical inference techniques that proposes an unvalidated unified evaluation metric called the Inference Efficiency Score.
-
Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
A survey of small language models that organizes known methods into taxonomies but adds no new models, data, or validated benchmarks.
Discussion (0). Sign in to comment.