Pith. sign in

REVIEW 7 cited by

Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.20306 v1 pith:DQNHWRUV submitted 2024-03-29 cs.AI cs.ARcs.DC

classification cs.AIcs.ARcs.DC
keywords energymodelsinferencellmscenterdataforefrontknobs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requirements of modern LLMs, more and more top-of-the-line GPUs are being deployed to serve these models. Energy availability has come to the forefront as the biggest challenge for data center expansion to serve these models. In this paper, we present the trade-offs brought up by making energy efficiency the primary goal of LLM serving under performance SLOs. We show that depending on the inputs, the model, and the service-level agreements, there are several knobs available to the LLM inference provider to use for being energy efficient. We characterize the impact of these knobs on the latency, throughput, as well as the energy. By exploring these trade-offs, we offer valuable insights into optimizing energy usage without compromising on performance, thereby paving the way for sustainable and cost-effective LLM deployment in data center environments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

    cs.DC 2026-08 conditional novelty 6.0 of 10

    AFlex combines operator-level disaggregation with per-operator DVFS to cut LLM serving energy per token by up to 49% without violating P90 TTFT/TPOT SLOs.

  2. Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations

    cs.SE 2026-07 conditional novelty 6.0 of 10

    vLLM attention kernel and prefix caching drive model- and task-dependent energy and latency effects, with no universal best config, and can unexpectedly shift measured accuracy.

  3. ELASTIC: Event-Tracking Data Synchronization in Soccer Without Annotated Event Locations

    cs.DB 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims ELASTIC outperforms prior soccer event-tracking synchronizers on 2,134 annotated events, but the submitted full text is a different manuscript, leaving the central claim unverifiable.

  4. GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs

    cs.HC 2025-05 conditional novelty 5.0 of 10

    An eco-feedback browser extension for ChatGPT raises user awareness of energy and water use, but a nine-participant study finds limited effect on query frequency.

  5. Is There a Case for Conversation Optimized Tokenizers in Large Language Models?

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Retraining LLM tokenizers on chatbot conversation data reduces token counts by 5-10% on conversational text with minimal impact on general text.

  6. Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    A survey of LLM routing and hierarchical inference techniques that proposes an unvalidated unified evaluation metric called the Inference Efficiency Score.

  7. Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation

    cs.CL 2025-05 unverdicted novelty 2.0 of 10

    A survey of small language models that organizes known methods into taxonomies but adds no new models, data, or validated benchmarks.

Pith tools