Introduces a quantization-enabled demand response framework for LLM data centers that maps precision levels to power parameters and achieves 34.3% cost reduction in case studies while maintaining token volume.
Systematic characterization of LLM quantization: A performance, energy, and quality perspective,
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Small LLMs under 2B parameters achieve better economic break-even, energy efficiency, and hardware density than larger models on legacy GPUs for industrial tasks.
Layered prefill replaces token-chunked prefill with layer-group interleaving in MoE models, cutting TTFT by up to 70%, end-to-end latency by 41%, and per-token energy by 22% while preserving stall-free TBT.
Across four edge platforms, NPU/GPU accelerators improve LLM energy efficiency by up to ~40x over CPU-only, and volume-normalized throughput favors the tiny M5Stack over the fastest Jetson.
citing papers explorer
-
From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads
Introduces a quantization-enabled demand response framework for LLM data centers that maps precision levels to power parameters and achieves 34.3% cost reduction in case studies while maintaining token volume.
-
Are Large Language Models Economically Viable for Industry Deployment?
Small LLMs under 2B parameters achieve better economic break-even, energy efficiency, and hardware density than larger models on legacy GPUs for industrial tasks.
-
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
Layered prefill replaces token-chunked prefill with layer-group interleaving in MoE models, cutting TTFT by up to 70%, end-to-end latency by 41%, and per-token energy by 22% while preserving stall-free TBT.
-
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
Across four edge platforms, NPU/GPU accelerators improve LLM energy efficiency by up to ~40x over CPU-only, and volume-normalized throughput favors the tiny M5Stack over the fastest Jetson.