A hybrid analytical-ML framework predicts LLM inference latency and energy from architectural parameters, with MAPE below 5 percent on selected models and about 10 percent on a broad set.
Power hungry processing: Watts driving the cost of AI deployment? InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 85–99, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors
A hybrid analytical-ML framework predicts LLM inference latency and energy from architectural parameters, with MAPE below 5 percent on selected models and about 10 percent on a broad set.