REVIEW 3 major objections 4 minor 57 references
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read FinGraV reconstructs sub-millisecond GPU power profiles by stitching 1 ms averaged samples across hundreds of time-shifted runs, and shows that ignoring power-profile differentiation can cause up to 80% energy measurement error.
desk verdict Useful MI300X power measurement guidance and new component-level data, but the 'fine-grain' claim is undermined by unaddressed 1ms averaging; the profiles are blurred traces, not sub-ms truth. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the stitched power profile built from the MI300X's internal 1 ms power logger, where each sample is the average of instantaneous power readings over the preceding 1 ms. Three mechanisms carry the argument: execution time binning, which excludes outlier runs and groups runs whose execution times fall within a 2-5% margin; CPU-GPU time synchronization, which reads a GPU timestamp before kernel launch and benchmarks the read delay so each power log can be assigned a time of interest within the kernel; and power profile differentiation into SSE and SSP profiles, which separates the first stabilized execution from the later power-stabilized execution. Random delays before each run place the 1 ms averaging window at different points in the kernel, so stitching hundreds of runs produces a time-series view of average power across the kernel's duration.
What would settle it
Run a sub-millisecond GEMM on an MI300X while simultaneously logging with the internal 1 ms power logger and an independent high-frequency external power meter on the GPU power rails, then compare the FinGraV stitched profile to the external trace; if the shapes diverge by more than the reported error, the moving-average or cross-run stability assumption fails.
Extended reading notes
Core claim
The paper's central claim is that a GPU power logger with a 1 ms averaging window can still yield a fine-grain, sub-millisecond power profile, provided the same kernel is executed many times with random start delays and the resulting power samples are stitched together after careful CPU-GPU time synchronization, execution-time binning, and power-profile differentiation. On the AMD MI300X, FinGraV separates the steady-state execution (SSE) profile, the first execution whose time has stabilized after warm-up, from the steady-state power (SSP) profile, the later execution beyond which power no longer varies substantially. The paper reports that without this differentiation, power and energy measurements can be wrong by as much as 80%, depending on the relative magnitudes of kernel execution time and the power logger's averaging window. With the reconstructed profiles, the paper identifies which GPU sub-components dominate for compute-bound versus memory-bound kernels, and shows that kernels shorter than the averaging window inherit power from kernels that precede them.
Load-bearing premise
The load-bearing premise is that the internal 1 ms power logger really does return a moving average of instantaneous power over the preceding 1 ms, and that the GPU's power state is stable enough across repeated runs for stitching time-shifted samples to reconstruct one true kernel profile; the paper asserts this model but does not independently validate it.
Editorial extensions
If this is right
- For kernels shorter than the power logger's averaging window, a single power sample cannot represent the kernel's energy, so repeated-run stitching is necessary to recover the within-kernel power shape.
- Differentiating SSE from SSP is essential for accurate energy measurement: the paper reports up to 36% power error for a compute-bound 4K GEMM and up to 80% energy error across kernels, with the largest gap when kernel execution time is much shorter than the 1 ms averaging window.
- Kernels shorter than the averaging window inherit power from preceding kernels when executed in an interleaved fashion, so isolated executions are necessary to assess their true power draw.
- The profiles show that compute-heavy GEMMs are dominated by XCD power while bandwidth-bound communication and memory-bound GEMVs stress IOD and HBM, suggesting that complementary kernels could be co-scheduled to use available power headroom.
- The methodology extends to external power loggers such as amd-smi, with the resulting profile quality depending on the averaging window those loggers report.
Reading between the lines
- If the 1 ms moving-average model is exactly right, the same stitching recipe should transfer to any averaged power logger by scaling the number of runs roughly with the averaging window; a 10 ms logger would need about ten times as many runs for the same time resolution.
- The 80% error bound implies that published energy measurements that sample sub-millisecond kernels once per execution may be systematically biased, so kernel-efficiency comparisons that ignore this effect could rank kernels incorrectly.
- A direct test of the method would be to compare a FinGraV stitched profile against a high-bandwidth external power meter on the GPU power rails; if the internal logger's averaging is not a simple moving average, or if power state drifts across runs, the reconstructed profile would need correction.
- For kernels whose execution time is close to or larger than the averaging window, SSE and SSP profiles coincide, so the differentiation overhead may be unnecessary and FinGraV's main benefit is concentrated in the sub-millisecond regime.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FinGraV addresses the difficulty of obtaining fine-grain GPU power profiles for sub-millisecond to few-millisecond AI kernels on the AMD MI300X. The paper identifies four challenges (low native sampling frequency, CPU-GPU time synchronization, execution-time variation, and power variation across executions) and proposes a methodology that combines an internal 1ms power logger, timestamp synchronization, random delays across repeated runs, execution-time binning to discard outlier runs, and a distinction between steady-state execution (SSE) and steady-state power (SSP) profiles. The authors apply FinGraV to compute-bound and memory-bound GEMM/GEMV kernels and to RCCL communication kernels, report total and component-level (XCD, IOD, HBM) power profiles, and derive measurement guidance and optimization recommendations, including the claim that failing to differentiate power profiles can lead to power/energy measurement errors as high as 80%.
Significance. If the central accuracy claim holds, FinGraV would be a useful methodology for power profiling of short kernels on a state-of-the-art accelerator, and the component-level observations would provide actionable guidance for power-aware scheduling and hardware/software optimization. The paper is strongest in its careful enumeration of the practical pitfalls (C1-C4), the proposed synchronization and binning machinery, and the broad set of reported profiles across GEMM, GEMV, and collective communication kernels. However, the methodology's central claim, that the stitched profiles are fine-grain in time, rests on an averaging model that is asserted but not validated, and the evaluation is qualitative rather than quantitative. The headline 80% error figure is also not independently grounded. The work is therefore a promising methodology study whose load-bearing claims need significant additional support.
major comments (3)
- [IV-A (S1), IV-B (step 9), V-C1] Section IV-A states that each power sample from the internal logger is the average of multiple instantaneous power readings over the preceding 1ms. Under this model, a sample taken at offset tau from kernel start is the convolution of the kernel's instantaneous power trace with a 1ms boxcar window, not a pointwise sample of that trace at tau. Random delays between runs shift the phase of the boxcar window; they do not undo the averaging. For kernels in Table I with 25-50us execution times, the stitched SSP profile is therefore a smoothed, aggregate quantity, and the claim of a fine-grain sub-millisecond profile (abstract, Section V-C1) is not supported as stated. The paper should either deconvolve the known or estimated averaging kernel, or explicitly reframe all profiles as 1ms-window-smoothed average power and show that this smoothed quantity supports the subsequent insights.
- [V-C1, Table II] The headline 'measurement error as high as 80%' is presented as the spread between the SSE and SSP profiles, but both profiles are outputs of the same 1ms-averaging logger under different execution histories. Neither is independently established as the true kernel power or energy, so calling their difference a measurement error presumes a reference that is not defined. The paper should define an unambiguous reference (for example, the integral of a high-fidelity instantaneous power trace over a single kernel execution) and provide a quantitative comparison against it; otherwise the 80% figure is a statement about profile variation, not a validated error bound.
- [V-B] The evaluation of the FinGraV methodology is qualitative: it visually compares synchronized vs unsynchronized profiles, binned vs unbinned profiles, and 200-run vs 50-run profiles, and concludes that binning yields a profile 'more tuned to the true shape of power consumed.' No quantitative accuracy metric or independent ground truth is provided for any reconstructed profile. A validation experiment against a high-bandwidth external power meter, or against a synthetic power signal with a known shape, is needed to support the central accuracy claim and the measurement guidance in Table II.
minor comments (4)
- [IV-A (S4) vs IV-B (step 3)] The warm-up count is stated as 'typically three warm-up executions' in Section IV-A, whereas Section IV-B step 3 says to execute the kernel four times, noting that three executions sufficed for stabilization. Please make the recommended procedure and the empirical observation consistent.
- [Figure 9] The labels 'CB– >8K' and 'MB– >4K gemv' appear to be rendering errors for arrows such as 'CB→8K'; please fix the notation.
- [V-B] The footnote marker '12' after 'y-axis' appears malformed; the intended footnote markers should be rendered consistently as superscripts.
- [VII] The paper mentions in the related-work section that clock drift was observed and will be addressed in future work; since Section IV-A (S2) relies on timestamp synchronization, this limitation should be stated prominently in the methodology section so that readers understand the current accuracy limitations of the sync procedure.
Circularity Check
No significant circularity: FinGraV is an empirical measurement study whose error figures are measured differences between defined profiles, not derived from its own inputs.
full rationale
FinGraV does not derive predictions from fitted constants or from a self-citation chain. The SSE/SSP categories are empirical labels defined by observed stabilization of execution time and power, and the 'as high as 80%' figure is a measured relative difference between the two labeled profiles, not a value forced by the definitions themselves. No parameter fitted to data is subsequently renamed as a prediction; the polynomial regression in Section V-B is illustrative smoothing of the same measured data rather than an independent forecast. Self-citations such as [13] for GEMM/communication prevalence are background motivation and are not load-bearing for the measurement methodology. The strongest concern is the acknowledged averaging property of the 1ms GPU power logger (Section IV-A S1): each stitched sample is a moving average over the preceding 1ms, so the 'fine-grain' profiles are boxcar-convolved average-power traces, and true sub-millisecond temporal detail is not recovered without deconvolution. This is a validity and correctness limitation, not a circular reduction: the paper explicitly states that profiles provide 'a fine-grain time view of average power' and that averaging can impact the profiles, and it honestly notes that an instantaneous power sampler would be needed to assess kernels shorter than the averaging window. Because no load-bearing argument reduces to its own input by construction, the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Run count guidance =
400 for 25-50us kernels, 200 otherwise (Table I)
- Binning margin =
5% for <200us, 2% for larger (Table I)
- Warm-up execution count =
3 (typically, Section IV-B)
- Degree of regression for 50-run profile =
4 (Section V-B)
assumptions (4)
- domain assumption MI300X internal power logger reports average power over the last 1ms window (S1, Section IV-A).
- domain assumption CPU-GPU time synchronization can be accurately achieved by benchmarking the timestamp read delay (S2, Section IV-A).
- domain assumption Kernel execution time stabilizes after warmup and outlier runs can be discarded without biasing the power profile (S3, Section IV-A; step 6).
- domain assumption Power variations across repeated executions are primarily due to the logger's averaging window, not to uncontrolled DVFS or temperature dynamics (Section IV-A and V-C3).
Cite this review
Pith. "Pith review of FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights." pith.science (2026). https://pith.science/paper/2DCCO7YN
@misc{pith2026241212426,
author = {Pith},
title = {Pith review of: FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights},
year = {2026},
howpublished = {\url{https://pith.science/paper/2DCCO7YN}},
note = {Machine review of arXiv:2412.12426}
}
read the original abstract
Ubiquity of AI makes optimizing GPU power a priority as large GPU-based clusters are often employed to train and serve AI models. An important first step in optimizing GPU power consumption is high-fidelity and fine-grain power measurement of key AI computations on GPUs. To this end, we observe that as GPUs get more powerful, the resulting sub-millisecond to millisecond executions make fine-grain power analysis challenging. In this work, we first carefully identify the challenges in obtaining fine-grain GPU power profiles. To address these challenges, we devise FinGraV methodology where we employ execution time binning, careful CPU-GPU time synchronization, and power profile differentiation to collect fine-grain GPU power profiles across prominent AI computations and across a spectrum of scenarios. Using the said FinGraV power profiles, we provide both, guidance on accurate power measurement and, in-depth view of power consumption on state-of-the-art AMD Instinct MI300X. For the former, we highlight a methodology for power differentiation across executions. For the latter, we make several observations pertaining to GPU sub-component power consumption and GPU power proportionality across different scenarios. We believe that FinGraV unlocks both an accurate and a deeper view of power consumption of GPUs and opens up avenues for power optimization of these ubiquitous accelerators.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Introducing the AI Research Su- perCluster — Meta’s cutting-edge AI supercomputer for AI research,
Kevin Lee and Shubho Sengupta, “Introducing the AI Research Su- perCluster — Meta’s cutting-edge AI supercomputer for AI research,” https://ai.meta.com/blog/ai-rsc/, 2023
work page 2023
-
[2]
Microsoft announces new supercomputer, lays out vision for future AI work,
Jennifer Langston, “Microsoft announces new supercomputer, lays out vision for future AI work,” https://news.microsoft.com/source/features/ ai/openai-azure-supercomputer/, 2020
work page 2020
- [3]
-
[4]
POLCA: Power Oversubscription in LLM Cloud Providers,
P. Patel, E. Choukse, C. Zhang, ´I˜nigo Goiri, B. Warrier, N. Mahalingam, and R. Bianchini, “POLCA: Power Oversubscription in LLM Cloud Providers,” 2023. [Online]. Available: https://arxiv.org/abs/2308.12908
arXiv 2023
-
[5]
Towards improved power management in cloud gpus,
P. Patel, Z. Gong, S. Rizvi, E. Choukse, P. Misra, T. Anderson, and A. Sriraman, “Towards improved power management in cloud gpus,” Proceedings of the IEEE Computer Architecture Letters (CAL) , 2023
work page 2023
-
[6]
Z. Yang, K. Adamek, and W. Armour, “Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU’s Built-In Power Sensor,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC) , 2024
work page 2024
-
[7]
MI300X powers LLaMA405 at Meta,
“MI300X powers LLaMA405 at Meta,” https://www.linkedin.com/posts/lisasu-amd deep-partnership-with-industry-leaders-is-ugcPost-7261855624236285952-JUxo, 2024
work page 2024
-
[8]
A. Smith, E. Chapman, C. Patel, R. Swaminathan, J. Wuu, T. Huang, W. Jung, A. Kaganov, H. McIntyre, and R. Mangaser, “11.1 AMD Instinct™ MI300 Series Modular Chiplet Package – HPC and AI Accel- erator for Exa-Class Systems,” in Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC) , 2024
work page 2024
Show all 57 references
-
[9]
AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,
A. Smith, G. H. Loh, J. Wuu, S. Naffziger, T. Huang, H. McIntyre, R. Mangaser, W. Jung, and R. Swaminathan, “AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,” in Proceed- ings of the IEEE Symposium on VLSI Technology and Circuits (VLSI Technology an...
2024
-
[10]
The AMD CDNA ™ 3 architecture,
AMD, “The AMD CDNA ™ 3 architecture,” https://www.amd. com/content/dam/amd/en/documents/instinct-tech-docs/white-papers/ amd-cdna-3-white-paper.pdf, 2024
2024
-
[11]
Constraint-Driven Innovation,
James Hamilton, “Constraint-Driven Innovation,” https://mvdirona.com/ jrh/talksandpapers/JamesHamiltonCIDR2024.pdf, 2024
2024
-
[12]
How much electricity does an American home use?
, “How much electricity does an American home use?” https://www.eia. gov/tools/faqs/faq.php?id=97&t=3., 2024
2024
-
[13]
Tale of Two Cs: Computation vs. Communication Scaling for Future Transformers on Future Hardware,
S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair, “Tale of Two Cs: Computation vs. Communication Scaling for Future Transformers on Future Hardware,” in Proceedings of the IEEE International Symposium on Workload Characterization (IISWC) , 2023
2023
-
[14]
AMD SMI documentation,
AMD, “AMD SMI documentation,” https://rocm.docs.amd.com/projects/ amdsmi/en/latest/, 2024
2024
-
[15]
AMD ROCm ™ Software,
——, “AMD ROCm ™ Software,” https://www.amd.com/en/products/ software/rocm.html, 2024
2024
-
[16]
ROCm ™/rocBLAS: Next generation BLAS implementation for ROCm™ platform,
——, “ROCm ™/rocBLAS: Next generation BLAS implementation for ROCm™ platform,” https://github.com/ROCm/rocBLAS, 2024
2024
-
[17]
ROCm ™ Communication Collectives Library,
——, “ROCm ™ Communication Collectives Library,” https://github. com/ROCm/rccl, 2024
2024
-
[18]
NanoFlow: Towards Optimal Large Language Model Serving Throughput,
K. Zhu, Y . Zhao, L. Zhao, G. Zuo, Y . Gu, D. Xie, Y . Gao, Q. Xu, T. Tang, Z. Ye, K. Kamahori, C.-Y . Lin, S. Wang, A. Krishnamurthy, and B. Kasikci, “NanoFlow: Towards Optimal Large Language Model Serving Throughput,” 2024. [Online]. Available: https://arxiv.org/abs/2408.12757
2024 arXiv
-
[19]
System Management Interface SMIn,
NVIDIA, “System Management Interface SMIn,” https://developer. nvidia.com/system-management-interface, 2024
2024
-
[20]
Variorum,
Variorum, “Variorum,” https://variorum.readthedocs.io/en/latest/index. html, 2023
2023
-
[21]
Standardizing Power Monitoring and Control at Exascale,
R. E. Grant, M. Levenhagen, S. L. Olivier, D. DeBonis, K. T. Pedretti, and J. H. Laros III, “Standardizing Power Monitoring and Control at Exascale,” Computer, 2016
2016
-
[22]
PowerSensor 2: A Fast Power Mea- surement Tool,
J. W. Romein and B. Veenboer, “PowerSensor 2: A Fast Power Mea- surement Tool,” in Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2018
2018
-
[23]
Measuring GPU Power with the K20 Built-in Sensor,
M. Burtscher, I. Zecena, and Z. Zong, “Measuring GPU Power with the K20 Built-in Sensor,” in Proceedings of the Workshop on General Purpose Processing Using GPUs (GPGPU) , 2014
2014
-
[24]
Towards Accurate and Reliable Energy Measurement of NLP Models,
Q. Cao, A. Balasubramanian, and N. Balasubramanian, “Towards Accurate and Reliable Energy Measurement of NLP Models,” 2020. [Online]. Available: https://arxiv.org/abs/2010.05248
2020 arXiv
-
[25]
A Comparative Study of Techniques for Energy Predictive Modeling Using Performance Monitoring Counters on Modern Multicore CPUs,
A. Shahid, M. Fahad, R. R. Manumachu, and A. Lastovetsky, “A Comparative Study of Techniques for Energy Predictive Modeling Using Performance Monitoring Counters on Modern Multicore CPUs,” IEEE Access, 2020
2020
-
[26]
An experimental comparison of software-based power me- ters: focus on CPU and GPU,
M. Jay, V . Ostapenco, L. Lefevre, D. Trystram, A.-C. Orgerie, and B. Fichel, “An experimental comparison of software-based power me- ters: focus on CPU and GPU,” in Proceedings of the International Symposium on Cluster, Cloud and Internet Computing (CCGrid) , 2023
2023
-
[27]
AccelWattch: A Power Modeling Framework for Modern GPUs,
V . Kandiah, S. Peverelle, M. Khairy, J. Pan, A. Manjunath, T. G. Rogers, T. M. Aamodt, and N. Hardavellas, “AccelWattch: A Power Modeling Framework for Modern GPUs,” in Proceedings of the IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2021
2021
-
[28]
Understanding the Future of Energy Efficiency in Multi-Module GPUs,
A. Arunkumar, E. Bolotin, D. Nellans, and C.-J. Wu, “Understanding the Future of Energy Efficiency in Multi-Module GPUs,” in Proceed- ings of the International Symposium on High Performance Computer Architecture (HPCA), 2019
2019
-
[29]
Measuring and modeling on-chip interconnect power on real hardware,
V . Adhinarayanan, I. Paul, J. L. Greathouse, W. Huang, A. Pattnaik, and W.-c. Feng, “Measuring and modeling on-chip interconnect power on real hardware,” in Proceedings of the IEEE International Symposium on Workload Characterization (IISWC) , 2016
2016
-
[30]
Power and Performance Characterization and Modeling of GPU-Accelerated Systems,
Y . Abe, H. Sasaki, S. Kato, K. Inoue, M. Edahiro, and M. Peres, “Power and Performance Characterization and Modeling of GPU-Accelerated Systems,” in Proceedings of the IEEE 28th International Parallel and Distributed Processing Symposium (IPDPS) , 2014
2014
-
[31]
Online Power Estimation of Graphics Processing Units,
V . Adhinarayanan, B. Subramaniam, and W.-C. Feng, “Online Power Estimation of Graphics Processing Units,” in Proceedings of the IEEE/ACM International Symposium on Cluster, Cloud and Grid Com- puting (CCGrid), 2016
2016
-
[32]
GPGPU performance and power estimation using machine learning,
G. Wu, J. L. Greathouse, A. Lyashevsky, N. Jayasena, and D. Chiou, “GPGPU performance and power estimation using machine learning,” in Proceedings of the IEEE 21st International Symposium on High Performance Computer Architecture (HPCA) , 2015
2015
-
[33]
High-Resolution Power Profiling of GPU Functions Using Low-Resolution Measurement,
J. Lang and G. R ¨unger, “High-Resolution Power Profiling of GPU Functions Using Low-Resolution Measurement,” in Proceedings of the European Conference on Parallel Processing (Euro-Par) , 2013
2013
-
[34]
Optimizing performance-per-watt on GPUs in high performance com- puting: Temperature, frequency and voltage effects,
D. C. Price, M. A. Clark, B. R. Barsdell, R. Babich, and L. J. Greenhill, “Optimizing performance-per-watt on GPUs in high performance com- puting: Temperature, frequency and voltage effects,” Computer Science - Research and Development , 2015
2015
-
[35]
Benchmarking the Performance and Energy Efficiency of AI Acceler- ators for AI Training,
Y . Wang, Q. Wang, S. Shi, X. He, Z. Tang, K. Zhao, and X. Chu, “Benchmarking the Performance and Energy Efficiency of AI Acceler- ators for AI Training,” in Proceedings of the IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID), 2020
2020
-
[36]
Efficiency Near the Edge: Increasing the Energy Efficiency of FFTs on GPUs for Real-Time Edge Computing,
K. Ad ´amek, J. Novotn ´y, J. Thiyagalingam, and W. Armour, “Efficiency Near the Edge: Increasing the Energy Efficiency of FFTs on GPUs for Real-Time Edge Computing,” IEEE Access, 2021
2021
-
[37]
GPU-NEST: Characterizing Energy Efficiency of Multi-GPU Inference Servers,
A. Jahanshahi, H. Z. Sabzi, C. Lau, and D. Wong, “GPU-NEST: Characterizing Energy Efficiency of Multi-GPU Inference Servers,” IEEE Computer Architecture Letters , 2020
2020
-
[38]
Carbon Emissions and Large Neural Network Training,
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon Emissions and Large Neural Network Training,” 2021. [Online]. Available: https://arxiv.org/abs/2104.10350
2021 arXiv
-
[39]
Cutting the cost of pulsar astronomy: Saving time and energy when searching for binary pulsars using NVIDIA GPUs,
J. White, K. Adamek, and W. Armour, “Cutting the cost of pulsar astronomy: Saving time and energy when searching for binary pulsars using NVIDIA GPUs,” 2022. [Online]. Available: https://arxiv.org/abs/2211.13517
2022 arXiv
-
[40]
Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving,
J. Yu, J. Kim, and E. Seo, “Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA), 2023
2023
-
[41]
On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,
G. Schieffer, D. A. De Medeiros, J. Faj, A. Marathe, and I. Peng, “On the Rise of AMD Matrix Cores: Performance, Power Efficiency, and Programmability,” in Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2024
2024
-
[42]
A measurement study of GPU DVFS on energy conservation,
X. Mei, L. S. Yung, K. Zhao, and X. Chu, “A measurement study of GPU DVFS on energy conservation,” in Proceedings of the Workshop on Power-Aware Computing and Systems (HotPower) , 2013
2013
-
[43]
Comparing GPU Power and Frequency Capping: A Case Study with the MuMMI Workflow,
T. Patki, Z. Frye, H. Bhatia, F. Di Natale, J. Glosli, H. Ingolfsson, and B. Rountree, “Comparing GPU Power and Frequency Capping: A Case Study with the MuMMI Workflow,” in Proceedings of the IEEE/ACM Workflows in Support of Large-Scale Science (WORKS) , 2019
2019
-
[44]
The Impact of GPU DVFS on the Energy and Performance of Deep Learning: an Empirical Study,
Z. Tang, Y . Wang, Q. Wang, and X. Chu, “The Impact of GPU DVFS on the Energy and Performance of Deep Learning: an Empirical Study,” in Proceedings of the ACM International Conference on Future Energy Systems (e-Energy), 2019
2019
-
[45]
Performance/Energy Aware Optimiza- tion of Parallel Applications on GPUs Under Power Capping,
A. Krzywaniak and P. Czarnul, “Performance/Energy Aware Optimiza- tion of Parallel Applications on GPUs Under Power Capping,” in Parallel Processing and Applied Mathematics , 2020
2020
-
[46]
Input-Dependent Power Usage in GPUs,
T. Gregersen, P. Patel, and E. Choukse, “Input-Dependent Power Usage in GPUs,” 2024. [Online]. Available: https://arxiv.org/abs/2409.18324
2024 arXiv
-
[47]
Dynamic GPGPU Power Management Using Adaptive Model Predictive Control,
A. Majumdar, L. Piga, I. Paul, J. L. Greathouse, W. Huang, and D. H. Albonesi, “Dynamic GPGPU Power Management Using Adaptive Model Predictive Control,” in Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA), 2017
2017
-
[48]
Predict; Do not React for Enabling Efficient Fine Grain DVFS in GPUs,
S. Bharadwaj, S. Das, K. Mazumdar, B. Beckmann, and S. Kosonocky, “Predict; Do not React for Enabling Efficient Fine Grain DVFS in GPUs,” 2022. [Online]. Available: https://arxiv.org/abs/2205.00121
2022 arXiv
-
[49]
Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance Assurance,
Y . Zhang, Q. Wang, Z. Lin, P. Xu, and B. Wang, “Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance Assurance,” in Proceedings of the European Conference on Computer Systems (EuroSys) , 2024
2024
-
[50]
DRLCAP: Runtime GPU Frequency Capping With Deep Reinforce- ment Learning,
Y . Wang, M. Hao, H. He, W. Zhang, Q. Tang, X. Sun, and Z. Wang, “DRLCAP: Runtime GPU Frequency Capping With Deep Reinforce- ment Learning,” Proceedings of the IEEE Transactions on Sustainable Computing, 2024
2024
-
[51]
Going green: optimizing GPUs for energy efficiency through model-steered auto-tuning,
R. Schoonhoven, B. Veenboer, B. van Werkhoven, and K. J. Batenburg, “Going green: optimizing GPUs for energy efficiency through model-steered auto-tuning,” 2022. [Online]. Available: https: //arxiv.org/abs/2211.07260
2022 arXiv
-
[52]
Energy-Aware Tile Size Selection for Affine Programs on GPUs,
M. Jayaweera, M. Kong, Y . Wang, and D. Kaeli, “Energy-Aware Tile Size Selection for Affine Programs on GPUs,” in Proceedings of the IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 2024
2024
-
[53]
Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training,
J. You, J.-W. Chung, and M. Chowdhury, “Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training,” in Proceed- ings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2023
2023
-
[54]
Reducing Energy Bloat in Large Model Training,
J.-W. Chung, Y . Gu, I. Jang, L. Meng, N. Bansal, and M. Chowdhury, “Reducing Energy Bloat in Large Model Training,” in Proceedings of the ACM SIGOPS Symposium on Operating Systems Principles (SOSP) . ACM, 2024
2024
-
[55]
EnvPipe: Performance- preserving DNN training framework for saving energy,
S. Choi, I. Koo, J. Ahn, M. Jeon, and Y . Kwon, “EnvPipe: Performance- preserving DNN training framework for saving energy,” in Proceedings of the USENIX Annual Technical Conference (USENIX ATC) , 2023
2023
-
[56]
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency,
J. Stojkovic, C. Zhang, ´I˜nigo Goiri, J. Torrellas, and E. Choukse, “DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency,” 2024. [Online]. Available: https://arxiv.org/abs/ 2408.00741
2024
-
[57]
Splitwise: Efficient generative LLM inference using phase splitting,
P. Patel, E. Choukse, C. Zhang, A. Shah, ´I˜nigo Goiri, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative LLM inference using phase splitting,” 2024. [Online]. Available: https://arxiv.org/abs/2311.18677
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.