REVIEW 4 major objections 5 minor 55 references
THOR: A Generic Energy Estimation Approach for On-Device Training
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that a DNN's on-device training energy can be estimated as the sum of per-layer Gaussian-process predictions, cutting mean absolute percentage error from about 40% to around 10% across five devices.
desk verdict A sensible engineering method for on-device training energy estimation, with a load-bearing additivity assumption that is only weakly tested; deserves a referee but needs a direct subtractivity check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the layer-wise energy additivity and subtractivity assumption, stated in Section 3.2 and encoded in Eqs. (1) and (2), together with the Gaussian processes that fit each layer's energy as a smooth function of channel counts. A GP is a non-parametric probabilistic regressor whose predictions come with uncertainty, which the paper exploits to guide the choice of the next profiling point by maximum variance. The Matérn kernel with $\nu = 2.5$ supplies the GP's covariance, chosen to tolerate runtime optimization artifacts. The final estimator, Eq. (4), simply sums the per-layer GP predictions; profiling is a one-time cost per device and framework, and the paper reports most profiling and fitting runs finish within 20 minutes.
What would settle it
Build a two-convolutional-layer model and measure its training energy directly; then measure a one-layer model with the same first convolution and a one-layer model with the same second convolution, add them, and compare with the two-layer measurement over a range of channel counts on a device with operator fusion enabled. If the mismatch systematically exceeds the measurement noise, additivity fails and the subtractive profiling labels in Eqs. (1) and (2) are biased.
Extended reading notes
Core claim
The paper's central discovery candidate is that training energy obeys layer-wise additivity well enough to be instrumented: a model's total energy is the sum of its input, hidden, and output layer energies, and each layer's cost can be isolated by subtracting the other layers' costs from the measured total of a small variant network. Using this, THOR profiles one-, two-, and three-layer probe networks, fits a Gaussian process per layer type with the Matérn kernel, and predicts an unseen network's energy as $\hat{E}_{\mathrm{model}} = \hat{E}_{\mathrm{input}}(C_1) + \sum_{i=2}^{n-1} \hat{E}_{\mathrm{hidden}}(C_{i-1}, C_i) + \hat{E}_{\mathrm{output}}(C_{n-1})$. The paper's evaluation reports MAPE around 10% across LeNet-5, a 5-layer CNN, HAR, LSTM, Transformer, and ResNet variants on five devices, versus roughly 40% for FLOPs-based estimation. The paper acknowledges that the decomposition assumes sequential layer execution, that parallel-branch architectures such as GoogleNet and SqueezeNet fall outside the current treatment, and that estimates degrade when the framework version changes.
Load-bearing premise
The load-bearing premise is that a layer's training energy is independent of its neighbors, so the cost of any layer can be recovered by subtracting the other layers' costs from the measured total of a small probe network, and the whole model's energy is exactly the sum of these independent layer costs.
Editorial extensions
If this is right
- Energy-aware job schedulers can treat the per-layer GP sum as a cheap, differentiable budget oracle for on-device training.
- Energy-constrained pruning can use THOR as the objective: the case study hits a 50% energy budget while the FLOPs-based baseline misses it.
- One-time profiling per device and framework yields reusable layer models, so unseen architectures built from profiled layer types can be estimated without new measurements.
- Because the estimator tracks nonlinear layer behavior (plateaus and ridges) that FLOPs misses, it can detect slow spots or energy cliffs missed by compute-based proxies.
- The method's scope is sequential execution; the paper explicitly leaves parallel-branch models and framework-version drift as limitations.
Reading between the lines
- If additivity holds beyond the tested models, the same subtractive recipe could produce per-layer estimates for other resources, such as latency, memory bandwidth, or thermal load, extending the profiling idea beyond energy.
- The GP's predictive variance is a natural input for scheduling under probabilistic energy budgets, a use the paper hints at but does not develop.
- A direct stress test would be to profile GoogleNet or SqueezeNet-style multi-branch blocks; until that is done, the additivity claim is effectively established only for sequential architectures.
- The framework-version sensitivity reported (MAPE rising from 8% to 11% after an upgrade) suggests that a lightweight re-calibration rule could be derived from the GP's posterior variance, an extension the paper leaves implicit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes THOR, a per-layer Gaussian Process (GP) based energy estimator for on-device DNN training. The method assumes layer-wise energy additivity and subtractivity, profiles one-, two-, and three-layer probe networks to obtain per-layer energy labels, fits GP models to those labels, and then predicts total training energy as the sum of per-layer GP predictions (Eq. 4). The evaluation compares THOR against a FLOPs-based linear baseline on five devices (OPPO, iPhone, Xavier, TX2, Server) and several model families (LeNet-5, a 5-layer CNN, HAR, LSTM, Transformer, ResNet), reporting an average MAPE around 10% versus roughly 40% for the baseline, and a pruning case study on Xavier that claims a 50% energy reduction.
Significance. If the layer-wise additivity premise holds, THOR is a practically attractive approach: it is device-agnostic, uses active-learning-style GP profiling with uncertainty-based termination, and its per-layer decomposition could enable energy-aware scheduling and pruning. The breadth of the empirical study across heterogeneous devices and model types is a real strength, as is the use of a nonparametric model that does not require hardware-specific analytical simulators. However, the central additivity/subtractivity assumption is validated only qualitatively, the per-layer labels are derived by subtracting GP estimates rather than by direct measurement, and the main MAPE figures lack error bars. These issues currently limit confidence in the central claim and in the pruning case study.
major comments (4)
- [Sec. 3.2, Fig. 2, Eqs. (1)-(2)] The layer-wise energy additivity/subtractivity premise is load-bearing but is validated only visually on a CNN with identical stacked Conv2d layers. The paper's own Sec. 2.3 states that frameworks fuse operations such as Conv-BN-ReLU into a single kernel, making execution 'more like a black box'; under such fusion, per-layer energy is not independently measurable, so subtracting estimated energies of other layers from measured totals can produce biased hidden-layer labels. Please provide a direct validation of subtractivity on architectures that include BatchNorm fusion, residual connections, attention blocks, and recurrent cells, or otherwise demonstrate that the per-layer decomposition in Eqs. (1), (2), and (4) does not inherit a systematic bias from cross-layer coupling.
- [Eqs. (1)-(2), Sec. 4.2] The hidden- and input-layer labels are formed by subtracting GP-based estimates (\hat E_output and \hat E_input) from measured totals. Any systematic error in those GPs enters the training labels for the remaining GPs and propagates into the final sum in Eq. (4). The reported end-to-end MAPE can be accurate even when per-layer estimates are badly biased, because over- and under-estimates from different layers may cancel in the sum. The paper should report per-layer estimation errors (e.g., quantified MAPE/RMSE for the surfaces in Figs. 11-12), provide an error-propagation analysis, or compare the subtraction-derived labels against a direct per-layer measurement on at least one device.
- [Sec. 4.1, Fig. 8, Appendix A5.1] The main quantitative claim (MAPE around 10% versus roughly 40% for the FLOPs baseline) is presented as bar charts without error bars or confidence intervals, even though Appendix A5.1 states that each experiment was repeated three times and that the mean and standard error are reported. Without these values, the reader cannot judge whether the observed improvements are statistically significant, particularly for models with higher error such as HAR or LSTM on smartphones. Please add error bars to Figs. 8-10 and include a table with per-device, per-model means and standard errors.
- [Sec. 4.3, Fig. 13] The pruning case study claims that THOR reduces energy consumption by 50% while preserving accuracy, but it is unclear whether the reported 49.2% figure is a measured energy value or a THOR prediction. The claim also depends on per-layer GP gradients whose accuracy is not independently established. Please report the actually measured energy of the pruned model, specify the measurement procedure, and include multiple devices or models with variance across runs to support the generality of the claim.
minor comments (5)
- [Sec. 3.2] There is a typo in the sentence 'Based on the presented layer-wise energy additivity pf DNN'; it should read 'of DNN'. Similarly, Sec. 3.3 contains 'we use the the upper and lower bounds', which should be 'the upper and lower bounds'.
- [Fig. 2 caption] The caption says 'Energy consumption from NeuralPower estimation and from observation for a CNN', but the surrounding text describes an adapted training-phase profiling method. Please clarify exactly what is plotted, how NeuralPower was adapted to training, and how the linear trend supports additivity rather than merely showing a correlation.
- [Fig. 11] The subcaptions are garbled and repeated: the six panels are labeled '(a) H=W=42, N=10 (b) ... (a) Xavier, H=W=42 ...' in an overlapping way. Please relabel the panels uniquely by device and spatial size so the reader can map each surface to the correct setting.
- [References] The bibliography appears to contain a duplicated and largely unrelated block of references (e.g., Abelson et al. 1985, Baumgartner et al. 2001, Brachman and Schmolze 1985, Gottlob 1992, Levesque 1984, Nebel 2000) that are not cited in the text. Please remove this block and ensure every reference is cited and every citation appears in the reference list.
- [Abstract and Sec. 1] The phrase 'reduced the Mean Absolute Percentage Error (MAPE) by up to 30%' is ambiguous because MAPE is already a percentage. Since the reported improvement is from about 40% to about 10%, please state the change as an absolute percentage-point reduction or as a relative reduction to avoid confusion.
Circularity Check
No significant circularity: THOR is an empirical GP regression against external device measurements; the subtractive label construction is a decomposition assumption, not a self-fulfilling prediction.
full rationale
THOR's derivation chain is an empirical measurement pipeline. The claimed result, Eq. (4), sums GP models fitted to energy labels obtained by physically training 1-, 2-, and 3-layer probe networks on real devices (Sec. 3.2, Eqs. (1)-(2)); the ground-truth totals in the evaluation are external power measurements, not outputs of the model. The subtractive construction of hidden-layer labels (E_hidden = E_model - E_input_hat - E_output_hat) is a decomposition assumption, and it does create indirect coupling among the fitted GPs. However, that is a model-identification/robustness threat, not a circular reduction: the paper does not present the training profiles as predictions, and the reported MAPE is computed on randomly sampled, unseen channel configurations (Sec. 4.1) against independently measured totals, so the end-to-end estimate is not equal to its training labels by construction. The pruning case study similarly validates against a measured 20,000J baseline and a measured final energy of 49.2%. Self-citations in the paper (e.g., [Li et al., 2021], which shares co-author Miao Pan) are used only to contextualize prior energy assumptions and are not load-bearing for THOR's method. No uniqueness theorem, ansatz, or fitted parameter is imported from self-citation as a substitute for evidence, and no result in the paper is renamed from a known empirical pattern. Therefore the derivation chain is self-contained with respect to circularity, and the appropriate finding is no significant circularity (score 0).
Assumptions & free parameters
free parameters (5)
- Matérn kernel hyperparameters (length scale, signal variance, noise) =
not reported; learned from profiling data
- Matérn smoothness parameter nu =
2.5
- Profiling stopping thresholds (max points, variance threshold) =
variance threshold 5%; max points not specified
- Profiling iteration count =
500 batches
- Energy sampling interval =
100 ms on phones/Jetson; 50 Hz nvidia-smi on server
assumptions (6)
- domain assumption Total training energy is the sum of independent per-layer energies (additivity), and a layer's energy equals total minus the other layers' energies (subtractivity).
- domain assumption Layer energy depends only on channel counts and layer type for a fixed device, framework, input shape, and batch size, so a GP over channels can generalize to unseen widths.
- domain assumption Layers execute sequentially with negligible inter-layer effects.
- domain assumption Training time is positively correlated with training energy, so time uncertainty can be used as a surrogate for energy uncertainty during guided profiling.
- domain assumption Energy measured with the E = P-bar times T approximation at 100 ms sampling adequately represents true training energy.
- standard math Gaussian Process regression with the Matérn kernel is a valid model for layer-wise energy functions.
Cite this review
Pith. "Pith review of THOR: A Generic Energy Estimation Approach for On-Device Training." pith.science (2026). https://pith.science/paper/MCWSWXK6
@misc{pith2026250116397,
author = {Pith},
title = {Pith review of: THOR: A Generic Energy Estimation Approach for On-Device Training},
year = {2026},
howpublished = {\url{https://pith.science/paper/MCWSWXK6}},
note = {Machine review of arXiv:2501.16397}
}
read the original abstract
Battery-powered mobile devices (e.g., smartphones, AR/VR glasses, and various IoT devices) are increasingly being used for AI training due to their growing computational power and easy access to valuable, diverse, and real-time data. On-device training is highly energy-intensive, making accurate energy consumption estimation crucial for effective job scheduling and sustainable AI. However, the heterogeneity of devices and the complexity of models challenge the accuracy and generalizability of existing estimation methods. This paper proposes THOR, a generic approach for energy consumption estimation in deep neural network (DNN) training. First, we examine the layer-wise energy additivity property of DNNs and strategically partition the entire model into layers for fine-grained energy consumption profiling. Then, we fit Gaussian Process (GP) models to learn from layer-wise energy consumption measurements and estimate a DNN's overall energy consumption based on its layer-wise energy additivity property. We conduct extensive experiments with various types of models across different real-world platforms. The results demonstrate that THOR has effectively reduced the Mean Absolute Percentage Error (MAPE) by up to 30%. Moreover, THOR is applied in guiding energy-aware pruning, successfully reducing energy consumption by 50%, thereby further demonstrating its generality and potential.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
[Ahmad et al., 2015] Raja Wasim Ahmad, Abdullah Gani, Siti Hafizah Ab Hamid, Feng Xia, and Muhammad Shiraz. A Review on Mobile Application Energy Profiling: Tax- onomy, State-of-the-art, and Open Research Issues. Jour- nal of Network and Computer Applications , pages 42–59,
work page 2015
-
[4]
Neuralpower: Predict and Deploy Energy-efficient Convolutional Neural Networks
[Cai et al., 2017] Ermao Cai, Da-Cheng Juan, Dimitrios Sta- moulis, and Diana Marculescu. Neuralpower: Predict and Deploy Energy-efficient Convolutional Neural Networks. In ACML,
work page 2017
-
[10]
[Dai et al., 2019] Xiaoliang Dai, Peizhao Zhang, Bichen Wu, Hongxu Yin, Fei Sun, Yanghan Wang, Marat Dukhan, Yunqing Hu, Yiming Wu, Yangqing Jia, Peter Vajda, Matt Uyttendaele, and Niraj K. Jha. ChamNet: Towards Ef- ficient Network Design through Platform-Aware Model Adaptation. In CVPR,
work page 2019
-
[13]
HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients
[Diao et al., 2021] Enmao Diao, Jie Ding, and Vahid Tarokh. HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients. In ICLR,
work page 2021
-
[14]
Ember: Energy Management of Batteryless Event Detection Sensors with Deep Reinforcement Learning
[Fraternali et al., 2020] Francesco Fraternali, Bharathan Balaji, Dhiman Sengupta, Dezhi Hong, and Rajesh K Gupta. Ember: Energy Management of Batteryless Event Detection Sensors with Deep Reinforcement Learning. In SenSys,
work page 2020
-
[16]
Memory- Efficient DNN Training on Mobile Devices
[Gim and Ko, 2022] In Gim and JeongGil Ko. Memory- Efficient DNN Training on Mobile Devices. In MobiSys,
work page 2022
-
[17]
Chasing Carbon: The Elu- sive Environmental Footprint of Computing
[Gupta et al., 2021] Udit Gupta, Young Geun Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin S Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. Chasing Carbon: The Elu- sive Environmental Footprint of Computing. In HPCA,
work page 2021
-
[18]
Learning both weights and connections for efficient neural networks
[Han et al., 2015] Song Han, Jeff Pool, John Tran, and William J Dally. Learning both weights and connections for efficient neural networks. In NeurIPS,
work page 2015
Show all 55 references
-
[19]
Deep Residual Learning for Image Recognition
[He et al., 2016] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR,
2016
-
[20]
Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters
[Hu et al., 2021] Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen, and Tianwei Zhang. Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters. In SC,
2021
-
[21]
TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions
[Jia et al., 2019] Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions. In SOSP,
2019
-
[22]
Accel- Wattch: A Power Modeling Framework for Modern GPUs
[Kandiah et al., 2021] Vijay Kandiah, Scott Peverelle, Mah- moud Khairy, Junrui Pan, Amogh Manjunath, Timothy G Rogers, Tor M Aamodt, and Nikos Hardavellas. Accel- Wattch: A Power Modeling Framework for Modern GPUs. In MICRO,
2021
-
[23]
Gradient-Based Learning Applied to Document Recognition
[LeCun et al., 1998] Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-Based Learning Applied to Document Recognition. Proceedings of the IEEE, 86(11):2278–2324,
1998
-
[25]
To Talk or to Work: Flexible Communication Compression for Energy Efficient Feder- ated Learning over Heterogeneous Mobile Edge Devices
[Li et al., 2021] Liang Li, Dian Shi, Ronghui Hou, Hui Li, Miao Pan, and Zhu Han. To Talk or to Work: Flexible Communication Compression for Energy Efficient Feder- ated Learning over Heterogeneous Mobile Edge Devices. In INFOCOM,
2021
-
[26]
Revis- iting Random Channel Pruning for Neural Network Com- pression
[Li et al., 2022] Yawei Li, Kamil Adamczewski, Wen Li, Shuhang Gu, Radu Timofte, and Luc Van Gool. Revis- iting Random Channel Pruning for Neural Network Com- pression. In CVPR,
2022
-
[27]
Deep Learning Face Attributes in the Wild
[Liu et al., 2015] Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep Learning Face Attributes in the Wild. In ICCV,
2015
-
[28]
DARTS: Differentiable Architecture Search
[Liu et al., 2018] Hanxiao Liu, Karen Simonyan, and Yim- ing Yang. DARTS: Differentiable Architecture Search. In ICLR,
2018
-
[29]
Augur: Mod- eling the Resource Requirements of ConvNets on Mo- bile Devices
[Lu et al., 2019] Zongqing Lu, Swati Rallapalli, Kevin Chan, Shiliang Pu, and Thomas La Porta. Augur: Mod- eling the Resource Requirements of ConvNets on Mo- bile Devices. IEEE Transactions on Mobile Computing , 20:352–365,
2019
-
[30]
Clegg, Andrea Cavallaro, and Hamed Haddadi
[Malekzadeh et al., 2019] Mohammad Malekzadeh, Richard G. Clegg, Andrea Cavallaro, and Hamed Haddadi. Mobile Sensor Data Anonymization. In IoTDI,
2019
-
[31]
Distreal: Distributed Resource-Aware Learning in Heterogeneous Systems
[Martin Rapp and Ramin Khalili and Kilian Pfeiffer and Jörg Henkel, 2022] Martin Rapp and Ramin Khalili and Kilian Pfeiffer and Jörg Henkel. Distreal: Distributed Resource-Aware Learning in Heterogeneous Systems. In AAAI,
2022
-
[32]
[Mei et al., 2022] Linyan Mei, Huichu Liu, Tony F. Wu, H. Ekin Sumbul, Marian Verhelst, and Edith Beigné. A Uniform Latency Model for DNN Accelerators with Di- verse Architectures and Dataflows. In DATE,
2022
-
[33]
MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs
[Monil et al., 2020] Mohammad Alaul Haque Monil, Mehmet E Belviranli, Seyong Lee, Jeffrey S Vet- ter, and Allen D Malony. MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-Offs. In PACT,
2020
-
[34]
Efficient Federated Learning Algorithm for Re- source Allocation in Wireless IoT Networks
[Nguyen et al., 2020] Van-Dinh Nguyen, Shree Krishna Sharma, Thang X Vu, Symeon Chatzinotas, and Björn Ot- tersten. Efficient Federated Learning Algorithm for Re- source Allocation in Wireless IoT Networks. IEEE Inter- net of Things Journal, 8(5):3394–3409,
2020
-
[35]
Capuchin: Tensor-Based GPU Memory Manage- ment for Deep Learning
[Peng et al., 2020] Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian. Capuchin: Tensor-Based GPU Memory Manage- ment for Deep Learning. In ASPLOS,
2020
-
[36]
[Qasaimeh et al., 2019] Murad Qasaimeh, Kristof Denolf, Jack Lo, Kees Vissers, Joseph Zambreno, and Phillip H. Jones. Comparing Energy Efficiency of CPU, GPU and FPGA Implementations for Vision Kernels. In ICESS,
2019
-
[37]
[Qi et al., 2017] Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS,
2017
-
[38]
Delay and Energy Aware Task Schedul- ing Mechanism for Fog-Enabled IoT Applications: A Re- inforcement Learning Approach
[Raju and Mothku, 2023] Mekala Ratna Raju and Sai Kr- ishna Mothku. Delay and Energy Aware Task Schedul- ing Mechanism for Fog-Enabled IoT Applications: A Re- inforcement Learning Approach. Computer Networks , 224:109603,
2023
-
[39]
Riley, and Mikel Luján
[Rodrigues et al., 2018] Crefeda Faviola Rodrigues, Gra- ham D. Riley, and Mikel Luján. SyNERGY: An En- ergy Measurement and Prediction Framework for Convo- lutional Neural Networks on Jetson TX1. InPDPTA,
2018
-
[40]
DeLight: Adding Energy Dimension to Deep Neural Networks
[Rouhani et al., 2016] Bita Darvish Rouhani, Azalia Mirho- seini, and Farinaz Koushanfar. DeLight: Adding Energy Dimension to Deep Neural Networks. In ISLPED,
2016
-
[41]
Energy-Resilient Real-Time Scheduling
[Shirazi et al., 2023] Mahmoud Shirazi, Lothar Thiele, and Mehdi Kargahi. Energy-Resilient Real-Time Scheduling. IEEE Transactions on Computers, 72:69–81,
2023
-
[42]
Hyperpower: Power- and Memory-Constrained Hyper-Parameter Opti- mization for Neural Networks
[Stamoulis et al., 2018] Dimitrios Stamoulis, Ermao Cai, Da-Cheng Juan, and Diana Marculescu. Hyperpower: Power- and Memory-Constrained Hyper-Parameter Opti- mization for Neural Networks. In DATE,
2018
-
[43]
Energy and Policy Considerations for Modern Deep Learning Research
[Strubell et al., 2020] Emma Strubell, Ananya Ganesh, and Andrew McCallum. Energy and Policy Considerations for Modern Deep Learning Research. In AAAI,
2020
-
[44]
Generalized Latency Perfor- mance Estimation for Once-For-All Neural Architecture Search
[Syed and Srinivasan, 2021] Muhtadyuzzaman Syed and Arvind Akpuram Srinivasan. Generalized Latency Perfor- mance Estimation for Once-For-All Neural Architecture Search. arXiv preprint arXiv:2101.00732,
2021 arXiv
-
[45]
Efficient Processing of Deep Neu- ral Networks: A Tutorial and Survey
[Sze et al., 2017] Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S Emer. Efficient Processing of Deep Neu- ral Networks: A Tutorial and Survey. Proceedings of the IEEE, 105:2295–2329,
2017
-
[46]
Fed- erated Learning over Wireless Networks: Optimization Model Design and Analysis
[Tran et al., 2019] Nguyen H Tran, Wei Bao, Albert Zomaya, Minh NH Nguyen, and Choong Seon Hong. Fed- erated Learning over Wireless Networks: Optimization Model Design and Analysis. In INFOCOM,
2019
-
[47]
Attention Is All You Need
[Vaswani et al., 2017] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. In NeurIPS,
2017
-
[48]
Mixed-Criticality Scheduling of Energy-Harvesting Sys- tems
[Wang and Deng, 2022] Kankan Wang and Qingxu Deng. Mixed-Criticality Scheduling of Energy-Harvesting Sys- tems. In RTSS,
2022
-
[49]
Gaussian Processes for Ma- chine Learning, volume
[Williams and Rasmussen, 2006] Christopher KI Williams and Carl Edward Rasmussen. Gaussian Processes for Ma- chine Learning, volume
2006
-
[51]
Designing Energy-Efficient Convolutional Neu- ral Networks using Energy-Aware Pruning
[Yang et al., 2017] Tien-Ju Yang, Yu-Hsin Chen, and Vivi- enne Sze. Designing Energy-Efficient Convolutional Neu- ral Networks using Energy-Aware Pruning. In CVPR,
2017
-
[52]
Energy Efficient Federated Learning over Wireless Com- munication Networks
[Yang et al., 2020] Zhaohui Yang, Mingzhe Chen, Walid Saad, Choong Seon Hong, and Mohammad Shikh-Bahaei. Energy Efficient Federated Learning over Wireless Com- munication Networks. IEEE Transactions on Wireless Communications, 20(3):1935–1949,
2020
-
[53]
Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training
[You et al., 2023] Jie You, Jae-Won Chung, and Mosharaf Chowdhury. Zeus: Understanding and Optimizing GPU Energy Consumption of DNN Training. In NSDI,
2023
-
[54]
nn-Meter: Towards Accurate Latency Prediction of Deep-Learning Model Inference on Diverse Edge Devices
[Zhang et al., 2021] Li Lyna Zhang, Shihao Han, Jianyu Wei, Ningxin Zheng, Ting Cao, Yuqing Yang, and Yunxin Liu. nn-Meter: Towards Accurate Latency Prediction of Deep-Learning Model Inference on Diverse Edge Devices. In MobiSys,
2021
-
[55]
They may be battery-powered or even self-powered by harvesting their required energy from the environment
Appendix A1 More discussions on the Necessity of Accurate Estimation Mobile devices are often constrained by the limited power supply and computational resources compared to cloud-based devices. They may be battery-powered or even self-powered by harvesting their required ener...
2023
-
[1998]
Pruning Filters for Efficient ConvNets
[Li et al., 2017] Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. Pruning Filters for Efficient ConvNets. In ICLR,
2017
-
[2006]
Mandheling: Mixed-Precision On-Device DNN Training with DSP Offloading
[Xu et al., 2022] Daliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang, Yun Ma, Kang Huang, Gang Huang, Xin Jin, and Xuanzhe Liu. Mandheling: Mixed-Precision On-Device DNN Training with DSP Offloading. In Mobi- Com,
2022
-
[2009]
Tai- lorFL: Dual-Personalized Federated Learning under Sys- tem and Data Heterogeneity
[Deng et al., 2022] Yongheng Deng, Weining Chen, Ju Ren, Feng Lyu, Yang Liu, Yunxin Liu, and Yaoxue Zhang. Tai- lorFL: Dual-Personalized Federated Learning under Sys- tem and Data Heterogeneity. In SenSys,
2022
-
[2015]
Protean: An Energy-Efficient and Heterogeneous Platform for Adaptive and Hardware-Accelerated Battery- Free Computing
[Bakar et al., 2022] Abu Bakar, Rishabh Goel, Jasper de Winkel, Jason Huang, Saad Ahmed, Bashima Islam, Przemysław Pawełczak, Kasım Sinan Yıldırım, and Josiah Hester. Protean: An Energy-Efficient and Heterogeneous Platform for Adaptive and Hardware-Accelerated Battery- Free Co...
2022
-
[2016]
Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy
[Chen et al., 2018] Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In OSDI,
2018
-
[2017]
To- wards Ubiquitous Learning: A First Measurement of On- Device Training Performance
[Cai et al., 2021] Dongqi Cai, Qipeng Wang, Yuanqiang Liu, Yunxin Liu, Shangguang Wang, and Mengwei Xu. To- wards Ubiquitous Learning: A First Measurement of On- Device Training Performance. In EMDL Workshop,
2021
-
[2018]
Power-z testers,
[Chargerlab, 2023] Chargerlab. Power-z testers,
2023
-
[2019]
Imagenet: A large-scale hierarchical image database
[Deng et al., 2009] Jia Deng, Wei Dong, Richard Socher, Li- Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR,
2009
-
[2020]
Estimation of Energy Consumption in Machine Learning
[García-Martín et al., 2019] Eva García-Martín, Crefeda Faviola Rodrigues, Graham Riley, and Håkan Grahn. Estimation of Energy Consumption in Machine Learning. Journal of Parallel and Distributed Computing, pages 75–88,
2019
-
[2021]
Leaf: A benchmark for Federated Settings
[Caldas et al., 2018] Sebastian Caldas, Sai Meher Karthik Duddu, Peter Wu, Tian Li, Jakub Kone ˇcn`y, H Brendan McMahan, Virginia Smith, and Ameet Talwalkar. Leaf: A benchmark for Federated Settings. arXiv preprint arXiv:1812.01097,
2018 arXiv
-
[2022]
Performance Modeling of Computer Vision-based CNN on Edge GPUs
[Bouzidi et al., 2022] Halima Bouzidi, Hamza Ouarnoughi, Smail Niar, and Abdessamad Ait El Cadi. Performance Modeling of Computer Vision-based CNN on Edge GPUs. ACM Transactions on Embedded Computing Systems , 21(5):1–33,
2022
-
[2023]
Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neu- ral Networks
[Chen et al., 2016] Yu-Hsin Chen, Tushar Krishna, Joel S Emer, and Vivienne Sze. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neu- ral Networks. IEEE Journal of Solid-State Circuits , 52(1):127–138,
2016
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.