REVIEW 4 major objections 5 minor 77 references
eNPU claims that splitting an NPU into independently clocked components—frontend, systolic arrays, vector units, SRAM, HBM, and inter-chip links—and letting the compiler pick each component's voltage/frequency per operator can cut LLM servi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Component-level DVFS on NPUs, with pipeline refactoring and compiler-coordinated voltage/frequency selection, cuts LLM-serving energy by 25.8–35.2% at sub-4% area overhead in simulation.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A real architecture contribution with a believable energy-saving idea, but the headline numbers are projections from an unvalidated frequency-scaling model and the claimed co-optimization is actually fixed-schedule frequency selection. the 4 major comments →
Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that spatially fine-grained, component-level DVFS is practical on NPUs and delivers large energy savings during LLM serving. eNPU partitions the NPU core into separate V/f domains—frontend, systolic arrays, vector units, SRAM, HBM, and inter-chip interconnect—and adds lightweight asynchronous FIFOs, a shadow vector register file for local forwarding inside the vector-unit domain, and a vf.set ISA command that changes component frequencies in sub-microsecond timescales. A compiler-driven two-level greedy search first finds near-Pareto-optimal energy-delay plans per operator, then selects among those plans per request according to available SLO slack. Evaluated on producti
What carries the argument
The enabling mechanism is the component-level V/f domain: the NPU pipeline is refactored so that frontend, systolic arrays, vector units, SRAM, HBM, and ICI each form an independently clocked and voltage-scaled domain. Cross-domain communication uses carefully sized asynchronous FIFOs (minimum depth 10) to absorb synchronization round-trip delays, and a shadow vector register file in the VU domain forwards dependent values locally to avoid pipeline stalls. The vf.set instruction lets software set per-component frequency states in a single VLIW misc slot, and the ML compiler's two-level greedy search—operator-level Pareto exploration followed by request-level plan selection—co-optimizes instr
Load-bearing premise
The energy and SLO numbers rest on the authors' own reconstruction of the proprietary NPU compiler's VLIW scheduling and its frequency-dependent latency model; if that reconstruction diverges from the real compiler's optimized schedules when frequencies change, the reported savings are not established for actual TPU hardware.
What would settle it
Take one representative prefill operator, compile it with the reconstructed scheduler, and execute the resulting VLIW schedule on a real NPU or an FPGA implementation of the refactored pipeline at several component frequency states; if observed execution times and bottleneck components systematically differ from the cost model's predictions, the central quantitative claim is refuted.
If this is right
- If eNPU is right, LLM inference on NPUs can cut energy by roughly a quarter to a third without changing the model, the batch policy, or the hardware generation.
- Component-level DVFS beats whole-core DVFS by up to 9 percentage points at zero SLO slack, with the advantage holding as slack grows, because each component can be throttled to its own bottleneck.
- The savings compose with power gating, reaching up to 31.5% energy reduction when both techniques are used together.
- Per-operator DVFS is feasible in production because operators execute on microsecond timescales and vf.set acts in sub-microseconds, so frequency can track the changing bottleneck of each operator.
- The two-level search keeps the combinatorial space tractable: it explores tens of thousands of operator-level plans and a few thousand to a hundred thousand request-level plans instead of the full 10^7657 possible configurations.
Where Pith is reading between the lines
- Extension: the same Pareto-frontier machinery could be repurposed for power capping or overclocking within a power budget, since the operator-level energy-delay curve already spans both slower and faster configurations.
- Extension: if per-component DVFS becomes standard, cluster-level schedulers could treat energy as a first-class knob alongside auto-scaling, compounding off-peak savings by throttling the chips that remain active.
- Extension: future NPUs with larger vector units or standards like HBM4 that allow variable I/O voltages are likely to show larger gains than today's chips, because those are precisely the components eNPU can throttle independently.
- Extension: a direct silicon test—running eNPU-generated VLIW schedules on a prototype with per-component V/f domains and comparing measured per-operator latency against the cost model at several frequency points—would settle whether the quantitative claims transfer to real hardware.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes eNPU, a hardware/software co-design that enables component-level DVFS on NPUs. It partitions the NPU core into separate V/f domains for frontend, SA, VU, SRAM, HBM, and ICI, adds asynchronous FIFOs and a small shadow vector register file to reduce cross-domain synchronization overhead, and extends the ISA with a vf.set command. A compiler-driven two-level greedy search selects per-operator, per-component frequency plans under request-level SLO slack. The design is implemented on the open-source Coral NPU core and synthesized with ASAP7; evaluation uses an instruction-level performance model, a power model, and a cluster simulator validated against TPU/vLLM data, with production traces for dense and MoE LLMs. The headline result is 25.8–35.2% energy savings with a projected 3.45% area overhead on TPUv4 while preserving SLO targets.
Significance. The paper addresses a timely and important problem: energy-efficient LLM serving on NPUs. The hardware idea is well-motivated and the design is concrete; strengths include the open-source Coral implementation, the synthesis-based area/power estimates, and the validation of simulators against published TPU power and vLLM throughput. If the reported savings held on silicon, component-level DVFS would be a valuable complement to power gating. However, the evidence chain for the headline numbers is incomplete: the frequency-scaling performance model is unvalidated, and the claimed co-optimization of scheduling and DVFS is not implemented. These issues are correctable but currently limit the strength of the contribution.
major comments (4)
- [§3.3, Algorithm 1] The abstract and contributions claim that eNPU 'co-optimizes instruction scheduling and per-component V/f selection.' However, §3.3 states that the DVFS search runs 'after the compiler has finished all optimizations and generated the final VLIW instruction schedule for each operator,' and Algorithm 1 only enumerates frequency candidates and evaluates them with a cost model; it never regenerates the VLIW schedule, adjusts software pipelining, or reallocates shadow registers for the new latencies. The evaluated design is therefore a fixed schedule evaluated under frequency scaling, not the co-optimized system advertised. The paper should either implement the scheduling loop or re-scope the contribution claim.
- [§3.5] The instruction-level performance model drives the DVFS plan search, but it is validated only for static operator performance at peak frequency against real TPU chips. The frequency-dependent instruction latencies, async-FIFO round-trip delay (Eq. 1), and shadow-vector-register forwarding are not validated against any hardware, RTL simulation, or FPGA prototype at multiple V/f points. The Coral NPU prototype is synthesized but not run at varied V/f. If the latency-frequency model errs, both the selected DVFS plans and the reported 25.8–35.2% savings would change. Please add validation of the frequency-scaling behavior or explicitly frame the results as projections with a sensitivity analysis over model parameters.
- [Abstract; §2.3; §3.4] The abstract claims 'preserving strict SLO guarantees,' but the design targets 99% SLO satisfaction: §2.3 accepts that less than 1% of requests have negative slack, and §3.4's safeguard rule 2 locks to peak frequency only when more than 1% of requests violate SLO. This is not a strict 100% guarantee. The wording should be aligned with the actual SLO target throughout.
- [Abstract; §4.8] The headline range 25.8–35.2% is presented as a general result for LLM services, but the only end-to-end trace evaluation in §4.8 uses Llama3-70B; the text says 'We study Llama3-70B as an example' and asserts the insights apply to other models without showing results. Per-request savings for the other models appear in §4.2, but the service-level result is not reported for them. The abstract should either attribute the range to Llama3-70B or add end-to-end results for the other models.
minor comments (5)
- [Abstract; §1; §3.5] The abstract states '3.45% area overhead on a TPUv4 chip' without qualifier, while §1 and §3.5 describe this as a projection. Add 'projected' in the abstract.
- [§4 evaluation intro] In the enumerated evaluation claims, reference '§??' appears unresolved.
- [Table 5] The TPUv5p die area is proxied by NVIDIA A100's die area because TPUv5p is undisclosed. This proxy should be highlighted in the text, and the area-overhead projection's sensitivity to this assumption should be discussed.
- [§3.5] The description of the reconstructed scheduling passes is ambiguous: it is unclear whether the cost model re-schedules per candidate DVFS plan or evaluates the fixed XLA-generated schedule. Clarify to resolve the apparent inconsistency with §3.3.
- [§1] Duplicate 'are are' in the first sentence.
Circularity Check
No material circularity: energy-savings numbers come from externally calibrated component models, not from fitting the target result; only minor non-load-bearing self-citation remains.
full rationale
The central claim—component-level DVFS reduces LLM-serving energy by 25.8%–35.2%—does not reduce by construction to its inputs. The power model is built from RTL synthesis, public HBM/ICI datasheets, and published TDP, and is independently calibrated to within 4%/8% of TPUv3/TPUv4 chip TDP. The instruction-level performance model is validated against real TPUv4 chips for static operator performance; its frequency-dependent extensions (async-FIFO round-trip latency, Eq. 1–2; shadow-register forwarding) are analytical and explicitly stated, not fitted to the reported savings. The DVFS search uses this model as an objective and the evaluation reports the same model's output—a self-consistent co-simulation, not circularity, because the model parameters are externally derived rather than tuned to reproduce the headline numbers. The only self-citation is to the authors' prior ReGate work [65], used to supply TPUv5p chip data (§3.5) and as a complementary power-gating baseline (§4.7); it is not load-bearing for the DVFS mechanism or for any uniqueness/feasibility claim. A separate correctness risk should be noted: the abstract states that eNPU 'co-optimizes instruction scheduling and per-component V/f selection,' while §3.3 says the DVFS search runs after the compiler 'has finished all optimizations and generated the final VLIW instruction schedule.' This is an overclaim about the implemented algorithm, and the unvalidated frequency-scaling of the performance model is a validation gap, but neither constitutes a circular reduction of the result to its inputs.
Axiom & Free-Parameter Ledger
free parameters (6)
- shadow_vector_register_file_size =
4 registers
- frequency_step =
50 MHz, 34 states (0–1.7 GHz)
- slo_slack_compile_levels =
0%, 2%, 5%, 10%
- expert_capacity_factor =
worst-case default (adjustable)
- tpu_v5p_die_area_proxy =
826 mm² (A100)
- slo_target_multiplier =
5× single-request TTFT/TPOT
axioms (5)
- domain assumption P ∝ V² f holds across NPU components
- domain assumption Compiler cost model accurately predicts execution time under arbitrary V/f states
- domain assumption IVR transition delay (4 ns/54 mV) and conversion efficiencies from [10,34] transfer to NPU-scale domains
- domain assumption Production trace and 5× SLO target represent LLM service slack
- domain assumption 99% SLO satisfaction rate qualifies as strict SLO guarantees
invented entities (2)
-
Shadow vector register file
no independent evidence
-
vf.set VLIW instruction
no independent evidence
Cite this review
Pith. "Pith review of Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving." pith.science (2026). https://pith.science/paper/HCDZZ2FJ
@misc{pith2026260716473,
author = {Pith},
title = {Pith review of: Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving},
year = {2026},
howpublished = {\url{https://pith.science/paper/HCDZZ2FJ}},
note = {Machine review of arXiv:2607.16473}
}
abstract
As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting factor. Our study shows that using dynamic voltage and frequency scaling (DVFS) to exploit the service-level objective (SLO) slacks is a promising way to improve NPU energy efficiency for LLM services. And as tensor operators in LLMs exhibit diverse bottlenecks across NPU components, it is desirable to configure the frequency separately for each component to maximize their energy efficiency. In this paper, we develop eNPU that enables hardware and software support for spatially fine-grained, component-level DVFS on NPUs. eNPU refactors the NPU core pipeline to partition components into separate V/$f$ domains. It introduces lightweight cross-domain communication mechanisms to mitigate synchronization overheads across components, and extends the NPU ISA for sub-$\mu$s DVFS control. eNPU uses a compiler-driven two-level greedy search to co-optimize instruction scheduling and per-component V/$f$ selection under SLO constraints. We implement eNPU's pipeline design on an open-source NPU core to verify its functionality and evaluate the energy savings with a production-level NPU simulator with various LLMs using production traces. eNPU reduces energy consumption of LLM services by 25.8%--35.2% with 3.45% area overhead on a TPUv4 chip, while preserving strict SLO guarantees.
Figures
Reference graph
Works this paper leans on
-
[1]
2025. IEEE Standard for Design and Verification of Low-Power Energy-Aware Electronic Systems.IEEE Std 1801-2024 (Revision of IEEE Std 1801-2018)(2025), 1–614. https://doi.org/10.1109/IEEESTD.2025.10910082
arXiv 2025
-
[2]
Coral NPU
2026. Coral NPU. https://github.com/google-coral/coralnpu
2026
-
[3]
Abdelhafez, Christopher Zimmer, Sudharshan S
Hazem A. Abdelhafez, Christopher Zimmer, Sudharshan S. Vazhkudai, and Matei Ripeanu. 2019. AHEAD: A Tool for Projecting Next-Generation Hard- ware Enhancements on GPU-Accelerated Systems. In2019 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 583–592. https://doi.org/10.1109/IPDPSW.2019.00103
arXiv 2019
-
[4]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Tam- ing Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, Santa Clara, CA, 117–134. https://www.useni...
2024
-
[5]
Ionescu, Klaus E
Albert Alexandrov, Mihai F. Ionescu, Klaus E. Schauser, and Chris Scheiman
-
[6]
AMD. [n. d.]. AXI High Bandwidth Memory Controller LogiCORE IP Product Guide (PG276). https://docs.amd.com/r/en-US/pg276-axi-hbm
-
[7]
Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zachary DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael L...
-
[8]
AWS. 2025. Amazon EC2 Trn1/Trn1n Architecture. https://awsdocs- neuron.readthedocs-hosted.com/en/latest/about-neuron/arch/neuron- hardware/trn1-arch.html
2025
-
[9]
Omar Basit, Yunzhao Liu, Z Jonny Kong, and Y Charlie Hu. 2026. BiScale: Energy- Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS. arXiv preprint arXiv:2602.18755(2026). 14 Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving MICRO 2026, October 31–November 04, 2026, Athens, Greece
Pith/arXiv arXiv 2026
-
[10]
Beckmann, and Stephen Kosonocky
Srikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann, and Stephen Kosonocky. 2024. Predict; Don’t React for Enabling Efficient Fine- Grain DVFS in GPUs. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4(Vancouver, BC, Canada)(ASPLOS ’23). Association f...
arXiv 2024
-
[11]
Stephen Chen and Stephen F. Smith. 1999. Improving Genetic Algorithms by Search Space Reductions (with Applications to Flow Shop Scheduling). InProceed- ings of the Genetic and Evolutionary Computation Conference. Morgan Kaufmann, 135–140
1999
-
[12]
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. InProceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI’18). Carlsbad, CA
2018
-
[13]
ChipVerify. 2026. Introduction to UPF. https://chipverify.com/unified-power- format/introduction-to-upf. Accessed: June 2026
2026
-
[14]
Kihwan Choi, R. Soma, and M. Pedram. 2005. Fine-grained dynamic voltage and frequency scaling for precise energy and performance tradeoff based on the ratio of off-chip access to on-chip computation times.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems24, 1 (2005), 18–28. https://doi.org/10.1109/TCAD.2004.839485
arXiv 2005
-
[15]
Lawrence T. Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric. 2016. ASAP7: A 7-nm finFET predictive process design kit.Microelectronics Journal53 (2016), 105–115. https://doi.org/10.1016/j.mejo.2016.04.006
-
[16]
Tejas Dave. 2014. Clock Domain Crossing Techniques & Synchronizers. EDN Net- work. https://www.edn.com/synchronizer-techniques-for-multi-clock-domain- socs-fpgas/ Accessed: April 3, 2026
2014
-
[17]
DeepSeek-AI. 2025. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs.CL] https://arxiv.org/abs/2412.19437
Pith/arXiv arXiv 2025
-
[18]
DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Con- text Intelligence. Technical report and model card. https://huggingface.co/ deepseek-ai/DeepSeek-V4-Pro Accessed: Jun 2026
2026
-
[19]
Atul Dhamba and Anand V Kulkarni. [n. d.]. Design Considerations for High Bandwidth Memory Controller. https://www.design-reuse.com/articles/41186/ design-considerations-for-high-bandwidth-memory-controller.html
-
[20]
Stijn Eyerman and Lieven Eeckhout. 2011. Fine-grained DVFS using on-chip regulators.ACM Trans. Archit. Code Optim.8, 1, Article 1 (Feb. 2011), 24 pages. https://doi.org/10.1145/1952998.1952999
arXiv 2011
-
[21]
Google. 2023. XLA: Optimizing Compiler for Machine Learning. https://www. tensorflow.org/xla
2023
-
[22]
Google Cloud. 2026. TPU 8t and TPU 8i Technical Deep Dive. Google Cloud Blog. https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu- 8i-technical-deep-dive Accessed: Jun 2026
2026
-
[23]
Joao Guerreiro, Aleksandar Ilic, Nuno Roma, and Pedro Tomas. 2018. GPGPU Power Modeling for Multi-domain Voltage-Frequency Scaling. In2018 IEEE Inter- national Symposium on High Performance Computer Architecture (HPCA). 789–800. https://doi.org/10.1109/HPCA.2018.00072
arXiv 2018
-
[24]
Daniel Hackenberg, Robert Schöne, Thomas Ilsche, Daniel Molka, Joseph Schuchart, and Robin Geyer. 2015. An Energy Efficiency Feature Survey of the Intel Haswell Processor. InProceedings of the 2015 IEEE International Parallel and Distributed Processing Symposium Workshop (IPDPSW ’15). IEEE Computer Society, USA, 896–904. https://doi.org/10.1109/IPDPSW.2015.70
-
[25]
Shwai He, Weilin Cai, Jiayi Huang, and Ang Li. 2026. Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts. arXiv:2503.05066 [cs.LG] https://arxiv.org/abs/2503.05066
Pith/arXiv arXiv 2026
-
[26]
Chung-Hsing Hsu and Ulrich Kremer. 2003. The design, implementation, and evaluation of a compiler algorithm for CPU energy reduction.SIGPLAN Not.38, 5 (May 2003), 38–48. https://doi.org/10.1145/780822.781137
arXiv 2003
-
[27]
2025.Energy and AI
International Energy Agency. 2025.Energy and AI. Technical Report. Interna- tional Energy Agency, Paris. https://www.iea.org/reports/energy-and-ai
2025
-
[28]
Canturk Isci, Gilberto Contreras, and Margaret Martonosi. 2006. Live, Run- time Phase Monitoring and Prediction on Real Systems with Application to Dynamic Power Management. InProceedings of the 39th Annual IEEE/ACM Inter- national Symposium on Microarchitecture (MICRO 39). IEEE Computer Society, USA, 359–370. https://doi.org/10.1109/MICRO.2006.30
-
[29]
Behdad Jamadi, Meysam Sohani Darban, and Jeffrey S. Walling. 2026. Analysis of Edge Mismatch and Output Power Degradation in Cascoded Class-D Power Amplifiers Using Dual-Range Voltage Level Shifters. arXiv:2602.09820 [eess.SP] https://arxiv.org/abs/2602.09820
Pith/arXiv arXiv 2026
-
[30]
JEDEC. 2026. High Bandwidth Memory (HBM) DRAM. https://www.jedec.org/ standards-documents/docs/jesd235a
2026
-
[31]
JEDEC Solid State Technology Association. 2025. JEDEC and In- dustry Leaders Collaborate to Release JESD270-4 HBM4 Standard: Advancing Bandwidth, Efficiency, and Capacity for AI and HPC. https://www.jedec.org/news/pressreleases/jedec%C2%AE-and-industry- leaders-collaborate-release-jesd270-4-hbm4-standard-advancing
2025
-
[32]
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Clifford Young, Xiang Zhou, Zongwei Zhou, and David A Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings. InProceedings of the 50th Annual Inter...
arXiv 2023
-
[33]
Andreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis, and Dimitrios Soudris. 2025. throttll’em: Predictive gpu throttling for energy efficient llm inference serving. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 1363–1378
2025
-
[34]
Kim, Yi-Chun Shih, Kaushik Mazumdar, Rinkle Jain, Joseph F
Stephen T. Kim, Yi-Chun Shih, Kaushik Mazumdar, Rinkle Jain, Joseph F. Ryan, Carlos Tokunaga, Charles Augustine, Jaydeep P. Kulkarni, Krishnan Ravichan- dran, James W. Tschanz, Muhammad M. Khellah, and Vivek De. 2016. Enabling Wide Autonomous DVFS in a 22 nm Graphics Execution Core Using a Digitally Controlled Fully Integrated Voltage Regulator.IEEE Journ...
arXiv 2016
-
[35]
Gupta, Gu-Yeon Wei, and David Brooks
Wonyoung Kim, Meeta S. Gupta, Gu-Yeon Wei, and David Brooks. 2008. System level analysis of fast, per-core DVFS using on-chip switching regulators. In2008 IEEE 14th International Symposium on High Performance Computer Architecture. 123–134. https://doi.org/10.1109/hpca.2008.4658633
arXiv 2008
-
[36]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles(Koblenz, Germany)(SOSP ’23). Association for Computing Machinery, New York, N...
arXiv 2023
-
[37]
Aamodt, and Vijay Janapa Reddi
Jingwen Leng, Tayler Hetherington, Ahmed ElTantawy, Syed Gilani, Nam Sung Kim, Tor M. Aamodt, and Vijay Janapa Reddi. 2013. GPUWattch: enabling energy optimizations in GPGPUs. InProceedings of the 40th Annual International Symposium on Computer Architecture(Tel-Aviv, Israel)(ISCA ’13). Association for Computing Machinery, New York, NY, USA, 487–498. https...
arXiv 2013
-
[38]
Yueying Li, Zhanqiu Hu, Esha Choukse, Rodrigo Fonseca, G Edward Suh, and Udit Gupta. 2025. Ecoserve: Designing carbon-aware ai inference systems.arXiv preprint arXiv:2502.05043(2025)
Pith/arXiv arXiv 2025
-
[39]
Llama Team. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783
Pith/arXiv arXiv 2024
-
[40]
Ivan Miro Panades and Alain Greiner. 2007. Bi-Synchronous FIFO for Syn- chronous Circuit Communication Well Suited for Network-on-Chip in GALS Architectures. InFirst International Symposium on Networks-on-Chip (NOCS’07). 83–94. https://doi.org/10.1109/NOCS.2007.14
-
[41]
Thomas Norrie, Nishant Patil, Doe Hyun Yoon, George Kurian, Sheng Li, James Laudon, Cliff Young, Norman Jouppi, and David Patterson. 2021. The Design Process for Google’s Training Chips: TPUv2 and TPUv3.IEEE Micro41, 2 (2021), 56–63. https://doi.org/10.1109/MM.2021.3058217
arXiv 2021
-
[42]
NVIDIA Corporation. [n. d.]. NVIDIA A100 Tensor Core GPU. https://www. nvidia.com/en-us/data-center/a100/. Accessed: June 2026
2026
-
[43]
Pramesh Pandey, Noel Daniel Gundi, Koushik Chakraborty, and Sanghamitra Roy
-
[44]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient Generative LLM Infer- ence Using Phase Splitting. In2024 ACM/IEEE 51st Annual International Sym- posium on Computer Architecture (ISCA). 118–132. https://doi.org/10.1109/ ISCA59077.2024.00019
arXiv 2024
-
[45]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Brijesh Warrier, Nithish Mahalingam, and Ricardo Bianchini. 2023. POLCA: Power Oversub- scription in LLM Cloud Providers. arXiv:2308.12908 [cs.DC] https://arxiv.org/ abs/2308.12908
Pith/arXiv arXiv 2023
-
[46]
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew Kalbarczyk, Tamer Başar, and Ravishankar K Iyer
-
[47]
Qwen Team. 2025. Qwen3-Next-80B-A3B-Thinking. Hugging Face model card. https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking Accessed: Jun 2026
2025
-
[48]
Rambus. [n. d.]. HBM3E / HBM3 Controller IP. https://www.rambus.com/ interface-ip/hbm/hbm3-controller/
-
[49]
Tella Rajashekhar Reddy, Rohan Gandhi, Anjaly Parayil, Chaojie Zhang, Mike Shepperd, Liangcheng Yu, Jayashree Mohan, Srinivasan Iyengar, Shivkumar Kalyanaraman, Debopam Bhattacherjee, et al. 2025. AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron.arXiv preprint arXiv:2505.09989(2025)
Pith/arXiv arXiv 2025
-
[50]
In2024 USENIX Annual Technical Conference (USENIX ATC 24)
Power-aware deep learning model serving with{𝜇 -Serve}. In2024 USENIX Annual Technical Conference (USENIX ATC 24). 75–93
-
[51]
G. Semeraro, G. Magklis, R. Balasubramonian, D.H. Albonesi, S. Dwarkadas, and M.L. Scott. 2002. Energy-efficient processor design using multiple clock domains with dynamic voltage and frequency scaling. InProceedings Eighth International Symposium on High Performance Computer Architecture. 29–40. https://doi.org/10.1109/HPCA.2002.995696
arXiv 2002
-
[52]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism.arXiv preprint arXiv:1909.08053(2019)
Pith/arXiv arXiv 2019
-
[53]
Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Höl- zle, Stephen Stuart, and Amin Vahdat. 2015. Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Ne...
2015
-
[54]
H. Saputra, M. Kandemir, N. Vijaykrishnan, M. J. Irwin, J. S. Hu, C-H. Hsu, and U. Kremer. 2002. Energy-conscious compilation based on voltage scaling. In 15 MICRO 2026, October 31–November 04, 2026, Athens, Greece Yuqi Xue, Jerry Wu, Corey Yu, and Jian Huang Proceedings of the Joint Conference on Languages, Compilers and Tools for Embed- ded Systems: Sof...
arXiv 2002
-
[55]
Jovan Stojkovic, Chaojie Zhang, Inigo Goiri, Josep Torrellas, and Esha Choukse
-
[56]
Tianqi Tang, Sheng Li, Lifeng Nai, Norm Jouppi, and Yuan Xie. 2021. Neu- roMeter: An Integrated Power, Area, and Timing Modeling Framework for Machine Learning Accelerators Industry Track Paper. In2021 IEEE Interna- tional Symposium on High-Performance Computer Architecture (HPCA). 841–853. https://doi.org/10.1109/HPCA51647.2021.00075
arXiv 2021
-
[57]
The JAX Authors. [n. d.]. JAX: High performance array computing. https: //docs.jax.dev/en/latest/
-
[58]
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse, Haoran Qiu, Rodrigo Fonseca, Josep Torrellas, and Ricardo Bianchini. 2025. Tapas: Thermal-and power- aware scheduling for LLM inference in cloud platforms. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 1266–1281
2025
-
[59]
Amin Vahdat and Mark Lohmeyer. 2023. Enabling next-generation AI workloads: Announcing TPU v5p and AI Hypercomputer. https: //cloud.google.com/blog/products/ai-machine-learning/introducing-cloud-tpu- v5p-and-ai-hypercomputer
2023
-
[60]
Zibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Jingyi Zhang, Xiaoxin Xu, Jian He, Xiaoliang Wang, Wanchun Dou, Guihai Chen, and Chen Tian. 2025. Using Analytical Performance/Power Model and Fine- Grained DVFS to Enhance AI Accelerator Energy Efficiency. InProceedings of the 30th ACM International Conference on Architectural Su...
arXiv 2025
-
[61]
Mark Weiser, Brent Welch, Alan Demers, and Scott Shenker. 1994. Scheduling for reduced CPU energy. InProceedings of the 1st USENIX Conference on Operating Systems Design and Implementation(Monterey, California)(OSDI ’94). USENIX Association, USA, 2–es
1994
-
[62]
Qiang Wu, Philo Juang, Margaret Martonosi, and Douglas W. Clark. 2004. Formal online methods for voltage/frequency control in multiple clock domain micro- processors. InProceedings of the 11th International Conference on Architectural Support for Programming Languages and Operating Systems(Boston, MA, USA) (ASPLOS XI). Association for Computing Machinery,...
arXiv 2004
-
[63]
Patrick C. Toulme. 2025. From JAX to VLIW: Tracing a Computation Through the TPU Compiler Stack. https://patricktoulme.substack.com/p/from-jax-to-vliw- tracing-a-computation
2025
-
[64]
Fen Xie, Margaret Martonosi, and Sharad Malik. 2003. Compile-time dynamic voltage scaling settings: opportunities and limits. InProceedings of the ACM SIGPLAN 2003 Conference on Programming Language Design and Implementation (San Diego, California, USA)(PLDI ’03). Association for Computing Machinery, New York, NY, USA, 49–62. https://doi.org/10.1145/781131.781138
arXiv 2003
-
[65]
Yuqi Xue and Jian Huang. 2025. ReGate: Enabling Power Gating in Neural Processing Units. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture (MICRO ’25). Association for Computing Machinery, New York, NY, USA, 1160–1177. https://doi.org/10.1145/3725843.3756038
arXiv 2025
-
[66]
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. 2025. Gated Delta Networks: Improving Mamba2 with Delta Rule. InInternational Conference on Learning Rep- resentations. https://doi.org/10.48550/arXiv.2412.06464 arXiv:2412.06464 [cs.CL]
-
[67]
Jiahuan Yu, Aryan Taneja, Junfeng Lin, and Minjia Zhang. 2025. Voltanallm: Feedback-driven frequency control and state-space routing for energy-efficient llm serving.arXiv preprint arXiv:2509.04827(2025)
Pith/arXiv arXiv 2025
-
[68]
Qiang Wu, Margaret Martonosi, Douglas W. Clark, V. J. Reddi, Dan Connors, Youfeng Wu, Jin Lee, and David Brooks. 2005. A Dynamic Compilation Frame- work for Controlling Microprocessor Energy and Performance. InProceed- ings of the 38th Annual IEEE/ACM International Symposium on Microarchitec- ture(Barcelona, Spain)(MICRO 38). IEEE Computer Society, USA, 2...
-
[69]
Yong Zhao and Nobuo Sannomiya. 2001. An Improvement of Genetic Algorithms by Search Space Reductions in Solving Large-scale Flowshop Problems.IEEJ Transactions on Electronics, Information and Systems121, 6 (2001), 1010–1015. https://doi.org/10.1541/ieejeiss1987.121.6_1010
-
[70]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: efficient execution of structured language model programs. InProceedings of the 38th International Conference on Neural Information Processing Systems(Vancouver, BC, C...
2024
-
[71]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2025. DistServe: disaggregating prefill and decoding for goodput-optimized large language model serving. InProceedings of the 18th USENIX Conference on Operating Systems Design and Implementation(Santa Clara, CA, USA)(OSDI’24). USENIX Association, USA, Article...
2025
-
[72]
Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang, Aaron Wisner, Chris Lewis, and Henri Bahini. 2024. Resiliency at Scale: Managing Google’s TPUv4 Machine Learning Supercomputer. In21st USENIX Symposium on Networked Systems Design and Imp...
2024
-
[73]
Michael Zelikson, Kosta Luria, Lior Gil, Yuval Brown, Vadim Goldenbeg, Dor Kasif, Elias Hlees, and Alex Vinichuk. 2023. 14.3 A Digital Low-Dropout (LDO) Linear Regulator with Adaptive Transfer Function Featuring 125A/mm2 Power Density and Autonomous Bypass Mode. In2023 IEEE International Solid-State Circuits Conference (ISSCC). 230–232. https://doi.org/10...
arXiv 2023
-
[1995]
LogGP: incorporating long messages into the LogP model—one step closer towards a realistic model for parallel computation. InProceedings of the Seventh Annual ACM Symposium on Parallel Algorithms and Architectures(Santa Barbara, California, USA)(SPAA ’95). Association for Computing Machinery, New York, NY, USA, 95–105. https://doi.org/10.1145/215399.215427
-
[2021]
In2021 58th ACM/IEEE Design Automation Conference (DAC)
UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-Gating. In2021 58th ACM/IEEE Design Automation Conference (DAC). 325–330. https://doi.org/10.1109/DAC18074.2021.9586224
arXiv 2021
-
[2024]
Association for Computing Machinery, New York, NY, USA
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3620665.3640366
-
[2025]
In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency . In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE Computer Society, Los Alamitos, CA, USA, 1348–1362. https://doi.org/10.1109/HPCA61900.2025.00102
arXiv 2025
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.