Pith. sign in

REVIEW 3 major objections 7 minor 109 references

Energy-aware operation of HPC systems in Germany

T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper argues that no single technology can make supercomputing sustainable; eight German centers already run a combined portfolio of green power, liquid cooling, heat reuse, power capping, monitoring, scheduling, and application…

desk verdict Useful cross-site review of energy-efficient HPC operations in Germany; Table 1 needs more careful treatment before the paper becomes a reference. read the letter →

arxiv 2411.16204 v1 pith:P77WK6OS submitted 2024-11-25 cs.DC

classification cs.DC
keywords High-PerformanceComputingenergyefficiencydirectliquidcoolingpowercappingheatreusedatacentermonitoringschedulingGermany
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Supercomputers are growing faster than their per-watt efficiency improvements, and this paper argues that the only way German HPC centers can meet legal and economic pressure is to treat energy efficiency as a whole-system problem rather than a single technology choice. Drawing on production experience from eight major centers, it reports that compute hardware takes 78 to 86 percent of site power, infrastructure 7 to 15 percent, and storage 3 to 8 percent, and that successful sites combine green electricity, warm-water direct liquid cooling, heat reuse, power capping, monitoring, energy-aware scheduling, and application tuning. The authors present these measures as already implemented in production, not as research proposals, and conclude that no one measure alone can hit the efficiency targets. The reader should care because the same pressure is now reaching HPC centers worldwide, and the German sites provide a concrete portfolio of techniques that others can adopt.

What carries the argument

The object that carries the argument is the combined production efficiency stack: power supply choices, data-center cooling and heat-reuse infrastructure, heterogeneous compute hardware, storage, cluster-wide monitoring, scheduling and power management, and application-level optimization. The paper's method is the cross-site survey: tables of measured power shares, cooling parameters, heat-reuse factors, and monitoring-stack characteristics from eight production centers, read together with processor efficiency trends. The tables do the work of showing that the same portfolio appears across sites of different sizes, which is the basis for the claim that the portfolio is the cause of achievable efficiency rather than a site-specific accident.

What would settle it

An external audit using identical submetering at all surveyed centers would falsify the paper's core empirical claim if it found the compute share far outside the reported 78 to 86 percent band, or if a controlled comparison showed that a single measure alone met the legal efficiency targets at equal throughput.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a holistic, multi-layer efficiency portfolio is both necessary and already operational in large-scale HPC production. The evidence is a standardized cross-site comparison: despite system sizes spanning roughly 1.1 to 4.7 megawatts per site, the power breakdown is strikingly homogeneous—compute 78 to 86 percent, infrastructure 7 to 15 percent, storage 3 to 8 percent—and the centers already run direct liquid cooling, dynamic power limiting, job-aware monitoring, heat reuse (current energy reuse factors up to 20 percent, with future plans reaching higher), and energy-aware scheduling. The paper treats this convergence of practices as proof that the efficiency targets set by German law are reachable, but only when all these measures are pursued together.

Load-bearing premise

The argument assumes that the self-reported power breakdowns, cooling data, and monitoring intervals from the eight centers are consistent enough to compare, even though the paper notes that one site cannot separate storage power and that monitoring technologies and intervals differ across sites.

Editorial extensions

If this is right

  • If the holistic-approach conclusion is right, investments in any single category—say, only better cooling or only power capping—will fall short of the mandated efficiency targets.
  • Production sites can already achieve substantial efficiency gains: idle-node shutdown at one center cut annual energy use by about 25 percent, and warm-water cooling transfers more than 95 percent of system heat into water at several sites.
  • Future systems will likely become more heterogeneous and more tightly coupled to the electricity grid, because processor specialization offers the largest remaining per-watt gains and centers are exploring demand-response operation.
  • Heat reuse is the least mature layer, with current energy reuse factors of 0 to 20 percent, and its role will grow as legislation requires 10 to 20 percent reuse for new data centers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 78 to 86 percent compute-power share holds across sites, then the biggest energy lever is compute efficiency itself, and infrastructure savings—while legally necessary—have a smaller ceiling unless heat reuse turns waste heat into a usable product.
  • The paper's monitoring data imply that job-level energy accounting is close to becoming a standard operational tool; one testable extension is using historical power traces to predict and shift energy-intensive jobs to hours of cheap renewable electricity, a step the paper names as future research but does not claim to have demonstrated.
  • The trend toward lower cooling temperatures for denser GPU racks could reverse the gains from warm-water free cooling, so an open question is whether chip and packaging innovations can keep rack power density below the threshold where chillers return.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper is a multi-site survey/review of energy-efficiency strategies at eight German HPC centres (DKRZ, FAU, HLRS, JSC, KIT, LRZ, MPCDF, TUD). It covers electricity cost trends and legislation, infrastructure measures (power supply, data center design, cooling, heat reuse), system hardware (heterogeneity, storage, power management), monitoring, scheduling, and programming/algorithmic optimization, with four tables of site-reported data. The central conclusion is that no single measure suffices and that a holistic portfolio—green power, power capping, optimized cooling, heat recovery, hardware selection, monitoring, scheduling, and application optimization—is required, and that the participating sites are already implementing such portfolios in production.

Significance. If the survey is accurate, it is a valuable, current record of production experience across a large fraction of German national and regional HPC capacity (about 300 PFLOPS), covering systems from roughly 1 MW to 5 MW with concrete plans for exascale. Its strengths are the breadth of first-hand operational detail (DLC since 2008/2012 at MPCDF/JSC, heat-reuse projects with 2025 commissioning dates, monitoring stacks with up to 8M metrics), the explicit caveats in the text (e.g., MPCDF storage cannot be separated), and the absence of fitted parameters or circular derivations. The central holistic-approach claim is credible and consistent with the case studies.

major comments (3)
  1. [Section 2, Table 1] The text claims that despite varying system sizes the power-consumption breakdown is "quite homogeneous" (78-86% compute, 7-15% infrastructure, 3-8% storage), but the underlying measurements are not comparable across sites: MPCDF's storage power is included in compute (footnote 2), monitoring intervals and technologies vary by orders of magnitude (Table 4), and the allocation of UPS, chillers, pumps, and air-cooling equipment to "infrastructure" versus "compute" is not standardized. No averaging window or measurement uncertainty is reported, so the narrowness of the ranges could be an artifact of accounting boundaries. Because the paper's main holistic conclusion depends on the documented existence of measures in all categories rather than on these exact percentages, this is fixable by softening the "homogeneous" wording and adding a methodology caveat, but it should be addressed.
  2. [Section 3.4, Table 3] The reported Energy Reuse Factors (ERF) range from 0.009 to 0.20 (with planned values of 0.5 and 0.77), but the table does not specify a common measurement protocol: it is unclear whether heat is metered or estimated, over what 12-month reference period, and whether the ERF numerator is delivered heat or heat-pump input. As written, the numbers imply a cross-site comparability that the text does not establish. Please state the definition and measurement basis or explicitly label the values as site-specific approximations.
  3. [Section 8, Conclusions] The conclusion states that the reporting sites "are already implementing solutions in all these areas," but the evidence for green power supply is only general (Section 3.1), and Table 3 shows that some sites currently have no heat reuse. Rephrase to indicate that the solutions are implemented across the collective set of sites and to varying degrees, rather than implying that every site has deployed every category.
minor comments (7)
  1. [References] Reference 9 is incomplete: "Tech. rep., FixMe" is a placeholder and must be filled in before publication.
  2. [Section 4.1] The citation "NVIDIA (? )" is an incomplete placeholder and should be replaced with a proper reference or removed.
  3. [Author affiliations] "Bergische Universtät Wuppertal" should be "Bergische Universität Wuppertal."
  4. [Section 5.3] "highly customiszable" is a typo for "highly customisable."
  5. [Section 2] The text "This comes at the prize of higher power consumption" should read "at the price of higher power consumption," and the currency symbols in "e0.40 per kWh" and "e0.20 per kWh" appear garbled.
  6. [Table 3] The table header "Y ear" is a typo, and the "Since" column should be labeled more explicitly as the year since which heat reuse has been in operation.
  7. [Section 1] "20 machines ranging from ranking 21 ... to 446" should read "rank 21 ... to rank 446."

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper is a multi-site review whose conclusions rest on documented production deployments, not on a derivation chain that reduces to its own inputs.

full rationale

This paper is a descriptive review of energy-efficiency measures at eight German HPC centers. It makes no formal derivation, fits no parameters, and presents no predictive model; therefore none of the standard circularity patterns (self-definitional reasoning, fitted inputs renamed as predictions, ansatz smuggled in via citation, or uniqueness imported from authors) applies. The central conclusion—that no single measure suffices and a holistic portfolio is needed—is supported by site-specific evidence such as DLC adoption (Table 2), heat-reuse installations (Table 3), monitoring architectures (Table 4), and production scheduling/power-management examples in Sections 4.3 and 6. The paper does cite the authors' own tools and prior work (e.g., ClusterCockpit, DCDB, MetricQ, LLview, PowerSched, modular supercomputing architecture), but these citations are descriptive references to deployed systems, not load-bearing proofs of the review's claims. The most vulnerable part is the cross-site homogeneity statement built on Table 1, where the text itself notes that MPCDF storage is folded into compute and that monitoring intervals and technologies differ across sites; however, this is a data-comparability and correctness concern, not a circularity concern, because the percentages are self-reported measurements rather than predictions obtained from the conclusion. Accordingly, a low score of 1 reflects the presence of frequent but non-load-bearing self-citations, with no circular derivation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are introduced. The review rests on the unstated assumption that the self-reported site data are reliable and comparable, and that the efficiency metrics in Figure 1, computed from TDP and peak performance, are meaningful proxies. These are domain assumptions rather than deliberately introduced axioms.

assumptions (3)
  • domain assumption Self-reported measurements from the eight sites are accurate and comparable enough to support the aggregate claims in Tables 1 to 4.
    Invoked throughout Section 2 and the tables; the paper notes that storage power cannot be separated at MPCDF and that monitoring setups differ, so comparability is not fully established.
  • domain assumption FLOP/Watt efficiency computed from theoretical peak performance and TDP (Figure 1) is a meaningful proxy for real-world energy efficiency.
    Section 2 and the Figure 1 caption acknowledge that frequency types are not identical and cross-graph comparability is limited, yet the figure is used to conclude that specialization drives efficiency gains.
  • domain assumption The legal interpretation of the German Energy Efficiency Act and EU directives is correct.
    Sections 1 and 3.2 rely on these regulations to motivate the paper; no legal analysis is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Energy-aware operation of HPC systems in Germany." pith.science (2026). https://pith.science/paper/P77WK6OS

@misc{pith2026241116204,
  author       = {Pith},
  title        = {Pith review of: Energy-aware operation of HPC systems in Germany},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P77WK6OS}},
  note         = {Machine review of arXiv:2411.16204}
}
read the original abstract

High-Performance Computing (HPC) systems are among the most energy-intensive scientific facilities, with electric power consumption reaching and often exceeding 20 megawatts per installation. Unlike other major scientific infrastructures such as particle accelerators or high-intensity light sources, which are few around the world, the number and size of supercomputers are continuously increasing. Even if every new system generation is more energy efficient than the previous one, the overall growth in size of the HPC infrastructure, driven by a rising demand for computational capacity across all scientific disciplines, and especially by artificial intelligence workloads (AI), rapidly drives up the energy demand. This challenge is particularly significant for HPC centers in Germany, where high electricity costs, stringent national energy policies, and a strong commitment to environmental sustainability are key factors. This paper describes various state-of-the-art strategies and innovations employed to enhance the energy efficiency of HPC systems within the national context. Case studies from leading German HPC facilities illustrate the implementation of novel heterogeneous hardware architectures, advanced monitoring infrastructures, high-temperature cooling solutions, energy-aware scheduling, and dynamic power management, among other optimizations. By reviewing best practices and ongoing research, this paper aims to share valuable insight with the global HPC community, motivating the pursuit of more sustainable and energy-efficient HPC operations.

Figures

Figures reproduced from arXiv: 2411.16204 by the authors.

Figure 1
Figure 1. Energy efficiency for FP64 floating point throughput of a selection of CPUs (left) and GPUs (right). Determined with theoretical peak performance and TDP of one socket/GPU using the highest SKU of each generation. For CPUs, the frequency used to determine peak performance is the lowest frequency measured with a very hot benchmark. For GPUs, the base frequency is taken, assuming continued computations. For GPUs resul… view at source ↗
Figure 2
Figure 2. Average electricity price for new industrial consumers in Germany. Annual consumption 160 000 to 20 million kWh, medium-voltage supply. Data source (17). operated in such a way that at least 10 % of the waste heat is reused (this share grows to 20 % for data centers going into operation on July 1, 2028). The law also regulates air cooling temperatures, mandates the establishment of an energy management system, and c… view at source ↗
Figure 3
Figure 3. Typical components of a monitoring setup in HPC-Cluster environments not provided by the device itself but by the querying daemon. Some of the existing node collectors also support accessing external components via out-of-band plugins. Collection of monitoring data needs careful planning. The number of metrics available, particularly on compute nodes, can be overwhelming (65). For some metrics it may not always be c… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

109 extracted references · 53 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address annote author booktitle chapter doi edition editor eid howpublished institution journal key language month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := ...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter doi edition editor eid howpublished institution journal key language month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid...

  4. [4]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in "" FUNCTION format.date year ...

  5. [5]

    Adamidis, P., Pfister, E., Bockelmann, H., Zobel, D., Beismann, J.-O., and Jacob, M. (2024). The real challenges for climate and weather modelling on its way to sustained exascale performance: A case study using ICON (v2.6.6). Geoscientific Model Development Discussions DOI: 10.5194/gmd-2024-54 https://doi.org/10.5194/gmd-2024-54 gmd-2024-54

  6. [6]

    K., Amaro, E., Amit, N., Hunhoff, E., Yelam, A., and Zellweger, G

    Aguilera, M. K., Amaro, E., Amit, N., Hunhoff, E., Yelam, A., and Zellweger, G. (2023). Memory disaggregation: why now and what are the challenges. SIGOPS Oper. Syst. Rev. DOI: 10.1145/3606557.3606563 https://doi.org/10.1145/3606557.3606563 aguilera2023

  7. [7]

    Alappat, C., Hager, G., Schenk, O., and Wellein, G. (2023). Level-based blocking for sparse matrices: Sparse matrix-power-vector multiplication. IEEE Transactions on Parallel and Distributed Systems DOI: 10.1109/TPDS.2022.3223512 https://doi.org/10.1109/TPDS.2022.3223512 Alappat2023

  8. [8]

    Alappat, C., Thies, J., Hager, G., Fehske, H., and Wellein, G. (2024). Algebraic temporal blocking for sparse iterative solvers on multi-core CPUs . The International Journal of High Performance Computing Applications DOI: 10.1177/10943420241283828 https://doi.org/10.1177/10943420241283828 Alappat2024

Show all 109 references
  1. [9]

    Alpay, A., Soproni, B., W\" u nsche, H., and Heuveline, V. (2022). Exploring the possibility of a hipSYCL-based implementation of oneAPI . In Proceedings of the 10th International Workshop on OpenCL. DOI: 10.1145/3529538.3530005 https://doi.org/10.1145/3529538.3530005 adaptivecpp

  2. [10]

    [Dataset] AMD (2024 a ). HIP . https://github.com/ROCm/HIP. Accessed: 2024-10-12 hip

  3. [11]

    [Dataset] AMD (2024 b ). MlOpen . https://rocm.docs.amd.com/projects/MIOpen/en/latest/index.html. Accessed: 2024-10-10 mlopen

  4. [12]

    AMD EPYC Rome architecture

    [Dataset] AMD Corporation (2024). AMD EPYC Rome architecture . https://www.amd.com/system/files/documents/4th-gen-epyc-processor-architecture-white-paper.pdf. Accessed: 2024-09-12 amdEpyc

  5. [13]

    Emergence and Expansion of Liquid Cooling in Mainstream Data Centers

    American Society of Heating Refrigerating and Air-Conditioning Engineers (2021). Emergence and Expansion of Liquid Cooling in Mainstream Data Centers . Tech. rep., FixMe ashrae2021b

  6. [14]

    Ampere Computing processors

    [Dataset] Ampere Computing LLC (2024). Ampere Computing processors . https://amperecomputing.com/products/processors. Accessed: 2024-10-15 ampere

  7. [15]

    Apple M4 chip

    [Dataset] Apple (2024). Apple M4 chip . https://www.apple.com/newsroom/2024/05/apple-introduces-m4-chip. Accessed: 2024-10-15 apple

  8. [16]

    Auweter, A., Bode, A., Brehm, M., Brochard, L., Hammer, N., Huber, H., et al. (2014). A case study of energy aware scheduling on supermuc. In Proceedings of the 29th International Conference on Supercomputing (ISC 2014) (Springer-Verlag). DOI: 10.1007/978-3-319-07518-1\_25 htt...

  9. [17]

    Bates, N., Hsu, C.-H., Imam, N., Wilde, T., and Sartor, D. (2016). Re-examining HPC energy efficiency dashboard elements. In IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). DOI: 10.1109/IPDPSW.2016.184 https://doi.org/10.1109/IPDPSW.2016.18...

  10. [19]

    J., Sterling, T., Savarese, D., Dorband, J

    Becker, D. J., Sterling, T., Savarese, D., Dorband, J. E., Ranawak, U. A., and Packer, C. V. (1995). BEOWULF: A parallel workstation for scientific computation https://webhome.phy.duke.edu/ rgb/brahma/Resources/beowulf/papers/ICPP95/icpp95.html. In International Conference on ...

  11. [20]

    Browne, S., Dongarra, J., Garner, N., Ho, G., and Mucci, P. (2000). A portable programming interface for performance evaluation on modern processors. The International Journal of High Performance Computing Applications DOI: 10.1177/109434200001400303 https://dl.acm.org/doi/10....

  12. [21]

    [Dataset] Bundesverband der Energie- und Wasserwirtschaft e.V. (2024). Bdew-strompreisanalyse juli 2024 - analysis of electricity prices 2024 BDEW

  13. [22]

    Carabaño, J., Dios, F., Daneshtalab, M., and Ebrahimi, M. (2013). An exploration of heterogeneous systems. In 8th International Workshop on Reconfigurable and Communication-Centric Systems-on-Chip (ReCoSoC). DOI: 10.1109/ReCoSoC.2013.6581542 https://doi.org/10.1109/ReCoSoC.201...

  14. [23]

    Cerebras

    [Dataset] Cerebras (2024). Cerebras . https://cerebras.ai/. Accessed: 2024-10-15 cerebras

  15. [25]

    Curtis, R., Shedd, T., and Clark, E. B. (2023). Performance comparison of five data center server thermal management technologies. In 2023 39th Semiconductor Thermal Measurement, Modeling & Management Symposium (SEMI-THERM). DOI: 10.23919/SEMI-THERM59981.2023.10267908 https://...

  16. [26]

    Dietrich, R., Winkler, F., Knüpfer, A., and Nagel, W. (2020). Pika: Center-wide and job-aware cluster monitoring. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). doi:10.1109/CLUSTER49012.2020.00061 Dietrich2020_pika

  17. [27]

    Eastep, J., Sylvester, S., Cantalupo, C., Geltz, B., Ardanaz, F., Al-Rawi, A., et al. (2017). Global extensible open power manager: A vehicle for HPC community collaboration on co-designed energy management solutions. In High Performance Computing. DOI: 10.1007/978-3-319-58667...

  18. [28]

    Eitzinger, J., Gruber, T., Afzal, A., Zeiser, T., and Wellein, G. (2019). Clustercockpit – a web application for job-specific performance monitoring. In HPCMASPA 2019, the Workshop for Monitoring and Analysis for High Performance Computing Systems and Applications. DOI: 10.110...

  19. [29]

    AI-factories initiative from EuroHPC

    [Dataset] EuroHPC Joint Undertaking (2024 a ). AI-factories initiative from EuroHPC . https://eurohpc-ju.europa.eu/eurohpc-joint-undertaking-launches-ai-factories-calls-boost-european-leadership-trustworthy-ai-2024-09-10\_en. Accessed: 2024-09-11 AIfact

  20. [30]

    EuroHPC JU decission to fund RISC-V chiplet development

    [Dataset] EuroHPC Joint Undertaking (2024 b ). EuroHPC JU decission to fund RISC-V chiplet development . https://eurohpc-ju.europa.eu/document/download/38934e8b-30c9-4af6-b948-d73ab9107ce5\_en?filename=Decision Accessed: 2024-10-15 eurohpc-dare

  21. [31]

    EU Energy Efficiency Directive

    [Dataset] European Commission (2024). EU Energy Efficiency Directive . https://energy.ec.europa.eu/topics/energy-efficiency/energy-efficiency-targets-directive-and-rules/energy-efficiency-directive\_en. Accessed: 2024-10-14 EUenergyefficiency

  22. [32]

    EU supply chain directive: corporate sustainability due diligence and amending Directive

    [Dataset] European Parliament (2024). EU supply chain directive: corporate sustainability due diligence and amending Directive . https://eur-lex.europa.eu/eli/dir/2024/1760/oj. Accessed: 2024-10-09 EUsuplchain

  23. [33]

    Fujitsu A64Fx

    [Dataset] Fujitsu (2024). Fujitsu A64Fx . https://www.fujitsu.com/global/products/computing/servers/supercomputer/a64fx/. Accessed: 2024-10-15 fujitsu

  24. [34]

    Energy Efficiency Act in Germany

    [Dataset] German Federal Government (2024 a ). Energy Efficiency Act in Germany . https://www.bundesregierung.de/breg-en/federal-government/the-energy-efficiency-act-2184958. Accessed: 2024-09-12 eneffg

  25. [35]

    Gesetz zur Steigerung der Energieeffizienz in Deutschland (Energieeffizienzgesetz - EnEfG) - §4

    [Dataset] German Federal Government (2024 b ). Gesetz zur Steigerung der Energieeffizienz in Deutschland (Energieeffizienzgesetz - EnEfG) - §4 . https://www.gesetze-im-internet.de/enefg/BJNR1350B0023.html#BJNR1350B0023BJNG000400000. Accessed: 2024-09-30 enefG_dc

  26. [36]

    TPU (Cloud)

    [Dataset] Google (2024). TPU (Cloud) . https://cloud.google.com/tpu/docs. Accessed: 2024-10-15 tpu

  27. [37]

    Gouk, D., Kwon, M., Bae, H., Lee, S., and Jung, M. (2023). Memory pooling with CXL . IEEE Micro DOI: 10.1109/MM.2023.3237491 https://doi.org/10.1109/MM.2023.3237491 gouk2023

  28. [38]

    [Dataset] Grafana Labs (2024). Grafana . https://grafana.com. Accessed: 2024-10-25 grafana

  29. [39]

    Graphcore

    [Dataset] Graphcore (2024). Graphcore . https://www.graphcore.ai/. Accessed: 2024-10-15 graphcore

  30. [40]

    and Patterson, M

    Hackenberg, D. and Patterson, M. K. (2016). Evaluation of a new data center air-cooling architecture: The down-flow plenum. In 2016 15th IEEE Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronic Systems (ITherm). DOI: 10.1109/ITHERM.2016.7517576 http...

  31. [41]

    Hager, G., Treibig, J., Habich, J., and Wellein, G. (2013). Exploring performance and power properties of modern multicore chips via simple machine models. Concurrency Computat.: Pract. Exper. DOI: 10.1002/cpe.3180 https://doi.org/10.1002/cpe.3180 Hager2012

  32. [42]

    Hager, G., Wellein, G., and Treibig, J. (2010). LIKWID: A Lightweight Performance-Oriented Tool Suite for x86 Multicore Environments . In 2012 41st International Conference on Parallel Processing Workshops (IEEE Computer Society). DOI: 10.1109/ICPPW.2010.38 https://doi.ieeecom...

  33. [43]

    Herten, A. (2023). Many cores, many models: GPU programming model vs. vendor compatibility overview. In Proceedings of the SC '23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis. DOI: 10.1145/3624062.3624178 https://doi.o...

  34. [44]

    [Dataset] Herten, A., Achilles, S., Alvarez, D., Badwaik, J., Behle, E., Bode, M., et al. (2024). Application-Driven Exascale: The JUPITER Benchmark Suite . arXiv: 2408.17211 https://arxiv.org/abs/2408.17211 herten2024

  35. [45]

    Hofmann, J., Hager, G., and Fey, D. (2018). On the accuracy and usefulness of analytic energy models for contemporary multicore processors. In 33rd International Conference on High Performance Computing (ISC). DOI: 10.1007/978-3-319-92040-5\_2 https://doi.org/10.1007/978-3-319...

  36. [46]

    Huang, A. (2015). Moore's law is dying (and that could be good). IEEE Spectrum DOI: 10.1109/MSPEC.2015.7065418 https://doi.org/10.1109/MSPEC.2015.7065418 huang2015

  37. [47]

    and Steiner, S

    Hunold, S. and Steiner, S. (2020). Benchmarking julia’s communication performance: Is julia hpc ready or full hpc? In IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS). DOI: 10.1109/PMBS51919.2020.00008 https://doi.org/10.11...

  38. [48]

    Nagel, W

    Ilsche, T., Hackenberg, D., Schöne, R., Bielert, M., Höpfner, F., and E. Nagel, W. (2019). Metricq: A scalable infrastructure for processing high-resolution time series data. In IEEE/ACM Industry/University Joint International Workshop on Data-center Automation, Analytics, and...

  39. [49]

    Ilsche, T., Schrader, S., and Schöne, R. (2024). Optimizing idle power of HPC systems:practical insights and methods. In IEEE International Conference on Cluster Computing (CLUSTER) 2024_Ilsche_Idle

  40. [50]

    InfluxDB

    [Dataset] InfluxData (2024 a ). InfluxDB . https://www.influxdata.com. Accessed: 2024-10-25 influxdb

  41. [51]

    Kapacitor

    [Dataset] InfluxData (2024 b ). Kapacitor . https://www.influxdata.com/time-series-platform/kapacitor. Accessed: 2024-10-25 kapacitor

  42. [52]

    Telegraf

    [Dataset] InfluxData (2024 c ). Telegraf . https://www.influxdata.com/time-series-platform/telegraf. Accessed: 2024-10-25 telegraf

  43. [53]

    [Dataset] Intel (2024). OneMKL . https://www.intel.com/content/www/us/en/developer/tools/oneapi/onemkl.html. Accessed: 2024-10-10 mkl

  44. [54]

    u lich Supercomputing Centre

    [Dataset] J\" u lich Supercomputing Centre" (2024). J U elich S T orage cluster J U S T . https://www.fz-juelich.de/en/ias/jsc/systems/storage-systems/just. Accessed: 2024-10-11 JUST

  45. [55]

    LLview job monitoring

    [Dataset] J\" u lich Supercomputing Centre (2024). LLview job monitoring . https://www.fz-juelich.de/en/ias/jsc/services/user-support/software-tools/llview. Accessed: 2024-10-20 llview

  46. [56]

    [Dataset] Khronos (2024). SYCL . https://www.khronos.org/sycl/. Accessed: 2024-10-10 sycl

  47. [57]

    Koomey, J., Berard, S., Sanchez, M., and Wong, H. (2011). Implications of historical trends in the electrical efficiency of computing. IEEE Annals of the History of Computing DOI: 10.1109/MAHC.2010.28 https://doi.org/10.1109/MAHC.2010.28 koomey2011

  48. [58]

    (eds.) (2021)

    Kreuzer, A., Suarez, E., Eicker, N., and Lippert, T. (eds.) (2021). https://juser.fz-juelich.de/record/905738 P orting applications to a M odular S upercomputer - E xperiences from the DEEP - EST project . Schriften des Forschungszentrums Jülich IAS Series kreuzer2021

  49. [59]

    Lei a, R., Boesche, K., Hack, S., P\' e rard-Gayot, A., Membarth, R., Slusallek, P., et al. (2018). Anydsl: a partial evaluation framework for programming high-performance libraries. Proc. ACM Program. Lang. DOI: 10.1145/3276489 https://doi.org/10.1145/3276489 anydsl

  50. [60]

    Li, J., Michelogiannakis, G., Cook, B., Cooray, D., and Chen, Y. (2023). Analyzing resource utilization in an HPC system: A case study of NERSC’s perlmutter. In International Conference on High Performance Computing (ISC). DOI: 10.1007/978-3-031-32041-5\_16 https://doi.org/10....

  51. [61]

    Flux scheduler

    [Dataset] LLNL (2024). Flux scheduler . https://github.com/flux-framework. Accessed: 2024-10-20 flux

  52. [62]

    [Dataset] LLVM Foundation (2024). LLVM . https://llvm.org/. Accessed: 2024-10-15 llvm

  53. [63]

    Maloney, S., Suarez, E., Eicker, N., Guimarães, F., and Frings, W. (2024). Analyzing HPC monitoring data with a view towards efficient resource utilization. In 2024 IEEE 36th International Symposium on Computer Architecture and High Performance Computing ( SBAC-PAD ) (IEEE) ma...

  54. [64]

    Matsuoka, S. (2018). Cambrian explosion of computing and big data in the post-moore era. In Proceedings of the 27th International Symposium on High-Performance Parallel and Distributed Computing. DOI: 10.1145/3208040.3225055 https://doi.org/10.1145/3208040.3225055 matsuoka2018

  55. [65]

    Y., Glick, M., Dennison, L., et al

    Michelogiannakis, G., Klenk, B., Cook, B., Teh, M. Y., Glick, M., Dennison, L., et al. (2022). A case for intra-rack resource disaggregation in HPC DOI: 10.1145/3514245 https://doi.org/10.1145/3514245 michelogiannakis2022

  56. [66]

    Milojicic, D., Faraboschi, P., Dube, N., and Roweth, D. (2021). Future of HPC : Diversifying heterogeneity. In Design, Automation & Test in Europe Conference & Exhibition (DATE). DOI: 10.23919/DATE51398.2021.9474063 https://doi.org/10.23919/DATE51398.2021.9474063 milojicic2021

  57. [67]

    H., and Herten, A

    Nassyr, S., Mood, K. H., and Herten, A. (2023). Programmatically Reaching the Roof: Automated BLIS Kernel Generator for SVE and RVV. Tech. rep. DOI: 10.34734/FZJ-2023-03437 https://doi.org/10.34734/FZJ-2023-03437 nassyr2023programmatically

  58. [68]

    and Pleiter, D

    Nassyr, S. and Pleiter, D. (2024). Exploring processor micro-architectures optimised for BLAS3 micro-kernels. In Euro-Par: Parallel Processing. DOI: 10.1007/978-3-031-69766-1\_4 https://doi.org/10.1007/978-3-031-69766-1\_4 nassyreuropar

  59. [69]

    Netti, A., M\" u ller, M., Auweter, A., Guillen, C., Ott, M., Tafani, D., et al. (2019). From facility to application sensor data: modular, continuous and holistic monitoring with dcdb. In Proceedings of the International Conference for High Performance Computing, Networking, ...

  60. [70]

    [Dataset] NVIDIA (2024 a ). cuBLAS . https://developer.nvidia.com/cublas. Accessed: 2024-10-10 cublas

  61. [71]

    CUDA Toolkit

    [Dataset] NVIDIA (2024 b ). CUDA Toolkit . https://developer.nvidia.com/cuda-toolkit. Accessed: 2024-10-10 cuda

  62. [72]

    [Dataset] NVIDIA (2024 c ). NCCL . https://developer.nvidia.com/nccl. Accessed: 2024-10-10 nccl

  63. [73]

    [Dataset] OpenMP Architecture Review Board (2024). OpenMP . https://www.openmp.org/. Accessed: 2024-10-15 openmp

  64. [74]

    G., Groner, L., Ubbiali, S., Vogt, H., Madonna, A., Mariotti, K., et al

    Paredes, E. G., Groner, L., Ubbiali, S., Vogt, H., Madonna, A., Mariotti, K., et al. (2023). Gt4py: High performance stencils for weather and climate applications using python. arXiv preprint arXiv:2311.08322 gt4py

  65. [75]

    Peng, I., Pearce, R., and Gokhale, M. (2020). On the memory underutilization: Exploring disaggregated memory on HPC systems. In 32nd IEEE International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD). DOI: 10.1109/SBAC-PAD49847.2020.00034 https://d...

  66. [76]

    RISC-V Instruction Set Architecture specifications

    [Dataset] RISC-V Foundation (2024). RISC-V Instruction Set Architecture specifications . https://riscv.org/technical/specifications/. Accessed: 2024-10-15 riscv

  67. [77]

    Röhl, T., Treibig, J., Hager, G., and Wellein, G. (2014). Overhead analysis of performance counter measurements. In PSTI 2014, the Fifth International Workshop on Parallel Software Tools and Tool Infrastructures Gruber2014

  68. [78]

    Sato, M., Ishikawa, Y., Tomita, H., Kodama, Y., Odajima, T., Tsuji, M., et al. (2020). Co-design for a64fx manycore processor and ”fugaku”. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. DOI: 10.1109/SC41405.2020.00051 https...

  69. [79]

    Slurm scheduler

    [Dataset] SchedMD (2024). Slurm scheduler . https://slurm.schedmd.com/. Accessed: 2024-09-12 slurm

  70. [80]

    B., Trinitis, C., and Weidendorfer, J

    Schulz, M., Kranzlm\" u ller, D., Schulz, L. B., Trinitis, C., and Weidendorfer, J. (2021). On the inevitability of integrated HPC systems and how they will change HPC system operations. In Proceedings of the 11th International Symposium on Highly Efficient Accelerators and Re...

  71. [81]

    Shalf, J. (2020). The future of computing beyond M oore’s law. Phil. Trans. R. Soc. A DOI: 10.1098/rsta.2019.0061 https://doi.org/10.1098/rsta.2019.0061 Shalf_2020

  72. [82]

    Simmendinger, C., Marquardt, M., Mäder, J., and Schneider, R. (2024). Powersched - managing power consumption in overprovisioned systems. In IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops). DOI: 10.1109/CLUSTERWorkshops61563.2024.00012 https://...

  73. [83]

    Rhea chip

    [Dataset] SiPEARL (2024). Rhea chip . https://sipearl.com/wp-content/uploads/2024/05/PR\_SiPearl\_Rhea1\_EN.pdf. Accessed: 2024-10-15 sipearl

  74. [84]

    SpiNNaker2

    [Dataset] SpiNNcloud Systems (2024). SpiNNaker2 . https://spinncloud.com/portfolio/spinnaker2/. Accessed: 2024-10-15 spinnaker

  75. [85]

    and Reuter, K

    Stanisic, L. and Reuter, K. (2020). MPCDF HPC Performance Monitoring System: Enabling Insight via Job-Specific Analysis (Springer International Publishing). DOI: 10.1007/978-3-030-48340-1\_47 https://doi.org/10.1007/978-3-030-48340-1\_47 Stanisic2020_HPCMD

  76. [86]

    Bruttostromerzeugung in Deutschland: energy mix in Germany

    [Dataset] Statistisches Bundesamt (Destatis) (2024). Bruttostromerzeugung in Deutschland: energy mix in Germany . https://www.destatis.de/DE/Themen/Branchen-Unternehmen/Energie/Erzeugung/Tabellen/bruttostromerzeugung.html. Accessed: 2024-10-15 destatis_energy_24

  77. [87]

    (2024 a )

    [Dataset] Strohmaier, E., Dongarra, J., Simon, H., and Meuer, M. (2024 a ). Green500 . https://top500.org/lists/green500/. Accessed: 2024-09-11 green500

  78. [88]

    (2024 b )

    [Dataset] Strohmaier, E., Dongarra, J., Simon, H., and Meuer, M. (2024 b ). Top500 . https://top500.org. Accessed: 2024-09-12 top500

  79. [89]

    Suarez, E., Eicker, N., and Lippert, T. (2019). https://juser.fz-juelich.de/record/862856 M odular S upercomputing A rchitecture: from I dea to P roduction , vol. 3 of Contemporary High Performance Computing: From Petascale toward Exascale, chap. 9 suarez2019

  80. [90]

    Suarez, E., Eicker, N., Moschny, T., and Lippert, T. (2021). https://juser.fz-juelich.de/record/905854 C ritical A nalysis of the M odular S upercomputing A rchitecture (Forschungszentrum Jülich GmbH Zentralbibliothek, Verlag Jülich), vol. 48 of Schriften des Forschungszentrum...

  81. [91]

    Suggs, D., Subramony, M., and Bouvier, D. (2020). The AMD “Zen 2” processor. IEEE Micro DOI: 10.1109/MM.2020.2974217 https://doi.org/10.1109/MM.2020.2974217 suggs2020

  82. [92]

    TSMP manufacturing trends

    [Dataset] Taiwan Semiconductor Manufacturing Company (2024). TSMP manufacturing trends . https://www.tsmc.com/english/dedicatedFoundry/technology/logic. Accessed: 2024-09-11 tsmc

  83. [93]

    Teichgräber, J. M. R. (2022). Julia: A competitive high-level choice for performance portability in HPC? https://events.gwdg.de/event/243/contributions/509/attachments/145/186/julia-performance-portability-in-hpc-final.pdf In MPCDF Seminar (Performance) Portable Programming of...

  84. [94]

    ClusterCockpit Generic Datastructure Specification

    [Dataset] The ClusterCockpit Project (2024). ClusterCockpit Generic Datastructure Specification . https://github.com/ClusterCockpit/cc-specifications/tree/master/datastructures. Accessed: 2024-10-25 ClusterCockpitSpec

  85. [95]

    collectd - The system statistics collection daemon

    [Dataset] The collectd Project (2024). collectd - The system statistics collection daemon . https://www.collectd.org. Accessed: 2024-10-25 collectd

  86. [96]

    eeHPC project

    [Dataset] The eeHPC Consortium (2024). eeHPC project . https://eehpc.clustercockpit.org/. Accessed: 2024-10-25 eehpc

  87. [97]

    European Processor Inititative (EPI)

    [Dataset] The EPI Consortium (2024). European Processor Inititative (EPI) . https://www.european-processor-initiative.eu/. Accessed: 2024-10-15 epi

  88. [98]

    EUPEX project

    [Dataset] The EUPEX Consortium (2024). EUPEX project . https://eupex.eu/. Accessed: 2024-10-15 eupex

  89. [99]

    EUPILOT project

    [Dataset] The EUPILOT Consortium (2024). EUPILOT project . https://eupilot.eu/. Accessed: 2024-10-15 eupilot

  90. [100]

    Prometheus

    [Dataset] The Prometheus Project (2024). Prometheus . https://prometheus.io. Accessed: 2024-10-25 prometheus

  91. [101]

    and Altiparmak, N

    Tomes, E. and Altiparmak, N. (2017). A comparative study of HDD and SSD RAIDs’ impact on server energy consumption. In IEEE International Conference on Cluster Computing (CLUSTER). DOI: 10.1109/CLUSTER.2017.103 https://doi.org/10.1109/CLUSTER.2017.103 Tomes2017

  92. [102]

    R., Lebrun-Grandié, D., Arndt, D., Ciesko, J., Dang, V., Ellingwood, N., et al

    Trott, C. R., Lebrun-Grandié, D., Arndt, D., Ciesko, J., Dang, V., Ellingwood, N., et al. (2022). Kokkos 3: Programming model extensions for the exascale era. IEEE Transactions on Parallel and Distributed Systems DOI: 10.1109/TPDS.2021.3097283 https://doi.org/10.1109/TPDS.2021...

  93. [103]

    Tröpgen, H., Schöne, R., Ilsche, T., and Hackenberg, D. (2024). 16 years of SPEC power:an analysis of x86 energy efficiency trends. In IEEE International Conference on Cluster Computing (CLUSTER) Troepgen_2024_SPEC

  94. [104]

    [Dataset] UXL Foundation (2024). oneDNN . https://github.com/oneapi-src/oneDNN. Accessed: 2024-10-10 onednn

  95. [105]

    V an Z ee, F. G. and v an d e G eijn, R. A. (2015). BLIS : A framework for rapidly instantiating BLAS functionality. ACM Transactions on Mathematical Software DOI: 10.1145/2764454 https://doi.org/10.1145/2764454 blis

  96. [106]

    Wilde, T., Ott, M., Auweter, A., Meijer, I., Ruch, P., Hilger, M., et al. (2017). Coolmuc-2: A supercomputing cluster with heat recovery for adsorption cooling. In 2017 33rd Thermal Measurement, Modeling & Management Symposium (SEMI-THERM). DOI: 10.1109/SEMI-THERM.2017.7896917...

  97. [107]

    Williams, S., Waterman, A., and Patterson, D. (2009). Roofline: an insightful visual performance model for multicore architectures. Commun. ACM DOI: 10.1145/1498765.1498785 https://doi.org/10.1145/1498765.1498785 williams2009

  98. [108]

    Wittmann, M., Hager, G., Zeiser, T., Treibig, J., and Wellein, G. (2016). Chip-level and multi-node analysis of energy-optimized lattice Boltzmann CFD simulations. Concurrency and Computation: Practice and Experience DOI: 10.1002/cpe.3489 https://doi.org/10.1002/cpe.3489 Wittmann2016

  99. [109]

    and Di Napoli, E

    Wu, X. and Di Napoli, E. (2023). Advancing the distributed multi- GPU ChASE library through algorithm optimization and NCCL library. In Proceedings of the SC '23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis. DOI: 10.11...

  100. [110]

    Zenker, E., Worpitz, B., Widera, R., Huebl, A., Juckeland, G., Knüpfer, A., et al. (2016). Alpaka -- an abstraction library for parallel kernel acceleration. In IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). DOI: 10.1109/IPDPSW.2016.50 htt...

  101. [111]

    Zhao, D., Samsi, S., McDonald, J., Li, B., Bestor, D., Jones, M., et al. (2023). Sustainable supercomputing for AI : GPU power capping at HPC scale. SoCC ’23. DOI: 10.1145/3620678.3624793 https://doi.org/10.1145/3620678.3624793 Zhao_2023

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.