{"id":"dcf50df7-d3a1-43e1-af1d-24f70e9f514f","arxiv_id":"2411.16204","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"This review of eight German HPC centers concludes that a combination of cooling, scheduling, monitoring, and hardware choices is required to meaningfully reduce energy use.","lead":"This paper surveys how eight German supercomputing centers reduce electricity use through liquid cooling, monitoring, scheduling, and heat reuse. It gives operators elsewhere concrete practices and cost data for making high-performance computing more sustainable.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central holistic-approach claim is supported for the surveyed sites, but the underlying power-breakdown data are not comparable enough to justify cross-site 'homogeneity' assertions that the conclusion partly leans on.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test lands on the same weakest assumption: the comparability of self-reported power-breakdown data in Table 1. The paper's strongest claim is that a holistic combination of measures is necessary and is being implemented; that claim is well supported by the site-by-site descriptions of DLC, heat reuse, monitoring, scheduling, and power management, so I do not see an internal inconsistency or unsupported leap that would justify REJECT or a lower confidence. The most load-bearing numerical support is the homogeneity observation built on Table 1, because the paper explicitly uses it to argue that the energy problem looks similar across very different scales and hence that a general holistic strategy is appropriate. But the measurement methodology is heterogeneous: MPCDF cannot separate storage from compute; Table 4 shows monitoring systems with wildly different collection intervals and stacks; infrastructure share depends on whether cooling losses and UPS losses are charged to infrastructure or to the IT load; and no error bars or definitions are provided. These are exactly the conditions under which a narrow band of computed percentages could be methodological rather than physical. A concrete remedy is a common reporting protocol for a single reference period; this would settle whether the homogeneity claim reflects reality or accounting. That check is feasible because the data are only averages and the sites are the paper's authors. My verdict stays CONDITIONAL rather than UNCHANGED only in the sense that the requested revision (clarify or soften the Table-1 homogeneity claim) is already part of the reader's condition; I agree with the reader that the overall conclusion does not need to change. I therefore mark the recommended verdict as CONDITIONAL with agreement, to reflect that the concern is genuine but not fatal.","tokens_in":25788,"tokens_out":1793,"duration_ms":15486,"concrete_test":"Ask the eight sites to each report, for a common reference period and common boundary definitions, (a) total site power, (b) compute/IT load, (c) storage load (including file servers and storage network), (d) cooling-related power, (e) UPS/power-conversion losses, and (f) other infrastructure. If the recomputed infrastructure and storage shares still span roughly 7-15% and 3-8% with storage excluded from compute at every site, the homogeneity claim survives; if the ranges widen materially (e.g., storage becomes 0-12%), the homogeneity assertion in Section 2/Table 1 needs to be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is a review-level synthesis: no single measure suffices; a holistic portfolio of measures is required and is being implemented in production at eight German HPC centers. The evidence for implementation is credible and site-specific (DLC tables, heat-reuse table, monitoring table, scheduling examples). The weaker link is Table 1, which is used to assert that power breakdowns are 'quite homogeneous' across sites (78-86% compute, 7-15% infrastructure, 3-8% storage). These percentages are not apples-to-apples: MPCDF storage is folded into compute (footnote 2); monitoring intervals and technologies differ by orders of magnitude across sites (Table 4); infrastructure share depends on site-specific accounting choices (e.g., whether UPS, chillers, pumps, and air-cooling equipment are allocated to the IT load or to infrastructure); and no measurement uncertainty or averaging window is given. The 78-86% compute share range is narrow enough that such methodology differences could plausibly swamp the true cross-site differences, making the 'homogeneous' claim an artifact of how the boundaries between compute, infrastructure, and storage are drawn. However, this data-quality issue does not undermine the paper's main practical conclusion, because the conclusion does not depend on the exact percentages; it depends on the documented existence of a broad set of measures across all categories. The conditional acceptance is therefore appropriate, but the homogeneity statement should be softened or accompanied by a precise measurement protocol before the paper serves as a definitive reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a multi-site survey/review of energy-efficiency strategies at eight German HPC centres (DKRZ, FAU, HLRS, JSC, KIT, LRZ, MPCDF, TUD). It covers electricity cost trends and legislation, infrastructure measures (power supply, data center design, cooling, heat reuse), system hardware (heterogeneity, storage, power management), monitoring, scheduling, and programming/algorithmic optimization, with four tables of site-reported data. The central conclusion is that no single measure suffices and that a holistic portfolio—green power, power capping, optimized cooling, heat recovery, hardware selection, monitoring, scheduling, and application optimization—is required, and that the participating sites are already implementing such portfolios in production.","tokens_in":26014,"tokens_out":4986,"duration_ms":44312,"significance":"If the survey is accurate, it is a valuable, current record of production experience across a large fraction of German national and regional HPC capacity (about 300 PFLOPS), covering systems from roughly 1 MW to 5 MW with concrete plans for exascale. Its strengths are the breadth of first-hand operational detail (DLC since 2008/2012 at MPCDF/JSC, heat-reuse projects with 2025 commissioning dates, monitoring stacks with up to 8M metrics), the explicit caveats in the text (e.g., MPCDF storage cannot be separated), and the absence of fitted parameters or circular derivations. The central holistic-approach claim is credible and consistent with the case studies.","major_comments":[{"comment":"The text claims that despite varying system sizes the power-consumption breakdown is \"quite homogeneous\" (78-86% compute, 7-15% infrastructure, 3-8% storage), but the underlying measurements are not comparable across sites: MPCDF's storage power is included in compute (footnote 2), monitoring intervals and technologies vary by orders of magnitude (Table 4), and the allocation of UPS, chillers, pumps, and air-cooling equipment to \"infrastructure\" versus \"compute\" is not standardized. No averaging window or measurement uncertainty is reported, so the narrowness of the ranges could be an artifact of accounting boundaries. Because the paper's main holistic conclusion depends on the documented existence of measures in all categories rather than on these exact percentages, this is fixable by softening the \"homogeneous\" wording and adding a methodology caveat, but it should be addressed.","section":"Section 2, Table 1"},{"comment":"The reported Energy Reuse Factors (ERF) range from 0.009 to 0.20 (with planned values of 0.5 and 0.77), but the table does not specify a common measurement protocol: it is unclear whether heat is metered or estimated, over what 12-month reference period, and whether the ERF numerator is delivered heat or heat-pump input. As written, the numbers imply a cross-site comparability that the text does not establish. Please state the definition and measurement basis or explicitly label the values as site-specific approximations.","section":"Section 3.4, Table 3"},{"comment":"The conclusion states that the reporting sites \"are already implementing solutions in all these areas,\" but the evidence for green power supply is only general (Section 3.1), and Table 3 shows that some sites currently have no heat reuse. Rephrase to indicate that the solutions are implemented across the collective set of sites and to varying degrees, rather than implying that every site has deployed every category.","section":"Section 8, Conclusions"}],"minor_comments":[{"comment":"Reference 9 is incomplete: \"Tech. rep., FixMe\" is a placeholder and must be filled in before publication.","section":"References"},{"comment":"The citation \"NVIDIA (? )\" is an incomplete placeholder and should be replaced with a proper reference or removed.","section":"Section 4.1"},{"comment":"\"Bergische Universtät Wuppertal\" should be \"Bergische Universität Wuppertal.\"","section":"Author affiliations"},{"comment":"\"highly customiszable\" is a typo for \"highly customisable.\"","section":"Section 5.3"},{"comment":"The text \"This comes at the prize of higher power consumption\" should read \"at the price of higher power consumption,\" and the currency symbols in \"e0.40 per kWh\" and \"e0.20 per kWh\" appear garbled.","section":"Section 2"},{"comment":"The table header \"Y ear\" is a typo, and the \"Since\" column should be labeled more explicitly as the year since which heat reuse has been in operation.","section":"Table 3"},{"comment":"\"20 machines ranging from ranking 21 ... to 446\" should read \"rank 21 ... to rank 446.\"","section":"Section 1"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript reads as a community survey rather than an original research result, which is appropriate for the target venue. The self-citations are frequent but legitimate given that the authors describe their own production systems. The main fixes needed are the Table 1 comparability caveat, the Table 3 measurement-basis statement, and completion of the placeholder references."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a review, not a research contribution, but it is a genuinely useful one. The new value is the consolidated, self-reported picture from eight German centers: power breakdowns (Table 1), DLC adoption and temperatures (Table 2), heat reuse (Table 3), monitoring stacks (Table 4), plus production experiences with power capping, scheduling, and idle-node shutdown. I know of no other single document that puts these side by side. The central claim—that no single measure suffices and that a holistic portfolio is needed—is supported by the case studies and is not controversial; the paper does not try to prove something circular, and self-citations are appropriate in a multi-center survey of the authors' own systems.\n\nSoft spots: Table 1 is the weakest link. The 78–86% compute share range is narrow, but the measurements are not apples-to-apples. MPCDF's storage is folded into compute, monitoring intervals and collection technologies differ (Table 4), and sites likely draw the line between 'infrastructure' and 'compute' differently. With no error bars or averaging windows, the 'quite homogeneous' claim is more fragile than the prose suggests. The paper's main conclusion does not depend on those exact percentages, so this is not fatal, but the homogeneity statement should be softened or backed by a precise measurement protocol before people cite Table 1 as definitive. There are also drafting artifacts: a 'FixMe' reference, 'NVIDIA (?)', empty acknowledgments, and figure data not yet public. Minor, but a journal should ask for cleanup.\n\nWho it is for: HPC operators, facilities people, procurement, and policymakers. The individual techniques are established in the literature, so specialists will not learn much that is new, but the cross-site comparison and production experience are valuable. I'd send it to peer review rather than desk reject; it deserves a referee who will push on the measurement accounting. With a modest revision—especially Table 1 methodology—it can serve as a solid reference.","headline":"Useful cross-site review of energy-efficient HPC operations in Germany; Table 1 needs more careful treatment before the paper becomes a reference.","tokens_in":26644,"tokens_out":2174,"would_cite":true,"duration_ms":172082,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that no single technology can make supercomputing sustainable; eight German centers already run a combined portfolio of green power, liquid cooling, heat reuse, power capping, monitoring, scheduling, and application…","keywords":["High-Performance Computing","energy efficiency","direct liquid cooling","power capping","heat reuse","data center monitoring","scheduling","Germany"],"falsifier":"An external audit using identical submetering at all surveyed centers would falsify the paper's core empirical claim if it found the compute share far outside the reported 78 to 86 percent band, or if a controlled comparison showed that a single measure alone met the legal efficiency targets at equal throughput.","tokens_in":25589,"feed_emoji":"⚡","tokens_out":8142,"duration_ms":95909,"temperature":0.7,"pith_summary":"Supercomputers are growing faster than their per-watt efficiency improvements, and this paper argues that the only way German HPC centers can meet legal and economic pressure is to treat energy efficiency as a whole-system problem rather than a single technology choice. Drawing on production experience from eight major centers, it reports that compute hardware takes 78 to 86 percent of site power, infrastructure 7 to 15 percent, and storage 3 to 8 percent, and that successful sites combine green electricity, warm-water direct liquid cooling, heat reuse, power capping, monitoring, energy-aware scheduling, and application tuning. The authors present these measures as already implemented in production, not as research proposals, and conclude that no one measure alone can hit the efficiency targets. The reader should care because the same pressure is now reaching HPC centers worldwide, and the German sites provide a concrete portfolio of techniques that others can adopt.","feed_headline":"No single fix makes supercomputers green, eight German centers find","feed_subtitle":"A survey of production experience says cooling, capping, heat reuse, monitoring, and scheduling must work together.","key_machinery":"The object that carries the argument is the combined production efficiency stack: power supply choices, data-center cooling and heat-reuse infrastructure, heterogeneous compute hardware, storage, cluster-wide monitoring, scheduling and power management, and application-level optimization. The paper's method is the cross-site survey: tables of measured power shares, cooling parameters, heat-reuse factors, and monitoring-stack characteristics from eight production centers, read together with processor efficiency trends. The tables do the work of showing that the same portfolio appears across sites of different sizes, which is the basis for the claim that the portfolio is the cause of achievable efficiency rather than a site-specific accident.","core_discovery":"On the paper's own terms, the central discovery is that a holistic, multi-layer efficiency portfolio is both necessary and already operational in large-scale HPC production. The evidence is a standardized cross-site comparison: despite system sizes spanning roughly 1.1 to 4.7 megawatts per site, the power breakdown is strikingly homogeneous—compute 78 to 86 percent, infrastructure 7 to 15 percent, storage 3 to 8 percent—and the centers already run direct liquid cooling, dynamic power limiting, job-aware monitoring, heat reuse (current energy reuse factors up to 20 percent, with future plans reaching higher), and energy-aware scheduling. The paper treats this convergence of practices as proof that the efficiency targets set by German law are reachable, but only when all these measures are pursued together.","pith_inferences":["If the 78 to 86 percent compute-power share holds across sites, then the biggest energy lever is compute efficiency itself, and infrastructure savings—while legally necessary—have a smaller ceiling unless heat reuse turns waste heat into a usable product.","The paper's monitoring data imply that job-level energy accounting is close to becoming a standard operational tool; one testable extension is using historical power traces to predict and shift energy-intensive jobs to hours of cheap renewable electricity, a step the paper names as future research but does not claim to have demonstrated.","The trend toward lower cooling temperatures for denser GPU racks could reverse the gains from warm-water free cooling, so an open question is whether chip and packaging innovations can keep rack power density below the threshold where chillers return."],"forward_implications":["If the holistic-approach conclusion is right, investments in any single category—say, only better cooling or only power capping—will fall short of the mandated efficiency targets.","Production sites can already achieve substantial efficiency gains: idle-node shutdown at one center cut annual energy use by about 25 percent, and warm-water cooling transfers more than 95 percent of system heat into water at several sites.","Future systems will likely become more heterogeneous and more tightly coupled to the electricity grid, because processor specialization offers the largest remaining per-watt gains and centers are exploring demand-response operation.","Heat reuse is the least mature layer, with current energy reuse factors of 0 to 20 percent, and its role will grow as legislation requires 10 to 20 percent reuse for new data centers."],"supporting_citations":[{"why":"Positions the surveyed systems in the global performance ranking and documents the growth in system size.","marker":"(84)"},{"why":"Supplies the claim that computational efficiency per watt improves with each generation.","marker":"(83)"},{"why":"Defines the German legal targets for power usage effectiveness and waste-heat reuse that motivate the measures.","marker":"(31)"},{"why":"Provides the industrial electricity price data used in the cost-trend analysis.","marker":"(17)"},{"why":"Gives an example of a comprehensive production monitoring stack used at a large HPC center.","marker":"(65)"},{"why":"Shows a runtime energy-management system applied to adjust processor frequencies during job execution.","marker":"(20)"},{"why":"Provides a production power scheduler that enforces dynamic job-level power caps in overprovisioned systems.","marker":"(78)"},{"why":"Supports the claim that power capping at HPC scale can reduce energy consumption.","marker":"(107)"},{"why":"Offers a case study of waste-heat recovery used for cooling, informing the heat-reuse discussion.","marker":"(102)"}],"fun_headline_variants":["Supercomputer efficiency: one fix won't cut it, German study shows","German HPC centers reveal recipe: combine cooling, capping, and more","Efficiency in practice: German supercomputers blend many measures","No silver bullet for green supercomputers, German experience says","Holistic efficiency: German HPC sites show multi-pronged path"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the self-reported power breakdowns, cooling data, and monitoring intervals from the eight centers are consistent enough to compare, even though the paper notes that one site cannot separate storage power and that monitoring technologies and intervals differ across sites.","fun_headline_variants_meta":{"raw":{"variants":["Supercomputer efficiency: one fix won't cut it, German study shows","German HPC centers reveal recipe: combine cooling, capping, and more","Efficiency in practice: German supercomputers blend many measures","No silver bullet for green supercomputers, German experience says","Holistic efficiency: German HPC sites show multi-pronged path"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2241,"prompt_tokens":928,"completion_tokens":1313,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1220}},"tokens_in":544,"tokens_out":1313,"duration_ms":9974,"temperature":1.0,"reasoning_tokens":1220,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:22:22.922443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An external audit using identical submetering at all surveyed centers would falsify the paper's core empirical claim if it found the compute share far outside the reported 78 to 86 percent band, or if a controlled comparison showed that a single measure alone met the legal efficiency targets at equal throughput.","supporting_citations":[],"review_version":1}