{"id":"2a162e7d-709d-44ee-abad-ebebc8e0005b","arxiv_id":"2505.22864","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"NRP operates as a single stretched Kubernetes cluster across more than 75 sites, providing GPUs and CPUs to researchers at over 50 campuses.","lead":"This paper describes the National Research Platform (NRP), a distributed Kubernetes cluster spanning 75 locations, and reports its 2024 usage of 5M GPU-hours and 50M CPU core-hours. It matters because it shows how a stretched, multi-tenant Kubernetes platform can give researchers at 50 campuses access to shared GPUs and CPUs for AI and science.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 2024 usage figures and the §3.1 utilization gain both depend on an accounting pipeline whose metric definitions and validation are never described; reconstructing the totals and re-analyzing the reservation effect is needed to support the paper's central quantitative claims.","rationale":"The reader's weakest_assumption correctly isolates §3.1, and I agree that the 21.49% to 31.37% comparison is not an established effect. My stress-test read treats that as the clearest symptom of a broader load-bearing condition: the quantitative basis of the central claim is the accounting pipeline described in §5, and no definition, validation, or error analysis is offered for any of the headline numbers (670 namespaces, 5M GPU-hours, 50M CPU core-hours, OSG 6.5M CPU hours). The §3.1 causal attribution is additionally confounded by the rising demand shown in Fig. 2 and by changes in the GPU pool, so the paper's operational lesson is not supported by the numbers as written. This is not a reason to doubt that the platform exists; the paper's descriptive architecture is coherent and self-consistent. But for a claim that rests on scale and measured effectiveness, the missing methodology is a correctable but currently unmet condition. Conditional acceptance—requiring the authors to publish query definitions and a controlled reanalysis—is proportionate; outright rejection would ignore the substantial descriptive value and the absence of any internal contradiction.","tokens_in":5758,"tokens_out":9852,"duration_ms":105001,"concrete_test":"Ask the authors to release the raw time series and dashboard/PromQL queries behind Fig. 2 and the 2024 totals, then independently recompute the 5M GPU-hour and 50M CPU-core-hour sums from raw accounting exports using a defined metric (e.g., DCGM device busy-time versus allocated time). For §3.1, re-estimate the reservation-system effect with a regression that controls for month, GPU model mix, and node count. If the recomputed totals disagree materially, or if the adjusted utilization effect is not statistically significant, the quantitative basis of the central claim is not established.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing weakness is the absence of a stated measurement methodology for the paper's quantitative evidence, and the §3.1 utilization claim is the clearest case. The text reports that after introducing the A100 reservation system, 'average GPU utilization across the cluster improved significantly from 21.49% to 31.37%,' and attributes this to the reservation system. But the preceding sentence notes that demand for high-end GPUs was growing (Fig. 2), and the hardware pool was changing as new models arrived, so a simple before/after comparison cannot identify the reservation system's causal effect. 'Average utilization' is not defined (allocated request-time vs. measured device busy-time) and the denominator is unspecified. The same accounting pipeline described in §5 (Prometheus/Thanos with custom caching, query segmentation, and a Redis layer) is the sole source of the headline 2024 aggregates of 5M GPU-hours and 50M CPU core-hours, yet no validation or error analysis of that pipeline is reported. If these numbers are not reproducible, then even the central 'national-scale infrastructure' claim is only as strong as an unverifiable self-report.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes the National Research Platform (NRP), a distributed multi-tenant Kubernetes cluster that federates compute and storage nodes across more than 70 sites in the U.S. and internationally. It reports a cluster of 420+ nodes, 1,400+ GPUs, 28,000+ CPUs, and 161 TB of RAM, presenting the whole federation as a single batch cluster to users. The paper details the deployment toolchain (Ansible, NetBox), three integration layers (IPMI hardware-level, OS-level, and Kubernetes peering via Admiralty), user interfaces (JupyterHub, Coder, direct kubectl), storage (regionally distributed Ceph via Rook), security (Falco-based threat detection), accounting and monitoring (Prometheus/Thanos with custom caching, sFlow, PerfSONAR), and future scheduler work with YuniKorn. It also gives 2024 usage statistics: 670 namespaces, 5M GPU-hours, 50M CPU core-hours, and a claimed improvement in average GPU utilization from 21.49% to 31.37% after introducing an A100 reservation system.","tokens_in":6001,"tokens_out":5949,"duration_ms":58749,"significance":"If the operational claims are accurate, the NRP represents a notable and rare instance of a stretched, multi-tenant Kubernetes federation that gives researchers a unified cluster view across many autonomous sites. The paper's value is primarily descriptive: it documents a real system with specific engineering choices—layered integration, regional Ceph distribution, Falco-based intrusion detection, and a custom accounting layer—that are directly useful to operators of similar community infrastructures. The 2024 scale figures are striking and plausible given the NSF funding and prior Pacific Research Platform lineage, but they are self-reported and lack metric definitions; the utilization improvement claim is also methodologically under-specified. The paper includes a transparent disclosure that LLMs hosted on NRP assisted in editing, which is commendable. This is not a paper with machine-checked proofs or reproducible experiments, but as a practice/experience case study it has clear value to the PEARC audience.","major_comments":[{"comment":"The reported utilization improvement from 21.49% to 31.37% is presented without defining \"average GPU utilization\": it is not stated whether this is request-time from allocated reservations, device busy-time as measured by GPU telemetry, or something else, nor is the denominator or the measurement window specified. The surrounding text also notes that demand for high-end GPUs was growing (see Fig. 2) and that the hardware pool was changing as new models arrived, so a simple before/after comparison cannot identify the reservation system's causal effect. Please define the metric and its source, and either temper the claim that the improvement \"underscores the effectiveness of the reservation system\" or provide additional controls (e.g., matching by GPU type, time period, workload mix).","section":"Section 3.1"},{"comment":"The headline 2024 totals—5M GPU-hours, 50M CPU core-hours, and 670 namespaces—depend entirely on the accounting pipeline described in Section 5 as using Prometheus, Thanos, custom caching, query segmentation, and a Redis layer. No definition is given for how GPU-hours and core-hours are computed (e.g., allocated pod request-time multiplied by duration, versus measured busy-time), what query granularity is used, or how the pipeline is validated against ground truth. Without this information, the central scale claims are not interpretable or reproducible. Please add a short paragraph in Section 5 that states the exact metric definitions, the data sources, and any known limitations or approximations.","section":"Section 5 (and Section 1)"}],"minor_comments":[{"comment":"The abstract says \"over 75 locations\" while Section 2 says \"more than 70 sites\"; reconcile these counts or use a single approximate figure consistently.","section":"Abstract and Section 2"},{"comment":"\"The Network Research Platform\" should be \"The National Research Platform\" (the word \"Network\" is a typo).","section":"Section 2, first sentence"},{"comment":"\"Cloud Native Computing Federation\" should be \"Cloud Native Computing Foundation (CNCF)\".","section":"Section 4"},{"comment":"Reference [6] contains a capitalization typo: \"Red HAt\" should be \"Red Hat\".","section":"References"},{"comment":"The figure caption reports a 12-month total of 6,605,042 GPU hours from 412 research groups, while Section 1 reports 5M GPU-hours for calendar year 2024 from 670 namespaces; clarify the time window of the figure and how these numbers relate to the 2024 totals and the namespace/group counts.","section":"Figure 2 caption and Section 1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits PEARC's practice/experience scope, so I would not demand full independent validation of operational telemetry. However, the missing metric definitions in the accounting and utilization claims are substantial enough that authors should be asked to add a short methodology paragraph and to soften the causal language. The site-count discrepancy and typos are easily fixed. I would be comfortable with acceptance after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a credible practice-and-experience paper about the National Research Platform, a stretched Kubernetes cluster federating nodes across more than 70 sites. The genuinely new content is the operational description of NRP itself—how it integrates hardware at the IPMI, OS, and Kubernetes layers; how it manages inventory with NetBox and deployment with Ansible; how it uses Rook/Ceph for regional storage, Falco for container threat detection, and a custom Prometheus/Thanos/Redis accounting stack—plus the 2024 usage statistics. It is a direct continuation of the Pacific Research Platform, and the paper frames that lineage honestly.\n\nWhat it does well: the contrast with OSG and ACCESS is clear. NRP federates at the node level rather than overlaying batch schedulers, and the multi-tenancy story (670 namespaces, 525 campus researcher namespaces, 23 MSIs, EPSCoR states) is useful evidence of the platform's reach. The security section, though short, makes a concrete point about Falco being chosen after Prometheus overload alerts proved insufficient. The acknowledgment that the paper was edited with LLMs hosted on NRP is a nice honest footnote, not a flaw.\n\nSoft spots: the utilization improvement claim in Section 3.1—21.49% to 31.37%—is presented without defining what 'utilization' means or how it was measured. The stress-test note is correct that demand was growing and the hardware pool was changing, so a simple before/after comparison cannot isolate the reservation system's causal effect. This should be fixed with a definition and a caveat, but it is one metric, not the paper's backbone. The headline numbers (5M GPU-hours, 50M CPU core-hours) come from the accounting pipeline described in Section 5, and the paper does not report validation or error bars for that pipeline. For a PEARC practice paper, self-reported usage numbers are normal; a sentence on how the totals are derived would be enough. There is also a minor inconsistency: the abstract says 'over 75 locations' while Section 2 says 'more than 70 sites.'\n\nOverall, the central descriptive claim—that NRP operates as a distributed multi-tenant Kubernetes cluster at national scale and supports a real research community—holds up. The math is not the point; the operational evidence is. This paper deserves a serious referee and likely acceptance after minor revision. I would ask the authors to clarify the utilization metric and to add a brief note on the accounting pipeline's provenance and any known error margins.\n\nRecommendation: send it to peer review. It is exactly the kind of experience report PEARC is for.","headline":"A credible practice-and-experience paper on a real distributed Kubernetes platform; the usage numbers are self-reported and the §3.1 utilization gain is under-specified, but the architecture and operational detail carry the paper.","tokens_in":6537,"tokens_out":2559,"would_cite":true,"duration_ms":23526,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues for a 'stretched' single Kubernetes cluster as national scientific infrastructure, citing 670 namespaces, 5M GPU-hours, and 50M CPU core-hours in 2024.","keywords":["distributed computing","Kubernetes","multi-tenant cluster","scientific cyberinfrastructure","GPU allocation","high throughput computing","AI and ML workloads","federated infrastructure"],"falsifier":"Recompute average GPU utilization before and after the reservation rollout with one fixed formula, such as pod GPU-seconds divided by available GPU-seconds on the same node set; if the 21.49% to 31.37% gap disappears under that consistent measure, the reservation system's benefit is not established.","tokens_in":5616,"feed_emoji":"☸️","tokens_out":9556,"duration_ms":95140,"temperature":0.7,"pith_summary":"This paper reports on the National Research Platform, a Kubernetes installation that spans more than 75 U.S. and international sites and presents itself to users as a single batch cluster. The authors' central operating claim is that 'stretched' federation at the hardware, operating-system, and Kubernetes-peering levels is a viable way to build national-scale scientific cyberinfrastructure, especially for AI and machine learning. They point to 670 active research namespaces consuming 5M GPU-hours and 50M CPU core-hours in calendar year 2024, with more than three-quarters of the GPU usage coming from 525 individual campus-researcher namespaces across 50 campuses. The paper also describes the mechanisms that make the federation practical: an A100 GPU reservation system, regionally distributed storage, Kubernetes-integrated security monitoring, and preemptible background work from the high-throughput computing community. A sympathetic reader would take away that distributed, multi-tenant Kubernetes can serve researchers across the spectrum from large universities to community colleges.","feed_headline":"One Kubernetes cluster now serves 670 research groups across 75 sites","feed_subtitle":"In 2024 the platform logged 5M GPU-hours and 50M CPU core-hours for campus researchers.","key_machinery":"The central object is the stretched Kubernetes cluster itself, with everything else in service of making that federation manageable. Nodes join at one of three levels: hardware-level control through IPMI (server management hardware), operating-system-level onboarding via an automation tool that installs the container runtime, Kubernetes, and system settings, or peering with independently operated Kubernetes clusters. Multi-tenancy is enforced through namespaces, with users getting RBAC-scoped credentials from a federated identity system. Storage is provided by a distributed filesystem replicated across five U.S. regions, and security monitoring sits inside Kubernetes through a runtime threat-detection tool that captures process-level detail on pods. The A100 reservation system is the load-balancing mechanism that the paper credits with the utilization improvement.","core_discovery":"The central discovery is that a single Kubernetes cluster can be stretched across more than 75 administrative domains and still behave as one batch cluster from the user's point of view. As of spring 2025 the platform includes over 1,400 GPUs, 28,000 CPUs, and 161 TB of RAM spread across more than 420 nodes, with sites contributing hardware at the IPMI, operating-system, or Kubernetes-peering level. The paper offers 2024 accounting as evidence that this model works at national scale: 670 namespaces used 5M GPU-hours and 50M CPU core-hours, and 525 campus-researcher namespaces from 50 campuses consumed over three-quarters of the GPU-hours and two-thirds of the CPU-hours. It further reports that reserving A100 GPUs raised average cluster GPU utilization from 21.49% to 31.37%, that regionally replicated storage keeps data close to compute, and that preemptible work from a high-throughput federation contributed 6.5M CPU hours in 2024.","pith_inferences":["Editorial inference: a natural test of the stretched-cluster model is whether queue wait times and failure rates stay acceptable as sites and namespaces grow; the paper does not report queue-level quality-of-service data.","Editorial inference: if the utilization claim survives a fixed measurement definition, the reservation mechanism could be extended to other high-demand GPU models.","Editorial inference: the education-facing use, including community colleges, points to a workforce-training benefit that may grow faster than research workloads and become the platform's dominant justification."],"forward_implications":["If the platform performs as reported, a campus can contribute as little as one node and still draw on the full national resource pool.","The 525 individual campus-researcher namespaces from 50 campuses suggest the model spreads AI/ML capability beyond large research universities.","The reported jump in average GPU utilization after A100 reservations indicates scheduling policy can materially improve scarce-GPU use without new hardware.","Routing preemptible work from a high-throughput computing community through idle cycles contributed 6.5M CPU-hours in 2024 without displacing primary workloads."],"supporting_citations":[{"why":"Provides the multi-cluster peering mechanism that lets the platform integrate externally operated Kubernetes clusters.","marker":"[1]"},{"why":"Supplies the contrasting single-site resource model against which the paper positions the stretched-cluster approach.","marker":"[2]"},{"why":"Documents a campus research group using the platform with new hardware, an example of the adoption pattern the paper generalizes.","marker":"[7]"},{"why":"Gives the high-throughput federation model the platform extends and the source of preemptible jobs that used 6.5M CPU-hours in 2024.","marker":"[14]"},{"why":"Describes the earlier testbed from which the platform evolved, grounding its distributed-networking lineage.","marker":"[16]"},{"why":"Supplies the distributed filesystem backing the regionally replicated persistent storage that keeps data near compute.","marker":"[18]"}],"fun_headline_variants":["Single Kubernetes cluster spans 75 sites for 670 research groups","One stretched Kubernetes cluster: 5M GPU-hours in 2024","NRP: multi-tenant Kubernetes with 1,400 GPUs across 75 sites","Multi-site Kubernetes cluster powers 670 groups, 50M CPU-hours"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the reported rise in average GPU utilization from 21.49% to 31.37% is real and caused by the new A100 reservation system, but the paper never defines how utilization was measured.","fun_headline_variants_meta":{"raw":{"variants":["Single Kubernetes cluster spans 75 sites for 670 research groups","One stretched Kubernetes cluster: 5M GPU-hours in 2024","NRP: multi-tenant Kubernetes with 1,400 GPUs across 75 sites","Multi-site Kubernetes cluster powers 670 groups, 50M CPU-hours"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000738,"raw_usage":{"total_tokens":3262,"prompt_tokens":873,"completion_tokens":2389,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":2308}},"tokens_in":489,"tokens_out":2389,"duration_ms":15398,"temperature":1.0,"reasoning_tokens":2308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:57:40.461247+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute average GPU utilization before and after the reservation rollout with one fixed formula, such as pod GPU-seconds divided by available GPU-seconds on the same node set; if the 21.49% to 31.37% gap disappears under that consistent measure, the reservation system's benefit is not established.","supporting_citations":[{"cited_title":"2025.Multi-Cluster Kubernetes","cited_arxiv_id":null,"evidence_quote":"Provides the multi-cluster peering mechanism that lets the platform integrate externally operated Kubernetes clusters."},{"cited_title":"Adventures with Grace Hopper AI Super Chip and the National Research Platform","cited_arxiv_id":"2410.16487","evidence_quote":"Documents a campus research group using the platform with new hardware, an example of the adoption pattern the paper generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the distributed filesystem backing the regionally replicated persistent storage that keeps data near compute."}],"review_version":1}