{"id":"5939270f-4e16-4282-aa5a-c5be389dcfab","arxiv_id":"2507.22294","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proposes workflow templates and experiment management as key to simpler HPC benchmarking, but validates this only through the authors' own two tools.","lead":"This paper argues that reusable workflow templates and simple experiment-management tools can make HPC benchmarking much easier for scientists and students. It introduces the term 'benchmark carpentry' and compares two tools, Cloudmesh and SmartSim, to show they cover the same needs.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that the requirements are fundamental rests on an unverified independence premise and a circular validation: the requirements were derived from the very tools used as evidence of convergence.","rationale":"The reader's weakest_assumption correctly identifies the independence premise as the crucial unverified step, and the reader's rationale also notes the circularity of deriving requirements from the tools and then using those tools as validation. I agree with that assessment. My stress-test adds that the circularity is not merely a side issue: even if the tools were genuinely independent, the requirements list originates from the authors' own design experience, so the overlap only shows that two systems built from the same personal requirements happen to share those requirements. The independence claim is the only thing that could break the circle, and it is asserted rather than demonstrated. The paper offers no external comparison point despite the existence of many other workflow systems (Pegasus, Nextflow, Parsl, Snakemake, WfBench), and the related-work section is restricted to the authors' own publications. The proposed test—a provenance and cross-reference audit of the two repositories' histories—is a concrete, feasible way to check whether independence is real. If the test finds cross-references, the central claim collapses. If it finds none, the convergence argument would still be weaker than a direct external comparison, but at least the independence premise would be supported. Either way, the verdict of REJECT remains appropriate because the paper's central scientific claim is not adequately supported by the evidence presented; the software descriptions and use cases may still be valuable to practitioners, but that does not salvage the 'fundamental requirements' conclusion. I therefore recommend no change to the reader's verdict.","tokens_in":44227,"tokens_out":2814,"duration_ms":33804,"concrete_test":"Analyze the complete GitHub commit, issue, pull-request, and release histories of cloudmesh-cc/cloudmesh-ee and SmartSim (including all branches and tags) for any pre-2025 mention of the other project, of shared co-authors or contributors, or of common workflow predecessors (Karajan, Swift, Pegasus, Grid workflow, Globus). If no such cross-references exist, the independence claim gains support; if any exist, the 'completely independently developed' assertion in Section 3 is falsified and the central convergence argument loses its evidential force.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 6 and abstract) is that the requirements are fundamental because two 'completely independently developed' tools (Cloudmesh and SmartSim) converged on them. Yet Section 3 opens: 'The requirements discussed previously were distilled from our experience as principal developers of two workflow libraries Cloudmesh and SmartSim.' Thus the requirements were extracted from the same tools that are then presented as independent confirmation. The independence itself is asserted in Section 3 ('these two projects were developed completely independently without knowledge of the other until the writing of this paper') with no supporting evidence such as development timelines, issue histories, or absence of shared contributors. Both tools are developed by co-authors of this paper and are embedded in the same HPC workflow community with decades of shared prior art (e.g., Grid workflow systems, Pegasus, Swift, Karajan, MLCommons Science working group). The paper explicitly declines to compare with other workflow systems ('the authors are not privy to the software engineering and design history of those'), so no external baseline is offered. If the overlap stems from common lineage, shared community conventions, or the authors' prior collaborations rather than independent discovery, the convergence provides no external validation. This is the load-bearing step: without verified independence and without external comparison, the 'fundamental requirements' claim is circular and unsupported, regardless of whether the software itself is useful.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes workflow templates and \"benchmark carpentry\" as a way to make HPC benchmarks accessible to diverse scientific communities. It compiles a set of workflow requirements organized into compute systems, users, workflow specification, runtime, authentication, data management, and licensing, based on the authors' decades of experience and their participation in MLCommons Science. The paper's central claim is that these requirements are fundamental because two allegedly independent tools, Cloudmesh's Experiment Executor and Compute Coordinator and HPE's SmartSim, converge on the same abstractions. The paper also describes the OSMI benchmark, a cloud-provisioning plugin for AWS PCS with a cost equation, and several application use cases. A long appendix summarizes the co-authors' own prior work.","tokens_in":44439,"tokens_out":9336,"duration_ms":99760,"significance":"If the convergence claim were well founded, the paper would provide a useful requirements checklist for community HPC/AI benchmark workflows and a novel educational concept (benchmark carpentry). The paper has concrete strengths: the five-tier resource model, Eq. (1) for AWS cluster cost with cost tables, the explicitly FAIR-oriented data management discussion, and the public code repositories for Cloudmesh plugins and OSMI are checkable artifacts. However, the validation is circular and the independence premise is unverified, so the manuscript does not currently establish that the requirements are \"fundamental\" as claimed in the abstract and Section 6. The paper also ships no machine-checked proof or reproducible dataset for the index-equilibrium or education claims; the AWS cost arithmetic is internally consistent.","major_comments":[{"comment":"The requirements in Section 2 are stated to be \"distilled from our experience as principal developers of two workflow libraries Cloudmesh and SmartSim,\" yet the same two tools are then presented in Table 4 and the conclusion as independent confirmation that the requirements are fundamental. This is a circular validation: the evidence and the hypothesis share the same origin. To support the convergence claim, the authors need either to show that the requirements were derived before the tools' design decisions (e.g., from the MLCommons Science working group or from external sources) or to apply the requirement list to several workflow systems not authored by this paper's authors and demonstrate that the list predicts their feature sets.","section":"Section 3, opening paragraph; Section 6"},{"comment":"The claim that Cloudmesh and SmartSim \"were developed completely independently without knowledge of the other until the writing of this paper\" is asserted without supporting evidence. The manuscript itself notes that \"the authors are not privy to the software engineering and design history of those\" other workflow systems, so no external baseline is provided. Because both tools are authored by co-authors of this manuscript, the convergence between them could plausibly arise from shared community conventions, prior collaborations, or common lineage (e.g., the authors' own earlier workflow systems such as Karajan and Swift). Please provide concrete evidence of independence, such as development timelines, issue histories, or contributor lists, or alternatively add a systematic comparison of the requirements against external systems such as Pegasus, Nextflow, Parsl, and ExaWorks.","section":"Section 3, opening paragraph; Author Contributions"},{"comment":"The claim that using Cloudmesh EE reduces the on-ramp time from weeks-months to less than a day and lowers the required team from graduate students to a single undergraduate is supported only by an informal qualitative observation, with no sample size, no methodology, no control group, and no data. This claim is used in the abstract to support the educational benefits of \"benchmark carpentry.\" Please either provide a rigorous evaluation (e.g., a structured user study with numbers of participants and tasks) or reframe the claim as anecdotal and remove it from the abstract's central argument.","section":"Section 3.3.2, paragraph beginning \"While practically working with the system\""}],"minor_comments":[{"comment":"\"The increasing using of machine learning\" should be \"The increasing use of machine learning.\"","section":"Section 1, paragraph 4"},{"comment":"The tier labels are inconsistent (\"Tier-0\" vs. \"Tier 0\"); choose one convention for readability.","section":"Section 2.1"},{"comment":"The import line reads \"from smartsim impo rt Experiment\"; fix the typo.","section":"Section 3.2, Listing 1"},{"comment":"The output option \"-o\" appears twice; the second instance should presumably be \"-e\" for error output.","section":"Section 3.3.2, SLURM template example"},{"comment":"\"A WS PCS\" should be \"AWS PCS.\"","section":"Section 3.4, Table 1 heading"},{"comment":"\"undelaying respurces\" and \"direclty\" should be \"underlying resources\" and \"directly.\"","section":"Section 3.7"},{"comment":"\"facillitated\" should be \"facilitated.\"","section":"Section 4.1"},{"comment":"The claim \"Index Equilibrium is at about 7\" should cite the specific Top500 list version and date and provide the underlying data, since this number is used in the \"Minimal support for virtualization in the cloud\" implication.","section":"Section 2.1, Figure 1"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection because the weaknesses, while load-bearing, are in principle addressable: the authors could add external comparisons with non-author workflow systems, provide evidence for the independence claim, or systematically downgrade the \"fundamental requirements\" claim to an experience report. If they cannot do any of these, the paper would not be suitable for a research journal, since its contribution currently rests on an unverified premise. The manuscript's heavy reliance on the authors' own prior publications (Appendix A) also deserves editorial attention, though it may be appropriate in a special issue context."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's central claim — that two 'completely independently developed' tools landing on the same requirements proves those requirements are fundamental — doesn't hold up. Section 3 opens by admitting the requirements were 'distilled from our experience as principal developers of two workflow libraries Cloudmesh and SmartSim,' and the same two libraries are then presented as independent confirmation. That is circular, and the independence premise is asserted without evidence: no development timelines, no issue histories, no shared-contributor analysis. The paper explicitly declines to compare with other workflow systems, citing lack of access to their design history. Given that both tools have co-authors here embedded in the same HPC-workflow community — and that Appendix A documents von Laszewski's earlier Karajan/CoG Kit work already supporting iterative workflows — the 'remarkable convergence' is less surprising than claimed.\n\nNow the credit. The Section 2 requirements catalog is a genuinely useful synthesis: the five-tier machine model, queuing-policy workarounds, split-VPN and SSH access, FAIR data handling, licensing risk. Tables 3 and 4 give a clean side-by-side comparison. The AWS PCS cost arithmetic in Section 3.4 is internally consistent, and the paper is honest that it has not answered whether cloud is cheaper. Code ships on GitHub (cloudmesh-ee, cloudmesh-cc, cloudmesh-vpn, OSMI), which is reproducible evidence of the software's existence and shape.\n\nSoft spots beyond the circularity are minor but real. The education claim is a single anecdote — students took weeks to months without EE and under a day with it — with no sample size or methodology. The paper's own limitation notes are honest: it says the requirements are not complete and 'likely represent a necessary subset,' which cuts against the 'fundamental' language elsewhere.\n\nProportionate verdict: this is an experience report with a solid requirements synthesis and an unsupported convergence argument. The fix is reframing — drop 'fundamental,' present the catalog as lessons learned from two systems, supply independence evidence or drop the convergence claim, and either measure the education claim or cut it. The OSMI integration is the strongest concrete portability evidence and should stay.\n\nWho is this for? Practitioners and educators building HPC/AI benchmark workflows, and the workflow-systems community. The requirements checklist and comparison make it worth a serious referee, but it needs revision first. I'd bring it to a reading group as a case study in what convergence arguments require.\n\nRecommendation: send to peer review, but expect heavy revision.","headline":"The convergence claim is circular — the requirements were distilled from the same two tools offered as independent evidence — but this is a useful experience report with a solid requirements catalog that deserves peer review after reframing.","tokens_in":45002,"tokens_out":5608,"would_cite":false,"duration_ms":59754,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two independently built HPC tools converged on the same abstractions for running benchmark experiments, and the authors take the convergence as evidence that those abstractions are fundamental to emerging HPC and AI workflows.","keywords":["benchmarking","hyperparameter search","experiment execution","workflow templates","benchmark carpentry","HPC workflows","Cloudmesh","SmartSim"],"falsifier":"Audit the two projects' design histories: if they reveal that SmartSim's designers had access to Cloudmesh, to the first author's earlier loop-capable workflow systems, or to the same community venues before SmartSim's abstractions were fixed, the independence premise fails and the convergence argument collapses. The complementary check is to run the Section 2 requirement list against a third-party workflow engine built with no author involvement; an engine that satisfies most rows would confirm the list is fundamental, while one that fails most rows would indicate the list encodes local design choices.","tokens_in":44022,"feed_emoji":"🧩","tokens_out":11251,"duration_ms":113001,"temperature":0.7,"pith_summary":"This paper claims that a compact set of experiment-execution requirements—hyperparameter sweeps, batch-queue abstraction, reusable workflow templates, FAIR-compliant result reporting, and SSH-based access across heterogeneous machines—are fundamental to running benchmark workflows on today's HPC and AI systems. The supporting evidence is convergence: Cloudmesh's Experiment Executor and Compute Coordinator and the SmartSim toolkit were, the authors assert, developed independently, yet they landed on nearly identical abstractions, responsibilities, and terminology. The paper turns that overlap into a requirements list spanning compute systems, users, workflow specification, runtime, authentication, data management, and licensing, and it proposes 'benchmark carpentry' plus shared templates as the way to make benchmarking teachable and portable. If the convergence claim is right, the list is a durable checklist for building the next generation of benchmark workflow tools and for training the scientists who will use them.","feed_headline":"Two independent HPC tools converged on one benchmark recipe","feed_subtitle":"Their overlap becomes a durable checklist for running AI-era benchmark experiments on supercomputers.","key_machinery":"The central object is the experiment template: a compact specification—YAML in Cloudmesh, a Python driver script in SmartSim—that names the application, the hyperparameters to iterate over, and the resources to use, from which the tool generates concrete batch scripts for a workload manager. The evidentiary machinery is the cross-comparison of the two implementations: Table 3 catalogs which features each system provides, and Table 4 maps every Section 2 requirement onto the two systems, displaying the functional overlap that grounds the convergence claim. Both systems support gridsearch over parameters, cyclic as well as direct-acyclic-graph (DAG) execution, batch-queue integration, and template reuse, and this shared core is what the argument treats as fundamental.","core_discovery":"The paper's central discovery is that two separately built Python toolkits for experiment execution—a workflow that provisions data, runs an application, and varies hyperparameters across many single runs—supply the same core machinery: a way to define an experiment as a parameterized template, a generator that expands a Cartesian product of hyperparameters and hardware parameters into many individual batch jobs, an abstraction layer over workload managers, and a uniform scheme for collecting and reporting results. Cloudmesh expresses experiments in YAML and adds features for SSH-based federation, split-VPN access, and plugin extensibility, while SmartSim expresses experiments in Python driver scripts and adds an in-memory datastore for exchanging data between simulation and AI components; neither of these differences is needed for the overlap. The authors conclude that because two independent implementations converged, the requirements distilled in their Section 2 are not arbitrary design choices but map onto fundamental needs that emerging HPC and AI workflows will place on any experiment executor.","pith_inferences":["The independence premise is the vulnerable joint of the argument: the two lead authors co-author this paper, both systems' teams move in the same benchmark community, and the related-work section shows the Cloudmesh author's earlier workflow systems already supported loops and iterations, so the overlap could reflect a shared lineage rather than independent discovery. Checking the two projects' de","The paper explicitly declines to compare with the hundreds of other workflow engines because the authors are not privy to their design histories; a cheap external check of the 'fundamental' claim would therefore be to run the Table 4 requirement list against an unrelated third-party engine and see how many rows it satisfies.","The benchmark-carpentry report carries a testable educational prediction: students given templated experiments reached a working benchmark in under a day while untemplated efforts took weeks to months. A controlled classroom study with matched tasks could measure this gap directly.","The convergence claim could be deepened beyond feature tables by diffing the artifacts: running one benchmark, such as cloudmask or OSMI, through both systems on the same machine and comparing the generated batch scripts and result schemas would show whether the overlap sits at the API level or extends down to the emitted jobs."],"forward_implications":["New benchmark workflow tools can be checked against the Section 2 requirement list as a functional checklist; the paper presents the list as a necessary, though not complete, subset for emerging HPC and AI workflows.","Reusable templates cut the time to stand up a working benchmark: the paper reports students using templated experiments reached a working benchmark in under a day, where untemplated efforts consumed weeks to months and typically needed a graduate student.","The OSMI surrogate-inference benchmark can be driven through both Cloudmesh and SmartSim with no changes to either package, which the authors take as evidence that the distilled requirements are sufficient for a real hybrid simulation-AI workload.","A next-generation experiment executor should combine the unique strengths of both systems, namely Cloudmesh's plugin, federation, and split-VPN capabilities with SmartSim's in-memory datastore and built-in inference support.","Provisioning on-demand HPC clusters in the cloud, priced in the paper at roughly $4.18 per GPU-hour for an A100 example, adds a cost-estimation duty to experiment executors, because on-premise users traditionally never see the price of their allocation."],"supporting_citations":[{"why":"Provides the SmartSim library whose design history and feature set form one half of the convergence claim.","marker":"[42]"},{"why":"Provides the Cloudmesh Experiment Executor whose YAML templates and gridsearch generation form the other half.","marker":"[87]"},{"why":"Establishes the Cloudmesh Compute Coordinator workflow model and the cloudmask use case cited for template reusability.","marker":"[92]"},{"why":"Introduces benchmark carpentry and the educational experience from which the user and workflow requirements were distilled.","marker":"[95]"},{"why":"Supplies the six HPC/AI execution motifs that anchor the claim about emerging workflow requirements.","marker":"[19]"},{"why":"Provides the surrogate-model deployment studies that the OSMI benchmark was founded upon and that both systems run.","marker":"[21]"},{"why":"Defines the FAIR principles that the paper adopts for result reporting, templates, and the federated results repository.","marker":"[119]"},{"why":"Supplies the MLPerf Training HPC results used to project cloud costs for benchmark runs.","marker":"[58]"}],"fun_headline_variants":["Two HPC toolkits reveal the same core recipe for AI-era benchmarks","Benchmark carpentry: how two Python tools hit the same design","Converging toolkits map the essential shape of experiment execution","Two independent executors converge on one benchmark blueprint","HPC benchmark tools agree: experiment templates are the key"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two projects were genuinely independent—Section 3 opens by asserting they were developed without knowledge of each other until this paper was written—because only that asserted independence turns the overlap into evidence that the requirements are fundamental rather than shared community convention.","fun_headline_variants_meta":{"raw":{"variants":["Two HPC toolkits reveal the same core recipe for AI-era benchmarks","Benchmark carpentry: how two Python tools hit the same design","Converging toolkits map the essential shape of experiment execution","Two independent executors converge on one benchmark blueprint","HPC benchmark tools agree: experiment templates are the key"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3150,"prompt_tokens":857,"completion_tokens":2293,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":2206}},"tokens_in":473,"tokens_out":2293,"duration_ms":18120,"temperature":1.0,"reasoning_tokens":2206,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:50:48.812405+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Audit the two projects' design histories: if they reveal that SmartSim's designers had access to Cloudmesh, to the first author's earlier loop-capable workflow systems, or to the same community venues before SmartSim's abstractions were fixed, the independence premise fails and the convergence argument collapses. The complementary check is to run the Section 2 requirement list against a third-party workflow engine built with no author involvement; an engine that satisfies most rows would confirm the list is fundamental, while one that fails most rows would indicate the list encodes local design choices.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the surrogate-model deployment studies that the OSMI benchmark was founded upon and that both systems run."}],"review_version":1}