{"id":"251587f6-1d68-4da8-ba74-5e684679b29c","arxiv_id":"2506.09879","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Accounting for uncertainties in empirical scaling laws, plasma profiles, and impurities shifts the predicted optimal SPARC operating point away from the deterministic POPCON optimum.","lead":"This paper builds statistical POPCONs, maps of how likely a tokamak discharge is to meet its performance goals when uncertain physics inputs are varied. Using Monte Carlo sampling and Bayesian optimization, it finds a more robust operating point for the SPARC tokamak than traditional deterministic predictions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that the statistical optimum differs from the deterministic one is not yet statistically supported: the paper reports no uncertainty on the location of a deliberately broad, nearly flat success-rate maximum.","rationale":"I read the paper as a well-executed demonstration that optimizing expected success under parametric uncertainty can move the recommended SPARC operating point away from a deterministic optimum. The Monte Carlo machinery, the gradient-based profile form, and the multi-fidelity Bayesian workflow are appropriate, and the paper is candid about limitations. My concern is not with the machinery but with the statistical support for the headline location claim. Because the success-rate surface is deliberately broad and flat, the pointwise argmax is a fragile statistic. The authors compare a single white-star argmax to a single black-star deterministic optimum without reporting the variability of either, despite having the machinery to do so, since multiple seeds are already used in Figures 13-14. This is more load-bearing than the input-distribution caveat: it would matter even if the Table 1 distributions were perfect. The reader's conditional verdict already captures the need for more support; the proposed bootstrap or multi-seed test would either convert the conditional acceptance into a firm one or show that the claimed shift is within noise. I therefore recommend no change to the conditional verdict, but the test should be a condition for acceptance.","tokens_in":17916,"tokens_out":9839,"duration_ms":127802,"concrete_test":"Run 20 independent multi-fidelity Bayesian optimizations with different random seeds at the highest Monte Carlo fidelity (40,000 samples), recording the argmax in density-temperature space and the pointwise success rate at the deterministic optimum for each run. Compute the two-dimensional ensemble confidence ellipse of the argmax and the Monte Carlo standard error of the success-rate difference between the white-star and black-star points. If the confidence ellipse excludes the deterministic operating point and the success-rate difference exceeds its standard error by a factor of roughly three or more, the central claim is supported; if the ellipse contains the deterministic point, the 'different from deterministic' conclusion should be withdrawn or weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the white star in Figure 6, the argmax of the pointwise success rate, is 'relatively far' from the black star, the deterministic optimum. The load-bearing support for this claim is the location of an argmax of a success-rate surface that the paper itself describes as a broad, nearly maximal zone. For a flat objective, the argmax is ill-conditioned: Monte Carlo sampling noise, the Gaussian-process surrogate, and the random seed can shift the argmax substantially while changing the success rate by much less than the reported ~1% Monte Carlo error. The paper shows convergence curves in Figures 13-14 but never gives a confidence region or bootstrap spread for the final optimal density and temperature. Consequently, even if every input distribution in Table 1 were perfectly calibrated, the specific assertion that the optimal operating point differs from the deterministic prediction is not yet quantitatively supported; it may be a statistically insignificant displacement of a noisy argmax on a plateau. This concern is distinct from, and prior to, the input-distribution caveat: it concerns the statistical identifiability of the claimed optimum rather than the calibration of the input uncertainties.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces 'statistical POPCONs,' a Monte Carlo extension of conventional plasma operating contour analysis, and applies it to the SPARC primary reference discharge. Uncertainties in the confinement time multiplier H, the L-H threshold multiplier k_LH, profile parameters (Ti/Te, peaking factors, a/L_T), and impurity concentrations are propagated through the CFSPOPCON power-balance workflow. The pointwise success rate is defined as the fraction of Monte Carlo samples satisfying four criteria: Q>2, auxiliary power <25 MW, positive flattop time, and Greenwald fraction <1. A gradient-based profile parameterization replaces parabolic profiles, and a multi-fidelity Bayesian optimization workflow is developed to locate the operating point maximizing the success rate. The central finding is that this statistical optimum (white star in Fig. 6) differs from the deterministic optimum (black star), lying instead in a broad, nearly flat maximum that balances H-mode access, confinement, impurity dilution, and auxiliary power. The paper also scans distribution means and standard deviations, edge argon enrichment, magnetic field, plasma current, and auxiliary power, and identifies H, k_LH, and Ti/Te as the dominant sensitivity drivers.","tokens_in":18178,"tokens_out":7470,"duration_ms":88546,"significance":"If the central claim is made statistically robust, the paper provides a valuable methodological contribution: it converts a traditionally deterministic design tool into a probabilistic one, quantifies Monte Carlo resolution, benchmarks Bayesian optimization against brute force and Powell methods, and demonstrates a practical multi-fidelity speed-up. The use of open-source tools (CFSPOPCON, MITIM, BoTorch) and the explicit convergence checks in Figs. 5, 13, and 14 are strengths. The sensitivity rankings and the recommendation to avoid operating near the uncertain H-mode boundary are actionable for SPARC operational planning. However, the headline claim about a different optimal operating point needs additional statistical support before it can be accepted as stated.","major_comments":[{"comment":"The paper's central claim, stated in the abstract and Section 3, is that accounting for uncertainties leads to an optimal operating point (white star in Fig. 6) that differs from the deterministic prediction (black star). This claim is not yet quantitatively supported because the success-rate surface is described as a 'broad nearly maximal zone' (Section 3) and the paper gives no confidence region, bootstrap spread, or multi-seed distribution for the argmax location in the nominal case or in the Section 5 scans. Figures 13-14 show convergence of the averaged optimal density and temperature and a one-standard-deviation seed spread, but the reported optimal points in Section 5 appear to be outputs of a single high-fidelity run rather than a distribution. On a nearly flat objective, Monte Carlo noise, the Gaussian-process surrogate, and the random seed can move the argmax substantially while changing the success rate by less than the ~1% Monte Carlo error quoted in Section 2.4. Please either report the argmax with an uncertainty region (e.g., bootstrap over seeds or resamples) or replace the 'different optimal operating point' formulation with a robustly testable statement such as 'the deterministic optimum lies in a region where the pointwise success rate is significantly below the maximum'; the latter is directly supported by the success-rate surface.","section":"Section 3, Figure 6; Section 4, Figures 13-14"},{"comment":"The quantitative content of the sensitivity rankings (H, k_LH, Ti/Te as dominant drivers) and all absolute success rates are conditional on the chosen input distributions. The standard deviations for Ti/Te, nu_T, a/L_T, and W are not derived from regression errors of the relevant scalings but are 'motivated by' physics-based modeling or generic impurity uncertainty references, and all ten parameters are sampled independently. The authors do acknowledge the uncorrelated assumption in Section 2.2 and the Discussion, but the abstract and Section 6 present the findings without that conditionality. I would like to see either a small correlation-sensitivity test for the top three drivers (e.g., imposing a plausible correlation between H and k_LH or between Ti/Te and a/L_T) or an explicit caveat in the abstract and conclusions that the ranking and quantitative success rates are conditional on Table 1. This is load-bearing because the operational recommendation to avoid the H-mode boundary is grounded in the k_LH and H rankings.","section":"Section 2.2, Table 1"}],"minor_comments":[{"comment":"The sentence 'The automated enforcement of the 140 MW fusion power limit can be sen' contains a typo: 'sen' should be 'seen'.","section":"Section 3"},{"comment":"The caption states that both the low-fidelity and multi-fidelity cases switch from training to optimization at '10,000 point evaluations'; this is inconsistent with the described workflow of five initial training points plus five iterations and will confuse readers. It likely should be '10 point evaluations'.","section":"Figure 14 caption"},{"comment":"The phrase 'Note, a fusion power less than 140 MW is enforced...' should be 'Note that a fusion power less than 140 MW is enforced...' for grammatical clarity.","section":"Section 2.3"},{"comment":"The paper introduces the maximum pointwise success rate as a 'loosely proportional proxy' for the total success rate, but Section 6 later states that the pointwise rate is not directly proportional because the area of the successful region can change. The caveat should be stated where the proxy is first introduced, not only in the Discussion.","section":"Section 3 and Section 6"},{"comment":"The limitation that empirical models may not capture critical-gradient transport and that some temperatures in the statistical POPCONs 'may not be achievable at all' is important; I recommend restating this alongside the abstract's claim so that readers do not over-interpret the optimal operating point as directly realizable.","section":"Section 6"},{"comment":"The phrase 'The color squares correspond' should be 'The colored squares correspond'.","section":"Figure 8 caption"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the statistical identifiability of the argmax is valid and should be the primary focus of revision. The paper needs either a bootstrap/seed-spread uncertainty region for the optimal location or a softened central claim that focuses on the success-rate deficit at the deterministic optimum. The input-distribution conditionality should also be made prominent in the abstract. No concerns about novelty or scope; the methodological contribution is suitable for the journal once the statistical support is added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful methods paper. The authors fuse three things that already exist in pieces—POPCON analysis, Monte Carlo uncertainty propagation, and Bayesian optimization—into a coherent workflow for picking tokamak operating points under uncertainty. The statistical POPCON idea is a natural extension, and the gradient-based profile parameterization is a sensible improvement over parabolas. The multi-fidelity BO scheme is well benchmarked against brute force and Powell, and the Monte Carlo resolution convergence is explicitly quantified (Fig. 5). They also use the open-source CFSPOPCON tool and are candid about their caveats (uncorrelated parameters, density treated as controllable, the pointwise-versus-total success-rate issue). That is real substance, and the paper deserves referee time.\n\nThe main qualitative finding—accounting for uncertainty broadens the acceptable region and moves the optimal operating point away from the deterministic one—is credible under the stated assumptions. The stacked histograms showing which parameters drive failure (H, kLH, Ti/Te) are genuinely informative, and the sensitivity scans in Section 5 are physically sensible.\n\nWhere I push back is on the sharper claim that the white star in Fig. 6 is 'relatively far' from the deterministic black star. The authors themselves describe the success-rate maximum as a broad, nearly maximal zone. For a flat objective, the argmax is ill-conditioned; Monte Carlo noise, the GP surrogate, and the seed can shift it around without changing the success rate by more than the ~1% MC error. The paper never reports a confidence region or bootstrap spread on the optimal density and temperature. So the location difference may be real, but it is not statistically supported as stated. The fix is straightforward: run a few more seeds at high fidelity, or report a bootstrap or CRB-style interval on the argmax. That would not change the qualitative message, but it would turn the central claim from suggestive to quantitative.\n\nThe input distributions are also assumptions, not measurements—especially the gamma choices and the std devs from ref. [15]. The authors acknowledge this, but the sensitivity rankings could shift with different tails. That is a limitation, not a fatal flaw.\n\nFor a reader in the fusion operations/design community, this is worth a serious look. The lack of shipped code and data for the specific scans is a reproducibility gap, but not a reason to desk-reject. Send it to review, with a request for either error bars on the optimum location or a rephrased claim that the deterministic point sits in a risky, steep region while the statistical optimum is in a broad safe zone.","headline":"A practical, well-executed uncertainty workflow for SPARC operating scenarios, with a central qualitative result that holds up, though the exact location of the 'statistical optimum' is underdetermined by the flat success-rate surface.","tokens_in":18690,"tokens_out":1324,"would_cite":true,"duration_ms":17724,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Accounting for uncertainty shifts SPARC's optimal operating point to a broad success-rate maximum rather than the deterministic prediction.","keywords":["statistical POPCON","uncertainty quantification","SPARC tokamak","Monte Carlo analysis","Bayesian optimization","confinement scaling","L-H transition threshold","operating scenario design"],"falsifier":"Re-run the statistical POPCON using correlated Monte Carlo draws (joint covariance among $H$, $k_{LH}$, and $T_i/T_e$ from physics-based transport ensembles) or using the first measured SPARC shot-to-shot distributions; if the location of the maximum pointwise success rate shifts by more than the grid spacing, or if predicted high-success points fail systematically, the independent gamma and Gaussian assumptions are the deciding factor.","tokens_in":17699,"feed_emoji":"🧲","tokens_out":7621,"duration_ms":72930,"temperature":0.7,"pith_summary":"This paper argues that the operating point of a next-step tokamak should be chosen by maximizing the probability of meeting its goals under realistic model uncertainty, not just the nominal predicted performance. It builds statistical POPCONs by Monte Carlo sampling ten uncertain inputs—confinement factor $H$, L–H threshold factor $k_{LH}$, ion-to-electron temperature ratio, profile peaking and gradient parameters, and impurity concentrations—through SPARC's power-balance model, producing a pointwise success rate at each density–temperature point. The central finding is that this probabilistic view moves the optimum away from the deterministic POPCON prediction and into a broad maximum that balances H-mode access, confinement quality, impurity dilution, and the 25 MW auxiliary power limit. A multi-fidelity Bayesian optimization workflow finds these optima much faster than brute-force grids, enabling systematic scans of assumptions. A sympathetic reader should care because SPARC and similar devices need scenarios that succeed reliably, not just on paper.","feed_headline":"Uncertainty shifts SPARC's optimal operating point","feed_subtitle":"Error bars on confinement and threshold scalings favor a broad operating window over the sharp deterministic optimum.","key_machinery":"The central object is the statistical POPCON: a map over volume-averaged density and temperature in which each point is assigned the fraction of Monte Carlo samples that satisfy all success conditions ($Q>2$, $P_{\\rm aux}<25$ MW, nonzero flattop time, Greenwald fraction below one). Three mechanisms carry the argument. First, physically motivated gradient-based profile forms replace parabolic profiles, using density and temperature peaking plus a core inverse gradient scale length $a/L_T$ with a fixed pedestal width, so profiles do not demand unrealistically large logarithmic gradients. Second, Monte Carlo draws from gamma and Gaussian distributions propagate the table of ten uncertainties through the power-balance workflow. Third, a multi-fidelity Bayesian optimization routine—Gaussian-process surrogates with Matern kernel and logarithmic expected-improvement acquisition, starting at 2,000 Monte Carlo samples and rising to 40,000—locates the maximum pointwise success rate much faster than brute force, which is what makes the extensive assumption scans feasible.","core_discovery":"On the paper's own terms, the discovery is that uncertainty is not a small correction to POPCON-based scenario design: it changes the recommended operating point. Propagating the uncertainties of empirical scalings, profile shapes, and impurity concentrations through the SPARC power balance yields a broad plateau of high pointwise success rate whose maximum sits near the nominal Primary Reference Discharge location (when edge argon dilution is ignored) and quite far from the deterministic POPCON optimum that includes core dilution from argon. The dominant sensitivities are $H$, $k_{LH}$, and $T_i/T_e$; notably, only moderately high $H$ values are preferred, because too high a confinement factor lowers the predicted scrape-off-layer power below the L–H threshold and locks the sample into L-mode scaling. For the nominal assumptions the total success rate—the fraction of Monte Carlo samples that succeed somewhere in operating space—exceeds 50%, and the authors conclude that future devices should not rely on operating close to uncertain boundaries such as the H-mode transition.","pith_inferences":["Editorial inference: the independence assumption on the ten parameters is likely the weakest link; correlated draws (e.g., high $H$ accompanying low $k_{LH}$ through pedestal physics) could narrow or shift the success plateau in ways the current gamma and Gaussian sampling cannot show.","Editorial inference: the pointwise success rate is a proxy for the total success rate, and the proxy can mislead when the area of the successful region changes; future work should optimize the total success rate directly despite its higher cost.","Editorial inference: the same machinery could be turned into a real-time shot-planning tool that updates the success-rate map from measured confinement and L–H behavior after each SPARC discharge, which the authors name as future work.","Editorial inference: substituting physics-based transport surrogates for the empirical scalings inside the Monte Carlo loop would test whether the dominant-uncertainty ranking survives with more faithful profile and confinement correlations."],"forward_implications":["SPARC operators should aim for the broad high-success-rate plateau rather than the deterministic optimum, avoiding operation close to the H-mode transition boundary.","Reducing uncertainty in the confinement factor $H$ substantially raises the predicted success rate, whereas tightening $k_{LH}$ alone lowers it because it removes favorable low-threshold samples.","The full 25 MW of launched auxiliary power and high divertor argon enrichment are both worth more than small confinement improvements; reducing plasma current to avoid disruptions degrades success rapidly.","Total success rate exceeds 50% under nominal assumptions and is improved by attempting multiple shots, since different Monte Carlo samples succeed at different operating points.","The updated ITPA confinement scaling produces a statistical POPCON very similar to the ITER98y2-based one, so the main conclusions are not an artifact of that scaling choice."],"supporting_citations":[{"why":"defines SPARC device parameters, the Q>2 goal, and the 140 MW fusion-power limit that frame all success conditions","marker":"[1]"},{"why":"supplies the ITER98(y2) H-mode confinement scaling whose 15% error motivates the H uncertainty and drives fusion gain sensitivity","marker":"[8]"},{"why":"supplies the Martin L-H power threshold scaling used to decide H-mode versus L-mode confinement at each sample","marker":"[11]"},{"why":"is the open-source POPCON code used to solve the power-balance workflow at every operating point","marker":"[16]"},{"why":"defines the SPARC Primary Reference Discharge, providing nominal means and the gray-star comparison location","marker":"[17]"},{"why":"motivates the Ti/Te, peaking, and a/LT uncertainty ranges from physics-based gyrokinetic predictions","marker":"[15]"},{"why":"provides the density-peaking scaling on which the peaking offsets are based","marker":"[23]"},{"why":"supplies the experimental range of edge argon enrichment values used in the core-edge compatibility scan","marker":"[19]"},{"why":"provides the Bayesian optimization machinery (surrogate models and acquisition functions) used in the multi-fidelity workflow","marker":"[30]"},{"why":"supplies the logarithmic expected-improvement acquisition function used to select evaluation points","marker":"[33]"}],"fun_headline_variants":["Uncertainty shifts SPARC's recommended operating point","Error bars redraw SPARC's optimal operating window","Model uncertainty changes SPARC scenario predictions","Statistical POPCONs move SPARC's sweet spot","Uncertainty favors broad SPARC operating plateau"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the ten input distributions—their widths, shapes, and especially their independence—faithfully represent how SPARC's real confinement, threshold, profile, and impurity behavior will vary.","fun_headline_variants_meta":{"raw":{"variants":["Uncertainty shifts SPARC's recommended operating point","Error bars redraw SPARC's optimal operating window","Model uncertainty changes SPARC scenario predictions","Statistical POPCONs move SPARC's sweet spot","Uncertainty favors broad SPARC operating plateau"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000466,"raw_usage":{"total_tokens":2298,"prompt_tokens":892,"completion_tokens":1406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":1334}},"tokens_in":508,"tokens_out":1406,"duration_ms":10327,"temperature":1.0,"reasoning_tokens":1334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:38:03.233285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the statistical POPCON using correlated Monte Carlo draws (joint covariance among $H$, $k_{LH}$, and $T_i/T_e$ from physics-based transport ensembles) or using the first measured SPARC shot-to-shot distributions; if the location of the maximum pointwise success rate shifts by more than the grid spacing, or if predicted high-success points fail systematically, the independent gamma and Gaussian assumptions are the deciding factor.","supporting_citations":[{"cited_title":"Chapter 2: Plasma confinement and transport","cited_arxiv_id":null,"evidence_quote":"supplies the ITER98(y2) H-mode confinement scaling whose 15% error motivates the H uncertainty and drives fusion gain sensitivity"},{"cited_title":"cfs-energy/cfspopcon: v7.0.1","cited_arxiv_id":null,"evidence_quote":"is the open-source POPCON code used to solve the power-balance workflow at every operating point"},{"cited_title":"The SPARC Primary Reference Discharge defined by cfsPOPCON","cited_arxiv_id":"2311.05016","evidence_quote":"defines the SPARC Primary Reference Discharge, providing nominal means and the gray-star comparison location"},{"cited_title":"Nonlinear gyrokinetic predictions of SPARC burning plasma profiles enabled by surrogate modeling","cited_arxiv_id":null,"evidence_quote":"motivates the Ti/Te, peaking, and a/LT uncertainty ranges from physics-based gyrokinetic predictions"},{"cited_title":"Scaling of density peaking in H-mode plasmas based on a combined database of AUG and JET observations","cited_arxiv_id":null,"evidence_quote":"provides the density-peaking scaling on which the peaking offsets are based"},{"cited_title":"Unexpected improvements to expected improvement for bayesian opti- mization","cited_arxiv_id":null,"evidence_quote":"supplies the logarithmic expected-improvement acquisition function used to select evaluation points"}],"review_version":1}