{"id":"53244f76-0e85-4b20-8515-f8203be62af0","arxiv_id":"2608.11058","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Trained on nonlinear FAR3d simulations, Gaussian process and neural-network surrogates predict ITER energetic-particle transport fluxes with test R2 > 0.97 and a five to six order-of-magnitude speedup.","lead":"This paper builds two machine-learning models that predict how fast particles are lost from the ITER fusion reactor, based on data from expensive physics simulations. If the models hold up, fusion engineers can evaluate particle transport in seconds instead of hundreds of hours, speeding up reactor design.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-scenario generalization is untested: the four held-out profiles come from the same two FAR3d runs used for training, so the high test R2 may reflect interpolation along the training trajectories rather than predictive skill for new ITER plasma states.","rationale":"The reader's weakest assumption focused on the local-state uniqueness of Eq. 3, supported by the ad hoc 0.3 threshold and the 10-17% of high-variability states. That is a real concern and I agree with it. However, the more decisive issue is that the test set does not probe generalization beyond the two training simulations: all held-out profiles come from the same FAR3d runs as the training profiles, so temporal autocorrelation and dump-level global features can inflate apparent accuracy. Even a perfectly unique mapping within the sampled manifold would not establish the ability to replace transport evaluations in integrated modeling, which requires predictions for states not encountered during training. In this sense my concern is complementary to the reader's: the flux-variability analysis limits the validity of the mapping within the dataset, while the test-set construction limits the evidence that the mapping is useful outside it. The paper's own limitations section acknowledges that extension to multiple operating scenarios requires additional data, which is a sign of honest reporting. Because the authors claim only a proof of concept and hedge the integrated-modeling statement, the conditional verdict remains appropriate. No code or data release is provided, so independent reproduction is impossible; this does not by itself invalidate the results but supports keeping the verdict conditional rather than accept. If the leave-one-scenario-out test were to succeed, that would materially strengthen the paper; if it fails, the claim of integrated-modeling readiness would need to be reduced to within-scenario interpolation.","tokens_in":12887,"tokens_out":3418,"duration_ms":34903,"concrete_test":"Run leave-one-scenario-out cross-validation: retrain both the GP and hierarchical NN using only the reversed-shear data (11 profiles) and evaluate on all monotonic-shear profiles, then repeat with the roles reversed. Report R2, relative L2 error, and the fraction of held-out states with local-to-global flux variability above 0.3. If cross-scenario R2 falls materially below the in-run Table 3 values of ~0.97-0.98, the reported accuracy is interpolation within the training runs and the integrated-modeling-ready claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the mapping F in Eq. 3 is a well-defined local plasma-state-to-flux relation that transfers to new states encountered in integrated modeling. Two facts weaken this. First, the flux-variability analysis in Section 3 itself admits that 10% of beam and 17% of alpha states have local-to-global flux variability above 0.3, and these states are concentrated near the peak transport region (rho ~ 0.35) and at intermediate-to-high gradient strengths, where profile reconstruction errors are also largest (Section 5.2). The mapping is therefore not uniquely determined on a non-negligible, physically important subset of the sampled manifold. Second, and more load-bearing, the test set consists of four profiles drawn from the same two simulations as the training data. Output dumps are separated by roughly 14,000 timesteps within the same nonlinear runs, so test profiles lie on the same temporal trajectories as training dumps. For the GP, the profile-averaged densities are additional dump-level global inputs, further tying predictions to the specific simulation run. The reported test R2 values of 0.979/0.974 (GP) and 0.982/0.975 (NN) therefore demonstrate interpolation over time and radius within two known scenarios, not generalization to a new equilibrium, a different perturbation, or a substantially different plasma state. The paper is appropriately cautious about being a proof of concept for one steady-state scenario, but the abstract's claim that the surrogates are sufficiently accurate to be incorporated into integrated modeling workflows is not actually supported by the current test design.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops Gaussian process and hierarchical neural-network surrogate models that predict beam-ion and alpha-particle transport fluxes from nonlinear FAR3d gyrofluid simulations of an ITER steady-state scenario. The models use a seven-dimensional local plasma-state representation (radius, safety factor, magnetic shear, EP densities, and EP density gradients) to map to the two transport fluxes. A flux-variability analysis is used to argue that this local representation is sufficiently unique over most of the sampled feature space. The surrogates are trained on a dataset of about 9180 spatial samples derived from 20 radial profiles from two FAR3d simulations, with 16 profiles for training and 4 for testing. The reported test-set R² values are 0.9791/0.9737 (GP) and 0.9823/0.9752 (NN) for beam/alpha fluxes, and the evaluation time is reduced by five to six orders of magnitude relative to a full nonlinear simulation. The paper frames the result as a proof of concept for replacing repeated transport evaluations in integrated modeling workflows.","tokens_in":13205,"tokens_out":2930,"duration_ms":27693,"significance":"If the central claims hold, this would be a useful and novel proof of concept: the first ML surrogate for AE-driven energetic-particle transport, with explicit uncertainty quantification and a substantial speedup. The comparative study of GP and hierarchical NN uncertainty behavior is also informative. However, the significance is tempered by the narrow data basis: the surrogates are trained and tested on two simulations of one ITER steady-state scenario, and the test set consists of only four profiles drawn from those same two runs. The claimed predictive accuracy is therefore an interpolation result within the training manifold rather than demonstrated generalization to new equilibria, perturbations, or operating scenarios. The paper is appropriately cautious in its conclusions, but the abstract's statement that the surrogate is 'sufficiently accurate and computationally efficient to be incorporated into future integrated modeling workflows' goes beyond what the current evidence strictly supports.","major_comments":[{"comment":"The test set consists of four profiles taken from the same two FAR3d runs used for training, with output dumps separated by about 14,000 timesteps within the nonlinear saturated phase. Consequently, the test profiles lie on the same temporal trajectories as the training profiles, and the reported test R² values of 0.979–0.982 demonstrate interpolation along known simulation trajectories rather than predictive skill for a new plasma state. For the GP, the additional profile-averaged densities ⟨n_beam⟩ and ⟨n_alpha⟩ are dump-level global inputs, which further tie test predictions to the specific source run. The paper should either add a hold-out scenario evaluation (e.g., train on one q-profile and test on the other, or leave out a full simulation) or explicitly restrict the central claim to interpolation within the sampled manifold. As written, the evidence does not support the abstract's implication that the surrogates are ready for deployment in integrated modeling of unseen ITER states.","section":"§2 and §5.1 (Table 3)"},{"comment":"The flux-variability analysis is load-bearing for the uniqueness of the mapping in Eq. (3), but it leaves a non-negligible fraction of the data unexplained: approximately 10% of beam states and 17% of alpha states have local-to-global flux variability above the 0.3 threshold, concentrated near the peak transport region (ρ ≈ 0.35) and at intermediate-to-high gradient strengths. These are precisely the states that contribute most to transport, and the neighborhood radius of 0.3 is selected based on mean nearest-neighbor distances rather than derived from an objective criterion. The paper should quantify how the surrogate prediction errors on these high-variability states compare with errors on the rest of the dataset, and should discuss whether additional features (e.g., mode amplitudes or history-dependent variables) would be needed to make the mapping unique on this subset. Without this, the statement that the representation 'provides a sufficient unique parameterization over most of the sampled feature space' is too strong.","section":"§3 (Fig. 3 and surrounding text)"},{"comment":"The claim that 'neither model exhibits significant overfitting' is based on the similarity between training and testing L2 errors (Table 3). Because the test profiles come from the same two simulations as the training profiles, this comparison does not actually rule out overfitting to the simulation-specific trajectory. A more convincing check would be leave-one-simulation-out cross-validation or an evaluation on a simulation with a different initial condition or perturbation. The discussion of outlier profiles in §5.3 is useful, but it does not substitute for an independent test set.","section":"§5.2 and §5.3"},{"comment":"The hierarchical NN architecture is not fully specified. The high-ρ network uses 'the mean beam and alpha-particle transport fluxes predicted by the low-ρ network' as additional inputs, but it is unclear over which radial range or set of low-ρ points this mean is taken, and how the low-ρ network is applied when the high-ρ network is evaluated at inference time. This matters because the high-ρ network's inputs depend on the low-ρ network's outputs, creating a sequential dependency not described in enough detail for reproducibility. The manuscript should provide an explicit computational recipe for the hierarchical evaluation.","section":"§4.1"}],"minor_comments":[{"comment":"The text alternates between 'normal-shear' and 'monotonic q-profile' when referring to the second simulation. Please use a single consistent term throughout.","section":"§2"},{"comment":"The notation Γpred(ρ) and Γtrue(ρ) is introduced in Eq. (4) but the integration variable and the normalization are not fully defined; specify that ρ is the square-root of normalized toroidal flux and that the integral is over the radial domain.","section":"§5.2, Eq. (4)"},{"comment":"The 'distance correlation coefficient' is used to characterize the flux-gradient relationship, but no definition or reference is given. Please add a brief definition or citation.","section":"§5.3"},{"comment":"The figure caption says the 'average number of similar cross-profile states' is shown, but the vertical axis label and the text would be clearer if the units (count per radial bin) were made explicit.","section":"Figure 3(b)"},{"comment":"The sentence 'For the normal-shear simulations, we retained 9 output dumps' appears to refer to the monotonic q-profile case; please reconcile the terminology as noted above.","section":"§2, 'Training vs testing data split'"}],"recommendation":"major_revision","confidential_remarks":"The central proof-of-concept result is plausible and the paper is generally well written, but the evaluation design is the main weakness: all testing is performed on profiles from the same two simulations used for training, and the GP additionally uses dump-level average densities that further tie test predictions to the source runs. The flux-variability analysis also acknowledges a 10–17% subset of high-variability states in the most transport-relevant region. These issues are fixable within the manuscript's scope by reframing the claims as interpolation results and adding at least one stronger out-of-sample check (e.g., leave-one-simulation-out) or an explicit discussion of why such a check is not yet possible. I recommend major revision rather than rejection because the data generation and model comparison are otherwise sound and the paper is transparent about many limitations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: it's the first ML surrogate for AE-driven energetic-particle transport in an ITER steady-state scenario, and the basic engineering result holds—two standard methods (multitask GP, hierarchical NN) reproduce FAR3d transport fluxes on a small held-out set with R2 around 0.98 and cut evaluation cost by five to six orders of magnitude. That is a real, useful increment for integrated modeling, not a breakthrough, but not a dud either.\n\nWhat's genuinely good: the flux-variability analysis is a serious attempt to check whether the local plasma-state parameterization is well-posed, and the authors openly flag that 10–17% of states, concentrated near the transport peak, have local-to-global flux variability above 0.3. They also acknowledge the surrogate is limited to the saturated soft-MHD regime and one ITER scenario. The comparison between GP and NN uncertainty behavior is informative and not overclaimed. The citation pattern looks fine; they build on the FAR3d reference scenario and compare against the relevant reduced models (TGLF-EP, kick model, RBQ, Carlevaro).\n\nThe soft spots are the ones your stress-test flagged, and they are real. The test set is four profiles drawn from the same two FAR3d runs used for training. Output dumps are separated by roughly 14,000 timesteps within the same nonlinear trajectories, so the held-out profiles are interpolations along those trajectories, not new equilibria or perturbations. The abstract's claim that the surrogates are 'sufficiently accurate ... to be incorporated into integrated modeling workflows' is stronger than the evidence supports. Also, the GP uses profile-averaged densities as additional global inputs, which further ties predictions to the specific simulation run. The flux-variability threshold of 0.3 is ad hoc, and the 10–17% high-variability states sit exactly in the peak transport region. None of this kills the proof-of-concept, but it means the reported R2 values are not evidence of cross-scenario generalization.\n\nWho benefits: fusion modelers who want a fast transport closure for ITER scenario scans, and anyone building ML surrogates for nonlinear plasma simulations. It deserves a serious referee: the methods are reproducible in principle, the claims are mostly measured, and the limitations are stated. The referee should ask for a test on at least one simulation outside the training runs, ideally a different q-profile or perturbation, and a clearer statement that the current test is interpolation, not extrapolation.\n\nMy vote: send it to review, with the expectation that the authors tighten the generalization claims and add a genuinely held-out test if feasible. I'd cite it if I were working on EP transport surrogates.","headline":"First ML surrogate for AE-driven EP transport in ITER; solid proof-of-concept with a weak cross-scenario test.","tokens_in":13784,"tokens_out":1495,"would_cite":true,"duration_ms":12703,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trained on nonlinear FAR3d gyrofluid simulations of an ITER steady-state scenario, Gaussian-process and hierarchical neural-network surrogates predict energetic beam and alpha-particle transport fluxes with test $R^2$ around 0.98 while…","keywords":["energetic-particle transport","Alfvén eigenmodes","machine-learning surrogate","Gaussian process","neural network","ITER","FAR3d","integrated modeling"],"falsifier":"Run the trained GP and NN surrogates on new nonlinear FAR3d simulations of an ITER regime not represented in training data, such as a strongly bursting or large-scale relaxation case, and check whether test $R^2$ collapses or median relative $L^2$ error grows well beyond the 7–9% range; alternatively, repeat the flux-variability analysis with time-lagged features (e.g., the flux at the previous dump) and see whether the 10–17% of high-variability states near $\\rho \\sim 0.35$ become uniquely determined.","tokens_in":12667,"feed_emoji":"⚛️","tokens_out":8704,"duration_ms":70656,"temperature":0.7,"pith_summary":"This paper claims that machine-learning surrogates can replace expensive nonlinear simulations of energetic-particle transport in ITER design workflows. The authors train a multitask Gaussian process and a hierarchical neural network on nonlinear FAR3d simulations of Alfvén-eigenmode-driven transport of beam and alpha particles, using a seven-variable local plasma-state representation. Both models reproduce the simulated fluxes on unseen test profiles with coefficients of determination near 0.98 and evaluate all retained transport profiles in seconds instead of 158 wall-clock hours, a speedup of five to six orders of magnitude. If these findings hold, iterative reactor design and scenario optimization could afford repeated transport evaluations that currently require thousands of GPU-hours.","feed_headline":"Surrogates predict ITER particle transport 100,000x faster","feed_subtitle":"GP and neural-network models match simulated fluxes with test R² near 0.98, in seconds instead of 158 hours.","key_machinery":"The load-bearing object is the surrogate mapping $(\\Gamma_{\\rm beam}, \\Gamma_\\alpha) = F(\\rho, q, \\hat{s}, n_{\\rm beam}, n_\\alpha, \\nabla n_{\\rm beam}, \\nabla n_\\alpha)$, which states that an instantaneous local plasma state uniquely determines the transport flux. Two complementary learners realize this mapping: a multitask Gaussian process with a Matérn 3/2 ARD kernel that also receives profile-averaged beam and $\\alpha$ densities as global context, and a hierarchical neural network split at $\\rho = 0.5$, where a low-$\\rho$ network's predicted fluxes are fed as extra inputs to a high-$\\rho$ network. Uncertainty is carried by the GP posterior variance and by Monte-Carlo dropout in the NN. A flux-variability analysis, which compares fluxes among neighboring states within a Euclidean distance of 0.3 in the normalized feature space, tests whether the mapping is unique enough for the surrogate formulation to be valid.","core_discovery":"The central discovery is that the saturated-phase transport flux of energetic particles can, to a good approximation, be treated as a function of the local plasma state $(\\rho, q, \\hat{s}, n_{\\rm beam}, n_\\alpha, \\nabla n_{\\rm beam}, \\nabla n_\\alpha)$, and that both a multitask Gaussian process and a hierarchical neural network trained on these features from nonlinear FAR3d simulations predict beam and $\\alpha$ fluxes on test profiles with $R^2$ values of 0.979 and 0.974 (GP) and 0.982 and 0.975 (NN), respectively. The models reconstruct radial flux profiles with median relative $L^2$ errors of approximately 7–9% on unseen data and reduce the cost of evaluating a full set of transport profiles from 158 wall-clock hours to a few seconds. The authors demonstrate this for an ITER steady-state scenario covering both reversed-shear and monotonic-$q$ profiles, restricting training to the statistically stationary nonlinear saturated phase of the simulations where appreciable transport occurs.","pith_inferences":["The authors do not test whether the 10–17% of high-variability states near $\\rho \\sim 0.35$ become uniquely determined when time-lagged features are added; a direct test would show whether the current seven-variable representation is missing memory effects.","Because the flux-variability analysis flags the strongest transport and gradient regimes as least unique, integrated modeling should treat surrogate outputs there with wider uncertainty margins or resample with additional features.","Disagreement between GP posterior variance and MC-dropout uncertainty could serve as an out-of-distribution detector: when the two uncertainty estimates diverge sharply, a plasma state is likely beyond the training manifold.","The five-to-six-orders-of-magnitude speedup applies only to saturated-phase transport evaluation; using the surrogate to follow linear growth or bursting relaxation would place it outside its training regime."],"forward_implications":["Integrated modeling workflows could evaluate AE-driven energetic-particle transport repeatedly in seconds rather than hundreds of wall-clock hours, enabling design scans and uncertainty quantification for ITER and future burning-plasma devices.","The GP surrogate offers consistent global uncertainty estimates and slightly better reconstruction near the transport peak, while the NN surrogate's MC-dropout uncertainty distinguishes early- from late-saturation regimes, so the two models can serve complementary roles.","Because the surrogates learn the transport response directly from nonlinear simulation data, the approach transfers to other simulation codes and other devices as a proof of concept for reduced transport modeling.","The predictive uncertainty estimates provide a confidence measure when the surrogates are queried in plasma regimes sparsely represented in the training data.","The demonstrated accuracy supports extending the methodology to time-dependent transport prediction, with the long-term goal of accelerating nonlinear simulations themselves."],"supporting_citations":[{"why":"Supplies the nonlinear FAR3d simulation dataset of the ITER steady-state scenarios (reversed-shear and monotonic-q) from which the surrogate training and test data are drawn.","marker":"[10]"},{"why":"Documents the FAR3d gyro-fluid code that generated the nonlinear energetic-particle transport simulations used as ground truth.","marker":"[3]"},{"why":"Validation of the FAR3d simulation approach against D-T fusion plasma observations, underpinning the simulated transport as a reliable target.","marker":"[7]"},{"why":"Provides the ITER-relevant critical-gradient reduced model that the surrogate approach is positioned against as an alternative transport closure.","marker":"[13]"},{"why":"Supplies the Gaussian process regression framework, including the Matérn kernel and posterior predictive variance used by the GP surrogate.","marker":"[17]"},{"why":"Supplies the Monte-Carlo dropout method used to estimate predictive uncertainty for the neural-network surrogate.","marker":"[18]"}],"fun_headline_variants":["AI surrogates predict ITER transport 100,000x faster","Seconds vs 158 hours: ML surrogates for ITER","Machine learning cuts ITER transport time to seconds","Surrogates hit R² 0.98, speed up 100,000x in ITER","ITER transport predictions: from 158 hours to seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that similar values of the seven plasma-state variables always produce similar transport fluxes, with no dependence on history or on unmeasured quantities; if that fails, the surrogate's accuracy will not generalize beyond its training set.","fun_headline_variants_meta":{"raw":{"variants":["AI surrogates predict ITER transport 100,000x faster","Seconds vs 158 hours: ML surrogates for ITER","Machine learning cuts ITER transport time to seconds","Surrogates hit R² 0.98, speed up 100,000x in ITER","ITER transport predictions: from 158 hours to seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1464,"prompt_tokens":989,"completion_tokens":475,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":605,"tokens_out":475,"duration_ms":4873,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:07:53.700294+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained GP and NN surrogates on new nonlinear FAR3d simulations of an ITER regime not represented in training data, such as a strongly bursting or large-scale relaxation case, and check whether test $R^2$ collapses or median relative $L^2$ error grows well beyond the 7–9% range; alternatively, repeat the flux-variability analysis with time-lagged features (e.g., the flux at the previous dump) and see whether the 10–17% of high-variability states near $\\rho \\sim 0.35$ become uniquely determined.","supporting_citations":[{"cited_title":"Spong, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the nonlinear FAR3d simulation dataset of the ITER steady-state scenarios (reversed-shear and monotonic-q) from which the surrogate training and test data are drawn."},{"cited_title":"Varela, D","cited_arxiv_id":null,"evidence_quote":"Documents the FAR3d gyro-fluid code that generated the nonlinear energetic-particle transport simulations used as ground truth."},{"cited_title":"Solano, ˇZiga ˇStancar, Jacobo Varela, Matteo Baruzzo, Emily Belli, Phillip J","cited_arxiv_id":null,"evidence_quote":"Validation of the FAR3d simulation approach against D-T fusion plasma observations, underpinning the simulated transport as a reliable target."},{"cited_title":"Bass and R.E","cited_arxiv_id":null,"evidence_quote":"Provides the ITER-relevant critical-gradient reduced model that the surrogate approach is positioned against as an alternative transport closure."},{"cited_title":"Dropout as a bayesian approximation: Rep- resenting model uncertainty in deep learning","cited_arxiv_id":null,"evidence_quote":"Supplies the Monte-Carlo dropout method used to estimate predictive uncertainty for the neural-network surrogate."}],"review_version":1}