{"id":"9feabbd7-5952-48e9-9c24-a2a977c448d2","arxiv_id":"2507.03776","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Training ML models on 12.4M VMEC stellarator equilibria shows that two boundary coefficients, RBC1,0 and ZBS1,0, have little effect on quasisymmetry and quasi-isodynamicity, and provides surrogate predictors for these omnigenity metrics.","lead":"A database of 12.4 million simulated stellarator shapes was used to train machine learning models that predict how well each shape confines fusion alpha particles. The models show which geometric parameters matter most for omnigenity, potentially speeding up stellarator design.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The importance ranking is distribution-dependent and computed mostly on non-omnigenous equilibria; it has not been shown to hold in the near-omnigenous regime where the design conclusions apply.","rationale":"I read the paper in good faith: it builds a large VMEC-based database, releases code and data, and finds a consistent feature-importance pattern across LightGBM, LightGBM LSS, SHAP, and mutual information. This is real empirical work, and the R^2 values are credible for a smooth, low-dimensional regression problem on log-transformed targets. My concern is not that the calculations are wrong, but that the interpretation overreaches. The reader's weakest assumption was the representativeness of the restricted 8-mode vacuum family; my concern is closely related but more precise: even within that family, the importance ranking is computed over a population dominated by configurations far from omnigenity, so the ranking describes the broad sampled distribution rather than the near-omnigenous regime that the paper's design conclusions target. This is not an internal inconsistency, and the paper does flag some limitations in Section II, but the conclusions do not carry the necessary qualification. The concrete subset check would settle whether the ranking survives in the relevant regime. Because the reader already issued a CONDITIONAL verdict and this concern reinforces that condition rather than overturning the paper, I recommend leaving the verdict unchanged.","tokens_in":20418,"tokens_out":8307,"duration_ms":109358,"concrete_test":"Restrict the released database to configurations with f_QS below 1 and f_QI below 0.01 (or, if counts are too low, the label-1 converged subset) and rerun the exact LightGBM native-importance, SHAP, and mutual-information pipelines on this subset. If RBC1,0 and ZBS1,0 remain the lowest-ranked features in the near-omnigenous regime, the headline result is robust within the original 8-mode family; if the ordering changes or the gap closes, the global importance ranking is an artifact of averaging over a mostly non-omnigenous population and the conclusions should be reworded as conditional on the full database distribution. This check uses only existing data and the released code.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the inference from 'RBC1,0 and ZBS1,0 have low importance in models trained on the full database' to 'these coefficients are least important for omnigenity.' The database is generated with a uniform prior over eight Fourier coefficients (Section II, Eq. 4), and the targets f_QS and f_QI are computed for all converged equilibria. However, the resulting population is dominated by configurations far from omnigenity: of 12,421,004 configurations only 457,488 have f_QS below 10, and in the classification test set only 91,500 of 2.49 million instances satisfy 'converged and f_QS below 10' (Table III). Tree-importance and SHAP values are distribution-dependent summaries: they describe which inputs move predictions over the sampled population, not which inputs control the approach to omnigenity in the low-residual region where optimization actually operates. Section I states the goal is to quantify the role of each surface degree of freedom in the loss functions used to find precisely omnigenous designs, and the conclusions propose the models as surrogates for optimization. A ranking computed mostly on poor configurations does not establish that ranking for near-omnigenous designs, and Section II itself notes that precise omnigenity would require additional Fourier modes not present in this family. The central claim therefore needs to be read as a property of the sampled 8-mode distribution, not as a general statement about which geometric parameters influence omnigenity; the paper does not currently test the near-omnigenous regime.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper builds a large database of 12,421,004 VMEC vacuum stellarator equilibria, all with nfp=2, mpol=ntor=1, and R0=1, where the eight nonzero boundary Fourier coefficients are sampled uniformly in [-0.1, 0.1]. For each equilibrium the authors compute the quasisymmetry proxy f_QS (Eq. 5), the quasi-isodynamicity proxy f_QI (Eq. 6), and several auxiliary physics metrics. They then perform correlation analysis, outlier detection, PCA, and a supervised autoencoder, and train LightGBM, LightGBM LSS, and feed-forward neural network regressors to predict log f_QS and log f_QI. Feature-importance analyses (native gain, SHAP, and mutual information) are used to rank the eight boundary coefficients. The central claim is that RBC1,0 and ZBS1,0 are consistently the least important coefficients for omnigenity, while the m=1, n=±1 coefficients dominate, and that the FFNN predicts log f_QS and log f_QI on held-out data with R²=0.9909.","tokens_in":20762,"tokens_out":3502,"duration_ms":42578,"significance":"If the feature-importance result is robust, the paper would provide a useful, openly available database and fast surrogate models for stellarator design-space exploration, plus a concrete data-driven statement about which low-order boundary modes matter for quasisymmetry and quasi-isodynamicity. The strengths are the open database and code, the use of VMEC-generated targets (so the surrogate training is not circular), and the consistency of the least-important-pair finding across several independent methods. However, the significance is capped by the deliberately narrow 8-parameter, vacuum, nfp=2 boundary family and by the fact that the importance rankings are computed over a population dominated by far-from-omnigenous configurations; the paper's own text acknowledges that precise omnigenity requires additional boundary modes. The claimed insight therefore currently holds only for the sampled distribution, not for the near-omnigenous regime where the proposed design conclusions and optimization surrogates would be used.","major_comments":[{"comment":"The load-bearing conclusion that RBC1,0 and ZBS1,0 are the least important coefficients for omnigenity is drawn from importance measures computed on the full database, which is dominated by poor configurations: only 457,488 of 12,421,004 configurations have f_QS below 10, and in the classification test set only 91,500 of 2.49 million instances satisfy the converged-and-f_QS<10 condition (Section II and Table III). Tree-gain importance, SHAP values, and mutual information are distribution-dependent summaries: they describe how inputs move predictions over the sampled uniform prior, not which inputs control the approach to omnigenity in the low-residual region where optimization actually operates. To support the stated goal of quantifying the role of each surface degree of freedom in the loss functions used to find precisely omnigenous designs, the authors should recompute the importance rankings on a near-omnigenous subset (for example f_QS<10 and a corresponding f_QI threshold) and, if possible, compare with feature sensitivities for precisely omnigenous equilibria from the literature, such as Landreman-Paul quasisymmetric and Goodman et al. quasi-isodynamic configurations.","section":"Section VII and Figs. 13-17"},{"comment":"The FFNN SHAP analysis is performed with Kernel SHAP using a single-sample background drawn from the training data. With only one background sample, the conditional expectations required for Shapley-value computation are very crudely approximated, especially for eight correlated inputs, so the quantitative claim that RBC1,0 and ZBS1,0 have 'little to no impact' is not reliably established by this particular analysis. The authors should either use a sufficiently large background set (typically hundreds of samples), or corroborate the FFNN importance with an additional attribution method such as integrated gradients or permutation importance.","section":"Section VI, Figs. 16-17"},{"comment":"The database restricts the boundary to mpol=ntor=1 with nfp=2, and the text notes that precise omnigenity requires more Fourier modes. Consequently, the conclusion that RBC1,0 and ZBS1,0 are unimportant, and even the surrogate models themselves, are statements about this 8-parameter family rather than about magnetic geometry in general. The abstract and conclusions should carry this qualification explicitly, and the paper should either demonstrate that the rankings persist in an extended family or discuss how the omitted higher-order modes could couple to the modes studied here.","section":"Section II, Eq. (4)"}],"minor_comments":[{"comment":"There are typos in the omnigenity terminology: 'quasi-isodynamiticity' in the abstract and 'quasi-isodinamicity' in the introduction should be 'quasi-isodynamicity'.","section":"Abstract and Section I"},{"comment":"The text states that the classification model 'predicted VMEC convergence and quasisymmetry below 10, predicting 99.7% of the configurations with those conditions correctly,' but Table III reports recall of only 0.85 for label 1; 99.7% appears to refer to overall accuracy. This should be stated unambiguously.","section":"Section V, Table III"},{"comment":"The quantitative feature rankings differ between LightGBM, LightGBM LSS, and the FFNN SHAP results; only the least-important pair RBC1,0 and ZBS1,0 is consistent across methods. The text should explicitly acknowledge this disagreement and avoid implying that the models agree on the ordering of the dominant coefficients.","section":"Section V, Figs. 13-14"},{"comment":"The statement that quasisymmetry values 'can span up to 100 different orders of magnitude' should be clarified, since the caption of Table I lists quasisymmetry in [0,10] while the text discusses log-transformed values; the reader should know whether the reported range refers to the raw residual or its logarithm.","section":"Section III, Fig. 5"},{"comment":"The description of the FFNN architectures for the quasisymmetry and quasi-isodynamic models is identical, and both are reported to achieve R²=0.9909; the authors should confirm that these are indeed two distinct networks and report their train/validation/test splits consistently.","section":"Section VI"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid data-engineering and ML-application contribution, and the database/code availability is a real strength. My main concern is that the central interpretability claim is broader than what the experimental design can support: the importance rankings are averaged over a distribution that is almost entirely non-omnigenous, and the FFNN attribution uses a single background sample. These are fixable in revision by adding subset-restricted importance analyses and a more careful framing of the claims. I do not see grounds for rejection, because the underlying surrogate models and database are useful even if the geometric-insight claim needs to be narrowed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper builds a 12.4M-entry VMEC database of vacuum stellarator equilibria with eight boundary Fourier coefficients (mpol=ntor=1, nfp=2), computes f_QS and f_QI for every converged case, and then applies LightGBM, LightGBM LSS, and a feed-forward net to predict the residuals and rank feature importance. The new asset is the database itself, plus the consistent empirical finding that the two coefficients RBC1,0 and ZBS1,0 are the least important for both metrics across all the methods they tried, while the four helical coefficients RBC±1,1 and ZBS±1,1 dominate. That consistency across tree-based importance, SHAP, and mutual information is real evidence, not an artifact. The FFNN fits log f_QS and log f_QI on held-out data with R-squared near 0.99, which is respectable within this 8-parameter family.\n\nCredit where due: the paper is honest about scope. It states in Section II that precise omnigenity would require more modes than the eight used here, and the database is vacuum-only with two field periods. The authors frame the surrogates as screening tools for this family, not as generally applicable replacements for MHD solvers.\n\nThe soft spots are real but proportionate. The importance ranking is distribution-dependent: the database is dominated by poor configurations (only 457k of 12.4M have f_QS below 10), and SHAP/tree importance describe what moves predictions across that population, not which coefficients control the approach to omnigenity in the near-omnigenous regime where optimization actually operates. The paper does not validate the ranking on a known precisely-omnigenous equilibrium or on a dataset restricted to low f_QS. So the catchy conclusion that RBC1,0 and ZBS1,0 are unimportant should be read as 'unimportant in this sampled family, over this prior.' I also note the quantitative rankings disagree across models (LightGBM vs LSS vs FFNN produce different top-four orderings), though the least-important pair is stable. The single-sample background for SHAP is a rough approximation, and the paper flags it.\n\nNone of this is fatal. The consistent least-important finding is a useful empirical clue for stellarator designers, and the database is a resource that others can mine. The paper deserves peer review; a referee should ask for (1) a demonstration on a known near-omnigenous case, and (2) an analysis of how the ranking changes when training is restricted to low f_QS. I'd accept it with those requests.","headline":"Big new VMEC database and a robust empirical least-important finding, but the importance ranking is distribution-dependent and unverified near omnigenity.","tokens_in":21311,"tokens_out":2845,"would_cite":true,"duration_ms":30346,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["52.55.Hc"],"model":"deepseek-v4-flash","headline":"Across a 12-million-configuration stellarator database, four boundary Fourier coefficients dominate omnigenity while two are nearly irrelevant.","keywords":["stellarator","omnigenity","quasisymmetry","quasi-isodynamicity","boundary Fourier coefficients","machine learning surrogates","feature importance","ideal MHD equilibria"],"falsifier":"Run the same sampling procedure with $m_{pol}=n_{tor}=2$ (or with finite plasma pressure) and recompute the feature rankings; if $RBC_{1,0}$ or $ZBS_{1,0}$ rises out of the bottom two positions, the central ranking fails to transfer. A cheaper check is to take a published precisely omnigenous equilibrium, perturb $RBC_{1,0}$ and $ZBS_{1,0}$ by an amplitude comparable to perturbations of the claimed dominant coefficients, and measure whether $f_{QS}$, $f_{QI}$, or the radial drift degrades as much.","tokens_in":20166,"feed_emoji":"🧲","tokens_out":12821,"duration_ms":128144,"temperature":0.7,"pith_summary":"The paper asks which geometric features of a stellarator's plasma boundary actually control whether the device will confine fusion-born $\\alpha$ particles. To answer it, the authors generate a database of about 12.4 million vacuum stellarator equilibria, each described by eight boundary Fourier coefficients, and train several machine-learning models to predict two standard omnigenity metrics: quasisymmetry and quasi-isodynamicity. Their central finding is a ranking of the eight coefficients: the $m=1,n=0$ terms $RBC_{1,0}$ and $ZBS_{1,0}$ are consistently the least important for both metrics, while the $m=1,n=\\pm 1$ terms $RBC_{\\pm 1,1}$ and $ZBS_{\\pm 1,1}$ dominate, according to gradient-boosting feature importance, attribution values, and mutual information. The best surrogate, a feed-forward neural network, predicts the logarithm of both omnigenity metrics on held-out data with $R^2=0.9909$. If the ranking holds in realistic designs, it tells stellarator optimizers which boundary degrees of freedom deserve the most attention and gives them cheap surrogate models to speed up the search.","feed_headline":"Four shape terms control stellarator fusion-particle confinement","feed_subtitle":"A 12-million-configuration study ranks the boundary coefficients that decide quasisymmetry and quasi-isodynamicity.","key_machinery":"The load-bearing object is the Fourier representation of the plasma boundary, $R(\\theta,\\phi)=\\sum RBC_{m,n}\\cos(m\\theta-n_{fp}n\\phi)$ and $Z(\\theta,\\phi)=\\sum ZBS_{m,n}\\sin(m\\theta-n_{fp}n\\phi)$, truncated at $m_{pol}=n_{tor}=1$, which reduces the search space to eight coefficients with two field periods and major radius $R_0=1$. Each sampled boundary is fed to an ideal-MHD equilibrium solver that produces the magnetic field, from which the omnigenity proxies $f_{QS}$ and $f_{QI}$ and auxiliary quantities (rotational transform, mirror ratio, elongation, magnetic well, shear, inverse aspect ratio) are computed. The argument that some coefficients matter more than others is carried by three independent feature-attribution tools—gradient-boosting split counts, game-theoretic attribution values, and mutual information—plus a supervised autoencoder whose two-dimensional latent space displays a gradient of both metrics along one axis.","core_discovery":"The discovery the paper is trying to establish is that, within the eight-parameter family of two-field-period vacuum stellarators studied, omnigenity is governed by cross-section shaping coefficients with $|n|=1$ rather than by the two $n=0$ size-like coefficients. Concretely, the boundary Fourier coefficients $RBC_{1,1}$, $RBC_{-1,1}$, $ZBS_{1,1}$, and $ZBS_{-1,1}$ carry most of the predictive signal for both quasisymmetry and quasi-isodynamicity, while $RBC_{1,0}$ and $ZBS_{1,0}$ sit at the bottom of every importance ranking the paper computes. The paper further claims that these rankings are robust across method families—split-count importance in gradient-boosted trees, game-theoretic attribution, and mutual information—and that a feed-forward network trained on the same eight inputs reproduces both targets with $R^2=0.9909$. It concludes that the resulting models can act as fast surrogates for the ideal-MHD solver and that the supervised autoencoder's two-dimensional latent space offers a reduced design space aligned with omnigenity.","pith_inferences":["The least-importance of $RBC_{1,0}$ and $ZBS_{1,0}$ may partly be a scale artifact: in a high-aspect-ratio, vacuum, fixed-$R_0$ family these coefficients mostly set the overall cross-section size, which cancels in the dimensionless omnigenity proxies; at lower aspect ratio or finite plasma pressure, size-dependent equilibrium shifts could make them matter more.","A direct test of transferability would be to repeat the sampling with $m_{pol}=n_{tor}=2$ or with finite beta; if the same two coefficients stay at the bottom, the ranking is a property of omnigenity itself rather than of the restricted boundary representation.","Because the surrogate models are trained on the paper's random sample, they are most reliable where the sample is dense; using them as optimizers would require an active-learning loop that adds new ideal-MHD evaluations in underrepresented low-residual regions, where the reported errors were largest."],"forward_implications":["Stellarator optimizers in the same geometry family can concentrate search effort on the four $|n|=1$ coefficients, potentially reducing the effective dimension of the design problem from eight to four.","The trained feed-forward network and gradient-boosting regressors provide predictions of $\\log f_{QS}$ and $\\log f_{QI}$ accurate to $R^2 \\approx 0.99$ in this family, so they can screen candidate configurations without running the expensive ideal-MHD solver.","The classification model, with about 99% accuracy on the test set, can pre-filter boundaries that will either fail to converge or fail to reach quasisymmetry below a residual of 10.","Quasisymmetry and mirror ratio are strongly correlated in this database, so mirror ratio is a useful cheap proxy during early design.","The supervised autoencoder reduces the eight coefficients to two latent coordinates in which omnigenity varies monotonically along one axis, suggesting optimization could be performed in that two-dimensional space."],"supporting_citations":[{"why":"Defines the quasisymmetry proxy $f_{QS}$ and provides the precise-quasisymmetry targets against which the metric is calibrated.","marker":"[5]"},{"why":"Supplies the ideal-MHD equilibrium solver used to generate every database entry.","marker":"[9]"},{"why":"States the omnigenity conditions in Boozer coordinates that justify the $f_{QS}$ and $f_{QI}$ proxies.","marker":"[14]"},{"why":"Provides the gradient-boosting algorithm whose split-count feature importance yields one of the three rankings.","marker":"[19]"},{"why":"Provides the probabilistic gradient-boosting variant used for distributional prediction and for a second importance ranking.","marker":"[20]"},{"why":"Supplies the attribution-value method used for feature importance in the neural-network and probabilistic-model analyses.","marker":"[24]"},{"why":"Defines the quasi-isodynamic proxy $f_{QI}$ used as the second target.","marker":"[32]"},{"why":"Provides the mutual-information estimator that corroborates the importance ranking independently of the tree and attribution methods.","marker":"[52]"}],"fun_headline_variants":["Four boundary terms decide stellarator particle confinement","Which stellarator shapes trap fusion particles? Four coefficients","Data-mining 12M stellarators: four shape modes rule omnigenity","Four Fourier coefficients dominate stellarator confinement design","Stellarator geometry ranked: |n|=1 shaping beats size terms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The database only contains vacuum equilibria with two field periods, major radius one meter, and eight boundary Fourier coefficients drawn uniformly from $[-0.1,0.1]$, and the paper assumes this restricted family is representative enough that the feature-importance ranking carries over to realistic stellarators with many more modes and finite pressure.","fun_headline_variants_meta":{"raw":{"variants":["Four boundary terms decide stellarator particle confinement","Which stellarator shapes trap fusion particles? Four coefficients","Data-mining 12M stellarators: four shape modes rule omnigenity","Four Fourier coefficients dominate stellarator confinement design","Stellarator geometry ranked: |n|=1 shaping beats size terms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1290,"prompt_tokens":992,"completion_tokens":298,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":216}},"tokens_in":608,"tokens_out":298,"duration_ms":3567,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:03:48.055393+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same sampling procedure with $m_{pol}=n_{tor}=2$ (or with finite plasma pressure) and recompute the feature rankings; if $RBC_{1,0}$ or $ZBS_{1,0}$ rises out of the bottom two positions, the central ranking fails to transfer. A cheaper check is to take a published precisely omnigenous equilibrium, perturb $RBC_{1,0}$ and $ZBS_{1,0}$ by an amplitude comparable to perturbations of the claimed dominant coefficients, and measure whether $f_{QS}$, $f_{QI}$, or the radial drift degrades as much.","supporting_citations":[{"cited_title":"Helander","cited_arxiv_id":null,"evidence_quote":"Defines the quasisymmetry proxy $f_{QS}$ and provides the precise-quasisymmetry targets against which the metric is calibrated."},{"cited_title":"Jorge, A","cited_arxiv_id":null,"evidence_quote":"Supplies the ideal-MHD equilibrium solver used to generate every database entry."},{"cited_title":"Jorge, G","cited_arxiv_id":null,"evidence_quote":"States the omnigenity conditions in Boozer coordinates that justify the $f_{QS}$ and $f_{QI}$ proxies."},{"cited_title":"Landreman","cited_arxiv_id":null,"evidence_quote":"Provides the gradient-boosting algorithm whose split-count feature importance yields one of the three rankings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the attribution-value method used for feature importance in the neural-network and probabilistic-model analyses."},{"cited_title":"Brizard and T","cited_arxiv_id":null,"evidence_quote":"Defines the quasi-isodynamic proxy $f_{QI}$ used as the second target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the mutual-information estimator that corroborates the importance ranking independently of the tree and attribution methods."}],"review_version":1}