{"id":"b325d727-44a1-4d8f-a9df-84df78a9561d","arxiv_id":"2608.01344","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An LLM-controlled outer loop around the DESC solver improved finite-beta stellarator designs, boosting gate-valid configurations from 5 of 23 to 19 of 23, with median quasisymmetry error cut roughly in half.","lead":"This paper tests whether an AI language model can act as the manager of a stellarator design optimization loop, choosing which geometry parameters to tweak next in a sequence of physics simulations. On a set of 23 starting designs, the system raised the number meeting quality gates from 5 to 19 and reduced key error metrics, while logging every step as reusable data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Route-selection bias is the load-bearing flaw: Sec. 4.1 never defines how the 23 'completed' routes were selected, so the 5→19 gate-valid gain and paired medians may be artifacts of outcome-dependent completion.","rationale":"The paper is a proof of concept, not a comparative policy-efficiency study, and it explicitly scopes out baselines in Sec. 6.3. The most load-bearing assumption is therefore not the lack of a greedy/random baseline but the representativeness of the 23 selected routes. The reader's weakest_assumption identifies exactly this issue: route completion may be correlated with outcome. I agree. If the concern lands, the specific numerical claims in the abstract and Sec. 5.1 lose their force, because the paired medians would be conditional on the selection rule. If the concern does not land, the paper's evidence for the central claim is reasonably strong: deterministic DESC execution owns the physics, metrics are recomputed on a common grid, a 50-epoch route shows nonmonotone repair behavior, and 734 structured transitions are retained. The conditionality of the original verdict is appropriate: the issue is addressable from existing campaign logs, and the paper should be accepted only if the full-cohort reanalysis confirms the reported improvements. I therefore keep the reader's CONDITIONAL verdict rather than escalating to rejection.","tokens_in":10201,"tokens_out":3921,"duration_ms":43099,"concrete_test":"From the archived campaign, reconstruct the full set of all routes launched from source configurations under the same eight-epoch budget and physics contract. Apply the paper's deterministic evaluation and paired scoring to every such route, without excluding any route based on outcome or completion status. Report the total number of routes, gate-valid counts at input and output, and paired medians for QS RMS and max curvature on this full cohort. Also report the exact rule used to decide that a route was 'completed.' If the full-cohort medians are within ~10% of the reported values and the gate-valid output count remains 19, the selection concern is resolved. If the full-cohort improvements shrink or disappear, the headline numbers are artifacts of outcome-dependent route selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim rests on a paired comparison between 23 source configurations and 23 route outputs. The paper reports that gate-valid configurations rise from 5 to 19, median QS RMS falls from 2.39e-4 to 1.07e-4, and median max curvature falls from 62.56 to 33.00 m^-1. These numbers are only meaningful if the 23 routes are representative of the campaign under a common budget. Section 4.1 says only: \"We select 23 completed routes with a common eight-epoch budget for paired evaluation.\" It does not state how completion was determined, how many routes were launched, or whether the selection was made before inspecting outcomes. In an ongoing campaign, route completion can easily be correlated with outcome: routes that show early promise or reach gate-valid states may be allowed to continue to eight epochs, while stagnant or failing routes are closed early and excluded. The four incomplete routes are handled by using their closest-to-gate endpoints, but the omitted routes are never described. If the 23 are a post-hoc selection of successful or promising trajectories, the reported improvements are conditional on the outcome and cannot support the claim that agentic outer-loop control sustains finite-beta multi-objective search at campaign scale. The same selection also affects the 543 transition records from the short routes, since those records come from the same selected subset.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a bounded language-model agent that acts as an outer-loop controller for fixed-boundary DESC stellarator optimization. At each epoch the agent selects a parent equilibrium, Fourier-mode schedule, objective weights, and local solver budget, while a deterministic DESC layer owns the physics contract, evaluation, and acceptance. The manuscript reports a multi-start campaign on 23 common-budget routes: gate-valid configurations increase from 5 inputs to 19 outputs, median Boozer QS RMS drops from 2.39e-4 to 1.07e-4, and median max curvature drops from 62.56 to 33.00 m^-1. A separate 50-epoch route achieves a 9.10x QS reduction while repairing magnetic-well and curvature violations. The system also records 734 structured parent-action-outcome transitions. The central claim is that agentic outer-loop control can sustain finite-beta, multi-objective search and produce reusable optimization data.","tokens_in":10570,"tokens_out":7131,"duration_ms":72349,"significance":"If the quantitative claims are reliable, the paper makes a useful proof-of-concept contribution: it cleanly separates agent decisions from deterministic physics execution, explicitly labels transition records as confounded behavior-policy observations, and frames the work as a stepping stone to learned policies rather than as a final comparison. The description of the route graph and persistent provenance schema is a strength, as is the honest acknowledgment that matched-budget static/greedy/random controllers are still needed. The paper also openly reports that four short routes fail the terminal gate and that the long route misses its ambitious QS target. However, the central paired comparison currently rests on an underspecified route-selection protocol and lacks uncertainty quantification, so the headline improvements should be treated as provisional.","major_comments":[{"comment":"The central quantitative evidence is the paired comparison: gate-valid configurations rising from 5 to 19 inputs and median QS/curvature improvements. But Sec. 4.1 only says 'We select 23 completed routes with a common eight-epoch budget for paired evaluation.' It does not state how many routes were launched, what counts as 'completed,' whether routes could be closed early, or whether selection was outcome-blind. If completion correlates with progress or success, the reported improvements are conditional on a favorable subset and do not estimate the agent's average effect. The paper must report the full campaign funnel, define completion a priori, and provide an analysis that includes all started routes or a clearly outcome-blind subsample.","section":"Sec. 4.1 / Table 1 / Sec. 5.1"},{"comment":"The evaluation protocol for the four non-gate-valid routes is underspecified: they contribute 'the endpoint with minimum configured distance to the full gate,' but the distance measure is never defined as an equation or norm. Without this definition, part of the paired evaluation is not reproducible. In addition, Table 1 reports paired medians with no measures of dispersion, confidence intervals, or paired significance tests. The controller uses GPT-5.6-sol with high reasoning effort, and the paper does not state whether sampling is deterministic or whether routes were repeated/seeded. Please define the gate-distance metric and report distributions, paired differences, and bootstrap or repeated-seed uncertainties.","section":"Sec. 4.2 / Sec. 5.1"},{"comment":"The abstract and conclusion state that 'agentic outer-loop control can sustain finite-beta, multi-objective search,' yet the experiments include no comparison to a non-agentic baseline. The authors correctly note in Sec. 6.3 that matched-budget static, greedy, random, and memory-ablated controllers are required, but the headline causal attribution goes beyond that scoped claim. The improvement could plausibly arise from the deterministic DESC optimizer plus a simple continuation rule rather than from the agent's sequential reasoning. Please either add a minimal control (e.g., fixed-schedule or random-schedule DESC refinement from the same sources) or soften the causal framing to a capability demonstration.","section":"Abstract / Sec. 6.3 / Conclusion"}],"minor_comments":[{"comment":"Typo: 'foragentic' should be 'for agentic' in the abstract.","section":"Abstract / Title"},{"comment":"The expert-score weights alpha_j in Eq. (10) are stated to be fixed, but their values are not given. Provide them or refer to a code/data supplement, since they affect parent ranking.","section":"Sec. 4.2 / Eq. (10)"},{"comment":"The long-route selection criterion is not stated; if this route was chosen because it is successful or illustrative, label it as such so readers do not infer typical behavior. The R2_phase and p-value are useful but should be accompanied by a brief explanation of what variance is being explained.","section":"Sec. 5.3 / Fig. 5"},{"comment":"No data or code availability statement is provided. For a paper whose contribution includes a reusable transition corpus and provenance schema, make the dataset or a representative subset available, or state why it cannot be shared.","section":"General / Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the core idea is worth publishing once the evaluation protocol is clarified. The main risk is not technical fraud but that the headline numbers (5→19 gate-valid, median QS and curvature improvements) may reflect outcome-dependent route selection rather than a representative campaign effect. The authors' explicit limitations section gives me confidence that this can be addressed with additional reporting and reanalysis rather than requiring a fundamentally new study. I would encourage the editor to request the campaign funnel and a robustness analysis before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the arXiv:2608.01344 stellarator agent paper. The useful news: this is a genuinely new system in that literature. They put a bounded LLM in an outer loop over DESC, controlling parent choice, Fourier-mode schedule, objective weights, and solver budget, while DESC owns the physics contract and acceptance. That division is sensible. The contribution that matters is the transition corpus — every attempted local solve is recorded with parent, action, outcome, validity, and continuation decision. That is a real step toward making optimization runs reusable, and the authors are appropriately careful to label these observations as behavior-policy priors, not causal demonstrations.\n\nThe main results are the paired 23-route subset: gate-valid outputs go from 5 to 19, median QS RMS drops to 1.07e-4, max curvature to 33.00 m^-1. The long route is a nice controlled case: nonmonotone QS, deliberate well/curvature debt, repair, and an eventual 9.1x QS reduction. The figures are consistent with the text. So the proof of concept stands on its own terms: an agent can execute multi-stage finite-beta search and produce improved equilibria.\n\nSoft spots, in proportion. The biggest is that \"We select 23 completed routes\" is under-specified. We don't learn how many routes the campaign launched, how completion was decided, or whether completion was known before outcomes were inspected. If early-stopping or outcome-based selection happened, the 5→19 number is biased. That said, the paper describes a fixed eight-epoch budget and treats every route as contributing an endpoint, so my read is that all routes run to budget; still, the authors should state this explicitly and ideally report the full campaign count. This is a fixable reporting gap, not a hole in the concept.\n\nThe deeper limitation — which the authors themselves flag — is the absence of matched-budget baselines: random, greedy, or memory-ablated controllers. Without those, the paper shows the system can do the job, not that the agentic decisions are what make it work. For a proof of concept that is acceptable, but it should be said plainly in the title or framing.\n\nOther notes: no error bars (medians over 23 routes, some metrics tied to gates); only one long route; no data/prompt release; the R^2=0.31 action-phase regression is weak evidence but honestly reported. The citations look appropriate — DESC, QUASR, ConStellaration, Landreman et al. Self-citation is minimal.\n\nWho this is for: plasma physics / stellarator optimization crowd, and people building LLM agents for scientific control. A serious referee should get this; I'd want the selection rule clarified, baseline experiments sketched or included, and artifacts released before endorsement. But it deserves referee time.\n\nRecommendation: send to peer review as a proof-of-concept, conditional on tightening the evaluation narrative and the route-selection reporting.","headline":"A well-scoped proof of concept for LLM-driven stellarator route control; the improvements are plausible but the paper under-specifies route selection and defers the baseline that would make the central claim comparative.","tokens_in":10969,"tokens_out":4170,"would_cite":true,"duration_ms":38313,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["52.55.Hc"],"model":"deepseek-v4-flash","headline":"A language-model agent that picks each optimization step — parent, Fourier modes, objective weights, budget — can run stage-one stellarator search, lifting gate-valid finite-beta outputs from 5 to 19 and halving median quasisymmetry error.","keywords":["agentic optimization","stellarator design","finite-beta MHD equilibrium","quasisymmetry","multi-objective search","hyperparameter optimization","language-model agent","transition data"],"falsifier":"Run the same 23 sources with the same eight-epoch budget but replace the agent with a scripted scheduler — random parent choice with fixed objective weights, or a greedy rule that always continues the best gate-valid child — and compare gate-valid counts and median QS RMS. If a scripted controller matches the agent, the claimed agentic advantage is not the cause. Independently, re-run the four short routes that ended outside the terminal gate with longer budgets; if they converge, the 19-of-23 count was a budget effect. Also, publish the rule that selected the 23 'completed' routes: a rule tha","tokens_in":10133,"feed_emoji":"🧲","tokens_out":18495,"duration_ms":156850,"temperature":0.7,"pith_summary":"The paper tries to establish that the expert-driven 'outer loop' of stage-one stellarator design — choosing which parent equilibrium to refine, which boundary Fourier modes to activate, how to weigh competing objectives, and how long to let the local solver run — can be delegated to a bounded language-model agent that diagnoses the current equilibrium and proposes the next experiment, while the deterministic DESC solver keeps final authority over the physics. Stellarator design is a multi-objective inverse problem whose target metrics (quasisymmetry, magnetic well, rotational transform, force balance, aspect ratio, curvature) do not give a constructive map to a validated finite-$\\beta$ equilibrium, so outcomes hinge on exactly the choices the agent now controls. If the claim is right, the payoff is two-fold: design search becomes an executable, parallel campaign rather than a one-off expert session, and every attempted step is recorded as structured decision data. On a paired common-budget subset of 23 QUASR-derived sources, gate-valid configurations rise from five to nineteen, median Boozer QS RMS falls from $2.39\\times10^{-4}$ to $1.07\\times10^{-4}$, and median maximum curvature falls from $62.56$ to $33.00\\,\\mathrm{m}^{-1}$; one 50-epoch route achieves a $9.10\\times$ QS reduction while repairing magnetic-well and curvature defects.","feed_headline":"From 5 to 19 valid designs: AI agent steers stellarator search","feed_subtitle":"Only 5 of 23 finite-beta sources met all design gates; agent-run routes made 19 pass and halved quasisymmetry error.","key_machinery":"The carrying mechanism is the agent–harness split. The language-model planner proposes actions, but the deterministic DESC layer owns the physics contract; this separation means every policy decision becomes a declared, replayable numerical experiment. The action is the typed hyperparameter tuple $h=(K,w,\\eta,b)$ — Fourier-mode schedule, objective-weight vector, typed objective parameters, numerical budget — applied to a named parent $p$, with the budget part of the agent's visible state rather than an invisible conversation count. Every attempted solve is stored as transition evidence $e_i$ coupling parent metric state, declared action, child metrics, optimizer response, validity, and conti","core_discovery":"The paper claims a language-model agent can sustain finite-$\\beta$, multi-objective stage-one stellarator optimization, framing route control as state-dependent hyperparameter selection: the action $a=(p,h)$ — parent, Fourier-mode schedule, objective weights, budget — behaves differently on different parents. A bounded agent emits these from a grounded route state; deterministic DESC owns the physics. Evidence: a paired eight-epoch campaign over 23 QUASR-derived sources lifts gate-valid outputs from five to nineteen, cuts median Boozer QS RMS from $2.39\\times10^{-4}$ to $1.07\\times10^{-4}$, and median maximum curvature from $62.56$ to $33.00\\,\\mathrm{m}^{-1}$. A 50-epoch route reaches $4.157\\ti","pith_inferences":["The paper itself (Sec. 6.3) says matched-budget static, greedy, random, and memory-ablated controllers 'are required to quantify policy efficiency', so the strongest honest reading is that agentic routes beat unoptimized source inputs, not yet that the agent beats other outer-loop policies; a direct comparator run at equal budget would settle it.","The four short routes that remained outside the terminal gate, and the long route's miss of its $10^{-6}$ QS target, mark the current frontier; a plausible extension the paper does not make is that longer budgets or learned proposal priors trained on the transition corpus would close most of that gap — a claim the 734-record corpus is built to test.","The agent–harness split is portable: any staged inverse design with a validated solver and metric gates (coil design, optics, structural shape) admits the same outer-loop pattern, so the contribution is not stellarator-specific."],"forward_implications":["If the paired improvements hold, stage-one optimization can run as asynchronous multi-start campaigns: each completed route contributes a refined finite-beta configuration and its full lineage from the same solver budget.","Because rejected children, failed solves, and lost within-epoch comparisons are stored in one transition schema, the corpus can support failure prediction, action-conditioned surrogates, and offline policy evaluation as coverage grows.","The aggregate shift is joint — the gate structure rewards feasibility as much as quasisymmetry — so extending the workflow to other source families, field periods, and acceptance contracts should keep configurations comparable under a single evaluation contract.","The long route's nonmonotone path, including temporary QS regressions that buy curvature or well repairs, implies that a policy maximizing instantaneous quality would stop too early; route-level credit assignment is needed to reach the frontier."],"supporting_citations":[{"why":"Supplies the QUASR coil-vacuum source boundaries that seed all 23 campaign routes; the input side of the paired evaluation depends on this dataset.","marker":"[7]"},{"why":"The DESC stellarator code that performs every fixed-boundary finite-beta equilibrium solve — the deterministic executor the agent configures.","marker":"[8]"},{"why":"The DESC quasi-symmetry optimization module that executes each local boundary optimization under the agent-chosen mode schedule and weight vector.","marker":"[9]"},{"why":"Documents how equilibrium-optimization outcome depends on initialization, scalarization, and spectral resolution — the premise that makes route-level control a separate sequential decision problem.","marker":"[5]"},{"why":"Exponential spectral scaling for mode-dependent boundary optimization; motivates and supports the Fourier-mode schedule component of the agent's action.","marker":"[6]"},{"why":"Defines the magnetic-well and Mercier stability diagnostics used to form the dense-well gate and the long-route well profile.","marker":"[3]"},{"why":"Precise-quasisymmetry framework that underlies the Boozer QS error metric and the QS gates applied to every endpoint.","marker":"[10]"},{"why":"On-axis magnetic well and Mercier criterion for arbitrary stellarator geometries; anchors the dense magnetic-well profile diagnostic.","marker":"[11]"}],"fun_headline_variants":["AI agent triples valid stellarator designs, halves error","From 5 to 19 valid equilibria: agent steers stellarator search","Agentic search: 14 new stellarator designs pass gates","Agentic outer loop: stellarator designs jump from 5 to 19","Agentic optimization: error halved, curvature halved, designs tripled"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The evaluation assumes the 23 routes treated as 'completed' are a representative slice of the campaign — specifically, that deciding a route is 'completed' is not correlated with whether it reached a good endpoint; the paper does not state how those 23 routes were chosen, and if unpromising routes were stopped early and set aside, the 5-to-19 gate-valid count and the median improvements would both be overstated.","fun_headline_variants_meta":{"raw":{"variants":["AI agent triples valid stellarator designs, halves error","From 5 to 19 valid equilibria: agent steers stellarator search","Agentic search: 14 new stellarator designs pass gates","Agentic outer loop: stellarator designs jump from 5 to 19","Agentic optimization: error halved, curvature halved, designs tripled"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000521,"raw_usage":{"total_tokens":2425,"prompt_tokens":875,"completion_tokens":1550,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":1469}},"tokens_in":619,"tokens_out":1550,"duration_ms":12626,"temperature":1.0,"reasoning_tokens":1469,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:17:07.262036+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 23 sources with the same eight-epoch budget but replace the agent with a scripted scheduler — random parent choice with fixed objective weights, or a greedy rule that always continues the best gate-valid child — and compare gate-valid counts and median QS RMS. If a scripted controller matches the agent, the claimed agentic advantage is not the cause. Independently, re-run the four short routes that ended outside the terminal gate with longer budgets; if they converge, the 19-of-23 count was a budget effect. Also, publish the rule that selected the 23 'completed' routes: a rule tha","supporting_citations":[{"cited_title":"A com- prehensive exploration of quasisymmetric stellarators and their coil sets.Journal of Plasma Physics, 91(5):E128, 2025","cited_arxiv_id":null,"evidence_quote":"Supplies the QUASR coil-vacuum source boundaries that seed all 23 campaign routes; the input side of the paired evaluation depends on this dataset."},{"cited_title":"The DESC Stellarator Code Suite Part I: Quick and accurate equilibria computations","cited_arxiv_id":"2203.17173","evidence_quote":"The DESC stellarator code that performs every fixed-boundary finite-beta equilibrium solve — the deterministic executor the agent configures."},{"cited_title":"The DESC Stellarator Code Suite Part III: Quasi-symmetry optimization","cited_arxiv_id":"2204.00078","evidence_quote":"The DESC quasi-symmetry optimization module that executes each local boundary optimization under the agent-chosen mode schedule and weight vector."},{"cited_title":"Deflation Techniques for Stellarator Equilibrium and Optimization, 2026","cited_arxiv_id":null,"evidence_quote":"Documents how equilibrium-optimization outcome depends on initialization, scalarization, and spectral resolution — the premise that makes route-level control a separate sequential decision problem."},{"cited_title":"Exponential spectral scaling: robust and efficient stellarator boundary optimisa- tion via mode-dependent scaling.Journal of Plasma Physics, 92 (1):E15, 2026","cited_arxiv_id":null,"evidence_quote":"Exponential spectral scaling for mode-dependent boundary optimization; motivates and supports the Fourier-mode schedule component of the agent's action."},{"cited_title":"Magnetic well and Mercier stability of stellarators near the magnetic axis","cited_arxiv_id":"2006.14881","evidence_quote":"Defines the magnetic-well and Mercier stability diagnostics used to form the dense-well gate and the long-route well profile."},{"cited_title":"Magnetic fields with precise quasisymmetry for plasma confinement","cited_arxiv_id":"2108.03711","evidence_quote":"Precise-quasisymmetry framework that underlies the Boozer QS error metric and the QS gates applied to every endpoint."},{"cited_title":"The On-Axis Magnetic Well and Mercier's Criterion for Arbitrary Stellarator Geometries","cited_arxiv_id":"2011.07416","evidence_quote":"On-axis magnetic well and Mercier criterion for arbitrary stellarator geometries; anchors the dense magnetic-well profile diagnostic."}],"review_version":1}