{"id":"3fbf5178-d640-478b-8038-3539a87299a1","arxiv_id":"2504.17710","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A sequential, multimodal VAE maps TCV flat-top discharges to a 2D latent space where a calibrated disruption-risk variable separates operating regimes and disruption types.","lead":"This paper trains a multimodal variational autoencoder on about 1,600 TCV tokamak discharges to map plasma measurements into a two-dimensional latent space organized by operating regime and disruption risk. The latent space separates known disruption types and supports counterfactual-style studies of which signals differ between disrupting and non-disrupting discharges.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Disruption-time labels from DEFUSE are the load-bearing external input; no accuracy check is given, so risk calibration and terminal-state clustering could be corrupted by label error.","rationale":"I agree with the reader's weakest_assumption. The paper's own Footnote 4 is an in-scope limitation that the VAE evaluation never quantifies. All disruption-risk labels and terminal projections flow from tD; no independent validation is provided. The conditional verdict is appropriate: the method is plausible, the architecture is described in detail, and the qualitative physics align with known limits, but the central empirical claims need a label-accuracy check before full acceptance. My proposed test is single and decisive for this concern.","tokens_in":27184,"tokens_out":5490,"duration_ms":62856,"concrete_test":"Have a fusion expert (or a second independent algorithm) manually annotate thermal-quench onset tD for a stratified sample of about 100 disruptions spanning IBL, DL, NT and fast/slow events. Quantify DEFUSE vs manual tD offsets. Then: (a) recompute Eq. 29 labels with manual tD and retrain/evaluate Drisk, reporting test ECE; (b) recompute the terminal-state assignments in Figures 13/14 using manual tD. If test ECE remains near 0.07 and IBL/DL/NT still separate into distinct states, the concern is resolved. If calibration degrades or the cluster separation blurs, the central claim depends on uncertified tD.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative and qualitative claims -- the ECE values of 0.0307/0.0715 and the IBL/DL/NT terminal-state separation in Figures 13/14 -- assume the DEFUSE disruption onset times tD are accurate. tD enters twice: Eq. 29 converts tD into BCE training labels ytm with a 1 s ramp, and Section 5.3 places the 'last projection before disruption' 0-5 ms before tD. Section 3.1 states tD is computed automatically by DEFUSE, and Footnote 4 concedes the event-detection development is ongoing, but no per-shot accuracy or uncertainty for tD is reported. If tD is biased by disruption type, e.g. detecting a precursor rather than thermal-quench onset or triggering after the quench, then both the risk labels and the terminal-state assignments are corrupted in a type-dependent way. The held-out ECE does not remove this dependence because the same tD pipeline generates the test labels and the test terminal states. This is a label-quality assumption, not a criticism of the VAE itself, but it is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multimodal sequential variational autoencoder (VAE) for modeling TCV flat-top plasma discharges in a 2D latent space with a Gaussian mixture prior. The model is extended with a disruption-risk head trained on binary labels derived from DEFUSE disruption onset times, and the resulting risk variable is post hoc Platt-scaled to match an empirical disruption rate. The authors evaluate the latent space qualitatively and quantitatively: they report an expected calibration error of 0.0307 on training data and 0.0715 on held-out test data, show that latent states correlate with confinement states and disruption-related events, demonstrate that ITER Baseline Scenario, density limit, and negative triangularity disruptions separate into distinct terminal latent states, and run a proof-of-principle counterfactual analysis that identifies disruption-correlated parameters consistent with known physics.","tokens_in":27422,"tokens_out":5214,"duration_ms":52759,"significance":"If the central claims hold, this is a valuable contribution to disruption characterization and plasma state monitoring. The method is deliberately interpretable, in contrast to most black-box disruption predictors, and the qualitative findings are physically meaningful: low q95 near IBL disruptions, high Greenwald fraction near DL disruptions, and elevated vertical growth rate near NT disruptions. The paper is also unusually transparent about its limitations, including ongoing event-detection development (Footnote 4) and manual hyperparameter selection (Appendix A). Strengths include a clearly described dataset, careful attention to causal interpolation to avoid information leakage, and downstream analysis that produces falsifiable, physics-aligned predictions. However, the quantitative calibration claim is weakened by the post hoc Platt scaling and by the shared reliance on the same DEFUSE t_D labels for training labels, calibration targets, and terminal-state assignments, which limits the strength of the evidence for the headline ECE values.","major_comments":[{"comment":"The reported ECE values are computed after fitting Platt scaling on the same training set, and the empirical disruption rate eDrate used as the calibration target is defined using the same t_D that generates the y_tm training labels in Eq. (29). The training ECE of 0.0307 therefore measures how well a one-parameter calibrated curve fits the empirical rate on the data it was fit to, rather than how well D_risk predicts an independent quantity. Please report the ECE of the raw, uncalibrated D_risk before Platt scaling, and use a proper split procedure in which Platt scaling is fit only on the validation set and evaluated on the test set. This would make the test ECE of 0.0715 a cleaner generalization statistic.","section":"Section 5.2, Eq. (37), with Eqs. (29)-(30)"},{"comment":"The terminal-state assignments in the IBL/DL/NT clustering rely on the 'last projection before disruption,' defined as 0-5 ms before t_D (footnote 8). Since t_D is automatically computed by DEFUSE and no per-shot accuracy or uncertainty is reported, a type-dependent bias in t_D (e.g., detecting a precursor rather than the thermal quench onset) could create or distort the observed separation. Please add a sensitivity analysis that jitters t_D by at least ±10 ms and ideally ±50 ms and recomputes the state assignments, or report a per-disruption-type validation of the DEFUSE t_D against manual annotations.","section":"Section 5.3, Figures 13 and 14"},{"comment":"The final model configuration was selected by manual inspection of latent spaces, and the authors note that 'we are likely selecting suboptimal settings.' To rule out that the central qualitative results (multimodal separation and the disruption-type clustering) are artifacts of a particular configuration, please report quantitative latent-space quality metrics for the chosen model and for a small neighborhood of the scanned hyperparameters. The mutual information used in Figure A.1, the reconstruction error, and a cluster-separability measure would be appropriate; showing that the main conclusions are stable across these configurations would considerably strengthen the paper.","section":"Appendix A, Hyperparameter choice"}],"minor_comments":[{"comment":"The notation for the disruptivity alternates between \\hat{D}_disr in the text and D_disr in the figures and captions; please unify the notation.","section":"Section 5.2, Figures 9 and 10"},{"comment":"The threshold W1 ≥ 3 for selecting 'significant' feature differences is presented without supporting justification or multiple-testing correction; this is acceptable for a proof-of-principle, but the exploratory nature should be stated explicitly.","section":"Section 5.4"},{"comment":"The footnote states that early evaluations indicate the aggregate distributions of event detections are statistically meaningful; please clarify how this was assessed, since these detections are used as evaluation metadata.","section":"Section 3.1, Footnote 4"},{"comment":"The claim of 'a notable lack of estimates between 1/16 to 1/2' would be more informative if accompanied by the fraction of samples in each bin of the reliability diagram.","section":"Figure 8"},{"comment":"The sentence 'The output of the FNO layers is flattened and concatenated to the previous value of µ' is slightly misleading, since the architecture in Table A.5 uses a two-input MLP rather than a simple concatenation; please rephrase to match the implementation.","section":"Section 4.4, Table A.5"},{"comment":"Please define the domain of t_m explicitly and clarify that the ramp only applies for t_D - B ≤ t_m < t_D - A, with y_tm = 1 for t_m ≥ t_D - A, so that readers do not have to infer the piecewise structure.","section":"Section 4.3, Eq. (29)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its limitations, and the qualitative latent-space findings are compelling. The main risk is the heavy reliance on DEFUSE t_D labels, which enter the risk labels, the calibration target, and the terminal-state clustering; this is fixable with a sensitivity analysis and a cleaner calibration evaluation. I would not reject the paper on this basis, but the current evaluation of the headline ECE numbers is not yet fully convincing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful methods paper, not a breakthrough. It builds a sequential multimodal VAE for TCV flat-top discharges, and the latent space separates IBL, density-limit and negative-triangularity disruptions into distinct terminal states. That separation is the strongest result, and it lines up with known physics (low q95 for IBL, Greenwald fraction for DL, vertical growth rate for NT). The component planes and counterfactual-style analysis are nice downstream demonstrations.\n\nWhat's new: the specific combination—a residual (neural-ODE-like) encoder, a Gaussian-mixture prior with an explicit mode-coverage loss, and a learned Drisk head—is not in refs [39]/[40]. The paper is honest about what is fresh; those refs are cited properly. The dataset curation, including the stated 70.4% disruption bias and confidence-thresholded confinement labels, is also transparent.\n\nSoft spots: the quantitative risk calibration is weaker than it looks. Drisk is trained on labels built from tD, then Platt-scaled post hoc to the empirical disruption rate computed from the same tD. So the ECE numbers (0.0307 train, 0.0715 test) measure how well the post-hoc fit matches the rate, not an independent prediction. The reader flagged this; the critique lands. The stress-test note is also fair: no accuracy check is given for the DEFUSE tD values, and those times enter both the training labels and the terminal-state clustering. If tD is systematically early or late for one disruption type, Figures 13/14 could be corrupted in a type-dependent way. The paper should provide a sensitivity analysis or at least some validation of tD. That said, the qualitative physics alignment would likely survive modest label noise; this is not fatal for the main point, but it is for the headline ECE numbers.\n\nAlso worth noting: model selection was manual (Appendix A admits it), and no code or data is released. That limits reproducibility but does not invalidate the approach. The paper is also appropriately modest in its claims—it frames the counterfactual study as a proof of principle, not a full causal analysis.\n\nVerdict: a serious referee should engage with this. It is a solid contribution for people working on tokamak operational space and data-driven plasma state modeling. I'd ask for revisions that (1) address tD sensitivity, (2) clarify what the ECE does and does not show, and (3) state a code/data availability plan. I'd probably cite it.","headline":"Useful methods paper with a convincing latent-space separation of disruption types, but the quantitative risk calibration and the DEFUSE tD labels carry more weight than they should.","tokens_in":27978,"tokens_out":2146,"would_cite":true,"duration_ms":22644,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multimodal VAE maps TCV discharges into a latent space where a calibrated disruption risk tracks the empirical disruption rate.","keywords":["plasma disruption","tokamak","variational autoencoder","latent variable model","disruption risk","interpretable machine learning","TCV","counterfactual analysis"],"falsifier":"Take a held-out subset of TCV discharges and manually label the disruption onset time (start of the thermal quench) without reference to the automated detections, then retrain or re-evaluate the risk calibration with those labels; if the expected calibration error rises well above the reported test value or the three disruption families no longer fall into distinct terminal states, the central claim would be refuted.","tokens_in":26992,"feed_emoji":"⚡","tokens_out":3245,"duration_ms":35549,"temperature":0.7,"pith_summary":"This paper tries to establish that a multimodal variational autoencoder can turn tokamak discharge measurements into a low-dimensional, interpretable map of plasma operating regimes, with a calibrated disruption-risk coordinate that tracks how likely a discharge is to disrupt. The authors argue that such a representation complements black-box disruption predictors by making the proximity to disruption visible and separable across different disruption types. They demonstrate the claim on roughly 1,600 TCV flat-top discharges, showing that the learned risk variable matches the empirical disruption rate and that ITER baseline, density limit, and negative triangularity disruptions end in distinct latent states. A sympathetic reader would take the central claim to be that data-driven plasma-state monitoring can be both expressive and interpretable enough to support disruption characterization and downstream causal-style analyses.","feed_headline":"Latent risk map orders 1,600 TCV discharges by disruption proximity","feed_subtitle":"A calibrated VAE coordinate separates disruption types and matches the observed disruption rate to within a few percent.","key_machinery":"The central object is a two-dimensional latent state $z$ whose trajectory represents the plasma state over time, together with a learned disruption-risk map $D_{\\text{risk}}(z)\\in[0,1]$. The encoder is a residual sequential model that updates the latent mean from a sliding window of diagnostic signals, the prior is a Gaussian mixture with $K=8$ modes that encourages multimodal clustering, and the decoder maps each latent point back to the full signal vector so the space can be read off in physics quantities. The disruption-risk head is trained with binary cross-entropy labels that ramp to 1 during the second before the disruption time $t_D$, and the final risk values are recalibrated with Platt scaling; this risk term is what pulls disruptive and non-disruptive regions apart in the latent space.","core_discovery":"The central discovery is that a two-dimensional latent variable, learned by a sequential multimodal VAE with an added disruption-risk head, organizes the TCV flat-top operational space into smooth regions whose calibrated disruption risk $D_{\\text{risk}}$ closely follows the empirically observed disruption rate. After post-hoc Platt scaling, the expected calibration error is 0.0307 on the training set and 0.0715 on a held-out test set, and low-risk regions contain almost no disrupting discharges while risk rises exponentially near disruption. The same latent space separates known disruption families: projections of ITER baseline, density limit, and negative triangularity discharges start clustered together but terminate in different peaks of high-risk regions, and component planes show physically sensible correlations with $q_{95}$, Greenwald fraction, and vertical growth rate. The method also supports a proof-of-principle counterfactual analysis in which the most similar non-disrupting counterpart of a disrupting shot is found automatically, and the most distinct signal differences recover known disruption drivers such as MHD activity, density-limit proximity, and vertical instability.","pith_inferences":["Beyond the paper, the same latent-map construction could be extended to multi-machine databases, although device-specific time and spatial scales would need to be handled explicitly rather than letting each machine collapse into its own latent region.","The 5 ms latent timestep and slow-timescale focus mean the method is better suited to identifying global regime proximity than to resolving the fast chain-of-events immediately before a disruption; a dedicated early-warning comparison against existing predictors would be a natural next test.","If $D_{\\text{risk}}$ remains well calibrated under online projection of new discharges, it could be embedded in real-time control as a risk coordinate to steer trajectories away from high-risk regions, which the paper only sketches as future work.","A direct falsifiable extension would be to recompute the calibration error using independently labeled disruption-onset times rather than the automated detections, since the current risk labels inherit any error in those detections."],"forward_implications":["The calibrated $D_{\\text{risk}}$ variable can serve as a continuous, interpretable indicator of disruption rate across the flat-top operational space, with roughly 3% average calibration error on the training set and 7% on novel discharges.","The method automatically separates at least three distinct disruption families into different terminal latent states, meaning it can be used to label the type of disruption that a trajectory is heading toward.","Latent states correlate with plasma confinement states even though no confinement labels were used in training, so the learned map recovers physically meaningful operating regimes without supervision.","The counterfactual-style analysis identifies known disruption-related parameters for different scenarios, suggesting the latent space can automate the search for differences between disrupting and non-disrupting shots.","Because the decoder projects each latent point back to physics quantities, the same map can be reused to inspect which measured signals change as a discharge moves toward higher-risk regions."],"supporting_citations":[{"why":"Supplies the disruption-onset times $t_D$ and automated event detections used both to build the training labels and to populate the evaluation database.","marker":"[26]"},{"why":"Provides the automated confinement-state labels used to correlate latent states with L-mode, dithering, and H-mode behavior.","marker":"[32]"},{"why":"Introduces the variational autoencoder formulation that the method extends.","marker":"[20]"},{"why":"Establishes the stochastic backpropagation and reparameterization approach used for training the VAE.","marker":"[21]"},{"why":"Motivates the dynamical formulation of the encoder as latent trajectories over time.","marker":"[22]"},{"why":"Provides the Gaussian mixture variational autoencoder structure used for the multimodal prior and clustering.","marker":"[23]"},{"why":"Defines the ITER baseline scenario disruptions used as one of the three disruption-family validation cases.","marker":"[29]"},{"why":"Defines the density limit experiments used as the second disruption-family validation case.","marker":"[30]"},{"why":"Defines the negative triangularity scenarios used as the third disruption-family validation case and links them to vertical instability.","marker":"[31]"},{"why":"Provides the Platt scaling procedure used to calibrate the raw disruption-risk output to the empirical disruption rate.","marker":"[67]"}],"fun_headline_variants":["Latent VAE map orders 1,600 plasma shots by disruption risk","VAE latent space separates disruption types and predicts risk","Multimodal VAE yields interpretable disruption risk map for TCV","Latent coordinates order TCV shots by proximity to disruption","Counterfactual VAE analysis pinpoints disruption drivers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation depends on the automated disruption-onset times, event detections, and confinement-state labels being accurate enough to define both the training targets and the evaluation metadata; if those onset times are systematically wrong, the risk calibration and the terminal-state clustering would shift even if the VAE itself is sound.","fun_headline_variants_meta":{"raw":{"variants":["Latent VAE map orders 1,600 plasma shots by disruption risk","VAE latent space separates disruption types and predicts risk","Multimodal VAE yields interpretable disruption risk map for TCV","Latent coordinates order TCV shots by proximity to disruption","Counterfactual VAE analysis pinpoints disruption drivers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2677,"prompt_tokens":1062,"completion_tokens":1615,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":678,"completion_tokens_details":{"reasoning_tokens":1529}},"tokens_in":678,"tokens_out":1615,"duration_ms":10020,"temperature":1.0,"reasoning_tokens":1529,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:33:05.953681+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out subset of TCV discharges and manually label the disruption onset time (start of the thermal quench) without reference to the automated detections, then retrain or re-evaluate the risk calibration with those labels; if the expected calibration error rises well above the reported test value or the three disruption families no longer fall into distinct terminal states, the central claim would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Platt scaling procedure used to calibrate the raw disruption-risk output to the empirical disruption rate."}],"review_version":1}