Pith. sign in

REVIEW 4 major objections 5 minor 64 references

GyroSwin emulates full 5D gyrokinetic turbulence with a neural network, matching or beating reduced models on heat flux at 1/1000th the cost.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-04 10:57 UTC pith:5CDTSZHK

load-bearing objection First full-5D surrogate for gyrokinetic turbulence, with real out-of-sample flux gains over QL, but the headline cost/cascade claims outrun the evidence. the 4 major comments →

arxiv 2510.07314 v3 pith:5CDTSZHK submitted 2025-10-08 physics.plasm-ph cs.AIstat.ML

GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations

classification physics.plasm-ph cs.AIstat.ML
keywords gyrokinetic turbulence5D distribution functionneural surrogateplasma heat fluxzonal flowsSwin transformerfusion energyautoregressive rollout
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

GyroSwin is a neural network that directly learns the 5D distribution function of nonlinear gyrokinetic turbulence, evolving it snapshot by snapshot while predicting the 3D electrostatic potential and scalar heat flux in the same forward pass. The paper's central claim is that this 5D surrogate matches or beats quasilinear reduced-order models on time-averaged heat flux, reproduces the turbulent energy cascade and zonal-flow structure that linear-based models miss, and runs roughly three orders of magnitude faster than fully resolved nonlinear gyrokinetics. If correct, fusion design workflows that today rely on cheap but incomplete linear approximations could instead use a retrained surrogate that retains nonlinear physics. The authors also report favorable scaling with data and model size up to a billion parameters, suggesting the architecture can be extended to higher-fidelity regimes.

Core claim

GyroSwin's central claim is that a hierarchical vision transformer extended to five dimensions, trained with a combined loss on the distribution function f, the electrostatic potential φ, and the heat flux Q, can serve as a surrogate for the nonlinear gyrokinetic equation under the adiabatic-electron approximation. The model ingests one 5D snapshot of f (given as real and imaginary parts in spectral kx, ky, along-field coordinate s, parallel velocity v∥, and magnetic moment μ) together with operating parameters, and predicts the next snapshot plus φ and Q. Evaluated on unseen in-distribution and out-of-distribution parameter sets, the authors report that GyroSwin improves heat-flux RMSE over

What carries the argument

Three architectural ideas carry the argument. First, 5D shifted-window attention (5D W-MSA) restricts self-attention to small 5D windows and shifts the window grid between layers, making attention cost near-linear in the 5D resolution instead of quadratic, which is what makes a 5D transformer computationally feasible. Second, latent integrator and cross-attention modules replicate the physical integrals of Equation (2) in latent space: a pooling query contracts the 5D latent over velocity directions to obtain a 3D latent for the potential, and cross-attention passes information between the 5D and 3D branches, enabling multitask training on f, φ, and Q. Third, channelwise mode separation isol

Load-bearing premise

The saturated phase of turbulence is treated as a deterministic, Markovian map from one 5D snapshot to the next; if chaotic divergence breaks that assumption, the rollout-correlation and time-averaged flux results may not transfer beyond the training distribution.

What would settle it

Run GyroSwin autoregressively against two ground-truth nonlinear simulations that share the same operating parameters but different initial noise amplitudes (both within the trained range). If the model's per-snapshot Pearson correlation (τ=0.1) drops to near the correlation between the two ground-truth runs much earlier than its reported ~100-step rollouts, the deterministic snapshot-to-snapshot map is not tracking the true trajectory; the accurate time-averaged fluxes would then reflect ensemble statistics rather than a valid deterministic surrogate.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Replacing quasilinear reduced models: if the claims hold, GyroSwin can supply time-averaged heat flux and per-mode spectra for ion-scale adiabatic-electron turbulence at roughly one-thousandth the cost of a fully resolved nonlinear run, with nonlinear physics such as zonal flows included.
  • Self-consistent diagnostics: because f, φ, and Q are predicted together, flux spectra, turbulence intensity spectra, and zonal-flow profiles are all derived from the same predicted field, so they are mutually consistent rather than assembled from separate approximations.
  • Scalable path to higher fidelity: the monotone improvement from 90M to 1B parameters suggests the same architecture can absorb larger, higher-resolution datasets (e.g., kinetic electrons, multiple species) as they become available.
  • A learned saturation rule can outperform a hand-fitted one: the ablation in which GyroSwin is trained on linear simulations to predict nonlinear flux beats the fitted quasilinear saturation rule, implying that data-driven saturation may replace the free-parameter saturation rules in reduced transport models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because GyroSwin predicts the full 5D state, it could be used to synthesize dense turbulence statistics in regions of parameter space where nonlinear runs are too expensive to enumerate, effectively acting as an emulator for transport databases; the paper does not perform this synthesis.
  • The known chaotic, distributional nature of saturated turbulence suggests the next testable step is generative training (e.g., diffusion or score-based) over the saturated-phase ensemble, which the authors explicitly list as future work; such a model would trade pointwise rollout accuracy for calibrated ensemble statistics.
  • The linear-only saturation ablation points to a cheaper intermediate product: a network trained exclusively on linear simulations could calibrate quasilinear flux predictions across large parameter scans, avoiding the need to generate extensive nonlinear training data.
  • If the architecture transfers beyond the adiabatic-electron approximation, the same 5D→3D latent integration design could in principle couple ion and electron distribution functions to a shared potential field; this remains untested, as only adiabatic electrons are considered here.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. GyroSwin is a hierarchical Swin-Transformer-based UNet that operates directly on the 5D gyrokinetic distribution function (real-space x,y,s and velocity v_parallel, mu) and predicts, in a multitask fashion, the next-snapshot distribution function, the 3D electrostatic potential, and the scalar heat flux. It is trained on GKW simulations with adiabatic electrons, using Latin-hypercube sampling over four operating parameters. The paper reports autoregressive rollout correlation times, time-averaged heat-flux RMSE on held-out in-distribution and out-of-distribution simulations, flux/turbulence spectra Q(k_y) and W(k_y), and zonal-flow profiles, together with scaling experiments to a 998M-parameter model trained on 241 simulations. The central claim is that GyroSwin is the first scalable 5D neural surrogate that outperforms quasilinear reduced models in heat-flux prediction, captures the turbulent energy cascade, and reduces the cost of fully resolved nonlinear gyrokinetics by three orders of magnitude.

Significance. If the claims hold, GyroSwin would be a practically useful surrogate: it predicts 5D fields and derived physical diagnostics in an out-of-sample manner, in contrast to tabular regression surrogates that only predict scalar fluxes. The manuscript has genuine strengths: the held-out ID/OOD evaluation is real, the zonal-flow and spectrum diagnostics are computed from predicted fields rather than being directly supervised, ablations are provided for the main architectural components, code is released, and scaling to 1B parameters is demonstrated. The central claims, however, are currently stronger than the evidence: the headline comparison with QuaLiKiz is made at unequal training-data budgets, the 'three orders of magnitude' speedup is not cleanly supported by the reported numbers, and the 'turbulent energy cascade' claim rests on time-averaged spectra plus a low correlation-time threshold, while the paper's own appendices document substantial spectral and zonal-flow misses.

major comments (4)
  1. [Section 5, Tables 2 and 5, Appendix B] The headline QL comparison is at unequal training-data budgets. The QL row in Table 2 is fitted to the 48-simulation training set (Appendix B), whereas the GyroSwin rows in the 'Scaling' block use 241 simulations. Table 5 shows that fitting QL to 241 simulations reduces its ID RMSE from 89.53 to 40.74, so the apparent gap in Table 2 is partly a data-budget artifact. Also, on the 48-simulation budget, the parameter-only GPR and MLP baselines in Table 2 (ID RMSE 43.82 and 50.50, respectively) are notably better than GyroSwin (67.68), which contradicts an unqualified reading of 'outperforms widely used reduced numerics on heat flux prediction' in the abstract. Please include QL-241 in the main comparison, discuss the small-data regime explicitly, and qualify the abstract claim accordingly.
  2. [Abstract and Section 5, Table 3] The speed claim 'three orders of magnitude faster than GKW (4200 vs. 756 GFLOPs)' is internally inconsistent: 4200/756 ≈ 5.5, not 1000. Either the units are mislabeled or the comparison is between different workloads (e.g., a full GKW simulation versus one GyroSwin forward/rollout, with different snapshot averaging as described in Section 4). Because the 1000x speedup is a load-bearing claim in the abstract, it needs a precise definition and a reproducible measurement: per physical timestep advanced, including preprocessing and rollouts, and with the GKW cost counted over the same simulated time interval.
  3. [Sections 5–6, Appendices E and F] The evidence does not support the statement that GyroSwin 'captures the turbulent energy cascade.' The support consists of time-averaged spectra W(k_y) and Q(k_y) correlations (Table 3, Figures 9–10), but there is no direct cascade/energy-transfer diagnostic. Appendix F concedes that high-frequency spectral components are not captured, and Appendix E shows velocity-space overprediction after roughly ten rollout steps. Section 6 concedes the deterministic next-step model ignores the chaotic/distributional nature of turbulence. A deterministic L2-trained model can reproduce a time-mean flux while having the wrong conditional distribution. Please either add a spectral energy-transfer or cascade diagnostic, or replace 'captures the turbulent energy cascade' with a statement about low-frequency spectral reproduction. In addition, report the distribution of per-timestep Q(t) and compare agains
  4. [Section 5, Figures 4 and 11, Appendix F] The 'new capability: 5D zonal flow modelling' claim is overstated. The main text selects one favorable OOD case in Figure 4, while Appendix F states that on some test cases the zonal-flow profile is 'entirely off' and that amplitudes are overestimated and normalized for visualization. A single favorable example does not establish the capability. Please report quantitative errors (e.g., RMSE or correlation per ID/OOD case, with mean and spread) for the zonal-flow profile over all test simulations, and move the Appendix F caveat into the main-text discussion of this capability.
minor comments (5)
  1. [Section 3, Eq. (8); Appendix D] The multitask loss weights w_f, w_phi, and w_Q are never specified. These values affect the ablation and the final results; please report them and ideally a small sensitivity study.
  2. [Tables 2 and 3] The naming is inconsistent: Table 2 uses GyroSwinSmall/GyroSwinMedium/GyroSwin for the 241-simulation rows, while Table 3 uses GyroSwinSmall/GyroSwinMedium/GyroSwinLarge. Please use consistent names so the reader can identify the 998M-parameter model.
  3. [Abstract] The phrase 'first scalable 5D neural surrogate' is too broad. The present study considers local, adiabatic-electron, ion-scale gyrokinetics at a fixed resolution and in a restricted four-parameter region. Please qualify the novelty claim accordingly.
  4. [Throughout] There are numerous typos ('simualtions', 'appraoches', 'Plasmas', 'QuasiLinear' capitalization, etc.). A careful proofread is needed.
  5. [Section 4] The FNO baseline is a 3D FNO with velocity dimensions collapsed into channels, and PointNet/Transolver use subsampling. This should be stated more prominently in the main text so the reader does not interpret the baseline comparison as a comparison of equal-capability 5D models.

Circularity Check

0 steps flagged

No significant circularity: GyroSwin is a supervised surrogate with genuinely out-of-sample ID/OOD evaluation; disclosed QL baseline fitting and direct Q supervision do not reduce the claimed predictions to their inputs by construction.

full rationale

The claimed derivation chain is an empirical surrogate fit, not a derivation from first principles: GyroSwin is trained with a supervised next-step loss on GKW data (Eq. 8) and evaluated on simulations 'excluded from the training set' (Sec. 4). The central flux prediction is therefore a regression on held-out data, and the spectral/zonal diagnostics are computed from predicted fields rather than being separately fitted. The only parameters calibrated to nonlinear target fluxes are the QuaLiKiz saturation constant C=7.93 (Eq. 21, App. B) and the network weights themselves; the QL fit is a disclosed baseline, not a claimed prediction, and it is evaluated out-of-sample. Self-citations (Alkin et al. 2024a for correlation time; Bodnar et al. 2024 as weather background) are not load-bearing: the metric is restated in the text and the background claim does not justify GyroSwin's results. The admitted limitations in Sec. 6 and App. F (error accumulation, missed high frequencies, some zonal-flow profiles 'entirely off') weaken the evidence for the 'captures the cascade' claim, but they are evidential weaknesses, not circular reductions. No step equates a predicted output to a fitted input by construction.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The central contribution is an empirical fit: the trained network weights encode the gyrokinetic dynamics from GKW data. The ledger therefore lists the disclosed fit parameters (notably the QL saturation-rule C), the hand-set loss weights and architecture choices, and the domain assumptions (GKW as ground truth, saturated-phase-only data, deterministic Markov rollout, adiabatic electrons) that bound what the surrogate can claim. There are no invented physical entities; the only invented object is the architectural inductive bias of zonal-flow channel separation, which is validated by ablation but not derived.

free parameters (3)
  • QuaLiKiz saturation-rule normalization C = 7.93
    Equation (21) fits C by least squares to the nonlinear flux training set; this is the QL baseline's one free parameter, disclosed by the authors.
  • Multitask loss weights w_f, w_phi, w_Q = not reported
    Equation (8) defines the total loss with weighting factors; the values are not given in the text and are hand-chosen hyperparameters that affect the trade-off between f, phi, and Q accuracy.
  • Architecture hyperparameters (window size M, patch sizes, stage depths, channel widths) = partially in Section D
    GyroSwin's structure depends on these hand-chosen choices; ablations in Table 4 show component sensitivity, but the final configuration is not derived from first principles.
axioms (5)
  • domain assumption The gyrokinetic equation (Eq. 1) and its numerical solution by GKW provide the ground-truth dynamics for plasma turbulence.
    The surrogate is trained on GKW outputs; if GKW's flux-tube model or the adiabatic-electron approximation is wrong for the target regime, the surrogate inherits that error. Invoked throughout Sections 2 and 4.
  • domain assumption The saturated phase of the simulation (after discarding the first 80 snapshots) is representative of the turbulence relevant to time-averaged transport.
    Data Generation (Section 4) drops the linear phase; the surrogate never sees transient growth, so any application needing transient fluxes is outside scope. This is a stated modeling choice.
  • domain assumption The turbulent dynamics can be approximated by a deterministic map from one 5D snapshot to the next, with the 4 operating parameters and timestep as conditioning information.
    This is the autoregressive rollout premise; Section 6 concedes turbulence is chaotic/distributional, so the premise is only approximately valid and limits rollout length.
  • ad hoc to paper Separating the zonal-flow mode (ky=0) and adding real/imaginary parts as extra input channels improves generalization.
    Channelwise mode separation is an inductive bias specific to this architecture; validated by ablation in Table 4 but not derived from the gyrokinetic equation.
  • domain assumption The velocity-space dimension mu can be decoupled into channels with no loss of information.
    Section D states this is exact for adiabatic electrons; if coupling through collisions or the source term S were retained, the decoupling would be invalid.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations." pith.science (2026). https://pith.science/paper/5CDTSZHK

@misc{pith2026251007314,
  author       = {Pith},
  title        = {Pith review of: GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5CDTSZHK}},
  note         = {Machine review of arXiv:2510.07314}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Nuclear fusion plays a pivotal role in the quest for reliable and sustainable energy production. A major roadblock to viable fusion power is understanding plasma turbulence, which significantly impairs plasma confinement, and is vital for next-generation reactor design. Plasma turbulence is governed by the nonlinear gyrokinetic equation, which evolves a 5D distribution function over time. Due to its high computational cost, reduced-order models are often employed in practice to approximate turbulent transport of energy. However, they omit nonlinear effects unique to the full 5D dynamics. To tackle this, we introduce GyroSwin, the first scalable 5D neural surrogate that can model 5D nonlinear gyrokinetic simulations, thereby capturing the physical phenomena neglected by reduced models, while providing accurate estimates of turbulent heat transport. GyroSwin (i) extends hierarchical Vision Transformers to 5D, (ii) introduces cross-attention and integration modules for latent 3D$\leftrightarrow$5D interactions between electrostatic potential fields and the distribution function, and (iii) performs channelwise mode separation inspired by nonlinear physics. We demonstrate that GyroSwin outperforms widely used reduced numerics on heat flux prediction, captures the turbulent energy cascade, and reduces the cost of fully resolved nonlinear gyrokinetics by three orders of magnitude while remaining physically verifiable. GyroSwin shows promising scaling laws, tested up to one billion parameters, paving the way for scalable neural surrogates for gyrokinetic simulations of plasma turbulence.

Figures

Figures reproduced from arXiv: 2510.07314 by Fabian Paischer, Gianluca Galletti, Johannes Brandstetter, Lorenzo Zanisi, Naomi Carey, Paul Setinek, Stanislas Pamela, William Hornsby.

Figure 1
Figure 1. Figure 1: Left: GyroSwin models the 5D distribution function of nonlinear gyrokinetics and incor￾porates integration blocks to predict 3D electrostatic potential fields and scalar heat flux. Right: ROMs (quasilinear) solve a cartesian product of 2D modes in spectral space and 3D fields. Further￾more. They rely on saturation rules to approximate the nonlinear flux spectrum. et al., 2007; Staebler & Kinsey, 2010), are… view at source ↗
Figure 2
Figure 2. Figure 2: Left: GyroSwin receives the 5D distribution function as input and predicts the evolved 5D distribution function, as well as the respective 3D potential and heat flux at the next timestep. Right: Essential building blocks and integrator layers that enable multitask training. The latent 5D space is integrated over velocity space to obtain a latent 3D field for potential prediction via cross-attention. in des… view at source ↗
Figure 3
Figure 3. Figure 3: Scaling GyroSwin to ∼1B parameters trained on 241 simulations amounting to approxi￾mately 6TB of data. We show train/validation error for predicting the 5D distribution function (left) and the 3D electrostatic potential field (right). 10 1 10 0 ky 10 2 10 0 10 2 10 4 W(ky) 0 10 20 30 x 1.0 0.5 0.0 0.5 1.0 Amplitude GT GyroSwinLarge Quasilinear Transolver PointNet ViT [PITH_FULL_IMAGE:figures/full_fig_p009… view at source ↗
Figure 4
Figure 4. Figure 4: Left: W(ky) averaged over time and OOD simulations for different 5D neural surrogates. Competitors tend to underestimate while GyroSwin matches the spectrum well with a slight dis￾crepancy on higher frequencies. Right: Time-averaged zonal flow profile for a slice along s across radial coordinates x for a selected OOD simulation. GyroSwin captures the zonal flow profile. captured that contribute most to hea… view at source ↗
Figure 5
Figure 5. Figure 5: Distribution of input parameters sˆ, q, R/Ln, and R/Lt along with average heat flux Q¯. The sampled parameter space is evenly distributed. then the flattened patch dimension is embedded with a shallow MLP. Patch embedding, merging, and expansions are implemented as linear layers or MLPs. Furthermore, we add relative positional biases and condition all Swin layers on the 4 parameters as well as the current … view at source ↗
Figure 6
Figure 6. Figure 6: Visualization of the 5D distribution function of nonlinear gyrokinetics (ground truth, [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Comparison of 5D distribution functions for linear and nonlinear simulations (ground [PITH_FULL_IMAGE:figures/full_fig_p020_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Side-by-side comparison of autoregressive rollout predictions with GyroSwin compared [PITH_FULL_IMAGE:figures/full_fig_p022_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Comparison of W(ky) In-Distribution and Out-Of-Distribution. In addition to the time-averaged spectra, we provide visualizations for the time-averaged zonal flow profile for all ID simulations as well as OOD simulations for GyroSwinLarge. We observe that GyroSwinLargeusually tends to overestimate the amplitude therefore we normalize it for visualization purposes. The results for the ID and OOD test cases c… view at source ↗
Figure 10
Figure 10. Figure 10: Comparison of flux spectra In-Distribution and Out-Of-Distribution. [PITH_FULL_IMAGE:figures/full_fig_p023_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Comparison of zonal flow profiles In-Distribution and Out-Of-Distribution. [PITH_FULL_IMAGE:figures/full_fig_p023_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

64 extracted references · 16 canonical work pages

  1. [1]

    Accurate structure prediction of biomolecular interactions with alphafold 3

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, pp.\ 1--3, 2024

  2. [2]

    Universal physics transformers

    Benedikt Alkin, Andreas F \" u rst, Simon Schmid, Lukas Gruber, Markus Holzleitner, and Johannes Brandstetter. Universal physics transformers. CoRR, abs/2402.12365, 2024 a . doi:10.48550/ARXIV.2402.12365

  3. [3]

    Neuraldem-real-time simulation of industrial particulate flows

    Benedikt Alkin, Tobias Kronlachner, Samuele Papa, Stefan Pirker, Thomas Lichtenegger, and Johannes Brandstetter. Neuraldem-real-time simulation of industrial particulate flows. arXiv preprint arXiv:2411.09678, 2024 b

  4. [4]

    Accurate medium-range global weather forecasting with 3d neural networks

    Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. Accurate medium-range global weather forecasting with 3d neural networks. Nat., 619 0 (7970): 0 533--538, 2023. doi:10.1038/S41586-023-06185-3

  5. [5]

    C. K. Birdsall and A. B. Langdon . Plasma physics via computer simulation. New York: Taylor and Francis, first edition, 2005

  6. [6]

    Bruinsma, Ana Lucic, Megan Stanley, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

    Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Anna Vaughan, Jayesh K. Gupta, Kit Thambiratnam, Alex Archibald, Elizabeth Heider, Max Welling, Richard E. Turner, and Paris Perdikaris. Aurora: A foundation model of the atmosphere. CoRR, abs/2405.13063, 2024....

  7. [7]

    Bourdelle, X

    C. Bourdelle, X. Garbet, F. Imbeaux, A. Casati, N. Dubuit, R. Guirlet, and T. Parisot. A new gyrokinetic quasilinear transport model applied to particle transport in tokamak plasmas. Physics of Plasmas, 14 0 (11): 0 112501, 11 2007. ISSN 1070-664X. doi:10.1063/1.2800869

  8. [8]

    Bourdelle, A

    C. Bourdelle, A. Casati, X. Garbet, F. Imbeaux, J. Candy, F. Clairet, G. Dif-Pradalier, G. Falchetto, T. Gerbaud, V. Grandgirard, P. Hennequin, R. Sabot, Y. Sarazin, L. Vermare, and R. E. Waltz. Validity of quasi-linear transport model. In Proceedings of the 22nd IAEA Fusion Energy Conference, pp.\ 227, Vienna, Austria, 2008. International Atomic Energy A...

  9. [9]

    Core turbulent transport in tokamak plasmas: bridging theory and experiment with QuaLiKiz

    C Bourdelle, J Citrin, B Baiocchi, A Casati, P Cottier, X Garbet, and F Imbeaux and. Core turbulent transport in tokamak plasmas: bridging theory and experiment with QuaLiKiz . Plasma Physics and Controlled Fusion, 58 0 (1): 0 014036, December 2015. doi:10.1088/0741-3335/58/1/014036

  10. [10]

    Envisioning better benchmarks for machine learning pde solvers

    Johannes Brandstetter. Envisioning better benchmarks for machine learning pde solvers. Nature Machine Intelligence, pp.\ 1--2, 2024

  11. [11]

    Brunton, Bernd R

    Steven L. Brunton, Bernd R. Noack, and Petros Koumoutsakos. Machine learning for fluid mechanics. Annual Review of Fluid Mechanics, 52 0 (Volume 52, 2020): 0 477--508, 2020. ISSN 1545-4479. doi:https://doi.org/10.1146/annurev-fluid-010719-060214

  12. [12]

    o rler, O G\

    J Citrin, C Bourdelle, F J Casson, C Angioni, N Bonanomi, Y Camenen, X Garbet, L Garzotti, T G\" o rler, O G\" u rcan, F Koechl, F Imbeaux, O Linder, K van de Plassche, P Strand, and G Szepesi and. Tractable flux-driven temperature, density, and rotation profile evolution with the quasilinear gyrokinetic transport model QuaLiKiz . Plasma Physics and Contr...

  13. [13]

    Citrin, P

    J. Citrin, P. Trochim, T. Goerler, D. Pfau, K. L. van de Plassche, and F. Jenko. Fast transport simulations with higher-fidelity surrogate models for ITER . Physics of Plasmas, 30 0 (6), jun 2023. doi:10.1063/5.0136752

  14. [14]

    Torax: A fast and differentiable tokamak transport simulator in jax, 2024

    Jonathan Citrin, Ian Goodfellow, Akhil Raju, Jeremy Chen, Jonas Degrave, Craig Donner, Federico Felici, Philippe Hamel, Andrea Huber, Dmitry Nikulin, David Pfau, Brendan Tracey, Martin Riedmiller, and Pushmeet Kohli. Torax: A fast and differentiable tokamak transport simulator in jax, 2024

  15. [15]

    A. M. Dimits, G. Bateman, M. A. Beer, B. I. Cohen, W. Dorland, G. W. Hammett, C. Kim, J. E. Kinsey, M. Kotschenreuther, A. H. Kritz, L. L. Lao, J. Mandrekas, W. M. Nevins, S. E. Parker, A. J. Redd, D. E. Shumaker, R. Sydora, and J. Weiland. Comparisons and physics basis of tokamak transport models and turbulence simulations. Physics of Plasmas, 7 0 (3): 0...

  16. [16]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, ...

  17. [17]

    Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position

    Kunihiko Fukushima. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological cybernetics, 36 0 (4): 0 193--202, 1980

  18. [18]

    Multimodal convolutional neural networks for predicting evolution of gyrokinetic simulations

    Mitsuru Honda, Emi Narita, Shinya Maeyama, and Tomo-Hiko Watanabe. Multimodal convolutional neural networks for predicting evolution of gyrokinetic simulations. Contributions to Plasma Physics, 63 0 (5-6): 0 e202200137, 2023. doi:https://doi.org/10.1002/ctpp.202200137

  19. [19]

    A Hornsby, A

    W. A Hornsby, A. Gray, J. Buchanan, B. S. Patel, D. Kennedy, F. J. Casson, C. M. Roach, M. B. Lykkegaard, H. Nguyen, N. Papadimas, B. Fourcin, and J. Hart. Gaussian process regression models for the properties of micro-tearing modes in spherical tokamaks. Physics of Plasmas, 31 0 (1), jan 2024. ISSN 1089-7674. doi:10.1063/5.0174478

  20. [20]

    Itoh, S.-I

    K. Itoh, S.-I. Itoh, P. H. Diamond, T. S. Hahm, A. Fujisawa, G. R. Tynan, M. Yagi, and Y. Nagashima. Physics of zonal flowsa). Physics of Plasmas, 13 0 (5): 0 055502, 05 2006. ISSN 1070-664X. doi:10.1063/1.2178779

  21. [21]

    Perceiver: General perception with iterative attention

    Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Jo \ a o Carreira. Perceiver: General perception with iterative attention. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event , volume 139 of Proceedings of Machine Learning Resea...

  22. [22]

    Highly accurate protein structure prediction with alphafold

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \' dek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596 0 (7873): 0 583--589, 2021

  23. [23]

    Kiefer, C

    C.K. Kiefer, C. Angioni, G. Tardini, N. Bonanomi, B. Geiger, P. Mantica, T. Pütterich, E. Fable, P.A. Schneider, ASDEX Upgrade Team , EUROfusion MST1 Team , and JET Contributors . Validation of quasi-linear turbulent transport models against plasmas with dominant electron heating for the prediction of iter pfpo-1 plasmas. Nuclear Fusion, 61 0 (6): 0 06603...

  24. [24]

    Swift: Swin 4d fmri transformer

    Peter Kim, Junbeom Kwon, Sunghwan Joo, Sangyoon Bae, Donggyu Lee, Yoonho Jung, Shinjae Yoo, Jiook Cha, and Taesup Moon. Swift: Swin 4d fmri transformer. Advances in Neural Information Processing Systems, 36: 0 42015--42037, 2023

  25. [25]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , 2015

  26. [26]

    Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew M

    Nikola B. Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew M. Stuart, and Anima Anandkumar. Neural operator: Learning maps between function spaces with applications to pdes. J. Mach. Learn. Res., 24: 0 89:1--89:97, 2023

  27. [27]

    John A. Krommes. The gyrokinetic description of microturbulence in magnetized plasmas. Annual Review of Fluid Mechanics, 44 0 (Volume 44, 2012): 0 175--201, 2012. ISSN 1545-4479. doi:https://doi.org/10.1146/annurev-fluid-120710-101223

  28. [28]

    Kumar, Y

    N. Kumar, Y. Camenen, S. Benkadda, C. Bourdelle, A. Loarte, A.R. Polevoi, F. Widmer, and JET contributors. Turbulent transport driven by kinetic ballooning modes in the inner core of jet hybrid h-modes. Nuclear Fusion, 61 0 (3): 0 036005, jan 2021. doi:10.1088/1741-4326/abd09c

  29. [29]

    Fourcastnet: Accelerating global high-resolution weather forecasting using adaptive fourier neural operators

    Thorsten Kurth, Shashank Subramanian, Peter Harrington, Jaideep Pathak, Morteza Mardani, David Hall, Andrea Miele, Karthik Kashinath, and Anima Anandkumar. Fourcastnet: Accelerating global high-resolution weather forecasting using adaptive fourier neural operators. In Proceedings of the Platform for Advanced Scientific Computing Conference, PASC '23, New ...

  30. [30]

    Learning skillful medium-range global weather forecasting

    Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. Learning skillful medium-range global weather forecasting. Scien...

  31. [31]

    Boser, John S

    Yann LeCun, Bernhard E. Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, and Lawrence D. Jackel. Backpropagation Applied to Handwritten Zip Code Recognition . Neural Comput., 1 0 (4): 0 541--551, 1989. doi:10.1162/neco.1989.1.4.541

  32. [32]

    Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M

    Zongyi Li, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M. Stuart, and Anima Anandkumar. Neural operator: Graph kernel network for partial differential equations. CoRR, abs/2003.03485, 2020

  33. [33]

    Fourier neural operator for parametric partial differential equations, 2021

    Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations, 2021

  34. [34]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021 , pp.\ 9992--10002. IEEE , 2021. doi:10.1109/ICCV48922.2021.00986

  35. [35]

    Video swin transformer

    Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pp.\ 3192--3201. IEEE , 2022. doi:10.1109/CVPR52688.2022.00320

  36. [36]

    M. D. McKay, R. J. Beckman, and W. J. Conover. A comparison of three methods for selecting values of input variables in the analysis of output from a computer code. Technometrics, 21 0 (2): 0 239--245, May 1979. ISSN 0040-1706. doi:10.2307/1268522. OSTI 5236110

  37. [37]

    Batzner, Samuel S

    Amil Merchant, Simon L. Batzner, Samuel S. Schoenholz, Muratahan Aykol, Gowoon Cheon, and Ekin Dogus Cubuk. Scaling deep learning for materials discovery. Nat., 624 0 (7990): 0 80--85, 2023. doi:10.1038/S41586-023-06735-9

  38. [38]

    Van Mulders, F

    S. Van Mulders, F. Felici, O. Sauter, J. Citrin, A. Ho, M. Marin, and K.L. van de Plassche. Rapid optimization of stationary tokamak plasmas in RAPTOR : demonstration for the ITER hybrid scenario with neural network surrogate transport model QLKNN . Nuclear Fusion, 61 0 (8): 0 086019, July 2021. doi:10.1088/1741-4326/ac0d12

  39. [39]

    Narita, M

    E. Narita, M. Honda, S. Maeyama, and T.-H. Watanabe. Toward efficient runs of nonlinear gyrokinetic simulations assisted by a convolutional neural network model recognizing wavenumber-space images. Nuclear Fusion, 62 0 (8): 0 086037, jun 2022. doi:10.1088/1741-4326/ac70e8

  40. [40]

    Gupta, and Aditya Grover

    Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K. Gupta, and Aditya Grover. Climax: A foundation model for weather and climate. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , volume...

  41. [41]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 , pp.\ 4172--4182. IEEE , 2023. doi:10.1109/ICCV51070.2023.00387

  42. [42]

    Peeters, Y

    A.G. Peeters, Y. Camenen, F.J. Casson, W.A. Hornsby, A.P. Snodin, D. Strintzi, and G. Szepesi. The nonlinear gyro-kinetic flux tube code gkw. Computer Physics Communications, 180 0 (12): 0 2650--2672, 2009. ISSN 0010-4655. doi:https://doi.org/10.1016/j.cpc.2009.07.001. 40 YEARS OF CPC: A celebratory issue focused on quality software for high performance, ...

  43. [43]

    Courville

    Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. Film: Visual reasoning with a general conditioning layer. In Sheila A. McIlraith and Kilian Q. Weinberger (eds.), Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18),...

  44. [44]

    Pointnet: Deep learning on point sets for 3d classification and segmentation

    Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. arXiv preprint arXiv:1612.00593, 2016

  45. [45]

    Hamprecht, Yoshua Bengio, and Aaron C

    Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron C. Courville. On the spectral bias of neural networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA , volu...

  46. [46]

    Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian processes for machine learning. Adaptive computation and machine learning. MIT Press, 2006. ISBN 026218253X

  47. [47]

    U-net: Convolutional networks for biomedical image segmentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Nassir Navab, Joachim Hornegger, William M. Wells III, and Alejandro F. Frangi (eds.), Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 - 18th International Conference Munich, Germany, October 5 - 9, 2015, Proceed...

  48. [48]

    Simshift: A benchmark for adapting neural surrogates to distribution shifts, 2025

    Paul Setinek, Gianluca Galletti, Thomas Gross, Dominik Schnürer, Johannes Brandstetter, and Werner Zellinger. Simshift: A benchmark for adapting neural surrogates to distribution shifts, 2025

  49. [49]

    Staebler, C

    G. Staebler, C. Bourdelle, J. Citrin, and R. Waltz. Quasilinear theory and modelling of gyrokinetic turbulent transport in tokamaks. Nuclear Fusion, 64 0 (10): 0 103001, sep 2024. doi:10.1088/1741-4326/ad6ba5

  50. [50]

    G. M. Staebler and J. E. Kinsey. Electron collisions in the trapped gyro-landau fluid transport model. Physics of Plasmas, 17 0 (12), dec 2010. ISSN 1089-7674. doi:10.1063/1.3505308

  51. [51]

    G. M. Staebler, J. E. Kinsey, and R. E. Waltz. A theory-based transport model with comprehensive physics. Physics of Plasmas, 14 0 (5), may 2007. ISSN 1089-7674. doi:10.1063/1.2436852

  52. [52]

    Physics-based deep learning

    Nils Thuerey, Philipp Holl, Maximilian M \" u ller, Patrick Schnell, Felix Trost, and Kiwon Um. Physics-based deep learning. CoRR, abs/2109.05237, 2021

  53. [53]

    Factorized fourier neural operators, 2023

    Alasdair Tran, Alexander Mathews, Lexing Xie, and Cheng Soon Ong. Factorized fourier neural operators, 2023

  54. [54]

    The Particle-in-Cell Method, pp.\ 161--189

    David Tskhakaya. The Particle-in-Cell Method, pp.\ 161--189. Springer Berlin Heidelberg, Berlin, Heidelberg, 2008. ISBN 978-3-540-74686-7. doi:10.1007/978-3-540-74686-7_6

  55. [55]

    K. L. van de Plassche, J. Citrin, C. Bourdelle, Y. Camenen, F. J. Casson, V. I. Dagnelie, F. Felici, A. Ho, S. Van Mulders, and JET Contributors. Fast modeling of turbulent transport in fusion plasmas using neural networks. Physics of Plasmas, 27 0 (2): 0 022310, 02 2020. ISSN 1070-664X. doi:10.1063/1.5134126

  56. [56]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 30: Annual Conference o...

  57. [57]

    A high-fidelity surrogate model for the ion temperature gradient (itg) instability using a small expensive simulation dataset

    Chenguang Wan, Youngwoo Cho, Zhisong Qu, Yann Camenen, Robin Varennes, Kyungtak Lim, Kunpeng Li, Jiangang Li, Yanlong Li, and Xavier Garbet. A high-fidelity surrogate model for the ion temperature gradient (itg) instability using a small expensive simulation dataset. Nuclear Fusion, 65 0 (5): 0 054001, apr 2025. doi:10.1088/1741-4326/adc7c9

  58. [58]

    Factorized convolutional neural networks

    Min Wang, Baoyuan Liu, and Hassan Foroosh. Factorized convolutional neural networks. In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), pp.\ 545--553, 2017. doi:10.1109/ICCVW.2017.71

  59. [59]

    Transolver: A fast transformer solver for pdes on general geometries

    Haixu Wu, Huakun Luo, Haowen Wang, Jianmin Wang, and Mingsheng Long. Transolver: A fast transformer solver for pdes on general geometries. In International Conference on Machine Learning, 2024

  60. [60]

    Mattersim: A deep learning atomistic model across elements, temperatures and pressures

    Han Yang, Chenxi Hu, Yichi Zhou, Xixian Liu, Yu Shi, Jielan Li, Guanzhi Li, Zekun Chen, Shuizhou Chen, Claudio Zeni, Matthew Horton, Robert Pinsler, Andrew Fowler, Daniel Zügner, Tian Xie, Jake Smith, Lixin Sun, Qian Wang, Lingyu Kong, Chang Liu, Hongxia Hao, and Ziheng Lu. Mattersim: A deep learning atomistic model across elements, temperatures and press...

  61. [61]

    Zanisi, A

    L. Zanisi, A. Ho, J. Barr, T. Madula, J. Citrin, S. Pamela, J. Buchanan, F.J. Casson, and V. Gopakumar. Efficient training sets for surrogate models of tokamak turbulence with active deep ensembles. Nuclear Fusion, 64 0 (3): 0 036022, February 2024. ISSN 1741-4326. doi:10.1088/1741-4326/ad240d

  62. [62]

    A generative model for inorganic materials design

    Claudio Zeni, Robert Pinsler, Daniel Z \"u gner, Andrew Fowler, Matthew Horton, Xiang Fu, Zilong Wang, Aliaksandra Shysheya, Jonathan Crabb \'e , Shoko Ueda, et al. A generative model for inorganic materials design. Nature, pp.\ 1--3, 2025

  63. [63]

    Hofgard, Aria Mansouri Tehrani, Rui Wang, Ameya Daigavane, Montgomery Bohde, Jerry Kurtin, Qian Huang, Tuong Phung, Minkai Xu, Chaitanya K

    Xuan Zhang, Limei Wang, Jacob Helwig, Youzhi Luo, Cong Fu, Yaochen Xie, Meng Liu, Yuchao Lin, Zhao Xu, Keqiang Yan, Keir Adams, Maurice Weiler, Xiner Li, Tianfan Fu, Yucheng Wang, Haiyang Yu, Yuqing Xie, Xiang Fu, Alex Strasser, Shenglong Xu, Yi Liu, Yuanqi Du, Alexandra Saxton, Hongyi Ling, Hannah Lawrence, Hannes St \" a rk, Shurui Gui, Carl Edwards, Ni...

  64. [64]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.