REVIEW 2 major objections 5 minor 18 references
A transformer network that reads sequences of photonic-lantern core powers recovers low-order wavefronts better than a one-shot convolutional net, raising Strehl and cutting tip/tilt residuals in simulation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 03:31 UTC pith:YUKUJEB2
load-bearing objection Solid open-loop CNN-vs-transformer comparison on simulated HMS-PL data; title overclaims the low-SNR closed-loop problem the experiments never touch. the 2 major comments →
Overcoming the low signal-to-noise problem for hybrid mode-selective photonic lantern-based wavefront correction using machine learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When a hybrid mode-selective photonic lantern is used for simultaneous science injection and wavefront sensing, a transformer network that ingests a short temporal sequence of the sensing-core intensities recovers the dominant low-order wavefront structure more accurately than a convolutional network that sees only a single frame, reducing tip/tilt residual RMS to about 0.05 rad and raising mean Strehl ratio on temporal Von Kármán screens from roughly 0.25 to 0.36.
What carries the argument
The transformer neural network (TNN) that maps a sequence of previous lantern-core power vectors, via multi-head self-attention and positional encoding, onto an estimated pupil-plane wavefront to be applied to an upstream deformable mirror.
Load-bearing premise
Open-loop, noise-free simulations with a fixed lantern transfer matrix and no correction latency are enough to decide which network architecture will work for the real low-SNR closed-loop problem.
What would settle it
Train and run both networks on the same lantern with deliberately reduced sensing-core light (or added photon noise matching the intended science-core allocation) and measure whether the transformer still delivers higher closed-loop Strehl once latency and real deformable-mirror dynamics are included.
If this is right
- Temporal sequence models become the preferred architecture for lantern-based wavefront sensors when residual errors are temporally correlated.
- More light can be reserved for the science core while still obtaining usable tip/tilt and low-order correction, provided the sensing cores are read as a short time series.
- The same TNN approach can be extended to closed-loop predictive control for kernel-nulling instruments that rely on hybrid mode-selective lanterns.
- Network choice (one-shot CNN versus sequential TNN) should be driven by whether the residual wavefront is dominated by random or temporally evolving errors.
Where Pith is reading between the lines
- If the transformer advantage survives real photon noise and latency, hybrid lanterns can operate with even smaller sensing-core fractions than the six-core device simulated here.
- The same sequence model could be trained end-to-end with a differentiable lantern model to co-optimize core geometry and the reconstructor for a target science-core coupling.
- Because both networks act as low-pass filters, higher-order residual power may still require a conventional high-order sensor or a deeper temporal architecture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript compares a convolutional neural network (CNN) and a transformer neural network (TNN) for open-loop wavefront estimation from the five wavefront-sensing cores of a simulated 6-core hybrid mode-selective photonic lantern (HMS-PL), motivated by the Seidr/Asgard instrument at VLTI. Light is propagated from a monochromatic point source through either tip/tilt (100 nm RMS) or Von Kármán phase screens (r0 = 0.4 m), decomposed into LP modes via normalized inner products, and mapped through a fixed finite-difference transfer matrix to core intensities. Both networks are trained on 70k simulated input–output pairs; the TNN additionally receives Mt = 50 temporal frames when sequences are available. On held-out data the networks reduce tip/tilt residual RMS from ~0.5 rad to ~0.05–0.08 rad, and on temporal Von Kármán screens the TNN raises mean Strehl from ~0.25 to ~0.36 (versus ~0.29 for the one-shot CNN). The authors present this as a first step toward the low-SNR closed-loop problem that arises when light is preferentially reserved for the mode-selective science core.
Significance. If the architecture ranking survives realistic noise and closed-loop operation, the work would supply a concrete design path for HMS-PL wavefront sensing in high-contrast interferometry (kernel nulling, Seidr). The simulation pipeline is standard and fully specified (Von Kármán PSD, Zernike tip/tilt, Fraunhofer PSF, transfer matrix), residual RMS and Strehl PDFs are reported against known ground truth, and the temporal advantage of the TNN is shown clearly in Figures 4–6. The explicit framing as an open-loop first step and the provision of inference latencies are useful for the community. The central empirical claim is therefore of genuine instrumental interest, even though the title’s low-SNR closed-loop regime is not yet exercised.
major comments (2)
- Title, abstract and §1 frame the work as addressing the low-SNR trade-off that arises when light is maximized in the mode-selective core. Sections 2–4 and 3.3, however, report only noise-free, open-loop simulations with a fixed transfer matrix and no photon or detector noise; latency and closed-loop effects are explicitly deferred to “ongoing work.” The architecture ranking (TNN > CNN on temporal data) is therefore demonstrated only in a high-SNR regime and does not yet constitute evidence that the same ranking, or any useful correction, survives the low-SNR conditions that motivate the paper. At minimum the manuscript should either (i) add a controlled noise study that scales the five WFS-core intensities and injects realistic photon/read noise, or (ii) substantially revise the title/abstract framing so that the claimed contribution matches the open-loop, noise-free experiments actually
- Section 4.2 and Figure 5 report mean Strehl rising only from ~0.25 to ~0.36 (TNN) or ~0.29 (CNN) on Von Kármán screens. Both networks act as low-pass filters (Figure 6). Without a quantitative link between residual Strehl and mode-selective core injection efficiency (or kernel-nulling contrast) for Seidr, it remains unclear whether the demonstrated correction is instrumentally useful. A short calculation or simulation of residual light in the LP01 core after the estimated correction would make the practical significance of the ~0.1 Strehl gain concrete.
minor comments (5)
- Equation (1): the normalization of the inner product is written with a product of two L2 norms under a single square root; a brief statement that the LP modes are taken as already normalized (or an explicit check) would remove ambiguity.
- Table 1 and Figure 1: the transfer matrix is central; stating the wavelength, mesh resolution and any validation against analytic LP modes would strengthen reproducibility.
- Section 3.3: inference latencies are given for an RTX 4090 without TensorRT; a one-sentence note on expected closed-loop frame rates for Seidr would help readers judge real-time feasibility.
- Figure 4 caption and text: residual RMS values are quoted to two decimals; adding the standard deviation of the residual (or a box-plot) would make the scatter visible.
- Typographical: “Von Kármán” is inconsistently accented; “Maréchal” appears as “Maréchal” and “Maréchal”; a few missing spaces after periods appear in the reference list.
Circularity Check
No circularity: NN performance metrics are computed on held-out simulated test sets against independent ground-truth wavefronts; self-citations supply only instrument background.
full rationale
The paper's load-bearing claims are empirical comparisons of CNN vs TNN residual RMS (tip/tilt) and Strehl ratio (Von Kármán) on 15 000 held-out test instances generated from the same forward model used for training (Eq. 1 + fixed transfer matrix + Zernike/Von Kármán screens). Training minimizes MSE on known input–output pairs; evaluation reports residual error against the known true wavefronts, not a quantity defined by the fit. No parameter is fitted to a subset and then re-presented as a prediction of a closely related observable. Self-citations (Norris et al. 2020/2022, Taras et al. 2024, etc.) establish the HMS-PL concept and Seidr context but are not invoked as uniqueness theorems or as the sole justification for the architecture ranking. The acknowledged open-loop, noise-free limitation is a scope gap, not a circular reduction. The derivation chain is therefore self-contained simulation + supervised learning + independent test-set metrics.
Axiom & Free-Parameter Ledger
free parameters (5)
- tip/tilt RMS amplitude =
100 nm
- Fried parameter r0 and outer scale L0 =
r0=0.4 m, L0=10 m
- temporal sequence length Mt =
50
- network hyperparameters (n_filt, n_dense, d_m, n_heads, dropout, etc.) =
see Tables 2–3
- wind speed for frozen-flow sequences =
10 m/s
axioms (5)
- domain assumption Fraunhofer diffraction and normalized inner-product decomposition into LP modes accurately map pupil-plane phase to lantern input coefficients.
- domain assumption A fixed, pre-computed 6-core transfer matrix fully describes the lantern’s linear mapping from LP modes to core fields.
- domain assumption Post-upstream-AO residuals at Seidr are dominated by tip/tilt that can be represented by the n=1 Zernike terms.
- ad hoc to paper Open-loop residual RMS and Strehl on noise-free data are informative proxies for closed-loop low-SNR performance.
- standard math Mean-squared-error training with Adam yields networks whose test-set residuals can be compared fairly.
read the original abstract
Hybrid mode-selective photonic lanterns transform an input complex point-spread function into several single-mode outputs, where a selected core feeds the fundamental mode to a photonic science instrument, while the remaining cores are used for wavefront sensing in a closed-loop adaptive optics system. A neural network maps the intensities of the wavefront sensing cores to an estimated wavefront correction, which is applied to an upstream deformable mirror. However, there exists a trade between maximizing the amount of light reserved for the photonic instrument and the reduced signal-to-noise ratios for the wavefront sensing cores. We explore wavefront correction for the Seidr instrument, a part of the Asgard Suite for the Very Large Telescope Interferometer. We evaluate different neural network architectures, comparing wavefront estimation performance for different wavefront error types, as a first step toward addressing the signal-to-noise trade-off.
Figures
Reference graph
Works this paper leans on
-
[1]
An all-photonic focal-plane wavefront sensor,
Norris, B. R., Wei, J., Betters, C. H., Wong, A., and Leon-Saval, S. G., “An all-photonic focal-plane wavefront sensor,”Nature Communications11(1), 5335 (2020)
2020
-
[2]
Optimal broad- band starlight injection into a single-mode fibre with integrated photonic wavefront sensing,
Norris, B., Betters, C., Wei, J., Yerolatsitis, S., Amezcua-Correa, R., and Leon-Saval, S., “Optimal broad- band starlight injection into a single-mode fibre with integrated photonic wavefront sensing,”Optics Ex- press30(19), 34908–34917 (2022)
2022
-
[3]
Kernel nulling at VLTI with photonic lanterns for optimal fibre injection,
Taras, A. K., Norris, B., Chhabra, S., Cvetojevic, N., Foriel, V., Ireland, M., Kraus, S., Leon-Saval, S., Martinache, F., Paul, J., et al., “Kernel nulling at VLTI with photonic lanterns for optimal fibre injection,” in [Optical and Infrared Interferometry and Imaging IX],13095, 242–250, SPIE (2024)
2024
-
[4]
Seidr update: photonic ‘black magic’ for high-contrast interferometry using kernel-nulling and photonic lanterns,
Long, N. K., Dahl, D. S., Betters, C. H., Bryant, J. J., Cvetojevic, N., Ireland, M. J., Kraus, S., Leon-Saval, S., Martinache, F., Martinod, M.-A., Norris, B., Paul, J., Rodziewicz-Ryan, A., Taras, A. K., Wei, J., and Tuthill, P. G., “Seidr update: photonic ‘black magic’ for high-contrast interferometry using kernel-nulling and photonic lanterns,” in [Op...
2026
-
[5]
High-angular resolution and high contrast observations from Y to L band at the Very Large Telescope Interferometer with the Asgard Instrumental suite,
Martinod, M.-A., Defr` ere, D., Ireland, M., Kraus, S., Martinache, F., Tuthill, P., Bigioli, A., Bouzerand, E., Bryant, J., Chhabra, S., et al., “High-angular resolution and high contrast observations from Y to L band at the Very Large Telescope Interferometer with the Asgard Instrumental suite,”Journal of Astronomical Telescopes, Instruments, and System...
2023
-
[6]
Heimdallr, Baldr, and Solarstein: designing the next generation of VLTI instruments in the Asgard suite,
Taras, A. K., Robertson, J. G., Allouche, F., Courtney-Barrer, B., Carter, J., Crous, F., Cvetojevic, N., Ireland, M., Lagarde, S., Martinache, F., et al., “Heimdallr, Baldr, and Solarstein: designing the next generation of VLTI instruments in the Asgard suite,”Applied Optics63(14), D41–D49 (2024)
2024
-
[7]
Baldr: a Zernike wavefront sensor for VLTI/Asgard,
Courtney-Barrer, B., Robertson, G., Taras, A., Bernard, J. T., McGuinness, G., Crous, F., Tuthill, P., N’Diaye, M., Langford, C., Cvetojevic, N., et al., “Baldr: a Zernike wavefront sensor for VLTI/Asgard,” in [Adaptive Optics Systems IX],13097, 362–375, SPIE (2024)
2024
-
[8]
Asgard/NOTT: L-band nulling interferometry at the VLTI. II. Warm optical design and injection system,
Garreau, G., Bigioli, A., Laugier, R., Raskin, G., Morren, J., Berger, J.-P., Dandumont, C., Goldsmith, H.- D. K., Gross, S., Ireland, M., et al., “Asgard/NOTT: L-band nulling interferometry at the VLTI. II. Warm optical design and injection system,”Journal of Astronomical Telescopes, Instruments, and Systems10(1), 015002–015002 (2024)
2024
-
[9]
L., [Beam propagation method for design of optical waveguide devices], John Wiley & Sons (2015)
Pedrola, G. L., [Beam propagation method for design of optical waveguide devices], John Wiley & Sons (2015)
2015
-
[10]
Efficient nonuniform schemes for paraxial and wide-angle finite-difference beam propagation methods,
Shibayama, J., Matsubara, K., Sekiguchi, M., Yamauchi, J., and Nakano, H., “Efficient nonuniform schemes for paraxial and wide-angle finite-difference beam propagation methods,”Journal of Lightwave Technol- ogy17(4), 677–683 (1999)
1999
-
[11]
D., [Numerical simulation of optical wave propagation with examples in MATLAB], SPIE, Bellingham, Wash
Schmidt, J. D., [Numerical simulation of optical wave propagation with examples in MATLAB], SPIE, Bellingham, Wash. (2010)
2010
-
[12]
The spectrum of turbulence,
Taylor, G. I., “The spectrum of turbulence,”Proceedings of the Royal Society of London. Series A- Mathematical and Physical Sciences164(919), 476–490 (1938)
1938
-
[13]
Zernike polynomials and atmospheric turbulence,
Noll, R. J., “Zernike polynomials and atmospheric turbulence,”Journal of the Optical Society of Amer- ica66(3), 207–211 (1976)
1976
-
[14]
Learning the lantern: neural network applications to broadband photonic lantern modeling,
Sweeney, D., Norris, B. R., Tuthill, P., Scalzo, R., Wei, J., Betters, C. H., and Leon-Saval, S. G., “Learning the lantern: neural network applications to broadband photonic lantern modeling,”Journal of Astronomical Telescopes, Instruments, and Systems7(2), 028007 (2021)
2021
-
[15]
ImageNet classification with deep convolutional neural networks,
Krizhevsky, A., Sutskever, I., and Hinton, G. E., “ImageNet classification with deep convolutional neural networks,” in [Advances in Neural Information Processing Systems], Pereira, F., Burges, C., Bottou, L., and Weinberger, K., eds.,25, Curran Associates, Inc. (2012)
2012
-
[16]
Attention is all you need,
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I., “Attention is all you need,”Advances in Neural Information Processing Systems30(2017)
2017
-
[17]
Adam: A method for stochastic optimization,
Kingma, D. P. and Ba, J., “Adam: A method for stochastic optimization,” in [3rd International Conference on Learning Representations],arXiv preprint arXiv:1412.6980(2015)
Pith/arXiv arXiv 2015
-
[18]
Strehl ratio for primary aberrations in terms of their aberration variance,
Mahajan, V. N., “Strehl ratio for primary aberrations in terms of their aberration variance,”Journal of the Optical Society of America73(6), 860–861 (1983)
1983
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.