REVIEW 4 major objections 5 minor 38 references
A neural network approach for line detection in complex atomic emission spectra measured by high-resolution Fourier transform spectroscopy
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A neural network trained on simulated Fourier spectra can detect weak, blended, and ringing-distorted emission lines that standard peak finding misses, and its new detections identify two previously unidentifiable Ni II levels.
desk verdict A serious, well-scoped ML-for-spectroscopy paper whose core claim survives contact with the stress-test note: the Ni II levels are anchored by Ritz-consistent human lines, not just 3-S/N features, but the missing code/data and lack of an FDR estimate for the extra NN lines should be fixed before the method becomes a standard tool. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the pointwise class-1 probability $p_1$ produced by a two-layer bidirectional LSTM that encodes overlapping spectrum chunks, concatenates the normalised wavenumber to the LSTM output, and decodes through a fully connected network with PReLU activations. The training data are simulated FT spectra: line wavenumbers are uniform random at density $\rho_\text{line}=10\rho_{10}$; Doppler and pressure widths are drawn from distributions fitted to the experimental spectrum; $S/N$s follow a linear log-$S/N$ extrapolation from line counts above $S/N=10$; and each simulated line is convolved with the FT instrumental profile by multiplying the interferogram with the apodisation function $A(x)$ of Eq. 2, including finite aperture, zero-padding, and cosine-bell apodisation. Focal loss (Eq. 3) makes training focus on the rare line-centre points and hard cases. This machinery lets the network learn line shapes, blends, noise, and resolution ringing from known ground truth, so the model can flag weak lines in the real spectrum without inheriting the incompleteness of human line lists.
What would settle it
Take the nine Ni spectra, inject synthetic Voigt lines with known wavenumbers and $S/N$ between 2 and 6 into quiet regions and into blend and ringing regions, train the pipeline exactly as described, and measure recovery at the known wavenumbers. If the model misses most injected lines below $S/N=5$, or flags ringing artefacts as lines at a rate comparable to the old peak finder, the simulation-fidelity premise is falsified; this is a test the paper's data and code can run directly.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that a bidirectional LSTM-FCNN, trained per experimental spectrum on simulations whose line positions, $S/N$ distributions, and instrumental profiles are tuned to that spectrum, transfers to experimental data and outperforms the conventional peak-finding routines on exactly the cases that defeat them. In nine Ni-He hollow-cathode FT spectra covering 1800–70,000 cm$^{-1}$ the model listed 13,400 lines versus 6,700 in the human list; in the dense Nd-Ar Penning spectrum it listed 9,400 versus 7,000. In both cases about 95% of the human lines were matched within 0.05 cm$^{-1}$, and the unmatched remainder were predominantly weak lines and centre-of-gravity wavenumbers from isotope and hyperfine structure. Four lines found only by the model—three of them at $S/N = 3$—removed the ambiguity in assigning two Ni II levels, $3d^8(^3F_4)6f[2]_{3/2}$ at $134,261.8946 \pm 0.0081$ cm$^{-1}$ and $3d^8(^3F_4)6f[1]_{3/2}$ at $134,249.5264 \pm 0.0054$ cm$^{-1}$, which the previous analysis had concluded were unidentifiable from the Ni spectra.
Load-bearing premise
The model learns only what the simulated spectra contain, so the load-bearing premise is that spectra built from uniform random line positions, an $S/N$ distribution extrapolated linearly from lines above $S/N=10$, Gaussian white noise, $S/N_\min=3$, a 20% blend-masking rule, and a line density of $10\rho_{10}$ are faithful enough to real FT spectra that a network trained on them detects real lines correctly. The authors themselves note in Section 6 that unmodelled isotope shifts, hyperfine structure, phase errors, and self-absorption limit the approach.
Editorial extensions
If this is right
- About 95% of the lines in the human-produced Ni and Nd line lists are recovered by the model within 0.05 cm$^{-1}$, so existing analyses carry over and only need verification.
- Line lists become substantially more complete: for the nine Ni spectra the model detected 13,400 lines against 6,700 in the human list, and for the dense Nd spectrum 9,400 against 7,000, with the extra lines mostly weak or blended.
- The $p_1$ probability curve acts as a de-noised spectrum in which weak and blended features appear as peaks, providing a new visual aid for manual fitting and for decisions about which lines are real.
- Blended components that the conventional peak finder misses, such as the Nd III line at 26,140.225 cm$^{-1}$ in the wing of a strong line, are detected, and resolution-limited ringing is learned so that real lines are not confused with ringing maxima.
- Four weak lines visible only in the neural-network line list make the identification of two Ni II levels unambiguous: $3d^8(^3F_4)6f[2]_{3/2}$ at 134,261.8946 cm$^{-1}$ and $3d^8(^3F_4)6f[1]_{3/2}$ at 134,249.5264 cm$^{-1}$.
Reading between the lines
- An injection-recovery test on the same experimental spectra—adding synthetic lines of known wavenumber and $S/N$ near the detection limit and measuring recovery—would give a ground-truth measure of performance that the paper's comparison with human lists cannot provide.
- The main systematic gap is isotope and hyperfine structure: the model misses most of the human-list centre-of-gravity wavenumbers, so extending the simulator to generate such components is the most direct route to closing that gap.
- If per-spectrum training on simulations becomes routine, the long-term bottleneck shifts from months of manual line identification to accurate simulation of the light source and instrument, since the model can only learn what is put into the simulator.
- A testable extension of the paper's stated expectation is to add self-absorption and phase-error models to the simulator and check whether the same architecture then detects lines in optically thick or infrared spectra without extra false positives.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents a supervised line-detection method for high-resolution Fourier transform (FT) atomic emission spectra. A bidirectional LSTM encodes spectrum chunks and a fully connected decoder classifies each point as a line centre or not. Training spectra are simulated from per-experiment estimates of line widths, line densities, S/N distributions, noise, and the instrumental line profile. Models are trained separately for each experimental spectrum and evaluated on nine Ni-He hollow-cathode spectra and one Nd-Ar Penning-discharge spectrum. The NN lists recover 96% of the human-made Ni lines and 94% of the human-made Nd lines within a 0.05 cm^-1 tolerance, while adding thousands of additional candidates, 1,900 of which match known Ni II Ritz wavenumbers. A following energy-level analysis assigns two Ni II 3d^8(3F4)6f levels, with several supporting lines detected only by the NN at S/N = 3. The paper also reports improved detection of blends and of lines masked by instrumental resolution-limited ringing.
Significance. If validated, the method would be a valuable tool for line-list compilation in complex d- and f-shell spectra, potentially reducing months of manual effort and improving completeness near the noise limit. The use of simulated training data with known labels avoids the circularity of training on imperfect human line lists, and the evaluation across multiple instruments, elements, and line densities is a genuine strength. The energy-level demonstration is, however, the weakest link: the newly claimed levels rest on a small number of S/N = 3 detections with no false-discovery control, and the authors themselves state in §5.1 that many of the 1,900 Ritz matches 'could be chance coincidences.' The paper also makes no public code or data available, which limits independent verification of the simulation pipeline.
major comments (4)
- [§5.1, Table 2] The claim that the two Ni II levels are 'confidently identified' is not supported by a statistical control. The paper reports 6,700 NN-only Ni lines, of which 1,900 match known Ni II Ritz wavenumbers, but it explicitly states that many of these matches could be chance coincidences and gives no estimate of the expected number of chance matches within the 0.05 cm^-1 tolerance. Table 2 then uses three NN-only lines at S/N = 3, exactly the simulation threshold S/N_min, together with a few human-detected lines, to define the levels. With an unstated number of candidate upper levels searched, the probability that random 3-S/N features align into a consistent three-line pattern is not quantified. Please add a chance-coincidence or false-discovery-rate control (e.g., Monte Carlo searches over random level candidates, or matching against shifted Ritz lists) and report the number of level searches performed.
- [§4.3, §5.2] No experimental false-positive rate is measured for the NN output. The precision on simulated test data is 0.1–0.5 at p_th = 0.5, which the paper interprets as adjacent-point duplicates; whether isolated noise features are also produced is not reported. The comparison with Xgremlin at smin = 3 in §5.2 shows total line counts (1,300 vs 13,600 for the NiHeEH spectrum), but total counts do not separate true weak lines from false positives. A fixed false-positive operating point, or a labeled validation set derived from independent energy-level analyses, is needed to support the conclusion in §6 that the NN 'decrease[s] the number of false line detections.'
- [§3.3, §3.6] The simulation pipeline depends on several hand-tuned parameters: rho_line = 10 x rho_10, S/N_min = 3, the 20% blend-masking ratio, and p_th = 0.5 in postprocessing. Because each model is retrained on a spectrum simulated from parameters estimated from that same experimental spectrum, the reported weak-line gains could be sensitive to these choices. I ask for a sensitivity analysis—for example, varying rho_line by factors of 0.5 and 2 and S/N_min between 2 and 5—while reporting human-list recall and Ritz-confirmation counts at each setting. This would show whether the central result is robust or an artifact of the chosen simulation thresholds.
- [§5.4] The resolution-ringing evaluation is performed on a spectrum artificially degraded by reducing the maximum OPD by a factor of four, not on an independently measured spectrum with native resolution-limited ringing. Since the central novelty includes handling instrumental ringing, a demonstration on a real under-resolved archive spectrum (e.g., a Kitt Peak spectrum) would make the claim more general. As it stands, the section establishes that the model can learn the simulated apodization model, but not necessarily that it will reject unmodeled phase and profile errors in real ringing.
minor comments (5)
- [Figs. 2, 7] Axis labels such as 'Wavenumber (cm 1)' should read 'cm^-1'; the missing superscript appears in several figures.
- [Eq. (2)] Please define the normalization of the sinc function explicitly; as written, sinc(y) = sin(y)/y is used with arguments involving sigma*Omega*x/2 and pi, but the reader cannot verify the aperture factor without a stated convention.
- [§4.3] The sentence on precision states that false positives 'almost always occurred adjacent to true positives'; this is important, but the subsequent claim that precision is 'partially representative of uncertainties' should be quantified, for example by reporting the fraction of false positives that lie within one line width of a true positive.
- [§5.1] The rounding rule ('all numbers of lines quoted above 1000 are rounded to the nearest 100') conflicts with quoted values such as 1,900 and 6,700; please state which numbers are exact and which are rounded.
- [Data and code availability] The data/code availability statement promises GitHub availability 'after publication' but gives no repository identifier; for a methods paper, please include a permanent archive such as a Zenodo DOI or provide full code and data as supplementary material.
Circularity Check
No significant circularity: the network detections and Ni II level identifications are not equivalent to the simulation or fitting inputs.
full rationale
The central pipeline is not circular by construction. The LSTM-FCNN is trained on simulated spectra with randomly sampled line positions (Section 3.3) and known labels; the experimental spectrum is then fed to the trained network, so the detected line positions are outputs of a learned map, not values that were used to define the training target. Per-spectrum simulation parameters (Gw/Lw, rho10, S/N distribution) are estimated from the same experimental spectra (Sections 3.2-3.3), introducing a mild domain-adaptation dependence, but this does not make detection of an individual line equivalent to an input: the network cannot reproduce an experimental line unless that line is present in the input spectrum. The two Ni II levels (Table 2) are optimized with LOPT from NN-detected transition wavenumbers together with previously known lower levels; the level energy is an output of the fit, and the lower levels come from independent prior analyses [7,8]. The residuals in Table 2 are fit residuals, not independent predictions, but the identification claim rests on consistency across multiple transitions rather than on a definitional identity. The authors explicitly flag the main non-circular weakness: 'many of them could be chance coincidences' (Section 5.1), i.e. false positives among weak Ritz matches are unquantified; that is a validation/statistics concern, not a circularity. Self-citations [7,8,11,12] supply spectra and prior line lists but do not justify the NN method or the level identification through an unverified uniqueness claim.
Assumptions & free parameters
free parameters (6)
- Line density multiplier rho_line = 10 x rho_10 =
10 x rho_10, with rho_10 estimated from each spectrum
- S/N_min detection threshold =
3
- Blend masking ratio =
20%
- Probability threshold pth =
0.5
- Doppler width scatter sigma =
0.3
- Wavenumber match tolerance =
0.05 cm^-1
assumptions (5)
- domain assumption Spectral lines in the experimental FT spectra are well approximated by Voigt profiles convolved with the instrumental line profile of Eq. 2.
- domain assumption Spectra simulated with uniformly random line positions, a power-law-like S/N distribution, and Gaussian white noise capture the line-detection-relevant statistics of real spectra.
- ad hoc to paper The 20% blend masking threshold and S/N_min=3 reflect human detectability and therefore produce valid training labels.
- domain assumption Human-made line lists and Ritz wavenumber lists are appropriate evaluation ground truth at the 0.05 cm^-1 tolerance.
- domain assumption The chosen BiLSTM-FCNN architecture and focal loss can represent the line detection mapping from simulated spectra.
Cite this review
Pith. "Pith review of A neural network approach for line detection in complex atomic emission spectra measured by high-resolution Fourier transform spectroscopy." pith.science (2026). https://pith.science/paper/OUKNMFRY
@misc{pith2026250112276,
author = {Pith},
title = {Pith review of: A neural network approach for line detection in complex atomic emission spectra measured by high-resolution Fourier transform spectroscopy},
year = {2026},
howpublished = {\url{https://pith.science/paper/OUKNMFRY}},
note = {Machine review of arXiv:2501.12276}
}
abstract
The atomic spectra and structure of the open d- and f-shell elements are extremely complex, where tens of thousands of transitions between fine structure energy levels can be observed as spectral lines across the infrared and UV per species. Energy level quantum properties and transition wavenumbers of these elements underpins almost all spectroscopic plasma diagnostic investigations, with prominent demands from astronomy and fusion research. Despite their importance, these fundamental data are incomplete for many species. A major limitation for the analyses of emission spectra of the open d- and f-shell elements is the amount of time and human resource required to extract transition wavenumbers and intensities from the spectra. Here, the spectral line detection problem is approached by encoding the spectrum point-wise using bidirectional Long Short-Term Memory networks, where transition wavenumber positions are decoded by a fully connected neural network. The model was trained using simulated atomic spectra and evaluated against experimental Fourier transform spectra of Ni ($Z=28$) covering 1800-70,000 cm$^{-1}$ (5555-143 nm) and Nd ($Z=60$) covering 25,369-32,485 cm$^{-1}$ (394-308 nm), measured under a variety of experimental set-ups. Improvements over conventional methods in line detection were evident, particularly for spectral lines that are noisy, blended, and/or distorted by instrumental spectral resolution-limited ringing. In evaluating model performance, a brief energy level analysis of Ni II using lines newly detected by the neural networks has led to the confident identification of two Ni II levels, $3\text{d}^8$$(^3\text{F}_4)6\text{f} [2]_{3/2}$ at 134,261.8946 $\pm$ 0.0081 cm$^{-1}$ and $3\text{d}^8$$(^3\text{F}_4)6\text{f} [1]_{3/2}$ at 134,249.5264 $\pm$ 0.0054 cm$^{-1}$, previously concluded to be unidentifiable using previously analysed Ni spectra.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Johansson S 1996 Physica Scripta 1996 7
work page 1996
-
[2]
von Hellermann M G, Bertschinger G, Biel W, Giroud C, Jaspers R, Jup´ en C, Marchuk O, O’Mullane M, Summers H, Whiteford A et al. 2005 Physica Scripta 2005 19
work page 2005
-
[3]
Atkins J F and Baranov P V 2013 Nature Editorial 503 437
work page 2013
-
[4]
2021 Astronomy & Astrophysics645 A106
Heiter U, Lind K, Bergemann M, Asplund M, Mikolaitis ˇS, Barklem P S, Masseron T, de Laverny P, Magrini L, Edvardsson B et al. 2021 Astronomy & Astrophysics645 A106
work page 2021
-
[5]
Cowan J J, Sneden C, Lawler J E, Aprahamian A, Wiescher M, Langanke K, Mart ´ ınez-Pinedo G and Thielemann F K 2021 Reviews of Modern Physics93 015002
work page 2021
-
[6]
Cowan R D 1981 The Theory of Atomic Structure and Spectra(Berkeley and Los Angeles, CA, USA: Univ. of California Press)
work page 1981
-
[7]
Clear C P, Pickering J C, Nave G, Uylings P and Raassen T 2022 The Astrophysical Journal Supplement Series 261 35
work page 2022
-
[8]
Clear C P, Pickering J C, Nave G, Uylings P and Raassen T 2023 The Astrophysical Journal Supplement Series 269 36 A neural network approach for FT atomic spectra line detection 24
work page 2023
Show all 38 references
-
[9]
Lawler J E, Schmidt J R and Den Hartog E A 2022 Journal of Quantitative Radiative Transfer 289 108283
2022
-
[10]
Lawler J E, Schmidt J R and Den Hartog E A 2022 The Astrophysical Journal Supplement Series 258 27
2022
-
[11]
Ding M, Ryabtsev A N, Kononov E Y, Ryabchikova T, Clear C P, Concepcion F and Pickering J C 2024 Astronomy & Astrophysics684 A149
2024
-
[12]
Ding M, Ryabtsev A N, Kononov E Y, Ryabchikova T and Pickering J C 2024 Astronomy & Astrophysics 692 A33
2024
-
[13]
Nave G, Griesmann U, Brault J and Abrams M 2015 Xgremlin: Interferograms and spectra from Fourier transform spectrometers analysis Astrophysics Source Code Library, record ascl:1511.004
2015
-
[14]
Li D W, Hansen A L, Yuan C, Bruschweiler-Li L and Br¨ uschweiler R 2021Nature Communications 12 5229
-
[15]
2023 Journal of Magnetic Resonance347 107357
Schmid N, Bruderer S, Paruzzo F, Fischetti G, Toscano G, Graf D, Fey M, Henrici A, Ziebart V, Heitmann B et al. 2023 Journal of Magnetic Resonance347 107357
2023
-
[16]
Hochreiter S and Schmidhuber J 1997 Neural Computation 9 1735–1780
1997
-
[17]
Lin T Y, Goyal P, Girshick R, He K and Doll´ ar P 2017 Focal loss for dense object detection 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italypp 2999–3007
2017
-
[18]
Davis S P, Abrams M C and Brault J W 2001 Fourier Transform Spectrometry(San Diego, CA, USA: Academic Press)
2001
-
[19]
Thorne A, Litz´ en U and Johansson S 1999 Spectrophysics: Principles and Applications(Berlin, Germany: Springer)
1999
-
[20]
Danzmann K, G¨ unther M, Fischer J, Kock M and K¨ uhne M 1988Applied Optics 27 4947–4951
-
[21]
Finley D S, Jelinsky P, Bowyer S and Malina R F 1986 A penning discharge source for extreme ultraviolet calibration SPIE 30th Annual Technical Symposium, X-Ray Calibration: Techniques, Sources, and Detectors, San Diego, CA, USAed Rockett P D and Lee P pp 6–10
1986
-
[22]
Heise C, Hollandt J, Kling R, Kock M and K¨ uhne M 1994 Applied Optics 33 5111–5117
1994
-
[23]
Huber P J 1964 The Annals of Mathematical Statistics35 73 – 101
1964
-
[24]
Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M and Duchesnay E 2011 Journal of Machine Learning Research12 2825–2830
2011
-
[25]
Rosenblatt M 1956 The Annals of Mathematical Statistics27 832–837
1956
-
[26]
Parzen E 1962 The Annals of Mathematical Statistics33 1065–1076
1962
-
[27]
Learner R 1982 Journal of Physics B: Atomic and Molecular Physics15 L891
1982
-
[28]
Learner R C, Thorne A P, Wynne-Jones I, Brault J W and Abrams M C 1995 Journal of the Optical Society of America A12 2165–2171
1995
-
[29]
Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, Kopf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J and Chintala S 2019 Pytorch: An imperative style, high-performance deep lear...
2019
-
[30]
He K, Zhang X, Ren S and Sun J 2015 Delving deep into rectifiers: Surpassing human-level performance on imagenet classification 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chilepp 1026–1034
2015
-
[31]
Kingma D P and Ba J 2015 Adam: A method for stochastic optimization 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USAed Bengio Y and LeCun Y
2015
-
[32]
Kurucz R L 2017 Canadian Journal of Physics95 825–827
2017
-
[33]
Blaise J, Wyart J F, Hoekstra R and Kruiver P 1971 Journal of the Optical Society of America 61 1335–1342
1971
-
[34]
Blaise J, Wyart J F, Djerad M T and Ahmed Z B 1984 Physica Scripta 29 119 A neural network approach for FT atomic spectra line detection 25
1984
-
[35]
Kramida A 2011 Computer Physics Communications182 419–434
2011
-
[36]
Section A, Physics and Chemistry 74 801
Shenstone A 1970 Journal of Research of the National Bureau of Standards. Section A, Physics and Chemistry 74 801
1970
-
[37]
Hill F, Branston D and Erdwurm W 1997 The National Solar Observatory Digital Library American Astronomical Society Solar Physics Division Meeting 28pp 2–72
1997
-
[38]
National Solar Observatory (US) 2024 NSO Historical Archive: McMath-Pierce Facility accessed: [December 2024] URL https://nispdata.nso.edu/ftp/FTS_cdrom/
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.