{"id":"76ee446a-0928-4b5d-bd2e-099c8a7632ef","arxiv_id":"2501.12276","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A neural network trained only on simulated spectra detects weak, blended, and ringing-distorted lines in experimental FT atomic spectra, recovering roughly 95% of human-found lines and enabling two new Ni II level identifications.","lead":"This paper trains a bidirectional LSTM neural network on simulated spectra, then uses it to locate spectral lines in high-resolution Fourier transform emission spectra of nickel and neodymium. The network recovered about 95 percent of the lines humans had found, found many additional weak lines, and enabled identification of two previously unconfirmed Ni II energy levels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed physical payoff rests on unvalidated weak-line detections: no false-discovery-rate estimate is given for the newly added S/N≈3 lines, so the 'confident' Ni II level identifications may be chance alignments.","rationale":"I retain the reader's CONDITIONAL verdict: the method is plausible and the evaluation is serious, but the physical payoff requires an independent false-positive control that the paper does not provide. The reader's weakest_assumption concerns simulation fidelity; my concern is adjacent but distinct. Even if the simulations are faithful, the paper's real-spectrum evaluation lacks a false-discovery-rate estimate, and the two newly identified levels rest on a handful of S/N=3 lines whose chance alignment probability is not assessed. This does not warrant rejection, because the paper presents direct comparisons with Xgremlin, high recall against human lists, and a plausible energy-level analysis. It does warrant a concrete statistical validation before the 'confident identification' claim is accepted as established.","tokens_in":17750,"tokens_out":5731,"duration_ms":64281,"concrete_test":"Run a Monte Carlo false-line control for the Ni II level search. Take the nine NN line lists, keep all human lines, and replace the NN-only detections with the same number of randomly placed lines at the same measured S/N and wavenumber distribution. Repeat the exact level-search/LOPT procedure used in §5.2 many times (e.g. 1000 resamples), recording how often a candidate level with ≥3 supporting lines and residuals within the tabulated ±0.0x cm^-1 tolerance appears. If spurious levels occur at comparable rate, the claimed confidence is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: the detector improves on Xgremlin, and the detections enabled confident identification of two Ni II levels. The second part is load-bearing: if the newly detected weak lines used for those levels are mostly noise or chance coincidences, the headline improvement is not established because more detections without a false-positive bound is not necessarily better. The paper's only large-scale validation is overlap with human lists (96% recall) and Ritz matching of 1,900 of the ~6,700 extra NN lines. The authors themselves state that 'many of them could be chance coincidences' (§5.1), yet no false-discovery-rate calculation is provided. The two claimed levels then depend on four NN-only lines with S/N = 3 (Table 2), exactly at the simulation threshold S/N_min = 3, in a search over a large set of previously unidentified upper levels. With thousands of candidate levels and a dense Ritz line list, the probability that a few random 3-S/N features align to form a 'level' is not quantified. Without such a control, 'confident identification' is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a supervised line-detection method for high-resolution Fourier transform (FT) atomic emission spectra. A bidirectional LSTM encodes spectrum chunks and a fully connected decoder classifies each point as a line centre or not. Training spectra are simulated from per-experiment estimates of line widths, line densities, S/N distributions, noise, and the instrumental line profile. Models are trained separately for each experimental spectrum and evaluated on nine Ni-He hollow-cathode spectra and one Nd-Ar Penning-discharge spectrum. The NN lists recover 96% of the human-made Ni lines and 94% of the human-made Nd lines within a 0.05 cm^-1 tolerance, while adding thousands of additional candidates, 1,900 of which match known Ni II Ritz wavenumbers. A following energy-level analysis assigns two Ni II 3d^8(3F4)6f levels, with several supporting lines detected only by the NN at S/N = 3. The paper also reports improved detection of blends and of lines masked by instrumental resolution-limited ringing.","tokens_in":18022,"tokens_out":6335,"duration_ms":65705,"significance":"If validated, the method would be a valuable tool for line-list compilation in complex d- and f-shell spectra, potentially reducing months of manual effort and improving completeness near the noise limit. The use of simulated training data with known labels avoids the circularity of training on imperfect human line lists, and the evaluation across multiple instruments, elements, and line densities is a genuine strength. The energy-level demonstration is, however, the weakest link: the newly claimed levels rest on a small number of S/N = 3 detections with no false-discovery control, and the authors themselves state in §5.1 that many of the 1,900 Ritz matches 'could be chance coincidences.' The paper also makes no public code or data available, which limits independent verification of the simulation pipeline.","major_comments":[{"comment":"The claim that the two Ni II levels are 'confidently identified' is not supported by a statistical control. The paper reports 6,700 NN-only Ni lines, of which 1,900 match known Ni II Ritz wavenumbers, but it explicitly states that many of these matches could be chance coincidences and gives no estimate of the expected number of chance matches within the 0.05 cm^-1 tolerance. Table 2 then uses three NN-only lines at S/N = 3, exactly the simulation threshold S/N_min, together with a few human-detected lines, to define the levels. With an unstated number of candidate upper levels searched, the probability that random 3-S/N features align into a consistent three-line pattern is not quantified. Please add a chance-coincidence or false-discovery-rate control (e.g., Monte Carlo searches over random level candidates, or matching against shifted Ritz lists) and report the number of level searches performed.","section":"§5.1, Table 2"},{"comment":"No experimental false-positive rate is measured for the NN output. The precision on simulated test data is 0.1–0.5 at p_th = 0.5, which the paper interprets as adjacent-point duplicates; whether isolated noise features are also produced is not reported. The comparison with Xgremlin at smin = 3 in §5.2 shows total line counts (1,300 vs 13,600 for the NiHeEH spectrum), but total counts do not separate true weak lines from false positives. A fixed false-positive operating point, or a labeled validation set derived from independent energy-level analyses, is needed to support the conclusion in §6 that the NN 'decrease[s] the number of false line detections.'","section":"§4.3, §5.2"},{"comment":"The simulation pipeline depends on several hand-tuned parameters: rho_line = 10 x rho_10, S/N_min = 3, the 20% blend-masking ratio, and p_th = 0.5 in postprocessing. Because each model is retrained on a spectrum simulated from parameters estimated from that same experimental spectrum, the reported weak-line gains could be sensitive to these choices. I ask for a sensitivity analysis—for example, varying rho_line by factors of 0.5 and 2 and S/N_min between 2 and 5—while reporting human-list recall and Ritz-confirmation counts at each setting. This would show whether the central result is robust or an artifact of the chosen simulation thresholds.","section":"§3.3, §3.6"},{"comment":"The resolution-ringing evaluation is performed on a spectrum artificially degraded by reducing the maximum OPD by a factor of four, not on an independently measured spectrum with native resolution-limited ringing. Since the central novelty includes handling instrumental ringing, a demonstration on a real under-resolved archive spectrum (e.g., a Kitt Peak spectrum) would make the claim more general. As it stands, the section establishes that the model can learn the simulated apodization model, but not necessarily that it will reject unmodeled phase and profile errors in real ringing.","section":"§5.4"}],"minor_comments":[{"comment":"Axis labels such as 'Wavenumber (cm 1)' should read 'cm^-1'; the missing superscript appears in several figures.","section":"Figs. 2, 7"},{"comment":"Please define the normalization of the sinc function explicitly; as written, sinc(y) = sin(y)/y is used with arguments involving sigma*Omega*x/2 and pi, but the reader cannot verify the aperture factor without a stated convention.","section":"Eq. (2)"},{"comment":"The sentence on precision states that false positives 'almost always occurred adjacent to true positives'; this is important, but the subsequent claim that precision is 'partially representative of uncertainties' should be quantified, for example by reporting the fraction of false positives that lie within one line width of a true positive.","section":"§4.3"},{"comment":"The rounding rule ('all numbers of lines quoted above 1000 are rounded to the nearest 100') conflicts with quoted values such as 1,900 and 6,700; please state which numbers are exact and which are rounded.","section":"§5.1"},{"comment":"The data/code availability statement promises GitHub availability 'after publication' but gives no repository identifier; for a methods paper, please include a permanent archive such as a Zenodo DOI or provide full code and data as supplementary material.","section":"Data and code availability"}],"recommendation":"major_revision","confidential_remarks":"To the editor: This is a promising methods paper for the atomic spectroscopy community. The methodological core—supervised detection trained on simulation—is sound in principle, but the paper's headline physical result (two Ni II levels) currently rests on weak-line detections without a false-discovery control. I would recommend major revision with emphasis on a chance-coincidence analysis and a sensitivity study of the simulation parameters. The absence of public code and data is a concern for a methods paper, though not disqualifying if the statistical controls are added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a solid, honest paper that does something genuinely new. It applies a bidirectional LSTM-FCNN to line detection in high-resolution FT atomic emission spectra, trains on simulated spectra, and shows real gains over Xgremlin for weak, blended, and ringing-distorted lines. The headline numbers — 96% recall of the human Ni line list, 94% for Nd — are meaningful, and the blend and ringing sections are concrete rather than hand-waving. The two newly identified Ni II levels are a modest but real payoff; they were previously flagged as unidentifiable from the same spectra.\n\nOn the stress-test concern: I think it partly misses the mark. Each of the two levels in Table 2 is anchored by at least two human-detected lines, and the NN-only lines are checked against Ritz wavenumbers predicted from the level energy optimized with those human lines. That is a much stronger check than \"a few 3-S/N features lined up by chance.\" The paper does not provide a false-discovery-rate estimate for the thousands of extra NN lines, and it openly says many could be chance coincidences, so the broader claim about line-list completeness is softer than the level-identification claim. But the level-identification claim itself holds up reasonably well.\n\nThe real soft spots are the usual ones: no public code or data (the GitHub promise is post-publication), and a pile of hand-tuned simulation parameters — line-density multiplier, S/N threshold, blend masking ratio, 20% masking rule, probability threshold. These are chosen by trial and error and by inspection, which is fine for a proof of concept but makes the method hard to reproduce or transfer. The comparison to Xgremlin also deserves a careful referee: setting smin=3 makes Xgremlin look bad by design, even though the paper is transparent about that choice.\n\nI also want to credit the authors for flagging their own limitations. Section 6 explicitly names unmodeled isotope shifts, hyperfine structure, phase errors, and self-absorption as constraints, and Section 5.1 admits the 1900 Ritz-matched extra lines include many chance coincidences. That is honest engagement with the evidence.\n\nBottom line: this is a sound contribution for atomic spectroscopists, especially those who build line lists for open d- and f-shell elements. It deserves a serious referee. My recommendation: send it to peer review, and require the authors to release the code and data, add a false-discovery or synthetic-injection validation for the extra lines, and justify the key simulation thresholds with sensitivity checks. Those are fixable revisions, not fatal flaws.","headline":"A serious, well-scoped ML-for-spectroscopy paper whose core claim survives contact with the stress-test note: the Ni II levels are anchored by Ritz-consistent human lines, not just 3-S/N features, but the missing code/data and lack of an FDR estimate for the extra NN lines should be fixed before the method becomes a standard tool.","tokens_in":18515,"tokens_out":1486,"would_cite":true,"duration_ms":19066,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["07.05.Mh","07.60.-j","32.30.-r","32.70.-n"],"model":"deepseek-v4-flash","headline":"A neural network trained on simulated Fourier spectra can detect weak, blended, and ringing-distorted emission lines that standard peak finding misses, and its new detections identify two previously unidentifiable Ni II levels.","keywords":["atomic emission spectra","Fourier transform spectroscopy","spectral line detection","bidirectional LSTM","spectrum simulation","instrumental ringing","nickel energy levels","neodymium spectra"],"falsifier":"Take the nine Ni spectra, inject synthetic Voigt lines with known wavenumbers and $S/N$ between 2 and 6 into quiet regions and into blend and ringing regions, train the pipeline exactly as described, and measure recovery at the known wavenumbers. If the model misses most injected lines below $S/N=5$, or flags ringing artefacts as lines at a rate comparable to the old peak finder, the simulation-fidelity premise is falsified; this is a test the paper's data and code can run directly.","tokens_in":17569,"feed_emoji":"⚛️","tokens_out":12886,"duration_ms":119424,"temperature":0.7,"pith_summary":"The paper sets out to show that spectral line detection in high-resolution Fourier transform emission spectra of open-shell atoms can be done by a neural network trained only on simulated spectra, with no need for human-produced line lists as labels. The model, a bidirectional LSTM encoder with a fully connected decoder, classifies each spectrum point as being or not being a transition line centre. Tested on nine Ni spectra and one high-density Nd spectrum, it recovered about 95% of the lines in earlier human lists while finding thousands of additional weak and blended lines, including lines hidden in instrumental ringing. The practical consequence is that the bottleneck of atomic structure analysis—months of manual line-list construction—can be largely automated: the new detections led to confident assignment of two Ni II energy levels that had been previously judged unidentifiable.","feed_headline":"Neural network finds 6,700 spectral lines missed in human Ni analysis","feed_subtitle":"Trained on simulated Fourier spectra, it recovers weak and blended lines and confirms two Ni II energy levels.","key_machinery":"The load-bearing object is the pointwise class-1 probability $p_1$ produced by a two-layer bidirectional LSTM that encodes overlapping spectrum chunks, concatenates the normalised wavenumber to the LSTM output, and decodes through a fully connected network with PReLU activations. The training data are simulated FT spectra: line wavenumbers are uniform random at density $\\rho_\\text{line}=10\\rho_{10}$; Doppler and pressure widths are drawn from distributions fitted to the experimental spectrum; $S/N$s follow a linear log-$S/N$ extrapolation from line counts above $S/N=10$; and each simulated line is convolved with the FT instrumental profile by multiplying the interferogram with the apodisation function $A(x)$ of Eq. 2, including finite aperture, zero-padding, and cosine-bell apodisation. Focal loss (Eq. 3) makes training focus on the rare line-centre points and hard cases. This machinery lets the network learn line shapes, blends, noise, and resolution ringing from known ground truth, so the model can flag weak lines in the real spectrum without inheriting the incompleteness of human line lists.","core_discovery":"On the paper's own terms, the central claim is that a bidirectional LSTM-FCNN, trained per experimental spectrum on simulations whose line positions, $S/N$ distributions, and instrumental profiles are tuned to that spectrum, transfers to experimental data and outperforms the conventional peak-finding routines on exactly the cases that defeat them. In nine Ni-He hollow-cathode FT spectra covering 1800–70,000 cm$^{-1}$ the model listed 13,400 lines versus 6,700 in the human list; in the dense Nd-Ar Penning spectrum it listed 9,400 versus 7,000. In both cases about 95% of the human lines were matched within 0.05 cm$^{-1}$, and the unmatched remainder were predominantly weak lines and centre-of-gravity wavenumbers from isotope and hyperfine structure. Four lines found only by the model—three of them at $S/N = 3$—removed the ambiguity in assigning two Ni II levels, $3d^8(^3F_4)6f[2]_{3/2}$ at $134,261.8946 \\pm 0.0081$ cm$^{-1}$ and $3d^8(^3F_4)6f[1]_{3/2}$ at $134,249.5264 \\pm 0.0054$ cm$^{-1}$, which the previous analysis had concluded were unidentifiable from the Ni spectra.","pith_inferences":["An injection-recovery test on the same experimental spectra—adding synthetic lines of known wavenumber and $S/N$ near the detection limit and measuring recovery—would give a ground-truth measure of performance that the paper's comparison with human lists cannot provide.","The main systematic gap is isotope and hyperfine structure: the model misses most of the human-list centre-of-gravity wavenumbers, so extending the simulator to generate such components is the most direct route to closing that gap.","If per-spectrum training on simulations becomes routine, the long-term bottleneck shifts from months of manual line identification to accurate simulation of the light source and instrument, since the model can only learn what is put into the simulator.","A testable extension of the paper's stated expectation is to add self-absorption and phase-error models to the simulator and check whether the same architecture then detects lines in optically thick or infrared spectra without extra false positives."],"forward_implications":["About 95% of the lines in the human-produced Ni and Nd line lists are recovered by the model within 0.05 cm$^{-1}$, so existing analyses carry over and only need verification.","Line lists become substantially more complete: for the nine Ni spectra the model detected 13,400 lines against 6,700 in the human list, and for the dense Nd spectrum 9,400 against 7,000, with the extra lines mostly weak or blended.","The $p_1$ probability curve acts as a de-noised spectrum in which weak and blended features appear as peaks, providing a new visual aid for manual fitting and for decisions about which lines are real.","Blended components that the conventional peak finder misses, such as the Nd III line at 26,140.225 cm$^{-1}$ in the wing of a strong line, are detected, and resolution-limited ringing is learned so that real lines are not confused with ringing maxima.","Four weak lines visible only in the neural-network line list make the identification of two Ni II levels unambiguous: $3d^8(^3F_4)6f[2]_{3/2}$ at 134,261.8946 cm$^{-1}$ and $3d^8(^3F_4)6f[1]_{3/2}$ at 134,249.5264 cm$^{-1}$."],"supporting_citations":[{"why":"Supplies the nine Ni-He hollow-cathode FT spectra and the human Ni II line list against which the model's detections are compared.","marker":"[7]"},{"why":"Previous Ni II energy-level analysis that concluded the two levels were unidentifiable; the new NN lines are used to overturn that conclusion.","marker":"[8]"},{"why":"Provides the high-line-density Nd-Ar Penning FT spectrum and the human Nd III line list and blend classifications used in the blend evaluation.","marker":"[11]"},{"why":"Defines the conventional Xgremlin peak-finding routines that form the baseline the neural-network approach is compared with.","marker":"[13]"},{"why":"Earlier neural-network line-detection demonstrations on NMR spectra from which the LSTM-FCNN encoding idea is adapted.","marker":"[14, 15]"},{"why":"Introduces the LSTM architecture used as the bidirectional encoder for spectrum chunks.","marker":"[16]"},{"why":"Focal loss, used in Eq. 3 to handle the extreme imbalance between line-centre points and all other spectrum points.","marker":"[17]"},{"why":"Standard Fourier transform spectrometry theory behind the simulated instrumental line profile, zero-padding, apodisation, and white-noise model.","marker":"[18]"}],"fun_headline_variants":["Neural net finds 6,700 spectral lines missed by humans in Ni data","AI reveals 6,700 hidden lines in atomic spectra","Deep learning detects spectral lines humans missed","LSTM neural network uncovers two new Ni II levels","AI finds 6,700 spectral lines and two new Ni II levels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model learns only what the simulated spectra contain, so the load-bearing premise is that spectra built from uniform random line positions, an $S/N$ distribution extrapolated linearly from lines above $S/N=10$, Gaussian white noise, $S/N_\\min=3$, a 20% blend-masking rule, and a line density of $10\\rho_{10}$ are faithful enough to real FT spectra that a network trained on them detects real lines correctly. The authors themselves note in Section 6 that unmodelled isotope shifts, hyperfine structure, phase errors, and self-absorption limit the approach.","fun_headline_variants_meta":{"raw":{"variants":["Neural net finds 6,700 spectral lines missed by humans in Ni data","AI reveals 6,700 hidden lines in atomic spectra","Deep learning detects spectral lines humans missed","LSTM neural network uncovers two new Ni II levels","AI finds 6,700 spectral lines and two new Ni II levels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000372,"raw_usage":{"total_tokens":2137,"prompt_tokens":1239,"completion_tokens":898,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":855,"completion_tokens_details":{"reasoning_tokens":813}},"tokens_in":855,"tokens_out":898,"duration_ms":7782,"temperature":1.0,"reasoning_tokens":813,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:18:48.005163+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the nine Ni spectra, inject synthetic Voigt lines with known wavenumbers and $S/N$ between 2 and 6 into quiet regions and into blend and ringing regions, train the pipeline exactly as described, and measure recovery at the known wavenumbers. If the model misses most injected lines below $S/N=5$, or flags ringing artefacts as lines at a rate comparable to the old peak finder, the simulation-fidelity premise is falsified; this is a test the paper's data and code can run directly.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the nine Ni-He hollow-cathode FT spectra and the human Ni II line list against which the model's detections are compared."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Previous Ni II energy-level analysis that concluded the two levels were unidentifiable; the new NN lines are used to overturn that conclusion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the high-line-density Nd-Ar Penning FT spectrum and the human Nd III line list and blend classifications used in the blend evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the conventional Xgremlin peak-finding routines that form the baseline the neural-network approach is compared with."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Focal loss, used in Eq. 3 to handle the extreme imbalance between line-centre points and all other spectrum points."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Standard Fourier transform spectrometry theory behind the simulated instrumental line profile, zero-padding, apodisation, and white-noise model."}],"review_version":1}