REVIEW 3 major objections 5 minor 2 cited by
Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A lightweight autoencoder estimates quasar continua at roughly one percent error and transfers to unseen surveys without retraining.
desk verdict Solid mock-based benchmark, but the abstract misreports the headline numbers and the DESI tau_eff test overstates cross-survey generalization. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the autoencoder with a latent bottleneck: three dense layers compress each input spectrum into a low-dimensional code, and a decoder reconstructs the continuum from that code, with a masking layer that zeroes out pixels missing at the spectral edges. The decisive design choice is the input-output split — the network sees only the clean red part of the spectrum ($\lambda > 1216$ Å) and must predict the full continuum including the Lyman-$\alpha$ forest ($\lambda < 1216$ Å), forcing it to learn the continuum shape from an uncontaminated region and extrapolate into the absorbed one. The comparison machinery is the absolute fractional flux error (AFFE), the wavelength-averaged absolute fractional difference between predicted and true continuum, supplemented by covariate-shift tests in which the latent space is split in two and the network retrained on one half to see how predictions degrade outside the training distribution. The same architecture, re-optimized but not restructured, is applied to galaxy continua with the full spectrum (3500–5500 Å) as input.
What would settle it
Apply the mock-trained autoencoder to spectra where the true continuum is directly measurable — for instance, BAL quasars (excluded from training) observed at rest-frame wavelengths well redward of the Lyman-alpha line — and recompute the Lyman-alpha effective optical depth after applying the standard metal and optically-thick-absorber corrections that this paper omits; if the corrected values drift off the reference curve while a comparison network's values stay on it, the reported agreement is a coincidence of uncorrected systematics rather than continuum accuracy.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the simplest architecture wins: an optimized three-layer autoencoder reaches a median absolute fractional flux error (AFFE) of about 0.007–0.009 on mock WEAVE quasar spectra, versus 0.013 for the U-Net, with flatter wavelength-dependent bias than the published iQNet baseline. The network is trained on the relatively uncontaminated red side of the quasar spectrum ($\lambda > 1216$ Å) and asked to predict the full continuum from 1020 to 2000 Å, including the Lyman-$\alpha$ forest region where the true continuum is hidden behind dense absorption; passing the full spectrum as input does not materially improve the fit. Trained only on mocks, the autoencoder reproduces the redshift evolution of the Lyman-$\alpha$ effective optical depth, $\tau_{\rm eff}(z)$, in DESI Early Data Release quasars in agreement with earlier measurements (Becker et al. 2013; Turner et al. 2024), which the paper reads as evidence of genuine cross-survey generalization. For galaxies, a fresh optimization pass with the same architecture reaches a median AFFE of 0.014 against pPXF continua and reproduces the D4000n break in VIPERS and DESI data.
Load-bearing premise
Everything rests on one premise: the mock WEAVE spectra the network trains on — BOSS principal-component continuum shapes combined with simulated Lyman-alpha forest absorption — are representative enough of real quasar continua that a model which has only ever seen mocks can be trusted on DESI data, a premise the paper itself flags as imperfect when it shows the mock continua have systematically shallower slopes than the real spectra.
Editorial extensions
If this is right
- A WEAVE-trained quasar model can be deployed on DESI data with no retraining and still recover the known evolution of the Lyman-alpha effective optical depth, so cross-survey transfer of continuum estimators is practically achievable.
- Continuum fitting at the percent level becomes cheap enough to run on entire surveys: the paper estimates roughly 3 minutes to fit one million spectra with the autoencoder, versus about 40 minutes with the CNN.
- The same architecture family handles a very different continuum shape, galaxy spectra at 3500–5500 Å, after only a new optimization pass, with the autoencoder at median AFFE 0.014 and a working D4000n measurement.
- Extra architectural complexity buys nothing for this problem: the U-Net's deeper encoder-decoder with skip connections returns median AFFE 0.013, roughly double the autoencoder's error.
- A raw, uncorrected Lyman-alpha optical depth measurement computed from autoencoder continua lands within the scatter of the reference curve, indicating that percent-level continuum errors still propagate into a scientifically usable absorption measurement.
Reading between the lines
- Because the reported Lyman-alpha optical depth recovery skips the standard corrections for metal absorption and optically thick absorbers, both of which shift the measurement in the same direction, my read is that part of the agreement with the reference curve may reflect compensating errors; a corrected comparison would be the sharper test.
- The model's blind spot is precisely the quasars it cannot see: broad absorption line (BAL) quasars are removed from the mocks, so the natural next test is whether fine-tuning on a few hundred BAL spectra lets the same architecture track their heavily distorted continua.
- Figure 2 shows the mock continua have systematically shallower slopes than real DESI spectra, so the reported transfer accuracy is achieved despite a known distributional mismatch; enriching the mock library with a wider spread of continuum slopes should push survey-transfer errors lower.
- The bottleneck code is a compressed, survey-agnostic description of the continuum shape, so the same network could plausibly double as a continuum prior for absorption-line analyses, an extension the paper does not test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript compares three deep-learning architectures (an autoencoder, a CNN, and a U-Net) for automated quasar continuum estimation in the rest-frame range 1020–2000 Å. The models are trained on WEAVE mock quasar spectra built from BOSS PCA continuum shapes and hydrodynamical Lyα forest skewers, and are benchmarked against the published iQNet and LyCAN networks using the absolute fractional flux error (AFFE). The best-performing model is then applied without retraining to DESI EDR quasar spectra to measure the Lyα effective optical depth evolution, and, after retraining, the same architectures are applied to galaxy spectra from VIPERS and DESI to measure the D4000n break. The paper argues that the autoencoder performs as well as more complex architectures at much lower computational cost and generalizes across surveys.
Significance. If the reported results are correct, the paper provides a useful comparative benchmark for continuum-fitting methods needed for WEAVE and other large spectroscopic surveys. The study design is sound in its main components: training on realistic mocks, evaluating on held-out mock spectra, applying the trained model to independent DESI data, and comparing with two published baselines. The extensive appendices on magnitude cuts, S/N dependence, and redshift/SNR biases are valuable, and the public Data availability of the underlying surveys is a strength. However, the headline claim about the autoencoder being the best architecture is contradicted by the paper's own figures, and the DESI Lyα optical-depth test is presented in a way that does not independently establish unbiased continuum prediction on real data. These issues need to be resolved before the paper's central conclusions can be accepted.
major comments (3)
- [Abstract; Section 8; Figs. 4 and 10] The abstract states that 'the autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars.' This value is not the autoencoder's median AFFE in Fig. 4: the autoencoder red-part and full-spectrum medians are 0.007 and 0.008, while 0.009 is the median for iQNet. More importantly, Fig. 4 shows that the CNN trained on the red part achieves a median AFFE of 0.006, which is lower than any autoencoder result. Fig. 10, presented in the Summary, reports CNN red part at 0.009 and CNN full spectra at 0.006, while the autoencoders are 0.008 and 0.007; this contradicts Fig. 4, where CNN red part is 0.006 and CNN full spectra is 0.010. In either figure, the best quasar continuum model is a CNN, not the autoencoder. The abstract and the conclusions in Section 8 therefore misstate the main comparative result, and the discrepancy between Fig. 4 and Fig. 10 must be resolved.
- [Section 5.2, Eq. (4), Fig. 8] The comparison between the autoencoder's raw, uncorrected τeff and the corrected literature curves is not a valid test of continuum unbiasedness. The text explicitly states that no correction for optically thick absorbers or metals is applied, while the Turner et al. (2024) values are 'bias-corrected measurements corrected for metal line absorption... and optically thick absorbers,' and the Becker et al. (2013) fit is likewise a corrected measurement. An unbiased continuum should yield a raw τeff higher than the corrected curves by roughly the metal plus LLS opacity; the fact that the raw measurement lies on the corrected curves implies a systematic underestimate of the true DESI continuum in the Lyα forest. The covariate-shift exercise in Section 5.1 is internal to the mock distribution and does not constrain real-data bias, especially because Fig. 2 shows a slope offset between the WEAVE mocks and DESI. In addition, Section 5.2 defines a mean transmission in Eq. (4) but then reports the median transmitted flux, which is not the same statistic as the literature values with which it is compared. As presented, the DESI τeff test does not independently establish that the mock-trained continuum is unbiased on real data; the authors should either apply the corrections, compare with raw uncorrected literature measurements, or explicitly reframe the test as a consistency check that does not by itself demonstrate zero continuum bias.
- [Section 7, Fig. 9] The VIPERS D4000n validation is partially circular and should be identified as such. According to Section 2.2, the NN training labels for galaxies are the pPXF stellar-continuum fits; the left-hand panel of Fig. 9 compares the NN-derived D4000n against D4000n computed from the same pPXF continua. Agreement in that panel largely confirms that the network has learned the pPXF labels, not that it generalizes to an independent estimate. The DESI/Redrock comparison in the right-hand panel is the genuinely independent test, since Redrock is a separate pipeline. The text should explicitly state that the VIPERS comparison is a reproducibility check, not independent generalization evidence, and should not use the VIPERS correlation coefficient as support for the generalization claim.
minor comments (5)
- [Throughout] There are numerous typographical and spacing issues, including 'di fferent' (multiple instances), 'untractable' (Section 1), 'versitile', 'datatsets', 'availaility', 'modfications', 'res-frame', and 'WEA VE' with a stray space; these should be corrected in a careful proofreading pass.
- [Section 2.3, Fig. 2] The sentence 'WEA VE mocks generally overlap with the DESI data for the R parameter' refers to the flux ratio FR defined in the same paragraph; the notation should be consistent.
- [Section 4.1] The cross-reference 'Sect.??' is unresolved; the intended section on galaxy generalization should be cited explicitly.
- [Section 5.2] The text says 'we computed the median transmitted flux ⟨f⟩' after defining the mean transmission in Eq. (4); the statistical choice should be stated upfront and justified, and the estimator should match the quantity compared with the literature.
- [Appendix A, Fig. A.1] The caption and text use 'res-frame' and 'functional of the rest-frame wavelength'; these should be corrected to 'rest-frame' and 'as a function of'.
Circularity Check
No significant circularity in the quasar continuum chain; the only reduction-by-construction element is the VIPERS galaxy D4000n comparison, where the pPXF continuum is both the training label and the reference.
-
fitted input called prediction
[Sections 2.2 and 7, Fig. 9 left panel.]
"Having a good measurement of the true continuum as a label during the training of the NN model is essential. We estimate the continuum of the VIPERS galaxy spectra by modeling them (as done in Pistis et al. 2024) using the penalized pixel fitting code (pPXF, Cappellari & Emsellem 2004; Cappellari 2017, 2023) ... Fig. 9 shows the comparison of the values for the D4000n break calculated over the continua predicted by the autoencoder and the continua estimated using pPXF for VIPERS and Redrock for DESI."
For VIPERS, the pPXF continuum is the training label (Section 2.2) and also the reference used to compute the D4000n index compared in Fig. 9. The agreement therefore measures how well the autoencoder reproduces its own label-generator, not whether the galaxy continuum is independently correct on VIPERS data. This is a partial, secondary circularity confined to the galaxy branch; the corresponding DESI/Redrock comparison is an independent external test.
full rationale
The central quasar result is self-contained: the models are trained on WEAVE mock spectra whose true continua come from BOSS PCA shapes and hydrodynamical Ly-alpha forest skewers, and the DESI tau_eff measurement is an external generalization test on real data. Section 5.2 explicitly labels the tau_eff estimate as a raw consistency check with no metal or optically-thick absorber corrections, which is a correctness caveat rather than a circular step. The covariate-shift exercise in Section 5.1 is internal to the mock distribution, but it is not presented as an external validation. The only step that reduces by construction is the VIPERS D4000n comparison, where the pPXF fit serves as both training label and reference; this does not affect the main quasar claim. The Pistis et al. (2024) citation is a methodological pointer, not a load-bearing self-citation. Overall, the scores reflect one partial, secondary circularity while the primary derivation remains independent.
Assumptions & free parameters
free parameters (4)
- Autoencoder hyperparameters =
Dense 512-128-1024 for WEAVE red part (Table D.1)
- CNN hyperparameters =
Table D.2
- U-Net hyperparameters =
Table D.3
- r-band magnitude cut =
r < 21
assumptions (3)
- domain assumption WEAVE mock spectra reproduce the statistical properties of real quasar spectra (continuum from BOSS PCA, Ly-alpha forest from hydro simulations).
- domain assumption The red side (1216 to 2000 Angstrom) of quasar spectra is sufficiently uncontaminated to predict the blue-side continuum.
- domain assumption pPXF stellar fits provide a reliable ground-truth continuum for galaxy spectra.
Cite this review
Pith. "Pith review of Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures." pith.science (2026). https://pith.science/paper/N4NSIXLT
@misc{pith2026250510976,
author = {Pith},
title = {Pith review of: Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures},
year = {2026},
howpublished = {\url{https://pith.science/paper/N4NSIXLT}},
note = {Machine review of arXiv:2505.10976}
}
abstract
Context. Ongoing and upcoming large spectroscopic surveys are drastically increasing the number of observed quasar spectra, requiring the development of fast and accurate automated methods to estimate spectral continua. Aims. This study evaluates the performance of three neural networks (NN) - an autoencoder, a convolutional NN (CNN), and a U-Net - in predicting quasar continua within the rest-frame wavelength range of $1020~\text{\AA}$ to $2000~\text{\AA}$. The ability to generalize and predict galaxy continua within the range of $3500~\text{\AA}$ to $5500~\text{\AA}$ is also tested. Methods. The performance of these architectures is evaluated using the absolute fractional flux error (AFFE) on a library of mock quasar spectra for the WEAVE survey, and on real data from the Early Data Release observations of the Dark Energy Spectroscopic Instrument (DESI) and the VIMOS Public Extragalactic Redshift Survey (VIPERS). Results. The autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars. The best model also effectively recovers the Ly$\alpha$ optical depth evolution in DESI quasar spectra. With minimal optimization, the same architectures can be generalized to the galaxy case, with the autoencoder reaching a median AFFE of 0.014 and reproducing the D4000n break in DESI and VIPERS galaxies.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 2 Pith papers
-
Forecast for the detectability of patchy hydrogen reionization in WEAVE-QSO measurements of the Lyman-$\alpha$ forest power spectrum at redshift $z \geq 4$
A mock-observation forecast finds that WEAVE-QSO measurements of the Lyman-alpha forest power spectrum at z≥4 can detect the relic large-scale power from patchy hydrogen reionization at 4.5 sigma.
-
Uncertainty-Aware Deep Learning for the Ly$\alpha$ Forest: CNN-Based Absorber Detection and Characterization
A sliding-window CNN recovers Lyα absorber locations and Voigt parameters from spectra, reproducing CDDF and b–N relations on mocks and, more weakly, on UVES data.
Reference graph
Works this paper leans on
-
[1]
Abazajian, K. N., Adelman-McCarthy, J. K., Agüeros, M. A., et al. 2009, ApJS, 182, 543 Abraham, V ., Deville, J., & Kinariwala, G. 2024, Research Notes of the Ameri- can Astronomical Society, 8, 46 Aguirre, A., Schaye, J., & Theuns, T. 2002, ApJ, 576, 1 Ambrosch, M., Guiglion, G., Mikolaitis, Š., et al. 2023, A&A, 672, A46 Angthopo, J., Granett, B. R., La...
arXiv 2009
-
[3]
Pool Size: 2 Conv (2010,
work page 2010
-
[6]
Mask value: 0 Masking (2000,
work page 2000
-
[9]
Kernel Size: 2 Conv (2000,
work page 2000
-
[11]
Kernel Size: 5 Conv (2000,
work page 2000
-
[12]
Kernel Size: 1 Conv (2000,
work page 2000
-
[13]
Kernel Size: 1 Notes. The activation function (Act. value in the table) used is the rectified linear unit (ReLU) function. Article number, page 20 of 23 Francesco Pistis et al.: Automated quasar continuum estimation using neural networks Appendix E: Fit example: autoencoders and CNNs Fig. E.1 shows examples of spectra for the case of giving to the optimiz...
work page 2000
-
[32]
Kernel Size: 5 Act.: ReLU MaxPooling (2010,
work page 2010
Show all 10 references
-
[162]
Drop rate: 0.2 Flatten 14720 Dense 512 Act.: ReLU Dense 1024 Act.: ReLU Dense 5000 Act.: ReLU VIPERS Input 2000 Masking 2000 Mask value: 0 Conv (2000,
2000
-
[256]
The activation functions (Act
Pool Size: 2 Flatten 14080 Dense 384 Act.: ReLU Dense 1024 Act.: ReLU Dense 2000 Act.: ReLU Notes. The activation functions (Act. value in the table) used are the rectified linear unit (ReLU), exponential linear unit (ELU), and linear activation. Article number, page 19 of 23 ...
2000
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.