Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A lightweight autoencoder estimates quasar continua at roughly one percent error and transfers to unseen surveys without retraining.

desk verdict Solid mock-based benchmark, but the abstract misreports the headline numbers and the DESI tau_eff test overstates cross-survey generalization. read the letter →

arxiv 2505.10976 v1 pith:N4NSIXLT submitted 2025-05-16 astro-ph.GA astro-ph.IM

classification astro-ph.GAastro-ph.IM
keywords quasarcontinuumestimationLyman-alphaforestneuralnetworksautoencoderconvolutionalnetworkU-Netintergalacticmediumspectroscopicsurveys
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a lightweight autoencoder — a neural network that compresses a spectrum through a narrow bottleneck and then reconstructs it — can estimate quasar continua at sub-percent accuracy and, trained only on simulated spectra, can be applied without retraining to real spectra from a different survey. Comparing an autoencoder, a convolutional network, and a U-Net on mock spectra built for the WEAVE survey, the authors find the autoencoder clearly outperforms the U-Net and matches the CNN while consuming roughly an order of magnitude less compute. Applied unchanged to DESI Early Data Release quasars, the autoencoder reproduces the redshift evolution of the Lyman-alpha effective optical depth found by earlier dedicated measurements, and with modest retuning the same architecture fits galaxy continua and recovers the D4000n break. If the paper is right, a single small network can serve as a fast, survey-agnostic continuum estimator for the million-spectrum surveys now underway.

What carries the argument

The load-bearing object is the autoencoder with a latent bottleneck: three dense layers compress each input spectrum into a low-dimensional code, and a decoder reconstructs the continuum from that code, with a masking layer that zeroes out pixels missing at the spectral edges. The decisive design choice is the input-output split — the network sees only the clean red part of the spectrum ($\lambda > 1216$ Å) and must predict the full continuum including the Lyman-$\alpha$ forest ($\lambda < 1216$ Å), forcing it to learn the continuum shape from an uncontaminated region and extrapolate into the absorbed one. The comparison machinery is the absolute fractional flux error (AFFE), the wavelength-averaged absolute fractional difference between predicted and true continuum, supplemented by covariate-shift tests in which the latent space is split in two and the network retrained on one half to see how predictions degrade outside the training distribution. The same architecture, re-optimized but not restructured, is applied to galaxy continua with the full spectrum (3500–5500 Å) as input.

What would settle it

Apply the mock-trained autoencoder to spectra where the true continuum is directly measurable — for instance, BAL quasars (excluded from training) observed at rest-frame wavelengths well redward of the Lyman-alpha line — and recompute the Lyman-alpha effective optical depth after applying the standard metal and optically-thick-absorber corrections that this paper omits; if the corrected values drift off the reference curve while a comparison network's values stay on it, the reported agreement is a coincidence of uncorrected systematics rather than continuum accuracy.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the simplest architecture wins: an optimized three-layer autoencoder reaches a median absolute fractional flux error (AFFE) of about 0.007–0.009 on mock WEAVE quasar spectra, versus 0.013 for the U-Net, with flatter wavelength-dependent bias than the published iQNet baseline. The network is trained on the relatively uncontaminated red side of the quasar spectrum ($\lambda > 1216$ Å) and asked to predict the full continuum from 1020 to 2000 Å, including the Lyman-$\alpha$ forest region where the true continuum is hidden behind dense absorption; passing the full spectrum as input does not materially improve the fit. Trained only on mocks, the autoencoder reproduces the redshift evolution of the Lyman-$\alpha$ effective optical depth, $\tau_{\rm eff}(z)$, in DESI Early Data Release quasars in agreement with earlier measurements (Becker et al. 2013; Turner et al. 2024), which the paper reads as evidence of genuine cross-survey generalization. For galaxies, a fresh optimization pass with the same architecture reaches a median AFFE of 0.014 against pPXF continua and reproduces the D4000n break in VIPERS and DESI data.

Load-bearing premise

Everything rests on one premise: the mock WEAVE spectra the network trains on — BOSS principal-component continuum shapes combined with simulated Lyman-alpha forest absorption — are representative enough of real quasar continua that a model which has only ever seen mocks can be trusted on DESI data, a premise the paper itself flags as imperfect when it shows the mock continua have systematically shallower slopes than the real spectra.

Editorial extensions

If this is right

  • A WEAVE-trained quasar model can be deployed on DESI data with no retraining and still recover the known evolution of the Lyman-alpha effective optical depth, so cross-survey transfer of continuum estimators is practically achievable.
  • Continuum fitting at the percent level becomes cheap enough to run on entire surveys: the paper estimates roughly 3 minutes to fit one million spectra with the autoencoder, versus about 40 minutes with the CNN.
  • The same architecture family handles a very different continuum shape, galaxy spectra at 3500–5500 Å, after only a new optimization pass, with the autoencoder at median AFFE 0.014 and a working D4000n measurement.
  • Extra architectural complexity buys nothing for this problem: the U-Net's deeper encoder-decoder with skip connections returns median AFFE 0.013, roughly double the autoencoder's error.
  • A raw, uncorrected Lyman-alpha optical depth measurement computed from autoencoder continua lands within the scatter of the reference curve, indicating that percent-level continuum errors still propagate into a scientifically usable absorption measurement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the reported Lyman-alpha optical depth recovery skips the standard corrections for metal absorption and optically thick absorbers, both of which shift the measurement in the same direction, my read is that part of the agreement with the reference curve may reflect compensating errors; a corrected comparison would be the sharper test.
  • The model's blind spot is precisely the quasars it cannot see: broad absorption line (BAL) quasars are removed from the mocks, so the natural next test is whether fine-tuning on a few hundred BAL spectra lets the same architecture track their heavily distorted continua.
  • Figure 2 shows the mock continua have systematically shallower slopes than real DESI spectra, so the reported transfer accuracy is achieved despite a known distributional mismatch; enriching the mock library with a wider spread of continuum slopes should push survey-transfer errors lower.
  • The bottleneck code is a compressed, survey-agnostic description of the continuum shape, so the same network could plausibly double as a continuum prior for absorption-line analyses, an extension the paper does not test.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript compares three deep-learning architectures (an autoencoder, a CNN, and a U-Net) for automated quasar continuum estimation in the rest-frame range 1020–2000 Å. The models are trained on WEAVE mock quasar spectra built from BOSS PCA continuum shapes and hydrodynamical Lyα forest skewers, and are benchmarked against the published iQNet and LyCAN networks using the absolute fractional flux error (AFFE). The best-performing model is then applied without retraining to DESI EDR quasar spectra to measure the Lyα effective optical depth evolution, and, after retraining, the same architectures are applied to galaxy spectra from VIPERS and DESI to measure the D4000n break. The paper argues that the autoencoder performs as well as more complex architectures at much lower computational cost and generalizes across surveys.

Significance. If the reported results are correct, the paper provides a useful comparative benchmark for continuum-fitting methods needed for WEAVE and other large spectroscopic surveys. The study design is sound in its main components: training on realistic mocks, evaluating on held-out mock spectra, applying the trained model to independent DESI data, and comparing with two published baselines. The extensive appendices on magnitude cuts, S/N dependence, and redshift/SNR biases are valuable, and the public Data availability of the underlying surveys is a strength. However, the headline claim about the autoencoder being the best architecture is contradicted by the paper's own figures, and the DESI Lyα optical-depth test is presented in a way that does not independently establish unbiased continuum prediction on real data. These issues need to be resolved before the paper's central conclusions can be accepted.

major comments (3)
  1. [Abstract; Section 8; Figs. 4 and 10] The abstract states that 'the autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars.' This value is not the autoencoder's median AFFE in Fig. 4: the autoencoder red-part and full-spectrum medians are 0.007 and 0.008, while 0.009 is the median for iQNet. More importantly, Fig. 4 shows that the CNN trained on the red part achieves a median AFFE of 0.006, which is lower than any autoencoder result. Fig. 10, presented in the Summary, reports CNN red part at 0.009 and CNN full spectra at 0.006, while the autoencoders are 0.008 and 0.007; this contradicts Fig. 4, where CNN red part is 0.006 and CNN full spectra is 0.010. In either figure, the best quasar continuum model is a CNN, not the autoencoder. The abstract and the conclusions in Section 8 therefore misstate the main comparative result, and the discrepancy between Fig. 4 and Fig. 10 must be resolved.
  2. [Section 5.2, Eq. (4), Fig. 8] The comparison between the autoencoder's raw, uncorrected τeff and the corrected literature curves is not a valid test of continuum unbiasedness. The text explicitly states that no correction for optically thick absorbers or metals is applied, while the Turner et al. (2024) values are 'bias-corrected measurements corrected for metal line absorption... and optically thick absorbers,' and the Becker et al. (2013) fit is likewise a corrected measurement. An unbiased continuum should yield a raw τeff higher than the corrected curves by roughly the metal plus LLS opacity; the fact that the raw measurement lies on the corrected curves implies a systematic underestimate of the true DESI continuum in the Lyα forest. The covariate-shift exercise in Section 5.1 is internal to the mock distribution and does not constrain real-data bias, especially because Fig. 2 shows a slope offset between the WEAVE mocks and DESI. In addition, Section 5.2 defines a mean transmission in Eq. (4) but then reports the median transmitted flux, which is not the same statistic as the literature values with which it is compared. As presented, the DESI τeff test does not independently establish that the mock-trained continuum is unbiased on real data; the authors should either apply the corrections, compare with raw uncorrected literature measurements, or explicitly reframe the test as a consistency check that does not by itself demonstrate zero continuum bias.
  3. [Section 7, Fig. 9] The VIPERS D4000n validation is partially circular and should be identified as such. According to Section 2.2, the NN training labels for galaxies are the pPXF stellar-continuum fits; the left-hand panel of Fig. 9 compares the NN-derived D4000n against D4000n computed from the same pPXF continua. Agreement in that panel largely confirms that the network has learned the pPXF labels, not that it generalizes to an independent estimate. The DESI/Redrock comparison in the right-hand panel is the genuinely independent test, since Redrock is a separate pipeline. The text should explicitly state that the VIPERS comparison is a reproducibility check, not independent generalization evidence, and should not use the VIPERS correlation coefficient as support for the generalization claim.
minor comments (5)
  1. [Throughout] There are numerous typographical and spacing issues, including 'di fferent' (multiple instances), 'untractable' (Section 1), 'versitile', 'datatsets', 'availaility', 'modfications', 'res-frame', and 'WEA VE' with a stray space; these should be corrected in a careful proofreading pass.
  2. [Section 2.3, Fig. 2] The sentence 'WEA VE mocks generally overlap with the DESI data for the R parameter' refers to the flux ratio FR defined in the same paragraph; the notation should be consistent.
  3. [Section 4.1] The cross-reference 'Sect.??' is unresolved; the intended section on galaxy generalization should be cited explicitly.
  4. [Section 5.2] The text says 'we computed the median transmitted flux ⟨f⟩' after defining the mean transmission in Eq. (4); the statistical choice should be stated upfront and justified, and the estimator should match the quantity compared with the literature.
  5. [Appendix A, Fig. A.1] The caption and text use 'res-frame' and 'functional of the rest-frame wavelength'; these should be corrected to 'rest-frame' and 'as a function of'.

Circularity Check

1 steps flagged · score 2.0 of 10

No significant circularity in the quasar continuum chain; the only reduction-by-construction element is the VIPERS galaxy D4000n comparison, where the pPXF continuum is both the training label and the reference.

  1. fitted input called prediction [Sections 2.2 and 7, Fig. 9 left panel.]
    "Having a good measurement of the true continuum as a label during the training of the NN model is essential. We estimate the continuum of the VIPERS galaxy spectra by modeling them (as done in Pistis et al. 2024) using the penalized pixel fitting code (pPXF, Cappellari & Emsellem 2004; Cappellari 2017, 2023) ... Fig. 9 shows the comparison of the values for the D4000n break calculated over the continua predicted by the autoencoder and the continua estimated using pPXF for VIPERS and Redrock for DESI."

    For VIPERS, the pPXF continuum is the training label (Section 2.2) and also the reference used to compute the D4000n index compared in Fig. 9. The agreement therefore measures how well the autoencoder reproduces its own label-generator, not whether the galaxy continuum is independently correct on VIPERS data. This is a partial, secondary circularity confined to the galaxy branch; the corresponding DESI/Redrock comparison is an independent external test.

full rationale

The central quasar result is self-contained: the models are trained on WEAVE mock spectra whose true continua come from BOSS PCA shapes and hydrodynamical Ly-alpha forest skewers, and the DESI tau_eff measurement is an external generalization test on real data. Section 5.2 explicitly labels the tau_eff estimate as a raw consistency check with no metal or optically-thick absorber corrections, which is a correctness caveat rather than a circular step. The covariate-shift exercise in Section 5.1 is internal to the mock distribution, but it is not presented as an external validation. The only step that reduces by construction is the VIPERS D4000n comparison, where the pPXF fit serves as both training label and reference; this does not affect the main quasar claim. The Pistis et al. (2024) citation is a methodological pointer, not a load-bearing self-citation. Overall, the scores reflect one partial, secondary circularity while the primary derivation remains independent.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The training pipeline rests on the realism of WEAVE mock spectra and on pPXF-derived galaxy labels. Hyperparameters were tuned on validation data via random and Bayesian search, and the magnitude selection r < 21 was chosen to reduce FFE bias. No new physical entities are introduced.

free parameters (4)
  • Autoencoder hyperparameters = Dense 512-128-1024 for WEAVE red part (Table D.1)
    Layer sizes and activations chosen by random search over 100 trials on validation data.
  • CNN hyperparameters = Table D.2
    Layer sizes and kernels chosen by Bayesian search over 50 trials on validation data.
  • U-Net hyperparameters = Table D.3
    Layer sizes and kernels chosen by Bayesian search over 50 trials on validation data.
  • r-band magnitude cut = r < 21
    Selected in Appendix A to reduce FFE bias while maximizing sample statistics.
assumptions (3)
  • domain assumption WEAVE mock spectra reproduce the statistical properties of real quasar spectra (continuum from BOSS PCA, Ly-alpha forest from hydro simulations).
    Central training and validation set is mocked; if mocks are unrepresentative, all performance numbers and the chosen model may not transfer. Section 2.1.
  • domain assumption The red side (1216 to 2000 Angstrom) of quasar spectra is sufficiently uncontaminated to predict the blue-side continuum.
    All red-part models map red wavelengths to the full continuum; if the red side is contaminated, predictions degrade. Sections 3.3 and 4.1.
  • domain assumption pPXF stellar fits provide a reliable ground-truth continuum for galaxy spectra.
    Galaxy training labels are pPXF fits (Section 2.2); galaxy AFFE is measured against pPXF, so errors in pPXF propagate into both labels and validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures." pith.science (2026). https://pith.science/paper/N4NSIXLT

@misc{pith2026250510976,
  author       = {Pith},
  title        = {Pith review of: Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N4NSIXLT}},
  note         = {Machine review of arXiv:2505.10976}
}
abstract

Context. Ongoing and upcoming large spectroscopic surveys are drastically increasing the number of observed quasar spectra, requiring the development of fast and accurate automated methods to estimate spectral continua. Aims. This study evaluates the performance of three neural networks (NN) - an autoencoder, a convolutional NN (CNN), and a U-Net - in predicting quasar continua within the rest-frame wavelength range of $1020~\text{\AA}$ to $2000~\text{\AA}$. The ability to generalize and predict galaxy continua within the range of $3500~\text{\AA}$ to $5500~\text{\AA}$ is also tested. Methods. The performance of these architectures is evaluated using the absolute fractional flux error (AFFE) on a library of mock quasar spectra for the WEAVE survey, and on real data from the Early Data Release observations of the Dark Energy Spectroscopic Instrument (DESI) and the VIMOS Public Extragalactic Redshift Survey (VIPERS). Results. The autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars. The best model also effectively recovers the Ly$\alpha$ optical depth evolution in DESI quasar spectra. With minimal optimization, the same architectures can be generalized to the galaxy case, with the autoencoder reaching a median AFFE of 0.014 and reproducing the D4000n break in DESI and VIPERS galaxies.

Figures

Figures reproduced from arXiv: 2505.10976 by the authors.

Figure 1
Figure 1. Histograms of the redshift (top panel) and S/N (bottom panel) for quasars (left column in purple for the WEAVE mock catalog and in green for DESI EDR) and galaxies (right panel in blue for the VIPERS and in green for DESI EDR). The unfilled histograms show the distributions of the parent samples, while the filled histograms show the distributions of the final samples after all the data selections. continuum shapes f… view at source ↗
Figure 2
Figure 2. Histograms of the ratio (top) and slope (bottom) of the quasar spectra between the rest-frame wavelengths λ = 1800 Å and λ = 1450 Å for the WEAVE mock catalog (purple) and the DESI EDR (green). The unfilled histograms show the distributions of the parent samples, while the filled histograms show the distributions of the final samples after all the data selections. at these specific cases. For this task, we use the p… view at source ↗
Figure 3
Figure 3. Graphical representation of the workflow. The dotted double ar￾rows represent recursive steps. edges of the spectra, adding a masking layer to the autoencoder to exclude this portion of the spectra. 3.1. Performance metrics To measure the quality of the fit, we use the AFFE as a fit￾goodness metric (see, e.g., Liu & Bordoloi 2021; Turner et al. 2024). The AFFE is defined as: AFFE = |δF| = Z λ2 λ1 [PITH_FULL_IMAGE:f… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Performance of NN predictions: histograms of the absolute fractional flux error (AFFE) for the autoencoders (top row) and CNN plus U-Net (bottom row) applied to the quasar datasets. The vertical lines show the median AFFE value for each histogram. 0 1 2 3 4 5 6 Normali…
Figure 5
Figure 5. Figure 5: Example quasar spectrum. The top panel of each plot shows the data, the true continuum, and the fits of different runs of the autoencoders (left) and CNNs (right). The bottom panel shows the FFE as a function of the rest-frame wavelength for different runs of the autoe…
Figure 6
Figure 6. Figure 6: Performance of NN predictions with wavelength: median fractional flux error (FFE) as a function of the wavelength for the autoencoders (left panel) and CNN (right panel). The shaded area shows the 16th and 84th percentiles of the distribution of the whole sample. The v…
Figure 7
Figure 7. Figure 7: Distribution of the spectra in the latent spaces of the autoencoders (upper row) and CNNs (bottom row) for the WEAVE training sample (purple) and the DESI final sample (green). 2.00 2.25 2.50 2.75 3.00 3.25 3.50 3.75 4.00 Redshift z 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 …
Figure 8
Figure 8. Figure 8: shows the evolution of our measurement of the op￾tical depth. Our simple approach can reproduce quite well the observed evolution in more bespoke studies (Becker et al. 2013; Turner et al. 2024), which implies a good degree of prediction of the quasar continuum in unse…
Figure 9
Figure 9. Figure 9: Comparison of the D4000n values estimated with the NN in this work and with pPXF (Pistis et al. 2024) (left panel) for VIPERS (in blue) and with Redrock (right panel) for DESI (in green). The red solid line shows the line y = x, the light-shaded area limited by dashed …
Figure 12
Figure 12. Figure 12: Visualization of the training and prediction times for the au￾toencoders, U-Net, and CNNs. Also for galaxies, we tested the ability of our autoencoder to make reliable predictions on unseen DESI galaxy spectra [PITH_FULL_IMAGE:figures/full_fig_p011_12.png]
Figure 11
Figure 11. Figure 11: Performance of NN predictions: histograms of the absolute fractional flux error (AFFE) for NNs applied to VIPERS galaxies for the novel architectures used in this work. ily restructuring the algorithms, the architectures used for the quasars undergo a new optimization…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Forecast for the detectability of patchy hydrogen reionization in WEAVE-QSO measurements of the Lyman-$\alpha$ forest power spectrum at redshift $z \geq 4$

    astro-ph.CO 2026-08 conditional novelty 6.0 of 10

    A mock-observation forecast finds that WEAVE-QSO measurements of the Lyman-alpha forest power spectrum at z≥4 can detect the relic large-scale power from patchy hydrogen reionization at 4.5 sigma.

  2. Uncertainty-Aware Deep Learning for the Ly$\alpha$ Forest: CNN-Based Absorber Detection and Characterization

    astro-ph.GA 2026-07 conditional novelty 5.5 of 10

    A sliding-window CNN recovers Lyα absorber locations and Voigt parameters from spectra, reproducing CDDF and b–N relations on mocks and, more weakly, on UVES data.

Reference graph

Works this paper leans on

10 extracted references · 9 canonical work pages · cited by 2 Pith papers

  1. [1]

    N., Adelman-McCarthy, J

    Abazajian, K. N., Adelman-McCarthy, J. K., Agüeros, M. A., et al. 2009, ApJS, 182, 543 Abraham, V ., Deville, J., & Kinariwala, G. 2024, Research Notes of the Ameri- can Astronomical Society, 8, 46 Aguirre, A., Schaye, J., & Theuns, T. 2002, ApJ, 576, 1 Ambrosch, M., Guiglion, G., Mikolaitis, Š., et al. 2023, A&A, 672, A46 Angthopo, J., Granett, B. R., La...

  2. [3]

    Pool Size: 2 Conv (2010,

  3. [6]

    Mask value: 0 Masking (2000,

  4. [9]

    Kernel Size: 2 Conv (2000,

  5. [11]

    Kernel Size: 5 Conv (2000,

  6. [12]

    Kernel Size: 1 Conv (2000,

  7. [13]

    The activation function (Act

    Kernel Size: 1 Notes. The activation function (Act. value in the table) used is the rectified linear unit (ReLU) function. Article number, page 20 of 23 Francesco Pistis et al.: Automated quasar continuum estimation using neural networks Appendix E: Fit example: autoencoders and CNNs Fig. E.1 shows examples of spectra for the case of giving to the optimiz...

  8. [32]

    Kernel Size: 5 Act.: ReLU MaxPooling (2010,

Show all 10 references
  1. [162]

    Drop rate: 0.2 Flatten 14720 Dense 512 Act.: ReLU Dense 1024 Act.: ReLU Dense 5000 Act.: ReLU VIPERS Input 2000 Masking 2000 Mask value: 0 Conv (2000,

  2. [256]

    The activation functions (Act

    Pool Size: 2 Flatten 14080 Dense 384 Act.: ReLU Dense 1024 Act.: ReLU Dense 2000 Act.: ReLU Notes. The activation functions (Act. value in the table) used are the rectified linear unit (ReLU), exponential linear unit (ELU), and linear activation. Article number, page 19 of 23 ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.