{"id":"560e5390-1281-4d7f-880d-0fb753764012","arxiv_id":"2505.10976","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A simple autoencoder estimates quasar and galaxy continua with roughly one percent median error and generalizes to unseen DESI spectra.","lead":"The authors trained three neural networks to estimate the smooth continuum under quasar and galaxy spectra, and found that a simple autoencoder is both accurate and fast. It reaches about one percent error on mock spectra and, applied to real DESI data, recovers known intergalactic hydrogen absorption.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The DESI tau_eff test compares uncorrected (possibly median) measurements with corrected literature values; a compensating continuum bias in the Ly-alpha forest would produce exactly this agreement, so the cross-survey transfer claim is not independently established.","rationale":"The central claim has two legs: accurate continuum reconstruction on WEAVE mocks and successful transfer to DESI. The mock leg is internally sound: true continua are known, the AFFE metric is well defined, and Appendix D gives enough architectural detail to reproduce the models. The transfer leg, however, rests almost entirely on the DESI tau_eff comparison, because no true continua are available for DESI quasars. That comparison is between an uncorrected raw estimator and corrected literature values, so the direction of the implied continuum bias is fixed: to match corrected curves, the autoencoder must underestimate the true continuum in the forest by about the metal+LLS opacity. The paper itself supplies reasons such a bias is plausible (Fig. 2 slope offset, Appendix A magnitude-dependent FFE), and the latent-space covariate-shift test is not a check against real data. The reader's conditional verdict already requires addressing this cluster of issues, and this stress-test does not move the verdict; it sharpens the reason the condition is necessary. The proposed Redrock control is cheap and decisive: DESI EDR already contains an independent PCA continuum, so the same estimator applied to it would reveal whether the autoencoder's raw curve is anomalously low. Credit is given for the explicit limitation statements, the detailed architecture tables, and the honest framing that the tau_eff test is not a new measurement; those features make the missing control the only serious obstacle to the paper's headline transfer claim.","tokens_in":29212,"tokens_out":12551,"duration_ms":131168,"concrete_test":"On the same DESI EDR quasar sample, wavelength range (1070-1160 Å), and redshift bins, recompute the raw tau_eff from Eq. (4) using the Redrock PCA continuum from the DESI EDR catalog as F_cont, and plot this raw curve together with the autoencoder raw curve and the corrected Becker/Turner curves. If Redrock's raw curve lies above the corrected curves by the expected metal+LLS offset while the autoencoder's raw curve lies on them, the autoencoder continuum is biased low in the forest and the transfer claim is not supported; if the two raw curves coincide, the compensating-bias concern is resolved. The same comparison should be done with the mean, not the median, of the transmission if the text's definition is followed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 validates cross-survey transfer with Fig. 8. The tau_eff used for the autoencoder is derived from Eq. (4) with no corrections; the text explicitly says 'We also do not apply any correction for optically thick absorbers (Becker et al. 2013) or metals (Schaye et al. 2003...)'. The Becker et al. and Turner et al. curves shown for comparison are corrected for metals and optically thick absorbers, as the text states for Turner et al. An unbiased continuum should therefore give a raw tau_eff above the corrected literature curves by the combined metal+LLS opacity. The raw autoencoder curve lying on the corrected curves implies the autoencoder systematically underestimates the true DESI continuum in the Ly-alpha forest by roughly that opacity. This plausible compensating error is not excluded by the paper: Fig. 2 shows a slope offset between the WEAVE mock continua used for training and DESI, and Appendix A shows the fractional flux error depends on the magnitude cut used to build the training sample. The covariate-shift exercise in Sect. 5.1 is internal to the mock distribution and does not constrain real-data bias. In addition, Sect. 5.2 first defines a mean transmission but then reports a median transmitted flux, so the estimator may not even be the same quantity as in the literature it is compared with. The DESI test, as presented, does not independently establish that the mock-trained continuum is unbiased on real data; the mock AFFE remains a self-consistent but not cross-survey validation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript compares three deep-learning architectures (an autoencoder, a CNN, and a U-Net) for automated quasar continuum estimation in the rest-frame range 1020–2000 Å. The models are trained on WEAVE mock quasar spectra built from BOSS PCA continuum shapes and hydrodynamical Lyα forest skewers, and are benchmarked against the published iQNet and LyCAN networks using the absolute fractional flux error (AFFE). The best-performing model is then applied without retraining to DESI EDR quasar spectra to measure the Lyα effective optical depth evolution, and, after retraining, the same architectures are applied to galaxy spectra from VIPERS and DESI to measure the D4000n break. The paper argues that the autoencoder performs as well as more complex architectures at much lower computational cost and generalizes across surveys.","tokens_in":29579,"tokens_out":6168,"duration_ms":60286,"significance":"If the reported results are correct, the paper provides a useful comparative benchmark for continuum-fitting methods needed for WEAVE and other large spectroscopic surveys. The study design is sound in its main components: training on realistic mocks, evaluating on held-out mock spectra, applying the trained model to independent DESI data, and comparing with two published baselines. The extensive appendices on magnitude cuts, S/N dependence, and redshift/SNR biases are valuable, and the public Data availability of the underlying surveys is a strength. However, the headline claim about the autoencoder being the best architecture is contradicted by the paper's own figures, and the DESI Lyα optical-depth test is presented in a way that does not independently establish unbiased continuum prediction on real data. These issues need to be resolved before the paper's central conclusions can be accepted.","major_comments":[{"comment":"The abstract states that 'the autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars.' This value is not the autoencoder's median AFFE in Fig. 4: the autoencoder red-part and full-spectrum medians are 0.007 and 0.008, while 0.009 is the median for iQNet. More importantly, Fig. 4 shows that the CNN trained on the red part achieves a median AFFE of 0.006, which is lower than any autoencoder result. Fig. 10, presented in the Summary, reports CNN red part at 0.009 and CNN full spectra at 0.006, while the autoencoders are 0.008 and 0.007; this contradicts Fig. 4, where CNN red part is 0.006 and CNN full spectra is 0.010. In either figure, the best quasar continuum model is a CNN, not the autoencoder. The abstract and the conclusions in Section 8 therefore misstate the main comparative result, and the discrepancy between Fig. 4 and Fig. 10 must be resolved.","section":"Abstract; Section 8; Figs. 4 and 10"},{"comment":"The comparison between the autoencoder's raw, uncorrected τeff and the corrected literature curves is not a valid test of continuum unbiasedness. The text explicitly states that no correction for optically thick absorbers or metals is applied, while the Turner et al. (2024) values are 'bias-corrected measurements corrected for metal line absorption... and optically thick absorbers,' and the Becker et al. (2013) fit is likewise a corrected measurement. An unbiased continuum should yield a raw τeff higher than the corrected curves by roughly the metal plus LLS opacity; the fact that the raw measurement lies on the corrected curves implies a systematic underestimate of the true DESI continuum in the Lyα forest. The covariate-shift exercise in Section 5.1 is internal to the mock distribution and does not constrain real-data bias, especially because Fig. 2 shows a slope offset between the WEAVE mocks and DESI. In addition, Section 5.2 defines a mean transmission in Eq. (4) but then reports the median transmitted flux, which is not the same statistic as the literature values with which it is compared. As presented, the DESI τeff test does not independently establish that the mock-trained continuum is unbiased on real data; the authors should either apply the corrections, compare with raw uncorrected literature measurements, or explicitly reframe the test as a consistency check that does not by itself demonstrate zero continuum bias.","section":"Section 5.2, Eq. (4), Fig. 8"},{"comment":"The VIPERS D4000n validation is partially circular and should be identified as such. According to Section 2.2, the NN training labels for galaxies are the pPXF stellar-continuum fits; the left-hand panel of Fig. 9 compares the NN-derived D4000n against D4000n computed from the same pPXF continua. Agreement in that panel largely confirms that the network has learned the pPXF labels, not that it generalizes to an independent estimate. The DESI/Redrock comparison in the right-hand panel is the genuinely independent test, since Redrock is a separate pipeline. The text should explicitly state that the VIPERS comparison is a reproducibility check, not independent generalization evidence, and should not use the VIPERS correlation coefficient as support for the generalization claim.","section":"Section 7, Fig. 9"}],"minor_comments":[{"comment":"There are numerous typographical and spacing issues, including 'di fferent' (multiple instances), 'untractable' (Section 1), 'versitile', 'datatsets', 'availaility', 'modfications', 'res-frame', and 'WEA VE' with a stray space; these should be corrected in a careful proofreading pass.","section":"Throughout"},{"comment":"The sentence 'WEA VE mocks generally overlap with the DESI data for the R parameter' refers to the flux ratio FR defined in the same paragraph; the notation should be consistent.","section":"Section 2.3, Fig. 2"},{"comment":"The cross-reference 'Sect.??' is unresolved; the intended section on galaxy generalization should be cited explicitly.","section":"Section 4.1"},{"comment":"The text says 'we computed the median transmitted flux ⟨f⟩' after defining the mean transmission in Eq. (4); the statistical choice should be stated upfront and justified, and the estimator should match the quantity compared with the literature.","section":"Section 5.2"},{"comment":"The caption and text use 'res-frame' and 'functional of the rest-frame wavelength'; these should be corrected to 'rest-frame' and 'as a function of'.","section":"Appendix A, Fig. A.1"}],"recommendation":"major_revision","confidential_remarks":"The discrepancy between Fig. 4 and Fig. 10 is a red flag: if either figure is correct, the paper's central claim that the autoencoder is the best quasar model changes. It would be worth asking the authors to provide the exact per-run median AFFE values in a table to eliminate ambiguity. The Section 5.2 comparison is also a substantive issue that likely requires re-analysis, not just rewording, and the VIPERS circularity should be acknowledged in the main text. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing: this is the first head-to-head comparison of autoencoder, CNN, and U-Net for quasar continuum fitting on WEAVE-specific mocks, with iQNet and LyCAN as baselines. The mock setup is well described, the appendices explore S/N and magnitude biases carefully, and the conclusion that a lightweight autoencoder reaches ~0.7–0.8% median AFFE at a fraction of the training cost of CNNs is credible and useful for the WEAVE survey. The galaxy generalization test with minimal retuning is a nice addition, even if it is explicitly not a new galaxy method.\n\nNow the soft spots. The abstract misstates the headline numbers: Fig. 4 gives the optimized autoencoder a median AFFE of 0.007–0.008, not 0.009 (that is iQNet's number), and the CNN trained on the red part reaches 0.006, beating the autoencoder. The abstract's \"autoencoder outperforms the U-Net\" is true, but it ignores the better CNN result. Easy fix, but it makes the abstract look careless.\n\nMore substantively, the DESI tau_eff test does not support the strength of the generalization claim. The paper computes a raw, uncorrected median optical depth and compares it with literature values that are corrected for metals and optically thick absorbers. An unbiased continuum should yield a raw tau_eff above the corrected curves (extra opacity from metals and LLS). That the raw curve lands on the corrected curves suggests a compensating continuum bias in the forest. The covariate-shift test in Sect. 5.1 is internal to the mock distribution and does not constrain real-data bias. Also, the text defines a mean transmission but then reports the median. The authors are transparent about not applying corrections, and they frame this as a sanity check, but the conclusion that the autoencoder shows 'excellent' cross-survey generalization is stronger than the evidence.\n\nThe galaxy validation on VIPERS is partly circular: the pPXF continua are the training labels, so the VIPERS D4000n comparison is a reproducibility check. The DESI D4000n comparison against Redrock is independent and helps. Finally, no code or data is released, which limits adoption for a methods paper.\n\nNone of this sinks the central result. The mock-based benchmark is solid, the comparison with iQNet and LyCAN is fair, and the paper is honest about many caveats (BAL removal, no skylines, magnitude-dependent biases). The tau_eff issue is the one that needs attention before the cross-survey claim is taken at face value.\n\nRecommendation: send it to a serious referee. With a corrected abstract, a reframed tau_eff test (or corrected values), and released code, it would be a useful reference for the WEAVE and DESI communities.","headline":"Solid mock-based benchmark, but the abstract misreports the headline numbers and the DESI tau_eff test overstates cross-survey generalization.","tokens_in":30166,"tokens_out":3785,"would_cite":true,"duration_ms":36453,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A lightweight autoencoder estimates quasar continua at roughly one percent error and transfers to unseen surveys without retraining.","keywords":["quasar continuum estimation","Lyman-alpha forest","neural networks","autoencoder","convolutional neural network","U-Net","intergalactic medium","spectroscopic surveys"],"falsifier":"Apply the mock-trained autoencoder to spectra where the true continuum is directly measurable — for instance, BAL quasars (excluded from training) observed at rest-frame wavelengths well redward of the Lyman-alpha line — and recompute the Lyman-alpha effective optical depth after applying the standard metal and optically-thick-absorber corrections that this paper omits; if the corrected values drift off the reference curve while a comparison network's values stay on it, the reported agreement is a coincidence of uncorrected systematics rather than continuum accuracy.","tokens_in":29062,"feed_emoji":"🔭","tokens_out":14884,"duration_ms":123968,"temperature":0.7,"pith_summary":"This paper claims that a lightweight autoencoder — a neural network that compresses a spectrum through a narrow bottleneck and then reconstructs it — can estimate quasar continua at sub-percent accuracy and, trained only on simulated spectra, can be applied without retraining to real spectra from a different survey. Comparing an autoencoder, a convolutional network, and a U-Net on mock spectra built for the WEAVE survey, the authors find the autoencoder clearly outperforms the U-Net and matches the CNN while consuming roughly an order of magnitude less compute. Applied unchanged to DESI Early Data Release quasars, the autoencoder reproduces the redshift evolution of the Lyman-alpha effective optical depth found by earlier dedicated measurements, and with modest retuning the same architecture fits galaxy continua and recovers the D4000n break. If the paper is right, a single small network can serve as a fast, survey-agnostic continuum estimator for the million-spectrum surveys now underway.","feed_headline":"Fits quasar continua to ~1% with a simple autoencoder","feed_subtitle":"Trained only on mock spectra, it reproduced the Lyman-alpha forest's measured opacity in DESI data.","key_machinery":"The load-bearing object is the autoencoder with a latent bottleneck: three dense layers compress each input spectrum into a low-dimensional code, and a decoder reconstructs the continuum from that code, with a masking layer that zeroes out pixels missing at the spectral edges. The decisive design choice is the input-output split — the network sees only the clean red part of the spectrum ($\\lambda > 1216$ Å) and must predict the full continuum including the Lyman-$\\alpha$ forest ($\\lambda < 1216$ Å), forcing it to learn the continuum shape from an uncontaminated region and extrapolate into the absorbed one. The comparison machinery is the absolute fractional flux error (AFFE), the wavelength-averaged absolute fractional difference between predicted and true continuum, supplemented by covariate-shift tests in which the latent space is split in two and the network retrained on one half to see how predictions degrade outside the training distribution. The same architecture, re-optimized but not restructured, is applied to galaxy continua with the full spectrum (3500–5500 Å) as input.","core_discovery":"On the paper's own terms, the central discovery is that the simplest architecture wins: an optimized three-layer autoencoder reaches a median absolute fractional flux error (AFFE) of about 0.007–0.009 on mock WEAVE quasar spectra, versus 0.013 for the U-Net, with flatter wavelength-dependent bias than the published iQNet baseline. The network is trained on the relatively uncontaminated red side of the quasar spectrum ($\\lambda > 1216$ Å) and asked to predict the full continuum from 1020 to 2000 Å, including the Lyman-$\\alpha$ forest region where the true continuum is hidden behind dense absorption; passing the full spectrum as input does not materially improve the fit. Trained only on mocks, the autoencoder reproduces the redshift evolution of the Lyman-$\\alpha$ effective optical depth, $\\tau_{\\rm eff}(z)$, in DESI Early Data Release quasars in agreement with earlier measurements (Becker et al. 2013; Turner et al. 2024), which the paper reads as evidence of genuine cross-survey generalization. For galaxies, a fresh optimization pass with the same architecture reaches a median AFFE of 0.014 against pPXF continua and reproduces the D4000n break in VIPERS and DESI data.","pith_inferences":["Because the reported Lyman-alpha optical depth recovery skips the standard corrections for metal absorption and optically thick absorbers, both of which shift the measurement in the same direction, my read is that part of the agreement with the reference curve may reflect compensating errors; a corrected comparison would be the sharper test.","The model's blind spot is precisely the quasars it cannot see: broad absorption line (BAL) quasars are removed from the mocks, so the natural next test is whether fine-tuning on a few hundred BAL spectra lets the same architecture track their heavily distorted continua.","Figure 2 shows the mock continua have systematically shallower slopes than real DESI spectra, so the reported transfer accuracy is achieved despite a known distributional mismatch; enriching the mock library with a wider spread of continuum slopes should push survey-transfer errors lower.","The bottleneck code is a compressed, survey-agnostic description of the continuum shape, so the same network could plausibly double as a continuum prior for absorption-line analyses, an extension the paper does not test."],"forward_implications":["A WEAVE-trained quasar model can be deployed on DESI data with no retraining and still recover the known evolution of the Lyman-alpha effective optical depth, so cross-survey transfer of continuum estimators is practically achievable.","Continuum fitting at the percent level becomes cheap enough to run on entire surveys: the paper estimates roughly 3 minutes to fit one million spectra with the autoencoder, versus about 40 minutes with the CNN.","The same architecture family handles a very different continuum shape, galaxy spectra at 3500–5500 Å, after only a new optimization pass, with the autoencoder at median AFFE 0.014 and a working D4000n measurement.","Extra architectural complexity buys nothing for this problem: the U-Net's deeper encoder-decoder with skip connections returns median AFFE 0.013, roughly double the autoencoder's error.","A raw, uncorrected Lyman-alpha optical depth measurement computed from autoencoder continua lands within the scatter of the reference curve, indicating that percent-level continuum errors still propagate into a scientifically usable absorption measurement."],"supporting_citations":[{"why":"Supplies the iQNet autoencoder architecture (encoder, bottleneck, decoder) that this paper takes as its starting point and re-optimizes.","marker":"Liu & Bordoloi (2021)"},{"why":"Supplies the LyCAN CNN architecture used as a baseline and the DESI-based effective optical depth measurement that the autoencoder result is compared against.","marker":"Turner et al. (2024)"},{"why":"Provides the BOSS principal components from which the continuum shapes of the WEAVE mock quasar spectra are drawn.","marker":"Pâris et al. (2011)"},{"why":"Provides the hydrodynamical simulation skewers used to insert Lyman-alpha forest absorption into the mock spectra.","marker":"Bolton et al. (2017)"},{"why":"The reference measurement of the Lyman-alpha effective optical depth evolution that the DESI-based result is compared against.","marker":"Becker et al. (2013)"},{"why":"The original U-Net architecture that the paper adapts to one-dimensional spectra.","marker":"Ronneberger et al. (2015)"},{"why":"Supplies the pPXF full-spectrum fitting code used to define the reference stellar continua for the VIPERS galaxy spectra.","marker":"Cappellari (2017, 2023)"},{"why":"The source of the real DESI Early Data Release quasar and galaxy spectra used for the generalization tests.","marker":"DESI Collaboration et al. (2022, 2024)"}],"fun_headline_variants":["Simple autoencoder beats U-Net on quasar continua","Neural net fits quasar continua to ~1% accuracy","Autoencoder nails quasar continua, generalizes to galaxies","Three-layer net predicts quasar continua from mocks alone","Simple net matches Lyman-alpha forest opacity from mock data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on one premise: the mock WEAVE spectra the network trains on — BOSS principal-component continuum shapes combined with simulated Lyman-alpha forest absorption — are representative enough of real quasar continua that a model which has only ever seen mocks can be trusted on DESI data, a premise the paper itself flags as imperfect when it shows the mock continua have systematically shallower slopes than the real spectra.","fun_headline_variants_meta":{"raw":{"variants":["Simple autoencoder beats U-Net on quasar continua","Neural net fits quasar continua to ~1% accuracy","Autoencoder nails quasar continua, generalizes to galaxies","Three-layer net predicts quasar continua from mocks alone","Simple net matches Lyman-alpha forest opacity from mock data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1404,"prompt_tokens":1082,"completion_tokens":322,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":240}},"tokens_in":698,"tokens_out":322,"duration_ms":3275,"temperature":1.0,"reasoning_tokens":240,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:59:29.909939+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the mock-trained autoencoder to spectra where the true continuum is directly measurable — for instance, BAL quasars (excluded from training) observed at rest-frame wavelengths well redward of the Lyman-alpha line — and recompute the Lyman-alpha effective optical depth after applying the standard metal and optically-thick-absorber corrections that this paper omits; if the corrected values drift off the reference curve while a comparison network's values stay on it, the reported agreement is a coincidence of uncorrected systematics rather than continuum accuracy.","supporting_citations":[],"review_version":1}