{"id":"ef1a29b2-cd91-4ee0-b6ab-eabfcd1b9eea","arxiv_id":"2505.01596","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Active learning with a self-organizing map outlier filter selects 6,700 DESI spectra, yielding QuasarNET weights that match eBOSS-trained performance with about one tenth the training data.","lead":"This paper retrains DESI's quasar-finding neural network on a small, carefully selected set of DESI spectra, using human labeling only for the most confusing objects. The new network matches the old one in quasar identification while using about a tenth of the training data, and it exposes a hidden systematic error in the network's redshift estimates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The completeness/purity comparison rests on a validation set that is a random split of the active-learning-selected sample; because that sample is deliberately enriched in high-entropy boundary spectra, the measured gain over eBOSS may not transfer to the 3.6-million-spectrum unlabeled pool.","rationale":"The reader's weakest assumption and my concern are the same: the 30% validation split of DESI AL2 is not representative of the full DESI spectral population, because the sample was constructed by active learning to concentrate on high-entropy, boundary objects. This is the load-bearing issue because the paper's headline claim is a comparative performance claim on that validation set, and the measured advantage could shrink or disappear on a representative sample. The paper has real strengths: the SOM-based outlier rejection is a genuine methodological addition, the bootstrapped dataset-size control is a thoughtful check, the code and data are released, and the repeat-exposure consistency test provides independent evidence of improved stability. However, none of these strengths replaces a truth-labeled, representative validation set. The paper itself strengthens the concern in Section 4.3.2 by admitting the training sample does not cover the entire redshift range, and Figure 7 shows the performance gain is not uniform across thresholds. I also note a secondary arithmetic inconsistency in Section 4.2: a total of 6698 spectra cannot split into 2010 validation and 6790 training spectra; this is not the central issue but suggests reporting care around the split. The proposed concrete test would settle whether the active-learning advantage persists on an unbiased sample. Since the reader's CONDITIONAL verdict already flags this issue and asks for addressable improvements, I do not recommend changing the verdict; UNCHANGED is appropriate.","tokens_in":18891,"tokens_out":4801,"duration_ms":54498,"concrete_test":"Construct an independent validation set by visually inspecting a random sample of roughly 1000-2000 objects from the same Guadalupe/DR1 unlabeled pool, stratified by redshift and signal-to-noise ratio, with no overlap with DESI AL2 and no entropy or SOM selection. Evaluate the DESI and eBOSS weights at the DESI confidence threshold of 0.95 and at neighboring thresholds on this sample. If DESI AL2 completeness and purity meet or exceed eBOSS within uncertainties, the central claim generalizes; if eBOSS is equal or better, the current validation was biased by active-learning selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Section 4.2, Figures 5-7) is that DESI AL2 meets or exceeds eBOSS DR12 completeness and purity at threshold 0.95 with less than one tenth of the training data. This is measured on a 30% validation split of DESI AL2, but DESI AL2 is not a random sample of DESI spectra: it is built by two active-learning rounds that deliberately select the 1000 highest-entropy spectra (Section 3, Eq. 3.2) after SOM outlier rejection. The validation set therefore overrepresents objects near QuasarNET's decision boundaries and underrepresents the bulk of the quasar population. Training on those boundary objects and then validating on a split of the same enriched distribution can inflate the apparent benefit of the active-learning data relative to what would be seen on a representative DESI sample. The paper's own Section 4.3.2 concedes that the training sample does not cover the entire redshift range, producing the redshift oscillations, and Figure 7 shows the improvement is threshold-specific. The repeat-exposure consistency test (Section 4.3.1) is independent but indirect: consistency is not accuracy, and the paper itself cautions that more consistent results need not be more correct. The bootstrapped dataset-size control (Figure 6) shows the added spectra help within the same biased distribution, but it does not establish representativeness of the full unlabeled pool. Thus the strongest claim is not established for the full DESI population until an independent, representative validation set is used.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes modifications to QuasarNET, a convolutional network used for quasar classification and redshift estimation in the DESI spectroscopic pipeline. The authors add padding and dropout to the architecture, then run an active learning loop that uses a 200-network bootstrap ensemble and a self-organizing-map outlier rejection step to select 2000 additional DESI spectra for visual inspection. The resulting 6698-spectrum DESI AL2 dataset is used to train a candidate weights file, which is compared with an eBOSS-trained weights file on a 30% validation split of DESI AL2 in completeness and purity, on repeat-exposure consistency over the unlabeled Guadalupe sample, and in redshift estimation. The paper reports comparable or better classification at the DESI threshold of 0.95 with less than a tenth of the eBOSS training data, improved consistency on repeated observations, and identifies box-edge oscillations in QuasarNET redshift estimates caused by incomplete redshift coverage in the active-learning-selected training sample.","tokens_in":19249,"tokens_out":6127,"duration_ms":59566,"significance":"The active learning pipeline is well motivated and the paper provides a useful practical demonstration: all plots are reproducible from the released code and data, the bootstrap control for dataset size is a valuable check, and the repeat-exposure consistency test is an informative indirect metric. The SOM-based outlier rejection is a sensible guard against labeling rare pathological spectra, and the discovery and characterization of the box-edge redshift oscillations is a concrete contribution to understanding QuasarNET. However, the central quantitative claim—that the DESI-trained weights match or beat the eBOSS weights on DESI data with 10% of the training data—is not yet established for the DESI population because the validation set is a split of the active-learning-selected sample rather than an independent random sample. The paper's own Section 4.3.2 shows that the selected sample does not cover the full redshift range. With an independent validation sample or a clearly restricted claim, the result would be a solid methods paper.","major_comments":[{"comment":"The headline comparison is measured on a 30% random split of DESI AL2, but DESI AL2 is built by two active-learning rounds that deliberately select the 1000 highest-entropy spectra after SOM outlier rejection, so it is not a random sample of the DESI spectral population. The validation split therefore overrepresents spectra near QuasarNET's decision boundaries and underrepresents outliers, and training on those boundary objects while validating on the same enriched distribution can inflate the measured completeness/purity gain relative to what would be observed on the full unlabeled pool. The bootstrapped size control in Fig. 6 shows the added spectra help within that distribution, but it does not establish representativeness. The authors should either validate on an independent set of DESI spectra not selected by active learning (for example, random or survey-targeted spectra with visual-inspection labels) or clearly restrict the central claim to the DESI AL2 validation distribution.","section":"Section 3, Eqs. (3.1)-(3.2); Section 4.2, Figs. 5-7"},{"comment":"The paper's own Section 4.3.2 and Conclusions concede that the final training sample does not cover the entire redshift range, producing redshift oscillations at the edges of QuasarNET's boxes, and that for DESI Year 3 processing the oscillations were returned to the eBOSS level by adding additional DESI training data cross-matched from eBOSS. This means the weights file evaluated in Section 4.2 is a candidate rather than the production weights, and the redshift pathology is not an irrelevant detail because QuasarNET redshifts drive the redrock rerun in the DESI pipeline. The article should state clearly which weights are being claimed as an improvement and should report the redshift-coverage limitation alongside the classification claim.","section":"Section 4.3.2; Conclusions"},{"comment":"The claim of meeting or exceeding eBOSS is threshold-specific: at lower confidence thresholds the eBOSS DR12 weights exceed the DESI AL2 weights in completeness with similar or better purity, as the authors state. Since the paper's title and abstract advertise improved quasar identification generally, the conclusions should carry an explicit caveat that the improvement is demonstrated only at the DESI nominal 0.95 threshold and on the AL2 validation split.","section":"Section 4.2, Fig. 7"},{"comment":"The repeat-exposure consistency results support the narrower claim that the DESI weights classify repeated observations more consistently, but consistency is not accuracy; the paper itself cautions that more consistent results need not be more correct. The abstract's wording 'more consistently classify objects in the same way' is appropriate, but it should not be summarized as improved classification accuracy on unlabeled data.","section":"Section 4.3.1, Figs. 8-9"}],"minor_comments":[{"comment":"The abstract in the paper text says 'achieve similar performance' while the arXiv abstract says 'meet or exceed'; these should be aligned, and the phrase 'systemic error' should be 'systematic error'.","section":"Abstract and Section 1"},{"comment":"The sentence 'DESI VI 2 outperforms the eBOSS weights file' appears to refer to DESI AL 2, not DESI VI 2; please correct the label.","section":"Section 4.2, after Fig. 7"},{"comment":"After removing the two corrupted spectra the total is 6698, but the text states that the validation set is 2010 spectra and 'the remaining 6790 are available for use as training data'; these numbers are inconsistent and should be reconciled.","section":"Section 4.2"},{"comment":"There is a duplicated word in 'Classification therefore only only uses the first 13 outputs'; remove the second 'only'.","section":"Section 1"},{"comment":"The dropout rate is reported as tuned empirically to 0.8, but no sensitivity analysis is shown; a brief statement of the range tested would help readers assess the robustness of the architecture choice.","section":"Section 2"},{"comment":"The figure and table numbering is clear, but the caption of Figure 5 refers to the eBOSS run as 'yellow' while the plot legend appears to use orange; please make the color references consistent.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central scientific issue is fixable: the authors need either an independent validation sample from the DESI spectral population or a clearly narrowed claim that the comparison applies only to the active-learning-selected distribution. Given the paper's own admission of redshift-coverage gaps, I do not see this as a reject; the methodology and reproducibility are solid strengths. One additional consideration is that the manuscript is submitted to JCAP but is largely an instrumentation/methods paper; it may fit better in an astro-ph.IM or astronomical methods venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a competent, useful pipeline paper: it retrains QuasarNET on DESI-only spectra using active learning, adds a SOM-based outlier rejection step, and in the process uncovers a genuine issue with QuasarNET's redshift estimates (box-edge oscillations). Code and data are on Zenodo, so the work is reproducible. Second, the headline quantitative claim—that the DESI-trained weights meet or exceed the eBOSS weights with less than a tenth of the training data—rests on a validation split of the same active-learning-selected sample, so as stated it does not generalize to the full DESI pool. That is a fixable flaw, not a fatal one.\n\nWhat is actually new: applying active learning to a 1D line-finder CNN rather than image classifiers, the bootstrap ensemble entropy, and the SOM outlier rejection. The visual inspection procedure is careful: three inspectors plus a merger, quality cuts, and a check that SOM-rejected outliers really are outliers. The bootstrap control in Figure 6 is a good check: it shows the added spectra help beyond naive dataset size, at least within the selected distribution. The redshift box-edge oscillation discovery is real and already fed into DESI Y3 processing.\n\nThe weak spot is Section 4.2. The validation set is a 30% split of DESI AL2, which is deliberately enriched in high-entropy, decision-boundary spectra. Training on boundary objects and validating on a split of the same distribution can inflate the apparent gain over the eBOSS model, which was trained on a different distribution. The paper itself concedes in Section 4.3.2 that the training sample does not cover the full redshift range, so the measured completeness and purity may not transfer to the 3.6-million-spectrum pool. The repeat-exposure consistency test is representative but indirect: consistency is not accuracy, as the authors themselves note. Figure 7 shows the improvement is threshold-specific; at lower confidence thresholds the eBOSS weights do better. There is also an abstract inconsistency: the arXiv abstract says 'meet or exceed,' the full-text abstract says 'similar performance.' That needs aligning.\n\nNone of this is fatal. The paper is transparent about limitations and the data release makes it reproducible. What it needs is an independent, representative validation set—a random DESI sample with visual-inspection labels, not a split of the active-learning-selected sample—to support the headline claim.\n\nBottom line: worth a serious referee. I would ask for the independent validation before accepting, but the paper deserves referee time and would be useful to anyone working on survey spectroscopy pipelines or active learning for spectral classification.","headline":"A solid, honest active-learning pipeline paper with a real redshift-discovery side effect, but the headline 'meet or exceed eBOSS with 10% of the data' needs an independent validation set before it can be believed at face value.","tokens_in":20075,"tokens_out":3513,"would_cite":true,"duration_ms":33029,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Active learning lets DESI train its quasar classifier on its own spectra and match the old BOSS-trained model with a tenth of the data, while uncovering a redshift bias.","keywords":["active learning","quasar classification","QuasarNET","DESI","self-organizing map","spectral classification","convolutional neural network","redshift estimation"],"falsifier":"Train the same modified QuasarNET on 6,700 spectra drawn uniformly at random from the same DESI visual-inspection pool, keeping the identical architecture, epochs, and validation split; if the random sample attains the same completeness and purity as the active-learned sample, then the gain attributed to active learning is not driven by which spectra were selected.","tokens_in":18715,"feed_emoji":"🔭","tokens_out":10133,"duration_ms":97029,"temperature":0.7,"pith_summary":"QuasarNET, the convolutional network that flags quasars inside DESI's spectral classification pipeline, was trained on spectra from the earlier BOSS survey and does not perfectly match DESI's data. This paper claims that an active-learning loop can rebuild it from DESI spectra alone: an ensemble of 200 bootstrapped networks scores millions of unlabeled spectra by confusion, a self-organizing map removes rare outliers, and human inspectors label the 1,000 most confusing surviving spectra in each of two rounds. The assembled DESI training set of about 6,700 labeled spectra meets or beats the previous BOSS-trained weights in completeness and purity on a validation split, with less than a tenth of the old training data, and classifies repeated exposures of the same object more consistently. The same comparison exposes a systematic error in QuasarNET's redshift estimates, traced to its box-based line-position outputs, which the paper uses to inform DESI Year 3 processing.","feed_headline":"Active learning retrains quasar finder on one tenth the data","feed_subtitle":"DESI's QuasarNET matches the old BOSS-trained model on 5,600 labeled spectra, with steadier repeat classifications.","key_machinery":"The machinery is an active-learning loop built around QuasarNET's line-finder outputs: for each of 13 wavelength boxes and each emission line, the network predicts a coarse confidence and a fine position estimate. A bootstrap ensemble of 200 QuasarNET copies, each trained on a resampled version of the current labeled set, classifies every unlabeled spectrum into three classes (NOT QSO, LOW-Z, HIGH-Z); disagreement defines an entropy $H = -\\sum_i C_i \\log_2 C_i$ that ranks spectra by confusion. A self-organizing map trained on the same rebinned spectra then rejects spectra in sparse cells (below the 15th percentile of cell counts), so the 1,000 highest-entropy surviving spectra sent to visual inspection are representative of the bulk quasar population rather than rare anomalies. The network itself is modified with zero padding, which permits two additional convolution layers, and dropout at rate 0.8 to improve generalization from the small training set.","core_discovery":"The central claim is that careful selection of training spectra, not raw data volume, is what lets a deep line-finder network adapt to a new spectrograph. After two active-learning iterations, the paper obtains a weights file trained on DESI visual-inspection labels (about 5,600 spectra selected by the algorithm, later grown to 6,700 by including inspected outliers) that matches or exceeds the previous eBOSS-trained weights at DESI's nominal confidence threshold, achieving comparable validation completeness and purity with less than one tenth of the training data. On the unlabeled DESI DR1 'Guadalupe' sample, the new weights reduce classification confusion between repeated exposures and lower the scatter of redshift estimates. In building this comparison the paper identifies a systematic redshift defect: QuasarNET's coarse and fine line-position estimates clump toward the centers of its 13 wavelength boxes, producing oscillations in redshift estimates at box edges that are more visible with the smaller DESI training sample; the paper reports that adding cross-matched training data for Year 3 mitigates the oscillations without changing the underlying architecture.","pith_inferences":["The same three-class entropy plus self-organizing-map outlier-rejection loop should transfer to any line-detection network in a large survey; the ingredients are an ensemble, a confusion score, and a density map over inputs, none of which are specific to quasars.","A natural extension the paper does not pursue is to replace the 13-box coarse/fine output with a continuous position regression or add a loss term that penalizes box-center clustering; that would attack the oscillation at its source rather than mitigating it with more training data.","If the validation numbers generalize, only a few thousand human labels were needed to retrain the classifier for a new instrument, suggesting visual-inspection budgets for future surveys can be reduced by orders of magnitude for similar retraining tasks.","The reported completeness and purity are measured on a holdout drawn from the same active-learning-selected pool; an independent test set drawn from the full survey, not from the selected pool, is the direct check on whether the gains carry to the 3.6 million unlabeled spectra."],"forward_implications":["A DESI-only training set of roughly 6,700 labeled spectra is sufficient to replace the BOSS-era weights at DESI's nominal confidence threshold, matching or exceeding old completeness and purity.","Repeated-exposure tests show the new weights classify the same object more consistently and with lower redshift scatter, which should make the DESI pipeline's quasar flags more stable across epochs.","The bootstrap experiment indicates that the active-learning selection, not just increased dataset size, expands the quasar space covered by training, so future retraining runs can expect the same procedure to beat naive data collection.","The box-edge redshift oscillations are a property of QuasarNET's 13-box line-position representation, so any future weights file for DESI will need either denser redshift coverage in the training sample or a change to the network output.","The completeness gain is specific to the 0.95 confidence threshold used during active learning; at lower thresholds the eBOSS weights retain higher completeness, so threshold choice is part of the method."],"supporting_citations":[{"why":"Defines the original QuasarNET convolutional architecture and 13-box line-finder outputs that this paper modifies and retrains.","marker":"[21]"},{"why":"Provides the eBOSS DR12 visual-inspection labels used to train the baseline weights file and to augment the first active-learning pass.","marker":"[23]"},{"why":"Supplies the 3,761 visually inspected DESI quasar spectra that seed the initial training set.","marker":"[27]"},{"why":"Supplies the 839 non-quasar DESI spectra added so the initial training set is not overly quasar-dense.","marker":"[28]"},{"why":"Supplies the active-learning uncertainty-sampling methodology the paper adapts to QuasarNET's line-finder outputs.","marker":"[29]"},{"why":"Defines self-organizing maps, the mechanism behind the novel outlier-rejection step.","marker":"[32]"},{"why":"Provides prior validation of the eBOSS weights' redshift estimates on DESI data, the baseline for the redshift consistency comparison.","marker":"[22]"}],"fun_headline_variants":["Active learning trims quasar training data by 90%","Smarter data selection beats volume for DESI quasar finder","QuasarNET retrained with 1/10th the data, matches old model","Active learning exposes redshift flaws in DESI quasar classifier","Five thousand smart labels outperform 1.3M quasar targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 30 percent of the active-learning-selected, visually inspected spectra held out for validation represents the full population of DESI spectra the classifier will encounter; the paper's own redshift analysis shows the training sample does not cover the entire redshift range, so completeness and purity measured on that holdout may not transfer to the 3.6 million unlabeled spectra.","fun_headline_variants_meta":{"raw":{"variants":["Active learning trims quasar training data by 90%","Smarter data selection beats volume for DESI quasar finder","QuasarNET retrained with 1/10th the data, matches old model","Active learning exposes redshift flaws in DESI quasar classifier","Five thousand smart labels outperform 1.3M quasar targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1465,"prompt_tokens":1012,"completion_tokens":453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":361}},"tokens_in":628,"tokens_out":453,"duration_ms":4354,"temperature":1.0,"reasoning_tokens":361,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:15:28.338875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same modified QuasarNET on 6,700 spectra drawn uniformly at random from the same DESI visual-inspection pool, keeping the identical architecture, epochs, and validation split; if the random sample attains the same completeness and purity as the active-learned sample, then the gain attributed to active learning is not driven by which spectra were selected.","supporting_citations":[{"cited_title":"Alexander, T.M","cited_arxiv_id":null,"evidence_quote":"Supplies the 3,761 visually inspected DESI quasar spectra that seed the initial training set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 839 non-quasar DESI spectra added so the initial training set is not overly quasar-dense."},{"cited_title":"Settles,Active Learning, Synthesis Lectures on Artificial Intelligence and Machine Learning, Springer International Publishing, Cham (2012)","cited_arxiv_id":null,"evidence_quote":"Supplies the active-learning uncertainty-sampling methodology the paper adapts to QuasarNET's line-finder outputs."},{"cited_title":"Kohonen,Self-organized formation of topologically correct feature maps, Biological Cybernetics 43 (1982) 59","cited_arxiv_id":null,"evidence_quote":"Defines self-organizing maps, the mechanism behind the novel outlier-rejection step."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides prior validation of the eBOSS weights' redshift estimates on DESI data, the baseline for the redshift consistency comparison."}],"review_version":1}