{"id":"73657a72-089c-4539-b8a9-eef99a1cdad1","arxiv_id":"2607.22836","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"1,029 cataclysmic variables (221 new) were identified from 98.97 million DESI spectra using a CNN trained on SDSS CVs.","lead":"Researchers used a machine-learning filter to scan 98.97 million spectra from the DESI survey and identified 1,029 cataclysmic variables (CVs), 221 of them new. The catalogue is the largest spectroscopically selected CV sample and includes rare types that photometric surveys miss.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CNN training set lacks absorption-dominated AM CVn spectra; §6 shows the prototype was missed, so the ~99% completeness claim is unsupported for unseen spectral classes.","rationale":"The reader's weakest_assumption identifies the representativeness of the CNN training set. I agree, and want to sharpen it: the paper itself provides a demonstrated false negative (§6, J1234+3737), so the question is not hypothetical. The completeness argument in §4.5 has a logical gap: a random sample drawn from the CNN-selected set can only estimate the false-positive rate (precision), not the false-negative rate (recall). The 97.2% recall figure in §4.5 is computed by checking spectra of CVs already discovered by other techniques; it cannot detect CVs missed by all techniques. The validation set in §4.3.4 is the only independent recall check, but it is drawn from a white-dwarf candidate sample and contains only two missed CVs, both with emission or normalization issues; no absorption-dominated AM CVn is reported. Therefore the recall for that morphology is unmeasured. This matters because the central claims include 'ten AM CVn systems' and the first absorption-dominated AM CVn prototype is known to be missed. A targeted search for AM CVn-like spectra in the unselected DESI set is the decisive test. Provided such a search returns no new systems, the completeness claim is grudgingly supported; otherwise the catalogue and space densities must be revised. I therefore recommend keeping the reader's CONDITIONAL verdict: the catalogue is likely real, but the completeness/population sub-claims need further evidence.","tokens_in":44523,"tokens_out":7070,"duration_ms":74950,"concrete_test":"Run a template-matching or line-index search over all 80,767,382 DESI science spectra for absorption-dominated AM CVn signatures (e.g., strong He I 4471/5876 absorption, weak or absent Balmer emission, blue continuum), using the known prototype J1234+3737 as a template, and visually inspect all candidates. If this search finds CVs that the CNN did not flag, the CNN recall is not uniform across spectral classes and the ~99% completeness claim fails; if it finds none, the completeness claim is supported for this class.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.5 claims the DESI CV sample is ≃99 per cent complete, based on the CNN flagging all CV spectra and on a zero-detection check of 9,954 CNN-selected, otherwise-uninspected spectra. That check bounds only false positives among selected spectra; it says nothing about false negatives among the 80,767,382 unselected science spectra. The paper's own §6 provides a concrete false negative: the CNN did not identify the AM CVn prototype J1234+3737 because its helium-absorption spectrum was not represented in the training set (Table 4, 304 SDSS CV spectra, all emission-line dominated). The other selection routes (coordinate cross-match with 16,277 known CVs, REF_CAT/Gaia filtering, redshift cut) will not recover an unknown, absorption-dominated CV that the CNN fails to flag. The §4.3.4 validation sample (580 spectra / 221 CVs) is drawn from a white-dwarf-targeted programme and evidently lacks such systems; it therefore cannot constrain the false-negative rate for this class. If additional absorption-dominated or otherwise morphologically novel CVs exist in the uninspected majority of DESI DR2, the number of CVs, the ten AM CVn systems, and the §10 space densities are all biased low, and the 99% completeness statement is wrong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a systematic search for cataclysmic variables in DESI DR2, analysing 80,767,382 science spectra. Three shortlisting routes are used: coordinate cross-match with 16,277 known/candidate CVs, a CNN trained on SDSS CV spectra, and a redshift cut; the resulting candidates are manually inspected with the aid of Gaia, GALEX, ZTF, and archival databases. The authors identify 1,029 CVs, of which 221 are new, including ten AM CVn systems and twelve members of a peculiar state-change class; they add 84 new or improved orbital periods and derive revised CV subtype space densities. The catalogue is supported by an independent validation set of 580 DESI CV spectra from a white-dwarf-targeted programme and by a detailed comparison with the Hou et al. (2026) DESI DR1 CV sample.","tokens_in":44841,"tokens_out":5861,"duration_ms":62697,"significance":"If the completeness and space-density claims hold, this is the largest spectroscopically selected CV sample to date, roughly two magnitudes deeper than SDSS, and would provide a largely outburst-unbiased sample for population studies. The paper is commendably transparent: it releases the full catalogue, per-CV spectra, and training/test lists; it documents an independent validation set with only two non-detections (with stated causes); and it explicitly re-examines and reclassifies 13 objects from Hou et al. (2026). The identification of twelve new peculiar state-change CVs and five 'highly evolved donor' candidates are interesting scientific results in their own right. These strengths are partly offset by an overstated completeness statement and by an apparent inconsistency in the completeness corrections used for the space-density estimates.","major_comments":[{"comment":"The claim 'the sample is ≃99 per cent complete' is not supported by the tests described. The random sample of 9,954 spectra is drawn from the 293,670 CNN-selected spectra that were not otherwise inspected, so it bounds false positives among the selected set, not false negatives among the 80,767,382 unselected science spectra. The 234-spectrum TARGETID check measures the CNN's flagging rate for spectra of CVs that were already identified by another route, not its sensitivity to unfamiliar CV spectral types. Section 6 supplies a concrete false negative: the AM CVn prototype J1234+3737 was missed because its helium-absorption spectrum was not represented in the training set. The validation sample of §4.3.4 is drawn from a white-dwarf-targeted programme and did not include such systems. Therefore the 99% completeness statement is unsubstantiated, and the §10 space densities—which depend on c","section":"§4.5, §6"},{"comment":"The 'perfect' confusion matrix in Fig. 2 is an optimistic selection artifact. The model was chosen among 200 retrainings on the basis of the test-set confusion matrix, so the quoted zero false-positive/zero false-negative result is not an honest out-of-sample evaluation. The independent validation set of 580 spectra in §4.3.4 is better evidence and should be reported as the primary sensitivity estimate. Please report the distribution of test metrics across the 200 splits, or use a validation split during model selection, and state the final CNN sensitivity using the held-out validation sample.","section":"§4.3.3, Fig. 2"},{"comment":"There is an internal inconsistency between the claimed completeness and the completeness corrections used in the space-density calculation. Section 4.5 states the sample is ≃99 per cent complete, but Section 10 adopts completeness values of 0.6 for short-period subtypes and 0.2 for long-period subtypes, and Eq. (2) applies the inverse of these as a correction factor to N_obs. These two statements cannot both describe the same selection pipeline. Clarify what the 99% statement refers to (e.g., contamination among the selected spectra) and justify the 0.6/0.2 values independently of that claim, or recalculate the space densities consistently.","section":"§10, Eq. (2)"}],"minor_comments":[{"comment":"'We found 1029 CVs among the DESI DR1 spectra' should read DR2; the analysis covers DESI DR2 (Section 4.1).","section":"§5.1"},{"comment":"The caption says the 150-pc 'All CVs' bar suffers from only seven CVs, whereas Table 7 lists N=8 for that row; please reconcile.","section":"Fig. 33 caption"},{"comment":"The term 'sample' is ambiguous in the completeness discussion: it is not clear whether '≃99 per cent complete' refers to the final CV list, the CNN-selected set, or the manually inspected set. Please define it explicitly.","section":"§4.5"},{"comment":"Swan et al. (submitted) is cited without a year or arXiv number; please give a fuller reference or state its status.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a strong catalogue paper whose main deliverable will be widely used. The referee's main concern is not the reality of the 1,029 CVs but the quantitative completeness and the derived space densities. The manuscript should separate what is measured (false-positive contamination of the selected set) from what is assumed (CNN sensitivity to all CV spectral types), and the space-density estimates should be made consistent with the completeness analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real and worth having: 1,029 CVs pulled from the full DESI DR2 spectral set, 221 new, ten AM CVn, 84 new or improved periods, and a dozen more members of the peculiar state-change class. That is a genuinely useful resource, and the bulk of the catalogue work — cross-matching, three-way shortlisting, manual inspection of 41,578 spectra, independent validation on 580 spectra — looks careful. The paper also does a nice job documenting why so many CVs were serendipitously targeted by galaxy and quasar surveys, which will be a useful reference for future DESI-based source searches.\n\nThe soft spots are real but concentrated in the completeness and space-density sections, not in the catalogue itself. The CNN is a replication of Inight et al. (2025), fine in principle, but the model was chosen out of 200 retrainings based on the test confusion matrix. That makes the 'perfect' test performance an optimistic selection artefact, not a clean performance estimate. More importantly, the '~99 per cent complete' claim in Section 4.5 is only supported by a zero-detection check of 9,954 CNN-flagged spectra. That check constrains false positives, not false negatives. The paper's own Section 6 gives a concrete counterexample: the AM CVn prototype J1234+3737 was missed because its helium-absorption spectrum was not represented in the training set. Other absorption-dominated or otherwise morphologically novel CVs could be sitting in the 80 million unselected science spectra. So the catalogue is likely a lower bound on DESI's true CV population, and the space-density estimates inherit that bias. The hand-set completeness values (0.6 for short-period, 0.2 for long-period) are also fairly ad hoc, though at least they are stated openly and the comparison figure is useful.\n\nThe circularity burden is real but not disqualifying. The method, the training labels, and the completeness reference all come from the authors' own SDSS work, and the CNN code and weights are not released. That limits independent reproduction. But the central catalogue claim — 1,029 CVs, 221 new — does not depend on the completeness or space-density sub-claims, and the catalogue is the main product.\n\nWould I referee it? Yes. The catalogue deserves to be in the literature, but the completeness and density sections need a serious revision: narrow the completeness claim to what the validation actually supports (e.g., completeness of the CNN among emission-dominated CVs), provide a separate discussion of the absorption-dominated gap, and ideally release the code and weights. A referee should push on those points, not on the existence of the CVs.","headline":"Big, useful CV catalogue from DESI DR2; the catalog itself is solid, but the ~99% completeness claim doesn't survive the paper's own missed-AM-CVn example.","tokens_in":45639,"tokens_out":976,"would_cite":true,"duration_ms":12702,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained on known, mostly emission-line CVs and applied to every DESI science spectrum finds 1,029 cataclysmic variables — 221 new, ten of them AM CVn.","keywords":["cataclysmic variables","white dwarf binaries","machine-learning classification","DESI spectroscopy","AM CVn systems","orbital period distribution","space density","spectroscopic completeness"],"falsifier":"Inspect a random sample of the roughly 238,000 spectra the CNN flagged but the paper never manually examined (those with Z > 0.01 and no catalogue counterpart). If CVs appear there at a rate well above the assumed ~1 per cent — or if retraining the CNN with absorption-dominated AM CVn spectra added to the training set causes it to flag many additional systems — the claimed ~99 per cent completeness and the space densities built on it would be overestimates.","tokens_in":44360,"feed_emoji":"🔭","tokens_out":29875,"duration_ms":228072,"temperature":0.7,"pith_summary":"The paper aims to measure the cataclysmic-variable (CV) population — close binaries in which a white dwarf accretes from a companion star — without the bias that has shaped the known census, where systems are mostly found because they erupted in outburst. It runs a convolutional neural network, trained on previously known CV spectra, over roughly 81 million usable spectra from the Dark Energy Spectroscopic Survey and manually inspects the shortlist, recovering 1,029 CVs: 808 already known and 221 new, including ten AM CVn systems (ultra-short-period binaries with hydrogen-poor donors). The authors argue the sample is about 99 per cent complete, roughly two magnitudes deeper than earlier spectroscopic surveys, and skewed toward intrinsically faint, short-period, low-accretion-rate systems — indeed, of the 171 new CVs with light-curve coverage, only 43 show outbursts. If the claims hold, this is the largest spectroscopically selected CV sample to date, the first to test CV population models against a largely variability-free census, and the basis for revised space densities of CV subtypes as well as a dozen new members of a rare class showing erratic state changes.","feed_headline":"1,029 cataclysmic variables found among 99 million spectra, 221 new","feed_subtitle":"A scan of the Dark Energy Spectroscopic Survey catches quiet binaries that rarely outburst — a nearly unbiased census.","key_machinery":"The load-bearing tool is a convolutional neural network (CNN) trained on 5,484 DESI spectra matched from SDSS samples: 623 known CVs plus galaxies, quasars, stars, white dwarfs, and detached white-dwarf binaries. Trained 200 times on random 80/20 splits, the best model is error-free on the held-out test set. Its job is to shrink 81 million spectra to 293,670 candidates; a catalogue cross-match and redshift cut (Z < 0.01185) then trim the list to ~41,500 spectra for human inspection. Completeness rests on inspecting random and targeted samples of discarded spectra, including the 1.6 < Z < 1.7 spike where DESI's redshift pipeline is fooled — yielding the ~99 per cent estimate.","core_discovery":"A CNN trained on known, mostly emission-line CVs and applied to every DESI science spectrum finds 1,029 cataclysmic variables — 221 new, ten of them AM CVn. The CNN flagged 293,670 candidates; cross-matching, a redshift cut, and visual inspection of 41,500 spectra yielded 1,014 CVs, with six more from validation. The CNN's per-spectrum reliability is 97.2 per cent and the sample completeness is claimed at ~99 per cent. Of the 171 new CVs with ZTF light curves, only 43 show outbursts — evidence that spectroscopic selection offsets the bias toward outbursting systems. Twelve new members of the peculiar state-change class and five evolved-donor CVs complete the catalogue.","pith_inferences":["If the ~99 per cent completeness holds, the implied conclusion is that the known CV census is far from finished at faint magnitudes: the serendipitous detection rate inside DESI's deep but non-CV-focused targeting is high enough that future wide-field spectroscopic surveys will keep uncovering substantial new populations.","The CNN's demonstrated blind spot — an absorption-dominated AM CVn prototype missed because the training set contained no such spectrum — suggests that other spectral classes absent from the training sample are silently underrepresented; true completeness for rare or unusual CVs may be below 99 per cent even if the overall claim survives.","The near-total reliance on serendipity implies a cheap, testable strategy: deliberately reserving survey fibers for white-dwarf-binary candidates would multiply CV yields, and a future far-UV space survey of the kind the paper highlights would make such targeting efficient.","The blurring of the period-gap edges in an unbiased sample sharpens a specific prediction-check for binary evolution models: models that reproduce a sharp gap in outburst-selected samples but not in deep spectroscopic samples would be consistent with these data, whereas models predicting a sharp gap here would not."],"forward_implications":["The DESI sample is the deepest CV census to date, about two magnitudes fainter than previous spectroscopic surveys, and it contains the largest fraction of short-period systems; the sharp edges of the 2–3 hour period gap fade in such untargeted samples.","Revised space densities for SU UMa and WZ Sge subtypes agree with earlier estimates, but the values for every subtype are lower bounds because 156 unclassified CVs — many of them non-outbursting, hence likely short-period — are not counted.","With twelve new members added to the eight previously known examples, roughly one per cent of all CVs belong to the peculiar state-change class, and their scatter across orbital period and HR-diagram position argues against a single evolutionary stage as the cause.","Almost all 1,029 CVs — 709 of them — were observed serendipitously by DESI's galaxy and quasar programs rather than by CV-specific targeting, showing that deep multi-object surveys harvest CVs as a by-product.","Five CVs whose spectra mimic F-type stars but whose absolute magnitudes are too faint for F-type donors are likely systems with highly evolved helium-core donors, a population previously found mainly by targeted variability searches."],"fun_headline_variants":["AI scan of 99M spectra finds 1,029 cataclysmic variables, 221 new","Spectroscopic census catches 1,029 cataclysmic variables, 221 never seen","Deep survey finds quiet binaries: 1,029 cataclysmic variables, 221 new","Machine learning uncovers 1,029 cataclysmic variables, 221 new"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The completeness argument stands or falls on the premise that the CNN's training spectra — previously known CVs, nearly all with emission lines — cover every spectral shape a CV can present in DESI; the paper itself shows this premise fails for at least one absorption-dominated AM CVn system that was absent from the training set.","fun_headline_variants_meta":{"raw":{"variants":["AI scan of 99M spectra finds 1,029 cataclysmic variables, 221 new","Spectroscopic census catches 1,029 cataclysmic variables, 221 never seen","Deep survey finds quiet binaries: 1,029 cataclysmic variables, 221 new","Machine learning uncovers 1,029 cataclysmic variables, 221 new"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001101,"raw_usage":{"total_tokens":4408,"prompt_tokens":700,"completion_tokens":3708,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":444,"completion_tokens_details":{"reasoning_tokens":3610}},"tokens_in":444,"tokens_out":3708,"duration_ms":22206,"temperature":1.0,"reasoning_tokens":3610,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:22:20.720066+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect a random sample of the roughly 238,000 spectra the CNN flagged but the paper never manually examined (those with Z > 0.01 and no catalogue counterpart). If CVs appear there at a rate well above the assumed ~1 per cent — or if retraining the CNN with absorption-dominated AM CVn spectra added to the training set causes it to flag many additional systems — the claimed ~99 per cent completeness and the space densities built on it would be overestimates.","supporting_citations":[],"review_version":1}