{"id":"134ef39f-ca2d-4895-be70-8f51c6b7e3e7","arxiv_id":"2502.09790","paper_version":4,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"ExoMiner++ is a deep learning classifier for TESS transit signals that adds five diagnostic branches, trains jointly on Kepler and TESS, and produces a public catalog of 7,330 planet candidates.","lead":"A NASA team extended the ExoMiner deep learning vetter with five new diagnostic inputs and applied it to 2-minute TESS data, classifying 7,330 of 147,568 unlabeled transit signals as planet candidates. The paper also releases a public catalog and 50 new candidate objects, aiming to focus follow-up telescope time on the most promising planets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Catalog precision may be far below headline: on the 293 TCEs that transitioned from uncertain to certain labels in 2024, precision at score>0.5 is 0.76 versus the reported 0.933; the 7,330-PC claim is not validated on the ambiguous unlabeled population.","rationale":"The reader's weakest-assumption (surrogate label quality in training/evaluation) is real, and the paper acknowledges it with several corrected examples and an honest Section 4.1 discussion. However, I see a more directly load-bearing issue: the performance metrics that support the headline catalog claim are computed on a labeled subset that excludes the very population (PC/APC/FA and other UNK TCEs) to which the catalog applies. The paper's own external validation in Section 7 quantifies this gap: precision 0.76 at score>0.5 on the 293 TCEs that gained certain labels during 2024, versus 0.933 on the labeled set. That discrepancy is not an artifact of label noise; it is evidence of distribution shift between the training/evaluation distribution and the catalog application distribution. Because the catalog is the paper's primary community contribution, the claim that these 7,330 TCEs are 'planet candidates' with 'excellent ranking quality' needs either additional validation on the ambiguous subset or a more cautious presentation (e.g., reporting expected precision or recommending a higher score threshold for follow-up). The paper is otherwise strong: the architecture is well-motivated, ablations are informative, code and catalogs are public, and the Kepler baseline plus small positive external validation provide genuine support. A conditional accept—requiring a precision estimate for the unlabeled/ambiguous population or a softened catalog claim—preserves the contribution while correcting the overreach. This is not a rejection: the model and catalog are likely useful, but the headline precision should not be implicitly attributed to the 7,330 list without evidence.","tokens_in":43545,"tokens_out":12035,"duration_ms":118655,"concrete_test":"Recompute precision at score>0.5 on the labeled TESS set after excluding NTPs (i.e., only KP/CP/BD/EB/FP labels) and also on only TCEs with an ExoFOP photometric or spectroscopic disposition; if precision drops below ~0.85, the headline is driven by easy NTPs. Separately, using Table 14's December 2024 TFOPWG dispositions, compute precision and recall at score>0.5 for all UNK TCEs with any certain Dec 2024 label (KP/CP/FP/BD), not just the 293 that changed from January 2024, and report the 95% confidence interval; if this precision is materially below 0.93, the abstract's 7,330-PC claim should be qualified or the catalog should present precision estimates for the unlabeled population.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline precision/recall (Table 6, 0.933/0.951) is measured on TCEs with certain surrogate labels (KP/CP/FP/EB/BD/NTP), explicitly excluding the ambiguous PC/APC/FA dispositions. Yet the central catalog claim—7,330 planet candidates among 147,568 unlabeled TCEs—is a prediction on exactly the excluded ambiguous population. The only direct holdout evidence from that population is Section 7's Jan-to-Dec 2024 comparison: among 293 TCEs whose labels changed from PC/APC/FA to CP/FP, the model achieves precision = 147/(147+47) = 0.76 and recall = 147/151 = 0.97. This is a large drop from the headline 0.933 precision and shows the model's discrimination is much weaker in the decision-relevant regime. The 293-sample subset is itself enriched for planets (prior ~51.5%), whereas the overall UNK pool almost certainly has a far lower planet fraction; under a lower prior, precision at the same threshold would drop further. The paper's manual review of 91 high-confidence potential CTOIs (rejecting 41) does not cover the bulk of the 7,330 PC list. Thus the abstract's '7,330 planet candidates' and the implied follow-up narrowing rest on an unverified transfer of metrics from an easy, NTP-dominated labeled set to a harder, ambiguous unlabeled set. This is a correctness risk for the main deliverable, not just a label-noise nuisance.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"ExoMiner++ is a deep-learning classifier for TESS 2-minute transit signals that extends the earlier ExoMiner architecture with five new input branches: transit-view unfolded flux, difference image, phase-folded momentum dump flags, a Lomb-Scargle periodogram, and a full-orbit folded flux trend. The model is trained with multi-source learning on Kepler plus TESS labels derived from ExoFOP, the Villanova EB catalog, and TESS-ExoClass NTPs, and is evaluated with 10-fold cross-validation using target-star splits. The authors report precision/recall of 0.933/0.951 and PR AUC 0.976 on the labeled TESS subset with TESS+Kepler aggregate scoring, show via ablations that most new branches improve performance, and release a vetting catalog in which 7,330 of 147,568 previously unlabeled TCEs are classified as planet candidates, including 50 newly introduced CTOIs.","tokens_in":43856,"tokens_out":4468,"duration_ms":45887,"significance":"If the catalog is as reliable as the headline metrics suggest, this is a substantial community resource: it narrows the TESS follow-up target list, provides confidence scores for ranking, and introduces new candidates. The paper has genuine strengths: the code is public, the cross-validation scheme splits on target stars rather than TCEs, the ablation study in Section 6.9 tests each new branch against a baseline, and the authors are unusually candid about label noise and ephemeris-matching failures in Section 4.1 and Section 6.7. The main deliverable is the 7,330-planet-candidate catalog, and the decision-relevant performance question is how well the model discriminates on the ambiguous, previously unlabeled population. The paper's own Section 7 evidence indicates that precision is substantially lower there than on the certain-label test set, so the catalog claim needs additional validation or careful qualification before it can be fully credited.","major_comments":[{"comment":"The decision-relevant precision for the catalog is not the Table 6 value of 0.933 on the certain-label TESS set, but the precision on TCEs that were unlabeled or uncertain in January 2024 and later resolved. The paper's own Section 7 calculation gives 147/(147+47)=0.76 precision and 0.97 recall on the 293 TCEs whose labels changed from PC/APC/FA to CP/FP between January and December 2024. That is a large drop from 0.933, and the 293-TCE subset is enriched in planets (prior about 0.5), whereas the full 147,568 UNK pool almost certainly has a much lower planet fraction. At a lower prior, precision at score>0.5 would fall further. Since the abstract's \"7,330 planet candidates\" and the implied follow-up prioritization rest on the UNK population, please report precision/recall and score-threshold curves directly for the label-transition subset, and either validate the catalog on a held-out ambiguous sample or substantially qualify the headline precision claim.","section":"Section 7 and Table 6"},{"comment":"The training and evaluation labels are surrogate labels from ExoFOP, the Villanova EB catalog, and TESS-ExoClass NTP triage. Section 6.7 gives explicit examples of known planets mislabeled as EBs (e.g., TIC 309792357 TCEs) and of planets mislabeled as NTPs due to failed ephemeris matching, and the authors acknowledge that the Villanova catalog has label noise. Because the same label-generation procedure produces both the training targets and the Table 6 evaluation labels, the reported 0.933/0.951 may be optimistically biased, and the 7,330 catalog classifications inherit that bias. Please add a quantitative label-noise robustness analysis, for example by re-evaluating on TCEs whose ExoFOP disposition changed between January and December 2024, or by reporting performance after removing Villanova/TEC labels that cannot be independently confirmed.","section":"Section 4.1 and Section 6.7, item 4"},{"comment":"The 50 newly introduced CTOIs are carefully vetted, but they are only a small subset of the 7,330 planet candidates. After the aggressive period matching, 427 unmatched TCEs remain and are reduced to 288 events and then to 91 high-confidence potential CTOIs by requiring consensus across all ten models and at least three observed transits; only these 91 are manually reviewed. The remaining thousands of PC classifications, including the 6,322 TCEs matched to existing TOIs and the 1,008 unmatched TCEs, do not receive the same SME scrutiny. Please state explicitly how many of the 7,330 are in each category and give separate expected-precision estimates for matched versus unmatched TCEs, so that the catalog's overall precision claim is not conflated with the well-vetted 50-CTOI list.","section":"Section 7 (catalog construction)"}],"minor_comments":[{"comment":"The abstract states 147,568 unlabeled TCEs, while Section 7 states 147,567 unlabeled TCEs. The numbers should be reconciled.","section":"Abstract vs. Section 7"},{"comment":"There are many typographical errors in the text and tables, including \"T able\", \"T ransit\", \"V alizadegan\", \"TYIC\" in Table 15, and \"MESS\" in figure axis labels. A careful proofread is needed.","section":"Tables and figures throughout"},{"comment":"The Figure 6 caption identifies TIC 82707763-1-S37 as TOI 1991.01, while the Figure 7 caption identifies the same TCE as TOI 1046.01 and refers to the target as TIC 309787037. The target identifier in the Figure 7 caption appears to be a copy-paste error and should be corrected.","section":"Figure 6 and Figure 7 captions"},{"comment":"The sentence \"There is no clearn pattern\" should read \"no clear pattern,\" and the surrounding discussion of MES and performance would benefit from a more direct presentation of the heatmap values.","section":"Section 6.4"},{"comment":"The description of multi-source learning says the simpler combining approach proved more effective, but the paper does not report the transfer-learning or fine-tuning comparison quantitatively. A sentence stating that those results are omitted for brevity would help the reader interpret the claim.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"The core methodological contribution is sound and the catalog is useful, but the main deliverable hinges on transferring metrics from a certain-label set to an ambiguous unlabeled set. The Section 7 label-transition analysis provides the most decision-relevant estimate (precision 0.76) and should be foregrounded in the revision, along with a clear statement of how the 7,330 count is distributed across matched, unmatched, and manually reviewed categories. The paper is not fatally flawed, but the headline claims need to be re-anchored to the actual validation evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: ExoMiner++ is a real step forward for automated TESS vetting. The architecture adds five diagnostic branches (unfolded flux, difference image, periodogram, flux trend, momentum dump), combines Kepler and TESS labels in training, and ships a public catalog with scores for 147k unlabeled TCEs plus 50 new CTOIs. The 10-fold cross-validation with target-star splits is sound, the ablation study shows most branches earn their keep, and the discussion of label noise and ephemeris-matching failures is unusually honest. Code and catalogs are public. That is concrete, reproducible work, and I would cite it.\n\nThe soft spot is the gap between the headline metrics and the catalog’s decision-relevant population. Table 6 reports precision/recall of 0.933/0.951 on the labeled subset (KP/CP/EB/FP/BD/NTP), which excludes the ambiguous PC/APC/FA dispositions. But the 7,330 planet candidates are predictions on exactly that ambiguous UNK population. The paper’s own Section 7 holdout—TCEs whose uncertain labels became certain between January and December 2024—gives precision 0.76 and recall 0.97 on 293 TCEs. That is not terrible, and it beats the TFOPWG baseline precision of 0.50, but it is not 0.933, and the sample is small and planet-enriched. The manual review of 91 potential CTOIs covers only a sliver of the 7,330. So the abstract’s “identifies 7,330 as planet candidates” and the phrase “excellent ranking quality” would overstate confidence if read as applying to the catalog as a whole. The paper does include the 0.76 number, but it belongs in the abstract’s neighborhood, not buried in Section 7.\n\nMinor issues: no error bars on catalog counts; the Section 6.4 experiment removing low-MES exoplanets is confusingly reported, and the momentum dump branch does not help on average, which the authors acknowledge. None of this overturns the contribution.\n\nThis paper is for anyone doing TESS follow-up or building automated vetting pipelines. It deserves a serious referee. After a revision that reframes the catalog precision honestly, it should be accepted: I would ask the authors to add a caveat to the abstract and state directly that the 7,330 count is based on a score threshold whose precision on the ambiguous population is about 0.76, not 0.933.","headline":"Useful model and public catalog, but the headline precision belongs to the easy labeled subset; the only holdout from the ambiguous catalog population shows 0.76 precision, so the 7,330-PC claim needs a caveat.","tokens_in":44564,"tokens_out":2633,"would_cite":true,"duration_ms":25704,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ExoMiner++ claims that a deep network fed with five new diagnostic inputs—difference images, periodograms, flux trends, unfolded flux, and momentum-dump flags—can classify TESS 2-minute transit signals and rank candidates so precisely…","keywords":["exoplanet transit classification","TESS","threshold crossing events","deep learning","false positive vetting","multi-source learning","planet candidate catalog","convolutional neural network"],"falsifier":"Follow up the 50 newly introduced CTOIs with radial-velocity and high-resolution imaging; if most turn out to be eclipsing binaries, background transits, or other false positives rather than planets, the claim that ExoMiner++ ranks the most likely candidates would be refuted.","tokens_in":43319,"feed_emoji":"🪐","tokens_out":8643,"duration_ms":75761,"temperature":0.7,"pith_summary":"ExoMiner++ aims to establish that a deep-learning model can vet TESS 2-minute transit signals as reliably as experienced human reviewers, despite TESS's short observation windows, blended light, and noisy labels. The key proposal is to feed the model the same diagnostics a human would consult, including difference images, periodograms, flux trends, unfolded flux, and spacecraft attitude events, and to train on a mix of clean Kepler labels and noisier TESS labels. The paper reports that this combination reaches precision 0.933 and recall 0.951 on the labeled TESS subset, and that the top 1,000 ranked signals are nearly all exoplanets. Applied to 147,568 unlabeled threshold-crossing events, the model classifies 7,330 as planet candidates, including 50 new community TESS Objects of Interest. If the claims hold, follow-up efforts can concentrate on a few thousand candidates rather than the full catalog, raising the yield of confirmed planets per unit of telescope time.","feed_headline":"Deep learning flags 7,330 planet candidates in TESS data","feed_subtitle":"New classifier ranks TESS transits so follow-up can skip ~140,000 false alarms and focus on the most likely planets.","key_machinery":"The engine is a multi-branch convolutional network whose branches mirror the diagnostic tests in the telescope's Data Validation report: phase-folded flux, odd/even flux, weak secondary, centroid motion, unfolded flux, difference image, periodogram, flux trend, and momentum-dump time series, plus stellar parameters and detection statistics as scalars. Each branch has its own feature extractor, and the three transit-view flux branches share low-level convolutional features; the model also receives the standard error of each phase bin as an extra channel. Two additional mechanisms drive much of the gain: multi-source learning, which mixes clean Kepler labels into the noisier TESS training set, and multi-sector aggregation, which assigns one score per physical event based on the longest sector run. Detrending was switched from a spline fit to a Savitzky-Golay filter, and the removed trend is fed back in as its own branch, letting the model see signals like ellipsoidal variations that occur at the orbital period timescale.","core_discovery":"The paper's central claim is that one convolutional architecture, ExoMiner++, can perform automated transit vetting for TESS at a level that narrows the follow-up search space while keeping few false positives at the top. For each TCE the model outputs a score between 0 and 1; a score above 0.5 counts as a planet candidate. On the labeled TESS dataset, the best configuration, trained on TESS plus Kepler and aggregated across sectors, reports precision 0.933, recall 0.951, PR AUC 0.976, and ROC AUC 0.998. Ranking quality is the headline: every one of the top 200 TCEs is an exoplanet, Precision@1000 is at least 0.99 across models, and Precision@3000 is 0.977 for the TESS+Kepler aggregate model. The paper argues this ranking transfers to unlabeled data: 7,330 of 147,568 unlabeled TCEs are classified as planet candidates, corresponding to 1,868 matched TOIs and 50 new CTOIs introduced by the authors. A temporal check adds support: of 151 TCEs promoted from uncertain TOI dispositions to confirmed planets between January and December 2024, ExoMiner++ scored 147 as planet candidates.","pith_inferences":["Because ExoMiner++ scores are not calibrated probabilities, the 0.5 cutoff is a working threshold, not a posterior; comparing scores across surveys or converting them into false-positive rates would require a separate calibration step.","The paper's label-noise examples cut both ways: if true planets are sometimes labeled as eclipsing binaries or non-transiting phenomena, the reported recall on TESS may be an underestimate, and the true precision of the 7,330-candidate catalog could be better than the labeled-set metrics suggest.","A clean prospective test would be to score each newly dispositioned TOI as the follow-up community updates its catalog over the next year, avoiding the circularity of evaluating on the same surrogate labels used for training.","The contaminated-aperture and nearby-background-transit subclasses remain the weakest spots, so adding known nearby-star positions and brightnesses as extra difference-image channels is a natural, testable extension in crowded fields."],"forward_implications":["Follow-up programs can target roughly 7,330 planet candidates instead of 147,568 unlabeled threshold-crossing events, with the top of the ranking consisting almost entirely of exoplanets.","TOIs with uncertain community dispositions can be re-ranked by model score; among 151 TOIs later confirmed as planets, 147 received candidate scores, suggesting the ranking can anticipate future confirmations.","The public catalog provides a confidence score per TCE and per TOI, allowing observers to schedule follow-up by score threshold and to revisit borderline cases.","The same architecture and preprocessing are directly portable to TESS full-frame-image data and to future transit surveys, since all branches consume standard diagnostics rather than mission-specific features.","If ranking quality persists, the overall yield of confirmed planets per observation hour should rise because telescope time will be spent on high-score candidates."],"supporting_citations":[{"why":"Supplies the original ExoMiner architecture and Kepler validation that ExoMiner++ extends with new branches and inputs.","marker":"Valizadegan et al. 2022"},{"why":"Defines the Data Validation diagnostic tests and scalar statistics that form the model's input branches.","marker":"Twicken et al. 2018a"},{"why":"Provides the ephemeris-matching procedure used to label TESS TCEs against community TOI, EB, and flux-triage catalogs.","marker":"Twicken et al. 2018b"},{"why":"Supplies the TESS EB catalog whose manual light-curve classifications provide a large share of non-planet training labels.","marker":"Prša et al. 2022"},{"why":"Describes the detection pipeline that produces the TCEs and Data Validation reports used as data and as the architectural template.","marker":"Jenkins et al. 2016"},{"why":"Describes the TESS mission whose 2-minute cadence data are the target of classification and the catalog.","marker":"Ricker et al. 2015"},{"why":"Establishes the earlier deep-learning transit classifier and spline detrending approach that ExoMiner++ modifies with Savitzky-Golay filtering and more diagnostic inputs.","marker":"Shallue & Vanderburg 2018"},{"why":"Provides the closest comparable TESS classifier for the performance discussion; its public dataset is matched to ExoMiner++'s to quantify overlap.","marker":"Tey et al. 2023"}],"fun_headline_variants":["Deep learning spots 7,330 TESS planet candidates","AI triage finds 7,330 planet candidates in TESS data","ExoMiner++ winnows TESS signals to 7,330 planet candidates","New AI catalog: 7,330 TESS planet candidates","Deep learning cuts TESS planet search to 7,330 candidates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training and evaluation labels for non-planets, mostly non-transiting phenomena from flux triage and eclipsing binaries from the EB catalog, are accurate enough to serve as ground truth, despite acknowledged noise and failed ephemeris matches.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning spots 7,330 TESS planet candidates","AI triage finds 7,330 planet candidates in TESS data","ExoMiner++ winnows TESS signals to 7,330 planet candidates","New AI catalog: 7,330 TESS planet candidates","Deep learning cuts TESS planet search to 7,330 candidates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001042,"raw_usage":{"total_tokens":4475,"prompt_tokens":1134,"completion_tokens":3341,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":750,"completion_tokens_details":{"reasoning_tokens":3247}},"tokens_in":750,"tokens_out":3341,"duration_ms":22363,"temperature":1.0,"reasoning_tokens":3247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:27:26.405749+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Follow up the 50 newly introduced CTOIs with radial-velocity and high-resolution imaging; if most turn out to be eclipsing binaries, background transits, or other false positives rather than planets, the claim that ExoMiner++ ranks the most likely candidates would be refuted.","supporting_citations":[],"review_version":1}