{"id":"d02594d5-c29f-493f-9d0a-913b4cd56c33","arxiv_id":"2509.10619","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LEO-Vetter automatically vets TESS transit signals with flux and pixel checks, reporting 91% completeness and 97% false-alarm reliability and producing a uniform catalog of 172 M dwarf planet candidates.","lead":"Astronomers built LEO-Vetter, a computer program that automatically checks TESS planet signals to separate real planets from noise, star activity, and eclipsing stars without human review. Running it on 200,000 M dwarf stars cut 20,000 signals down to 172 candidate planets, a step toward unbiased planet counts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 91% completeness and 97% reliability are computed on the same injected and scrambled data used to optimize thresholds (§5.6, §6.1); the FGK re-run (§8.2) shows reliability is dataset-sensitive, so an unbiased held-out estimate is needed before occurrence-rate claims.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing issue: the performance statistics are measured on the same simulated datasets used to tune thresholds. My reading of Sections 5.6 and 6.1 confirms that no held-out split or cross-validation is reported, and the FGK application in Section 8.2 shows that reliability is sensitive to the stellar sample, dropping from 97% on M dwarfs to 88% on FGK dwarfs with the same thresholds. This does not falsify the central claim, but it does mean the headline numbers are not yet demonstrated to transfer to new users' data. The independent anchors are real evidence: the 93% recovery of known TOIs, the three confirmed-planet failures being explained by low SNR and long-cadence data, and the public GitHub code all support the tool's basic functionality. However, TOI recovery is a biased sample of bright, manually vetted signals, and the FGK test lacks a tuned comparison, so neither fully settles the overfitting concern. A held-out validation split is a straightforward check that would either confirm the quoted numbers or require re-reporting them. The existing CONDITIONAL verdict remains appropriate: the paper is a substantial software contribution, but the occurrence-rate guarantee needs an unbiased performance estimate.","tokens_in":29988,"tokens_out":4610,"duration_ms":42485,"concrete_test":"Re-run the Section 5.6 optimization on a randomly chosen half of the 62,410 injected and 15,655 scrambled TCEs, splitting by star to avoid leakage, then compute C, E, and R_FA on the held-out half using the resulting thresholds. Repeat with several random seeds or use k-fold cross-validation. If the validation completeness or reliability is materially below 91% or 97%, the headline statistics are in-sample and should be replaced with the validation estimates, along with bootstrap uncertainties.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.6 uses differential evolution to choose pass-fail thresholds by minimizing sqrt((1-C)^2 + (1-R_FA)^2) on the injected and scrambled TCE sets, and Section 6.1 then quotes C = 91.00%, E = 99.91%, and R_FA = 96.97% from those same sets. This is in-sample evaluation: the optimizer can exploit noise in the specific simulated TCEs, so the reported performance need not transfer to new data. The paper reports no held-out split or cross-validation. The FGK test in Section 8.2 provides an out-of-sample hint: applying the M-dwarf-optimized thresholds to FGK dwarfs gives C = 90.86% but R_FA = 88.19% overall (98.64% for SNR > 12). Thus the 97% reliability is not robust across stellar populations, and even within M dwarfs the tuned value is likely optimistic. Since the central claim is that LEO-Vetter produces catalogs with characterized completeness and reliability for occurrence-rate calculations, the absence of an unbiased validation of the headline numbers is the load-bearing weakness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents LEO-Vetter, a fully automated Robovetter-inspired vetting pipeline for TESS transit candidates. The tool implements thirteen flux-level tests against noise/systematic false alarms, four flux-level tests against astrophysical false positives, and a pixel-level centroid offset test using difference images. As a demonstration, the authors apply LEO-Vetter to roughly 200,000 M-dwarf QLP light curves, obtain 18,424 observed TCEs plus 62,410 injected and 15,655 scrambled TCEs, and use a differential evolution optimizer to set pass-fail thresholds on the false-alarm tests. They report a completeness of 91.00% and a false-alarm reliability of 96.97% on these same data, and after pixel-level vetting they produce a uniformly vetted catalog of 172 planet candidates. They also apply the M-dwarf-tuned thresholds to 300,000 FGK dwarfs and report lower overall reliability (88.19%), which they interpret as demonstrating that the thresholds are applicable to FGK dwarfs. The paper claims that the vetter's characterized completeness and reliability make it suitable for TESS occurrence-rate calculations.","tokens_in":30240,"tokens_out":5405,"duration_ms":47327,"significance":"If the performance claims hold, LEO-Vetter would be a valuable public resource: it is the first fully automated TESS vetter that combines flux-level tests with pixel-level difference-image analysis, and it is designed specifically for uniform catalog production in support of demographic studies. The paper's strengths include the public release of the code on GitHub and Zenodo, the use of injection and scrambling experiments following the Kepler DR25 methodology, the large-scale M-dwarf demonstration producing a 172-candidate catalog, and the per-candidate false-positive probabilities, positional probabilities, and disposition scores. The central weakness is that the headline completeness and reliability numbers are in-sample estimates: thresholds were optimized on the same injected and scrambled TCE sets used to quote performance, and the FGK re-run in §8.2 shows that reliability is sensitive to the stellar sample. This load-bearing issue must be addressed before the occurrence-rate support claim is credible.","major_comments":[{"comment":"The headline performance figures are in-sample fitted values, not independent predictions. In §5.6 the differential evolution optimizer minimizes sqrt((1-C)^2 + (1-R_FA)^2) over the injected and scrambled TCE sets, and §6.1 then quotes C = 91.00%, E = 99.91%, and R_FA = 96.97% computed on those same sets. The paper reports no held-out split, cross-validation, or alternative test set for the M-dwarf pipeline. Because the optimizer can exploit noise in the specific simulated TCEs, the quoted completeness and reliability may be optimistic. The FGK re-run in §8.2 provides the natural out-of-sample test: applying the same thresholds to FGK dwarfs gives C = 90.86% but R_FA = 88.19% overall. This demonstrates that the 97% reliability is not robust across stellar populations and that the claim of applicability to FGKM dwarfs is only weakly supported. An unbiased evaluation (e.g., tuning on one half of the M-dwarf data and evaluating on the other half, or presenting the FGK result as the headline transferable reliability) is required before the paper can claim that LEO-Vetter produces catalogs with characterized completeness and reliability for occurrence rates.","section":"§5.6 and §6.1"},{"comment":"The reliability of the final 172-candidate catalog is not characterized. The 96.97% false-alarm reliability in §6.1 is computed after flux-level vetting only; the pixel-level centroid offset test then removes 153 of 325 flux-level candidates, yet no completeness or reliability estimate is provided for the pixel-level stage. The centroid threshold of Δθ < 15″ appears to be hand-set based on performance on known TOIs rather than validated with injected or simulated pixel-level tests, and the paper does not quantify how many true planets might be lost at this stage. Since the stated purpose of the tool is to support occurrence-rate calculations that require end-to-end reliability of the final catalog, the absence of a validated pixel-level performance estimate is a load-bearing gap.","section":"§6.3 and §7"},{"comment":"The cleaning of simulated TCEs introduces manual judgment into the benchmark sets used for the headline numbers. The authors removed 389 injected and 479 scrambled TCEs from light curves containing known TOIs, eclipsing binaries, or 'signals that we determined to be previously unknown planet candidates or eclipsing binaries manually identified while testing the vetter.' This manual pruning can bias both completeness and effectiveness in unpredictable directions and weakens the claim that C and R_FA are purely simulation-based quantities. The paper should justify that the removed TCEs are not systematically different from the retained ones, or present the performance metrics with and without this cleaning step.","section":"§5.5"}],"minor_comments":[{"comment":"The point estimates 91.00% and 96.97% are quoted without any uncertainty. Statistical uncertainties from the finite TCE samples are small, but systematic uncertainties from simulation fidelity and threshold tuning are not; a brief statement about the dominant uncertainty would help readers.","section":"Abstract and §6.1"},{"comment":"There is a typo: 'A value if C_i ≈ 0' should read 'A value of C_i ≈ 0'.","section":"§3.8"},{"comment":"The definition of P_max as 'half the length of the longest continuous stretch of consecutive sectors' is confusing; it would be clearer to say 'half the length of the longest continuous observing stretch within a sector or across consecutive sectors.'","section":"§5.2"},{"comment":"The sentence 'LEO-Vetter successfully classified 91.00% of injected planets as PCs' should say 'of injected TCEs recovered by the BLS search,' to match the definition of completeness in Eq. (14).","section":"§6.1"},{"comment":"The conclusion that 'the thresholds optimized for M dwarfs are also applicable to FGK dwarfs' is too strong given that R_FA drops from 96.97% to 88.19%; the SNR > 12 reliability of 98.64% should be reported more prominently if the authors intend this as the applicable regime.","section":"§8.2"},{"comment":"The disposition scores are derived from metric standard deviations measured on the same injected dataset used for threshold optimization; the text should note that these scores are therefore not an independent validation of the vetter's reliability.","section":"§7.3"}],"recommendation":"major_revision","confidential_remarks":"The in-sample evaluation is the core issue. The FGK re-run is a ready-made out-of-sample test; the authors should either present it as the primary validation or hold out a portion of the M-dwarf data for evaluation. The pixel-level stage also needs a validation strategy. I would support publication after these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a real software contribution that should be out in the world, and the headline performance numbers are in-sample fitted values that should not be quoted without a held-out check. The paper does exactly what it says: a fully automated, public, pure-Python vetter for TESS, modeled on the Kepler Robovetter, with flux-level tests (many ported from KDR25, some new) and pixel-level difference imaging. The M dwarf application—~200k light curves, ~20k TCEs down to 172 candidates, 13 new CTOIs—is a concrete demonstration. I believe the tool works and will be used. Credit where due: the code is on GitHub with a Zenodo DOI; the paper ships machine-readable tables; the comparison to known TOIs (93% recovery) and the careful explanation of the three confirmed planets that failed (all low-SNR, and all pass with 2-min cadence data) are honest and informative. The authors also clearly state that users can customize thresholds and that optimal thresholds depend on the dataset. The main weakness is exactly what the stress test flags: the 91% completeness and 97% reliability are computed on the same injected and scrambled TCE sets used to tune the pass-fail thresholds with differential evolution (§5.6, §6.1). This is in-sample evaluation. The optimizer can exploit noise in those specific simulated sets, so the quoted numbers are best-case fitted values, not predictions. There is no held-out split or cross-validation. The FGK re-run (§8.2) is an out-of-sample probe and shows the reliability is dataset-sensitive: 88% overall, 98.6% for SNR>12. That tells me the 97% M dwarf reliability is probably optimistic as stated. This matters because the stated purpose is supporting occurrence-rate calculations, which require characterized completeness and reliability. One fix would be to split the injected/scrambled data into tuning and evaluation subsets, or at minimum report standard errors on C and R_FA from bootstrap/resimulation. Just adding a sentence of caution is not enough. I also note the pixel-level completeness/reliability is not characterized at all, which is a smaller issue but relevant for NEB rejection. My recommendation: send this to a serious referee. The tool is useful, the paper is readable, and the methodological gap is fixable in revision. I would ask for a held-out validation and uncertainty quantification before accepting the headline numbers.","headline":"A genuinely useful, publicly shipped TESS vetter whose headline completeness/reliability numbers are in-sample fitted values; the tool deserves peer review, and the numbers need a held-out check.","tokens_in":30835,"tokens_out":2322,"would_cite":true,"duration_ms":19851,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LEO-Vetter claims that a fully automated, publicly available pipeline can vet TESS planet candidates with 91% completeness and 97% reliability against false alarms, reducing about 20,000 detections to 172 candidates.","keywords":["exoplanet vetting","TESS","transit candidates","planet occurrence rates","false alarms","eclipsing binaries","automated pipelines","pixel-level analysis"],"falsifier":"Apply LEO-Vetter with the published thresholds to an independent set of TESS light curves with known confirmed planets and known false positives, for example a different sector range or the 2-minute SPOC data, and compare the automated labels against the confirmed disposition. If the recovered fraction of confirmed planets is clearly below 91% or the false-alarm contamination of the candidate list clearly exceeds 3%, the reported completeness and reliability would be overstated.","tokens_in":29796,"feed_emoji":"🛰️","tokens_out":7744,"duration_ms":60602,"temperature":0.7,"pith_summary":"The paper tries to show that the bottleneck of manually vetting TESS transit candidates can be removed entirely. It presents LEO-Vetter, a public, fully automated pipeline that takes transit-like detections as input and labels each one as a planet candidate, an astrophysical false positive, or a noise/systematic false alarm. On simulated data the pipeline claims 91% completeness and 97% reliability against false alarms, and a demonstration search of about 200,000 M dwarfs reduces roughly 20,000 detections to 172 candidates. If these numbers hold, the result matters because a uniform, reproducible vetting step is precisely what TESS demographic calculations need to convert observed candidate counts into occurrence rates.","feed_headline":"Automated vetter cuts 20,000 TESS signals to 172 candidates","feed_subtitle":"A fully automated pipeline claims 91% completeness and 97% reliability for TESS planet catalogs.","key_machinery":"The engine is a suite of decision metrics computed for each transit-like detection and compared against pass-fail thresholds: signal-to-noise ratio, transit model fits, a sine-wave variability test, transit duration and asymmetry checks, depth mean-to-median consistency, per-transit signal-to-noise consistency, uniqueness tests on the folded light curve, individual transit vetting, and a pixel-level difference-image centroid analysis. The thresholds were set by an optimization routine that varies false-alarm-test thresholds to minimize $\\sqrt{(1-C)^2 + (1-R_{\\rm FA})^2}$, using injected transits to measure completeness and orbit-scrambled light curves to estimate false-alarm reliability. This metric-plus-threshold design is what lets the pipeline be fully automated and uniform.","core_discovery":"The paper claims that LEO-Vetter, a fully automated vetting pipeline for TESS transit signals, can replace the manual inspection currently used to build planet candidate catalogs without sacrificing the uniformity needed for demographic studies. On a test set of roughly 200,000 M dwarf light curves, it classifies 91.00% of injected transiting planets as planet candidates, rejects 99.91% of scrambled-light-curve false alarms, and reduces about 20,000 observed transit-like detections to 172 uniformly vetted candidates. The pipeline combines flux-level tests of transit shape, depth consistency, uniqueness, and signal-to-noise with a pixel-level centroid-offset test that flags signals coming from nearby stars. The authors conclude that this makes statistically grounded TESS occurrence-rate calculations possible because users can characterize both vetting completeness and false-alarm reliability from simulated data.","pith_inferences":["Because the pass-fail thresholds were optimized on the very injected and scrambled datasets used to measure performance, an independent test on a separate sector set or on 2-minute cadence data would be the natural check of whether 91% and 97% generalize to new observations.","The same architecture could be applied to other transit surveys with different bandpasses and cadence mixes; the limb-darkening assumptions and pixel-level code currently tied to TESS would need to be generalized first.","If the pipeline's completeness depends strongly on cadence, users analyzing only Prime Mission 30-minute data should re-derive thresholds rather than reuse the published ones, since the paper itself reports a Prime Mission completeness of 69.88%."],"forward_implications":["Users can turn their own TESS transit-like detections into a vetted planet candidate catalog without manual review, using the published thresholds as a starting point.","The injected and scrambled data products supplied with the pipeline let users measure vetting completeness and false-alarm reliability for their own sample, which is required input for occurrence-rate calculations.","The same thresholds applied to a 300,000-star FGK dwarf sample give comparable completeness (90.86%) and reliability (88.19% overall, 98.64% at signal-to-noise above 12), suggesting the M-dwarf tuning transfers to other dwarf populations.","Known TOIs that failed flux-level vetting recovered cleanly when 2-minute SPOC light curves were substituted for the longer-cadence FFI data, indicating performance improves with shorter cadence."],"supporting_citations":[{"why":"Defines the uniform catalog and reliability framework that LEO-Vetter follows and compares against.","marker":"Thompson et al. 2018"},{"why":"Supplies the effectiveness and false-alarm reliability equations used to compute LEO-Vetter's 97% reliability.","marker":"Bryson et al. 2020a"},{"why":"Provides the Model-Shift uniqueness test whose metrics become LEO-Vetter's MS1, MS2, MS3, and SHP tests.","marker":"Coughlin 2017"},{"why":"Introduces the transit asymmetry metric adopted as LEO-Vetter's ASYM test.","marker":"Eschen & Kunimoto 2024"},{"why":"Supplies false positive probabilities used to characterize astrophysical false positive reliability of the 172 candidates.","marker":"Giacalone et al. 2021"},{"why":"Demonstrates the TESS Faint Star Search whose Robovetter-inspired vetting motivates and contextualizes LEO-Vetter's design.","marker":"Kunimoto et al. 2022b"},{"why":"Provides the centroid-free Bayesian pixel probability analysis used to estimate on-target probability for each candidate.","marker":"Bryson & Kunimoto 2025"}],"fun_headline_variants":["Robot vetter finds 172 TESS planets from 20,000 signals","TESS vetting goes full-auto: 20k signals to 172 candidates","No human eyes needed: TESS planet vetting automated","LEO-Vetter: 20k TESS hits trimmed to 172 candidates","Automated TESS vetting: 91% completeness, 97% reliability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The performance numbers rest on the assumption that transits injected into real light curves and false alarms made by scrambling orbit and sector order faithfully represent the real planets and noise/systematic signals a user will encounter.","fun_headline_variants_meta":{"raw":{"variants":["Robot vetter finds 172 TESS planets from 20,000 signals","TESS vetting goes full-auto: 20k signals to 172 candidates","No human eyes needed: TESS planet vetting automated","LEO-Vetter: 20k TESS hits trimmed to 172 candidates","Automated TESS vetting: 91% completeness, 97% reliability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00037,"raw_usage":{"total_tokens":2016,"prompt_tokens":1013,"completion_tokens":1003,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":903}},"tokens_in":629,"tokens_out":1003,"duration_ms":8381,"temperature":1.0,"reasoning_tokens":903,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:53:43.528879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply LEO-Vetter with the published thresholds to an independent set of TESS light curves with known confirmed planets and known false positives, for example a different sector range or the 2-minute SPOC data, and compare the automated labels against the confirmed disposition. If the recovered fraction of confirmed planets is clearly below 91% or the false-alarm contamination of the candidate list clearly exceeds 3%, the reported completeness and reliability would be overstated.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Model-Shift uniqueness test whose metrics become LEO-Vetter's MS1, MS2, MS3, and SHP tests."},{"cited_title":"2025, Research Notes of the AAS, 9, 81, 10.3847/2515-5172/adcb3a","cited_arxiv_id":null,"evidence_quote":"Provides the centroid-free Bayesian pixel probability analysis used to estimate on-target probability for each candidate."}],"review_version":2}