{"id":"a83871d1-bd1a-4ff0-960e-381b0c3252fa","arxiv_id":"2607.23781","paper_version":3,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A uniform re-analysis of 461 M-dwarf TOI hosts recovers 85.5% of confirmed planets and yields 13 new transit candidates, including two in the optimistic habitable zone.","lead":"This survey re-analyzes TESS photometry for 461 M-dwarf planet-hosting stars with one uniform pipeline, recovering 85% of known planets and finding 13 new transit candidates. It provides a prioritized follow-up list with ephemerides for the strongest candidates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Threshold transfer from a 10-system validation set to 461 gapped/active hosts is the key unmeasured risk for the new-candidate list; a null-cohort false-positive test would settle it.","rationale":"The reader identified the same weakest assumption: frozen thresholds calibrated on a ten-system validation set with only two multiyear-gapped systems. I agree that this is the most load-bearing concern. The paper itself flags the limitation (Sect. 2.3 and limitation 7), and the survey is explicitly a follow-up prioritization product rather than a validated-planet catalog, so a moderate risk of threshold mismatch does not invalidate the core contribution. However, the concern is concrete and testable: the end-to-end false-positive rate in the survey regime has not been demonstrated, and the new-candidate list is the part of the paper most sensitive to it. I therefore keep the ACCEPT verdict unchanged, but recommend running the null-cohort test before heavy follow-up resources are committed to the 13 new candidates and 87 follow-up candidates.","tokens_in":32344,"tokens_out":6209,"duration_ms":97816,"concrete_test":"Build a null cohort of 461 synthetic light curves with the same per-host window functions and noise properties as the survey targets (e.g., block-shift scrambling of each host's real light curve, preserving per-sector window functions and power, or GP-realized noise with injected activity). Run the full pipeline — segmented search, event-time screen, TRICERATOPS, and Gaia filter — on this cohort and count how many synthetic candidates survive all stages. If the per-host false-positive rate is ≲1% (i.e., ≲5 survivors in 461), threshold transfer is supported; if it is several percent, the reported recovery rates and the 13 new candidates are likely contaminated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The pipeline's added components — segmented-search federation thresholds (SDE≥3.5, p<0.01, stacked 5σ, P>12 d jury) and the event-time screen's p_noclock pass/flag/fail cuts — were frozen on ten benchmark systems, of which only L 98-59 and TOI-700 lie in the multiyear-gapped regime (Sect. 2.3). The paper concedes this 'measures the method's safety rather than its yield.' The central survey products — 85.5% confirmed-planet recovery, 78.7% candidate recovery, and the 13 new transit candidates plus 87 follow-up candidates — depend on these thresholds transferring to 461 hosts spanning 1–45 sectors and widely varying activity. The end-to-end false-positive rate of the full detection+screen chain in this regime is not directly measured: the block-shift scramble test (0/42) applies only to the segmented search's federation+confirmation, and the synthetic clockless-train test applies only to the event-time screen in isolation. If the p<0.01 federation gate or the p_noclock null under-rejects on gapped/active hosts, the new-candidate list could contain window-function aliases or quasi-periodic stellar-variability trains that the two gapped validation systems did not exercise. This is a genuine load-bearing uncertainty, though not a demonstrated failure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies a previously developed transit-detection and statistical-validation pipeline uniformly to 461 active ExoFOP M-dwarf TOI hosts (M0–M6). The pipeline merges a full-baseline TLS search with a segmented, semi-coherent search across gapped multiyear baselines, adds an event-time coherence screen (p_noclock), TRICERATOPS vetting, and Gaia DR3 background/aperture verification. Frozen-threshold validation was performed on ten benchmark systems. The survey reports recovery of 165/193 in-range confirmed planets (85.5%) and 225/286 in-range planet candidates (78.7%), with no confirmed planet flagged as a false positive. It yields 87 follow-up grade planet candidates, 13 new transit candidates (0.7–2.0 R_Earth), including two new members of a near-resonant three-planet candidate system around TOI-6284, and period determinations for four cataloged single- or dual-transit TOIs (TOI-2433.01, TOI-4565.01, TOI-4353.01, TOI-6492.01). The authors explicitly state that the survey validates no new planet; all strong candidates are cleared of dominant false-positive channels by archival follow-up and await radial velocities.","tokens_in":32723,"tokens_out":7541,"duration_ms":73744,"significance":"If the results hold, the paper provides the first homogeneous end-to-end detection, vetting, and verification survey of the active ExoFOP M-dwarf TOI population, producing a prioritized follow-up list that is immediately usable. The external benchmarks are genuinely strong: 165/193 confirmed planets recovered against independent ExoFOP labels with zero terminal false positives; blind co-recovery of TOI-237 c, a planet independently confirmed by Timmermans et al. (2026); and out-of-sample predicted-transit recoveries for TOI-4565.01 (sector 104) and TOI-4556.02 (held-out sectors). The paper is unusually candid about its limitations, enumerating ten systematic effects, and it ships public code and machine-readable catalogs. The principal unmeasured risk is the transfer of frozen thresholds from a ten-system validation set, of which only two systems lie in the multiyear-gapped regime the segmented search targets, to the full 461-host sample spanning 1–45 sectors and widely varying activity. That risk directly affects the false-positive rate of the 13 new transit candidates, which are the survey's most novel product.","major_comments":[{"comment":"The frozen thresholds of the added components—segmented-search federation (SDE≥3.5, p<0.01, stacked 5σ, P>12 d jury) and the event-time screen pass/flag/fail cuts—were calibrated on ten systems, of which only L 98-59 and TOI-700 lie in the multiyear-gapped regime. The paper's own statement in Sect. 2.3 that “this validation measures the method’s safety rather than its yield” is candid, but it means the end-to-end false-positive rate of the full detection+screen chain is not directly measured on the survey's actual gapped/active population. The block-shift scramble test (0/42) covers only the benchmark light curves, and the synthetic clockless-train test covers only the event-time screen in isolation. Because the 13 new transit candidates and the 87 follow-up candidates are central survey products, I ask for a null-cohort test: run the full chain on a sample of M dwarfs without known TOIs","section":"Sect. 2.2.2"},{"comment":"The statement that “because the p<0.01 gate is applied across the full sample, it still admits of order ten chance groupings” is not quantified. Across 461 hosts, with up to 30 peaks per block and multiple blocks per host, the total number of period groupings tested is large, and the expected number of chance groupings depends on that total. Please report the total number of groupings tested, the expected number under the null, and the number removed by each subsequent stage (windowed confirmation, depth consistency, P>12 d jury). This would let the reader judge whether the 0/42 block-shift scramble is representative of the full survey and whether the combined false-positive rate is actually controlled.","section":"Sect. 2.2.2"},{"comment":"For TOI-6284.02 and TOI-6284.03, the tabulated NFPP values (5.8% and 19.2%) still include the wide Gaia neighbors that the Gemini speckle imaging and ground-based nearby-eclipsing-binary check are claimed to clear. If those neighbors are indeed cleared, the NFPP should be recomputed with them excluded, as was done for TOI-4565.01 (Sect. 5.1) using archival contrast curves. As presented, the Tier 1 assignment for these two candidates rests on an external check that is not reflected in the reported vetting statistic, making the vetting pipeline appear more conservative than it actually is. Please either provide the updated FPP/NFPP after a TRICERATOPS rerun with the contrast curves, or explicitly state why such a rerun was not performed and how the reported NFPP should be interpreted.","section":"Sect. 5.2, Table 9"}],"minor_comments":[{"comment":"Title contains a spacing typo: “F ollow-up” should be “Follow-up”.","section":"Title"},{"comment":"The definition of ω as the median |O−C| is given in the text, but the notation is reused in Eq. (1) without explicit restatement near the equation. Consider defining ω directly in the equation caption or immediately before Eq. (1).","section":"Sect. 2.2.3, Eq. (1)"},{"comment":"The text states that four transits over a 2,577 d baseline hold a coherent linear ephemeris, while Table 9 gives p_noclock = FLAG LOWSENS (fewer than five timeable events). This is internally consistent because four events are fewer than five, but the reader would benefit from an explicit sentence noting that the flag reflects the small number of events, not incoherence.","section":"Sect. 5.1, TOI-4353.01"},{"comment":"The sentence “its cataloged 731.8 d period lying outside the search range, is not counted among the planet candidates” is slightly ambiguous: TOI-6492.01 is a recovery with a determined period, but it is not counted in the 87 follow-up planet candidates. Clarify that it is listed as a recovery and excluded from the PC counts for that reason.","section":"Sect. 5.1, TOI-6492.01"}],"recommendation":"major_revision","confidential_remarks":"The paper is substantial and thoughtful, and the external validation benchmarks are strong. My concern is not with the confirmed-planet recovery statistics, which are well supported, but with the new-candidate list, which depends on thresholds calibrated on a very small gapped-multiyear validation set. The requested null-cohort test is feasible and would directly address the main unresolved risk. I would be willing to accept after that test is provided and the TOI-6284 NFPP inconsistency is resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nBottom line up front: this paper deserves a serious referee. It is the first uniform end-to-end detection+vetting+verification run over the full active M-dwarf TOI population, and it ships two genuinely new pieces: a segmented semi-coherent search that recovers signals a full-baseline TLS scan dilutes (L 98-59 d is the worked example), and an event-time coherence screen (p_noclock) that separates real transit clocks from window-stacked stellar variability. The recovery statistics are internally consistent (85.5% of in-range confirmed planets, 78.7% of candidates), and the 13 new candidates include a plausible near-resonant three-planet chain around TOI-6284 and two habitable-zone period determinations. The blind recovery of TOI-237 c, validated independently by Timmermans et al., and the out-of-sample confirmations for TOI-4565.01 and TOI-4556.02 are genuine external checks.\n\nThe paper is also unusually candid. The false-negative taxonomy (Table 5) is honest, the limitations list (Sect. 6.3) is substantive, and the survey validates no new planet, a scope limitation the author states plainly. Credit for publishing the catalogs and code.\n\nThe soft spot is exactly where the author points: the added thresholds (segmented-search federation gates, p_noclock cuts) were frozen on a ten-system validation set in which only L 98-59 and TOI-700 lie in the multiyear-gapped regime the segmented search was built for. Sect. 2.3 says this “measures the method’s safety rather than its yield.” True. But that means the survey’s yield numbers are the least-tested part of the pipeline. The per-stage null tests (block-shift scramble 0/42, synthetic clockless trains) are good but do not measure the end-to-end false-positive rate of the full detection+screen chain on gapped, active hosts. A null-cohort test—running the whole chain on a set of known-negative light curves (e.g., non-TOI M dwarfs with similar sector coverage) and counting how often a candidate survives—would close this. Its absence does not break the paper, but it should be the first question a referee asks.\n\nMinor caveat: none of the new candidates are validated by design; the conservative all-sector FPP screen ranks rather than validates, and only 11% of confirmed planets pass the strict thresholds. That is explained, but the practical value depends on follow-up. The citation pattern is appropriate, since the detection core is the author’s own prior work.\n\nVerdict: send to review. The threshold-transfer risk is real, testable, and acknowledged—the kind of thing a good referee can push on without rejecting the paper.","headline":"A solid, honest survey product with real new candidates and a genuinely useful screen; the main open risk is that the new pipeline stages were calibrated on only two multiyear-gapped systems, an acknowledged and testable limitation rather than a fatal flaw.","tokens_in":33186,"tokens_out":2663,"would_cite":true,"duration_ms":36639,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-04T03:29:20.978492+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":2}