{"id":"9e524cc5-99f6-4ac1-89f7-36258e5351cf","arxiv_id":"1908.03593","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Simulations show a 2-meter-class telescope with a low-dark-current H-band NIR camera, dithering among 5-7 L/T dwarfs per night for 360-480 nights, has over an 80% chance of detecting at least one Earth-sized transiting planet.","lead":"This paper simulates telescopes and cameras to find the best design for a ground-based search for Earth-sized planets around tiny, cool L and T dwarfs. It recommends a 2-meter telescope with an infrared camera, watching five to seven targets per night for three to four years, and predicts roughly two planet detections.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation omits systematics the paper itself measures; adding 0.2% flat-field and variability noise pushes typical detections below the 7.1σ threshold, so the 80% success rate is an idealized upper limit.","rationale":"The reader's weakest assumption identifies precisely the load-bearing weakness: the simulation's noise model is uncorrelated Gaussian, while Section 4 provides evidence that flat-fielding and LT variability contribute noise comparable to the block-level photon noise and to the transit depths of the planets that dominate the detection count. My quantitative scaling confirms that adding the quoted 0.2% flat-field and variability terms reduces the SNR of a median detected planet (1.51 R⊕) from about 11 to about 6.6, below the 7.1σ threshold. This directly undermines the headline success rate, not merely the smallest planets. The qualitative telescope/instrument/cadence recommendations are unaffected, and the paper is transparent about its assumptions, so CONDITIONAL remains the appropriate verdict. No internal inconsistency or circularity was found; the concern is purely that the numerical predictions are optimistic because measured systematics are excluded from the Monte Carlo.","tokens_in":19952,"tokens_out":8472,"duration_ms":81643,"concrete_test":"Modify the Section 3.1 simulation to add, per target and per binned block, an independent Gaussian systematic of 0.2% (flat-field term from Section 4.1) plus a sinusoidal variability term with amplitude drawn from the Radigan (2014) and Metchev et al. (2015) distributions and a period of 2–12 hours. Re-run the 2000 Monte Carlo realizations for the recommended configuration (2-m, 4-year, 6 targets/night, 5-minute cadence) and recompute the success rate and mean planet yield. If the success rate drops below 50% or the mean yield below ~1, the abstract's 80%-and-~2 claim is not supported once measured systematics are included.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—over 80% chance of at least one planet and ~2 planets on average—rests on the Section 3.1 Monte Carlo, which models noise only as uncorrelated Gaussian scatter from the CCD equation. The paper's own Section 4 identifies two systematics that operate at exactly the level of the block noise used for detection. For the median target (mH≈15, 2-m telescope), the 3-minute block noise is about 0.21% (scaling the 1-hour, 7.1σ minimum radius of 0.57 R⊕ from Figure 2 to 6 exposures per block). The flat-field test in Section 4.1 gives 0.2% RMS stability even with careful placement, and Section 4.2 reports that 80% of L3–L9.5 dwarfs show >0.2% variability (Metchev et al. 2015). Adding two such 0.2% terms in quadrature raises the block noise to ~0.35%. A typical detected planet has radius 1.51 R⊕ (Figure 8d), corresponding to a depth of ~2.3% around a 0.88 RJup host; its SNR falls from ~11 to ~6.6, below the 7.1σ threshold. Thus a large fraction of the detections that drive the 80% success rate would be lost, and the headline yield is not robust to systematics the authors themselves demonstrate. The paper acknowledges these limitations qualitatively but does not propagate them into the simulation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a design study for a ground-based transit survey targeting L and T dwarfs. It first simulates photometric precision in several optical and near-infrared bandpasses using the CCD equation, and concludes that a low-dark-current H-band NIR detector on a 2-meter-class telescope provides the best sensitivity to Earth-sized transits. It then builds a forward Monte Carlo survey around a catalog of 998 spectroscopically confirmed L/T dwarfs, injecting planets according to Kepler M dwarf occurrence rates, simulating weather losses, and testing different dithering strategies, survey durations, and telescope sizes. The central quantitative claims are that an optimal survey uses a 2-m telescope, 360-480 observing nights, and 5-7 targets per night with 5-10 minute dithering, yielding over an 80% chance of detecting at least one planet and roughly 2 planets on average. Section 4 then discusses practical limitations including dithering flat-field instability, L/T dwarf photometric variability, single-transit follow-up, and false positives.","tokens_in":1445,"tokens_out":1485,"duration_ms":64392,"significance":"If the quantitative yields were robust, this would be a genuinely useful design study: a single modest ground-based telescope could conduct the first thorough transit search of L/T dwarfs and likely find several Earth-sized candidates, with important implications for exoplanet demographics and JWST follow-up. The paper's strengths are the explicit comparison of detector architectures, the realistic target sample, the transparent forward-modeling framework, and an unusually honest Section 4 that measures and reports the principal systematics. The recommended instrument and telescope combination is sensible and likely robust. However, the headline success rates and expected yields are computed from a noise model that excludes the systematics quantified later in the paper, so the abstract's 'over 80%' and 'around 2 planets' should be read as idealized upper limits rather than expected survey outcomes.","major_comments":[{"comment":"The survey simulation models photometric noise only as uncorrelated Gaussian scatter from the CCD equation, and defines a detection as a single 7.1 sigma binned point. This is the load-bearing assumption for the abstract's claim of over 80% success and about 2 planets. The systematics the authors themselves quantify in Section 4 act at exactly this scale. For a median mH about 15 target on a 2-m telescope, the 3-minute block noise is about 0.21% when the 1-hour, 7.1 sigma minimum radius of 0.57 R_Earth from Figure 2 is scaled to the six-exposure blocks used in Section 3.2. Section 4.1 reports 0.2% RMS flat-field stability even with careful target placement, and Section 4.2 quotes greater than 0.2% variability for 80% of L3-L9.5 dwarfs (Metchev et al. 2015). Adding two such 0.2% terms in quadrature raises the block noise to about 0.35%; the median detected planet radius is 1.51 R_Earth (Figure 8d), corresponding to a transit depth of about 2.3% around a 0.88 R_Jup host, so the typical detection SNR falls from about 11 to 6.6, below the 7.1 sigma threshold. A large fraction of the detections that drive the claimed success rates would be lost. The qualitative caveat in Section 2.4 and the discussion in Section 4 do not replace propagating these systematics through the Monte Carlo.","section":"§3.1, §3.2, §4.1, §4.2"},{"comment":"The detection criterion is a single binned point that crosses 7.1 sigma below the baseline; the simulation does not require the candidate signal to appear in multiple blocks, on multiple nights, or with the periodicity expected of a transit. Section 4.3 correctly states that most planets would produce only a single transit during the five-night observing windows, and Section 4.2 notes that non-periodic variability can mimic transits in non-continuous photometry. Since the success rates in Figures 6 and 7 are built on this single-block criterion, the yields are optimistic even apart from the systematics issue raised above. The assertion that 7.1 sigma 'virtually guarantees zero non-astrophysical false positives' holds only under the Gaussian-noise assumption, which is not the regime the paper itself documents.","section":"§3.1, §3.2, §4.3"}],"minor_comments":[{"comment":"The text describes an observing run in May 2018 but dates the Mimir observation of 2MASS 1337 as UT 25 May 2016; please reconcile this discrepancy.","section":"§4.1"},{"comment":"The reference 'Udalksi et al. (2015)' is a typo for 'Udalski et al. (2015).'","section":"§1.1"},{"comment":"The caveat that the analysis neglects systematic noise would be more useful if accompanied by a quantitative estimate of how much the minimum detection radii increase under the systematics listed in Section 4.","section":"§2.4"},{"comment":"Please check the panel-by-panel caption descriptions against the figure panels; as printed, the ordering of the labels in the text and the figure is confusing.","section":"Figure 8"},{"comment":"The meaning of '360-480 observing nights' is clear only after reading Section 3; consider stating this as '3-4 years at about 120 usable nights per year' in the abstract.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for PASP and the design recommendation is likely to survive revision. My main concern is that the abstract's quantitative yield claims are not supported by the simulation, because the simulation excludes systematics that the paper itself measures. This is fixable within the manuscript's scope by including those systematics in the Monte Carlo or by explicitly reframing the headline numbers as idealized upper limits with a sensitivity analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is the first paper I know of that actually simulates detector choice and survey yield for transiting planets around L/T dwarfs, and the qualitative recommendation is sound. Second thing: the quantitative success rates in the abstract are optimistic, because the Monte Carlo treats noise as Gaussian while Section 4 documents two systematics at exactly the level of an Earth-sized transit.\n\nWhat is new and good: the detector comparison uses a real sample of L/T dwarfs and nearby reference stars, and it makes a clear case for H-band with a low-dark-current NIR detector on a 2-m class telescope. The survey simulation is transparent: it propagates Kepler M dwarf occurrence rates from Dressing & Charbonneau through a forward model of geometry, weather, and dithering cadence, and it optimizes over telescope size, survey duration, targets per night, and time per target. The Mimir dithering test in Section 4.1 is real data, not a cartoon: 3% jumps if target placement is ignored, 0.2% with careful placement. The authors also flag that occurrence rates may be higher around L/T dwarfs than M dwarfs. No circularity here; the yield is a straightforward propagation of an assumed input.\n\nThe soft spots are real and load-bearing. The headline numbers—over 80% chance of at least one planet, about 2 planets on average—come from Section 3.1, where photometric noise is only the CCD equation. Section 4 then tells you that flat-field stability is 0.2% with careful placement, and that a large fraction of L dwarfs show intrinsic variability above 0.2%. The stress-test arithmetic is persuasive: for a median target, block noise is about 0.21%, and adding 0.2% flat-field and 0.2% variability in quadrature brings it to roughly 0.35%. A typical detected 1.5 R_Earth planet then drops from about 11 sigma to 6.6 sigma, below the 7.1 threshold. So a substantial fraction of the simulated detections would not survive real systematics. The paper acknowledges these limitations qualitatively, but it does not propagate them into the simulation that produces the abstract numbers.\n\nAlso, \"detection\" here means one binned point crossing 7.1 sigma, not a confirmed planet. The authors are honest about follow-up being needed, but the abstract's \"detect around 2 planets\" reads as confirmed planets when it really means candidate transits.\n\nNone of this kills the qualitative guidance. H-band, 2-m, dithering with careful placement, and 3-4 years are probably the right design choices. But the simulation should have included the systematics, or the abstract should have labeled the yields as idealized.\n\nI would send this to peer review. It deserves serious referee time, and any group planning such a survey should read it. The parts to trust are the detector comparison, the median detection radii, and the cadence optimization. The parts to quote with caution are the 80% success rate and the expected number of planets.","headline":"Useful design study with a sound qualitative recommendation, but the headline yields ignore systematics the authors themselves measure, so treat the 80% success rate as an idealized upper limit.","tokens_in":20796,"tokens_out":2823,"would_cite":false,"duration_ms":31510,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 2-meter-class telescope with an H-band detector, dithering among 5-7 L and T dwarfs per night for 360-480 nights, has an over-80% chance of detecting at least one transiting Earth-sized planet.","keywords":["surveys","stars: brown dwarfs","planets and satellites: detection","L dwarfs","T dwarfs","transit survey","near-infrared photometry","dithering strategy"],"falsifier":"Observe a sample of roughly 30-50 L/T dwarfs with a 2-meter-class telescope in H band using the recommended five-night dithering cadence, and measure the distribution of binned lightcurve scatter and the frequency of greater-than-1% excursions; if a substantial tail appears from flat-field placement or intrinsic variability, the assumed 7.1-sigma sensitivity to 1% transits fails and the predicted over-80% survey success rate is not attainable.","tokens_in":19710,"feed_emoji":"🪐","tokens_out":9805,"duration_ms":93188,"temperature":0.7,"pith_summary":"L and T dwarfs, objects straddling the boundary between the lowest-mass stars, brown dwarfs, and planetary-mass bodies, have never been thoroughly searched for transiting planets. This paper argues that a dedicated ground-based survey can be the first to do so: it simulates photometry and full observing campaigns, and finds that a 2-meter-class telescope with an H-band near-infrared detector, dithering between several targets each night for 360-480 nights, has an over-80% chance of detecting at least one Earth-sized planet, with about two detections on average. The numbers matter because L/T dwarfs are only about Jupiter-sized, so an Earth-radius transit is a roughly 1% dip, deep enough for a modest ground-based telescope to catch. The paper also shows why red-optical detectors are the wrong tool and why H-band is the right window. If the survey design is right, the first transiting planets around brown-dwarf-mass objects, prime targets for atmospheric characterization, are within reach of a single ground-based telescope.","feed_headline":"2-meter NIR survey has 80% chance of finding planets around L/T dwarfs","feed_subtitle":"Simulations show a 3-4 year H-band dithering survey would net about two Earth-sized planets.","key_machinery":"Two linked pieces carry the argument. The first is the CCD-equation photometric model: signal-to-noise is target signal divided by the quadrature sum of Poisson noise from the target, sky, dark current, and read noise, applied to six detector concepts (red-optical z-prime, and high- and low-dark-current NIR J/H/Ks bands) with real reference-star fields around 132 L/T dwarfs. It shows that a low-dark-current H-band detector gives the lowest one-hour scatter because L/T dwarfs are brightest there, reaching a median 7.1-sigma minimum detectable radius of roughly 0.57 Earth radii on a 2-m telescope. The second is a Monte Carlo survey simulator: 998 observable L/T targets, synthetic planetary systems drawn from Kepler M dwarf occurrence rates, host masses and radii from evolutionary models, transit shapes injected with the BATMAN code, photometry binned into blocks set by the dithering cadence, and a planet counted as detected when a binned point falls 7.1 sigma below the baseline. The 7.1-sigma threshold is chosen, as in large-scale transit surveys, to make non-astrophysical false positives essentially impossible, but it also means the design relies on detecting a single deep block rather than a full multi-transit lightcurve.","core_discovery":"The paper claims that the first thorough transit search of L and T dwarfs can be carried out from the ground with a single modest telescope, and specifies the design that maximizes yield. On the basis of simulated photometry for real L/T targets and Monte Carlo surveys of synthetic planetary systems, it recommends a 2-meter-class telescope with a low-dark-current near-infrared camera observing in H band, 360-480 observing nights (about 120 per year with 30% lost to weather), five to seven targets per night with a dithering cadence of 5-10 minutes per target, and five nights per target group. Under the assumption that L/T dwarfs host planets at the Kepler-measured M dwarf rate, this design has over an 80% chance of detecting at least one planet and yields about two detections on average; a specific four-year, 2-m, six-targets-per-night configuration has a fitted Poisson mean of 1.88 detections. The paper argues these would typically be 1-2 Earth-radius planets on roughly 4-day orbits receiving about a third of Earth's insolation, making them attractive targets for atmospheric follow-up. It also stresses that the yield is conservative, since occurrence rates appear to rise toward later spectral types.","pith_inferences":["Because the simulations treat noise as uncorrelated Gaussian scatter, the stated 80% success rate is likely an upper bound; the paper's own dithering test shows flat-field placement jumps of over 3% when uncontrolled, and L/T variability affects 3-24% of targets at the 2% level, both comparable to a 1% Earth transit.","The block-based detection method means most candidates will be single-transit events with unknown periods; the paper notes that period estimation from a single transit in non-continuous photometry is unproven, so the realistic near-term product may be a candidate list requiring follow-up rather than confirmed planets.","If the occurrence-rate trend toward later spectral types holds, the same design could yield three to four planets rather than two, making a dedicated 2-m NIR survey competitive with space-based transit searches for the lowest-mass hosts.","The detector comparison suggests that adding an H-band camera to existing red-optical ultracool-dwarf surveys is the most direct way to extend them into the L/T regime, since z-prime photometry is insensitive to sub-Earth planets for half the sample."],"forward_implications":["A 1-m telescope is insufficient for most of the sample, while a 4-m telescope adds only marginal success over a 2-m, so the efficient niche is a 2-meter-class facility.","A survey longer than about four years yields diminishing returns because the brightest observable targets have already been scheduled for five nights, so the fifth year adds little.","Dithering between five and seven targets per night outperforms staring at one target, and 5-10 minutes per visit beats both shorter and longer cadences because it balances time baseline against phase coverage.","Typical detections will be sub-2-Earth-radius planets on short periods, receiving roughly a third of Earth's insolation, making them potentially habitable-zone objects suitable for atmospheric follow-up.","If occurrence rates rise toward later spectral types, as the paper argues, the expected number of detections roughly doubles."],"supporting_citations":[{"why":"Supplies the Kepler M dwarf planet occurrence rates used to seed synthetic planetary systems; the predicted yields scale directly with these rates.","marker":"Dressing & Charbonneau (2015)"},{"why":"Provides the MEarth survey's reference-lightcurve requirement and dithering strategy that the single-telescope design adapts.","marker":"Nutzman & Charbonneau (2008)"},{"why":"Supplies the 7.1-sigma detection threshold and the criterion of counting a binned data point below threshold as a detection.","marker":"Sullivan et al. (2015)"},{"why":"The BATMAN code injects realistic transit shapes with limb darkening into the simulated lightcurves.","marker":"Kreidberg (2015)"},{"why":"Provides the low-dark-current NIR detector properties and 41.5% throughput used for the recommended H-band photometry.","marker":"Clemens et al. (2007)"},{"why":"Evolutionary models assign each target a mass and radius from its age and effective temperature, setting the transit geometry and depth.","marker":"Baraffe et al. (2015)"},{"why":"Quantifies the fraction of L/T dwarfs with more than 2% variability, the host variability that can mimic or mask a transit in the proposed cadence.","marker":"Radigan (2014)"},{"why":"Quantifies low-amplitude variability in L3-L9.5 dwarfs, a systematic the Gaussian-noise simulations do not include.","marker":"Metchev et al. (2015)"},{"why":"Evidence that occurrence rates increase with later M spectral types, used to argue the M dwarf occurrence rates are conservative for L/T hosts.","marker":"Hardegree-Ullman et al. (2019)"}],"fun_headline_variants":["80% odds: 2-m NIR survey reveals ~2 Earth-sized L/T dwarf planets","Dithering H-band survey could find 2 Earths around L/T dwarfs","80% chance: 2-meter NIR hunt for L/T dwarf transits nets ~2 planets","Earth-size planets around L/T dwarfs: 2-m survey shows >80% odds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey-success numbers assume the photometry is limited by uncorrelated Gaussian noise from the CCD equation, with no component from L/T dwarf variability or imperfect dithering flat-fielding, even though the paper's own Section 4 shows those systematics can be as large as an Earth-sized transit.","fun_headline_variants_meta":{"raw":{"variants":["80% odds: 2-m NIR survey reveals ~2 Earth-sized L/T dwarf planets","Dithering H-band survey could find 2 Earths around L/T dwarfs","80% chance: 2-meter NIR hunt for L/T dwarf transits nets ~2 planets","Earth-size planets around L/T dwarfs: 2-m survey shows >80% odds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000826,"raw_usage":{"total_tokens":3669,"prompt_tokens":1061,"completion_tokens":2608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":2511}},"tokens_in":677,"tokens_out":2608,"duration_ms":18405,"temperature":1.0,"reasoning_tokens":2511,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:08:34.799534+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe a sample of roughly 30-50 L/T dwarfs with a 2-meter-class telescope in H band using the recommended five-night dithering cadence, and measure the distribution of binned lightcurve scatter and the frequency of greater-than-1% excursions; if a substantial tail appears from flat-field placement or intrinsic variability, the assumed 7.1-sigma sensitivity to 1% transits fails and the predicted over-80% survey success rate is not attainable.","supporting_citations":[{"cited_title":"W., Winn, J","cited_arxiv_id":null,"evidence_quote":"Supplies the 7.1-sigma detection threshold and the criterion of counting a binned data point below threshold as a detection."},{"cited_title":"P., Sarcia, D., Grabau, A","cited_arxiv_id":null,"evidence_quote":"Provides the low-dark-current NIR detector properties and 41.5% throughput used for the recommended H-band photometry."},{"cited_title":"A., Heinze, A., Apai, D","cited_arxiv_id":null,"evidence_quote":"Quantifies low-amplitude variability in L3-L9.5 dwarfs, a systematic the Gaussian-noise simulations do not include."},{"cited_title":"Kepler Planet Occurrence Rates for Mid-Type M Dwarfs as a Function of Spectral Type","cited_arxiv_id":"1905.05900","evidence_quote":"Evidence that occurrence rates increase with later M spectral types, used to argue the M dwarf occurrence rates are conservative for L/T hosts."}],"review_version":1}