{"id":"807b5984-14f6-41bc-a324-64e77ed1b77e","arxiv_id":"2608.06207","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Simulated JWST retrievals with stellar contamination show CO2 on TRAPPIST-1f is detectable in 10+10 transits, CH4 needs about 50+50, and H2O is not detected in up to 100+100 transits.","lead":"This paper simulates JWST observations of TRAPPIST-1f to test how starspot contamination affects molecule detection in the planet's atmosphere. It finds carbon dioxide could be confirmed in about 20 total transit observations (10 with NIRSpec plus 10 with MIRI), while methane needs about 100 and water remains undetectable even after 200.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Variable-star contamination is the load-bearing weakness: §3.2's sensitivity test never co-adds multiple transits with per-transit stellar parameters, so the abstract's 10/50-transit thresholds are only shown for a constant star and are presented without that caveat.","rationale":"The reader's weakest-assumption concern is the time-invariance of stellar contamination, and that is indeed the most load-bearing assumption for the headline transit-number predictions. My stress-test sharpens this: the paper's own sensitivity test in §3.2 is not structured to test visit-to-visit variability, because it either retrieves single transits (where atmospheric constraints are weak for any reason) or, in the 15+15 case, appears to keep stellar parameters fixed within the co-added set. Thus the published numbers remain valid only as a constant-star, best-case baseline. The authors are transparent about this in the body, and the retrieval framework, open-source code, and Zenodo data availability are genuine strengths. The appropriate verdict is therefore unchanged from the reader's CONDITIONAL: the abstract should carry the constant-star / noise-free caveat and the transit-number claims should be presented as upper-bound expectations rather than unconditional JWST predictions. The proposed concrete test would directly determine whether per-transit stellar variability moves the Bayes factors below the stated thresholds, settling whether the concern lands.","tokens_in":20666,"tokens_out":6502,"duration_ms":81473,"concrete_test":"From the released forward model and PandExo settings (Zenodo doi:10.5281/zenodo.20836651), generate one 10+10 and one 50+50 transit dataset in which spot/facula covering fractions, temperatures, and surface gravities are drawn per transit from the Table 1 priors (or a rotation-modulated version matching Lim et al. 2023), while the atmospheric model stays fixed. Retrieve with a sampler that allows per-transit stellar parameters (with atmospheric parameters shared), and recompute the CO2 and CH4 Bayes factors via nested-model comparison. If CO2 drops below B=150 at 10+10, or CH4 fails to reach B=3 at 50+50, the constant-star assumption is load-bearing for the abstract's transit-number predictions and the headline numbers must be labeled as a constant-star baseline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 assumes 'the model holds for each transit' and explicitly calls this a baseline; Section 4 repeats that the results depend on an unchanging star. The central quantitative predictions (CO2 B>150 at 10+10 transits, CH4 B>3 at 50+50 transits, H2O not detected to 100+100 transits) therefore inherit this assumption. The paper tries to address this in §3.2, but the sensitivity test does not land: five randomized stellar-parameter sets are each used to simulate a single transit and retrieved separately, and the resulting unconstrained atmospheres are as consistent with one-transit SNR being too low as with variable contamination being the problem. The subsequent 15+15 test uses 'the same five instances' but, as written, still keeps stellar parameters fixed across the co-added transits within each retrieval; it does not allow contamination to change between visits. Real TRAPPIST-1 observations cited by the paper (Lim et al. 2023, Piaulet-Ghorayeb et al. 2025, Espinoza et al. 2025) show visit-to-visit contamination changes, so the regime the thresholds must survive is precisely the regime the sensitivity test does not simulate. The abstract presents ~10 and ~50 transits without this caveat, although the body labels them a best-case baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the POSEIDON retrieval framework to simulate JWST transmission spectroscopy of TRAPPIST-1f with a CO2-rich, potentially habitable atmosphere, including stellar contamination from unocculted spots and faculae. The authors generate noiseless PandExo simulations for NIRSpec PRISM and MIRI LRS, co-adding 5+5 through 100+100 transits, and retrieve atmospheric and stellar parameters. They report that CO2 is strongly detected (B>150) at 10+10 transits, CH4 is weakly detected (B>3) at 50+50 transits, H2O is never detected, and that NIRSpec-only retrievals perform nearly as well as the two-instrument combination. The abstract frames these as concrete JWST observing-time predictions, while the body repeatedly labels the results a best-case baseline because the star is assumed unchanged between transits and the simulated data contain no Gaussian scatter.","tokens_in":20952,"tokens_out":5789,"duration_ms":65302,"significance":"If the results hold, they provide useful, quantitative guidance for JWST program design: a short program could confirm a CO2-rich atmosphere on TRAPPIST-1f, while methane and water constraints require far more time than earlier contamination-free estimates suggested. The study is fully reproducible: it uses the open-source POSEIDON package, simulates data with PandExo, and deposits forward-model spectra, simulated data, and retrieval samples on Zenodo. The systematic grid over transit number and instrument combination is a strength, and the paper is unusually transparent about its assumptions, explicitly labeling the constant-stellar-contamination and noiseless-data choices as a baseline. The main risk is that the headline quantitative thresholds inherit these assumptions and are presented in the abstract without the same caveats.","major_comments":[{"comment":"The central quantitative claims (CO2 strong evidence at ~10 transits, CH4 weak evidence at ~50 transits) are derived under the assumption that stellar contamination parameters are identical in every co-added transit. Section 2.1 states this directly and calls the scenario a baseline, but the abstract presents the thresholds without this qualifier. The sensitivity test in Section 3.2 does not close the gap: the five randomized stellar-parameter instances are simulated only as single transits, and the 15+15 test uses each randomized instance as a fixed stellar model rather than allowing contamination to vary between the co-added transits. Because the paper itself cites observations (Lim et al. 2023; Piaulet-Ghorayeb et al. 2025; Espinoza et al. 2025) showing visit-to-visit contamination changes, the regime most relevant to real JWST data is precisely the regime not simulated. The abstract and conclusions should either state that the transit-number predictions assume a time-invariant star, or the simulations should be extended to co-add transits with per-transit stellar parameters.","section":"§2.1, §3.2, Abstract"},{"comment":"The simulated data are generated without Gaussian scatter. Section 2.2 acknowledges this and calls the results a best-case baseline, but the abstract's '~10 transits' and '~50 transits' are quoted as unconditional findings. The Gaussian-scatter sensitivity test is restricted to the 5+5 case, and Figures C1 and C2 show that individual noise draws can produce erroneous posterior constraints (for example, O3 in the third column of Figure C1). The transit numbers that anchor the abstract's claims (10+10 and 50+50) are not tested against noise draws, so the robustness of the headline evidence levels to realistic noise remains unknown. Please either add noise-draw tests at the threshold transit counts or qualify the abstract explicitly.","section":"§2.2, §3.2, Abstract"},{"comment":"The single-transit randomized-stellar-parameter retrievals are interpreted as evidence about a changing star, but the experiment lacks a control. A single transit with fixed, correctly modeled stellar parameters may also produce unconstrained atmospheric posteriors simply because the signal-to-noise ratio is too low; no uncontaminated single-transit retrieval is shown for comparison. As written, the statement that the posterior distributions are 'completely unconstrained' therefore does not isolate the effect of stellar variability from the effect of low SNR. A control retrieval on a single-transit dataset with the original stellar parameters (or with no contamination) is needed to support the interpretation.","section":"§3.2"},{"comment":"The statement in Section 2.3 that 'the wavelength range MIRI operates in is safe from stellar contamination' conflicts with the model's own limitation, stated in Section 2.1, that contamination is only modeled up to 5.5 µm, and with the paper's later acknowledgment in Section 4 that Espinoza et al. (2025) report contamination affecting features beyond 3 µm. Since MIRI covers 5–15 µm, the claim that MIRI adds no contamination-related value is not established by these simulations. The conclusion that NIRSpec alone achieves similar results should be framed as contingent on the adopted contamination model, which applies only to wavelengths shortward of 5.5 µm.","section":"§2.3, §4"}],"minor_comments":[{"comment":"The Bayes factors for CO2 at 15+15 transits disagree between text and table: the text gives 2.38×10^4 for both uncontaminated and contaminated cases, while Table A1 lists 2.38×10^5 (uncontaminated) and 1.04×10^4 (contaminated). Please correct the discrepancy.","section":"Table A1 and §3.1"},{"comment":"The abstract says '~10 transits' and '~50 transits' while the body uses '10+10' and '50+50' transit notation. Define the notation at first use so the abstract's numbers are unambiguous.","section":"Abstract and §3.1"},{"comment":"There are several typographical errors, including 'significnace' (§2.3), 'uncontamintated' (Table A1), 'tempreature' (§2.3), and 'find find' (§1). A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The classification scale for Bayes factors (no/weak/moderate/strong evidence and 'detection') is introduced only in Section 3.1; consider defining it earlier, in Section 2.3, since the abstract refers to B>150 and B>3.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's body is transparent about the baseline nature of its assumptions, which is commendable, but the abstract and conclusions present the transit-number thresholds as more robust than the simulations support. The key fix is to align the abstract with the body's own caveats and to strengthen the sensitivity tests at the threshold transit counts. The paper is within scope for the journal and, with those revisions, would be a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid simulation study that gives the first retrieval-based transit-number predictions for TRAPPIST-1f that explicitly include stellar contamination. That alone makes it useful for JWST proposal planning. The NIRSpec-only comparison is also genuinely new and relevant. The work is honestly executed: POSEIDON and PandExo are standard tools, the forward model and retrieval sharing opacity tables is normal for this kind of detectability study, and they run sensitivity tests for Gaussian scatter and for randomized stellar parameters. The data and example notebooks are archived, which is real evidence and worth crediting.\n\nThe soft spot is the time-invariant stellar assumption. The body is explicit: Section 2.1 says the model holds for each transit and calls it a baseline, and Section 4 repeats that. But the sensitivity test in Section 3.2 does not actually co-add transits with per-transit stellar parameters. The single-transit retrievals with different stellar draws being unconstrained show that one transit has too little SNR, not that variable contamination is harmless. The 15+15 test still fixes the stellar parameters across the co-added transits within each retrieval. So the headline numbers—10+10 transits for CO2 at B>150, 50+50 for CH4 at B>3, no H2O by 100+100—are demonstrated only for a constant star. This is a real limitation, not a manufactured one, and the observations they cite (Lim et al., Piaulet-Ghorayeb et al., Espinoza et al.) show the star does change between visits.\n\nThe abstract overstates the result by dropping the '+N' qualifiers and the constant-star assumption. That should be fixed before publication. The body itself is mostly honest—it repeatedly says 'best-case baseline'—so this is a framing and abstract problem rather than a technical error.\n\nThis deserves peer review. The modeling is appropriate, the caveats are in the body, and the results are directly relevant to JWST time allocation. I would send it to review with a request to correct the abstract and, if possible, add a test that co-adds transits with varying stellar parameters, even a toy version, so the threshold claims are qualified by what happens when the star actually changes. The paper is aimed at people planning TRAPPIST-1f observations and anyone working on M-dwarf contamination in retrievals; I would cite it if I worked in that area.","headline":"Useful new transit-number predictions for TRAPPIST-1f with stellar contamination, but the headline thresholds are only demonstrated for a constant star and the abstract overstates them.","tokens_in":21488,"tokens_out":1998,"would_cite":true,"duration_ms":23767,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even with the host star's spots and faculae veiling every spectrum, about ten transits of JWST data give strong evidence for CO2 in TRAPPIST-1f's atmosphere, while methane needs about fifty and water stays undetectable.","keywords":["exoplanet atmospheres","transmission spectroscopy","stellar contamination","TRAPPIST-1f","atmospheric retrieval","JWST","Bayesian evidence","habitable zone"],"falsifier":"Two observations would settle it. Monitoring the star across several visits — for instance with the back-to-back TRAPPIST-1b proxy strategy the paper cites — would show whether spot and facula parameters vary from transit to transit; if they vary at the level Lim et al. (2023) report, the fixed-star model breaks and the paper's own sensitivity test shows the atmosphere then becomes completely unconstrained. Separately, measuring the contamination spectrum beyond 5.5 μm, where this paper's model switches to a constant scaling: Espinoza et al. (2025) already report stellar contamination past 3 μm for TRAPPIST-1e, and contamination reaching the 4.3 μm CO2 band would directly invalidate the ten-transit strong-evidence result.","tokens_in":20437,"feed_emoji":"🔭","tokens_out":19398,"duration_ms":166580,"temperature":0.7,"pith_summary":"This paper asks how much JWST observing time is needed to read the atmosphere of the habitable-zone planet TRAPPIST-1f when the host star's own spots and faculae contaminate every transmission spectrum. Earlier detectability forecasts for the TRAPPIST-1 planets assumed a quiet, pristine star, but the first JWST visits show the star's active regions imprint spectral features far larger than any plausible atmosphere. Simulating a CO2-rich habitable atmosphere with a worst-case contamination model, the paper finds that roughly ten transits per instrument (twenty total) yield strong Bayesian evidence ($B>150$) for CO2, that methane reaches only weak evidence ($B>3$) at about fifty transits per instrument, and that water leaves no detectable trace even at one hundred transits per instrument. If true, the result sets a concrete observing budget: a short JWST program can establish the CO2-rich atmosphere that is likely a prerequisite for temperate surface conditions, while the molecules most diagnostic of a living world would require observing time well beyond the scale of current large programs.","feed_headline":"Ten JWST transits can confirm CO2 on TRAPPIST-1f","feed_subtitle":"Even with starspot veiling, methane needs about 50 transits and water stays out of reach — NIRSpec alone suffices.","key_machinery":"The argument runs through the POSEIDON retrieval code — a parametric planetary-atmosphere and radiative-transfer model coupled to Bayesian model comparison — extended with a two-heterogeneity stellar contamination model: one cool starspot and one hot facula whose fractional coverages, temperatures, and surface gravities are set to the median values retrieved from actual JWST observations of TRAPPIST-1 by Lim et al. (2023). Synthetic observations are generated with PandExo for NIRSpec PRISM (0.6–5.3 μm) and MIRI LRS (5–15 μm), with single-transit errors scaled by $1/\\sqrt{N}$ to represent $N$ transits, and MultiNest nested sampling computes Bayesian evidences for a full retrieval versus retrievals with one molecule removed. The detection metric is the Bayes factor $B$, with the paper's thresholds at $B>3$ (weak), $B>150$ (strong), and $B>600$ (detection); the central comparison is therefore the change in evidence when CO2, CH4, or H2O is dropped from the retrieval model.","core_discovery":"The paper's central claim is that stellar contamination, when modeled accurately and assumed identical across transits, does not prevent JWST from detecting a CO2-rich atmosphere on TRAPPIST-1f, but it roughly doubles the observing time needed to see methane and leaves water out of reach entirely. Retrievals on simulated NIRSpec PRISM and MIRI LRS observations recover CO2 with strong evidence ($B\\approx 197$) at ten-plus-ten transits and a confident detection by fifteen-plus-fifteen, whereas CH4 shows no evidence at twenty-five-plus-twenty-five and only weak evidence ($B\\approx 5$) at fifty-plus-fifty, rising to strong evidence ($B\\approx 177$) at one hundred-plus-one hundred. H2O never exceeds a Bayes factor of about 0.9 in any contaminated retrieval, and explicit tests for O2 and O3 at the largest dataset likewise return no evidence. The paper further claims that leaving MIRI LRS out of the retrievals changes none of these conclusions, so NIRSpec PRISM alone delivers the same atmospheric constraints at half the observing cost, and that the contamination hides every spectral feature shortward of about 1.5 μm.","pith_inferences":["The transit counts are optimistic lower bounds, not operating budgets: the simulated data contain no Gaussian scatter, and the paper's own five-draw noise tests at five-plus-five transits show individual draws can shift posteriors or create spurious constraints, so real programs should plan extra margin.","If TRAPPIST-1f behaves like its siblings, the fixed-star assumption will fail — Lim et al. (2023) and Piaulet-Ghorayeb et al. (2025) already see contamination varying between visits of planets b and d — and the paper notes that allowing stellar parameters to float per transit is not computationally feasible with current tools; per-visit stellar modeling is the natural next step.","The MIRI-redundancy result points to a staged observing strategy the paper leaves implicit: confirm CO2 with a small NIRSpec campaign first, and commit hundreds of hours to chase methane only after the star's contamination behavior is understood well enough to model it.","If contamination extends past 5.5 μm as Espinoza et al. (2025) observe elsewhere in the system, the CO2 band at 4.3 μm — the paper's main detection channel — could itself be veiled, which would push the ten-transit threshold upward."],"forward_implications":["A short JWST program of about ten NIRSpec PRISM transits can return strong evidence for a CO2-rich atmosphere on TRAPPIST-1f, the kind of atmosphere the paper argues is a likely prerequisite for habitable surface conditions.","Methane evidence appears only after roughly fifty transits per instrument once contamination is included — about twice the observing time the same simulations predict for a pristine star.","Water is not retrievable even at one hundred transits per instrument, so a null result for H2O in any near-term campaign cannot be read as a dry planet.","MIRI LRS adds essentially nothing for CO2 or CH4, so the same atmospheric constraints are available from NIRSpec alone at half the observing time.","The contamination hides all spectral features shortward of about 1.5 μm, so the short-wavelength half of the PRISM band contributes little to atmospheric inference under this model."],"supporting_citations":[{"why":"Supplies the two-heterogeneity stellar contamination model (spot plus facula) and the median retrieved parameters used to build the contaminated forward spectra.","marker":"Lim et al. (2023)"},{"why":"Provides the CO2-rich, 5-bar habitable atmosphere model for TRAPPIST-1f that the retrievals are designed to recover.","marker":"Payne & Kaltenegger (2024)"},{"why":"The PandExo package generates the simulated NIRSpec and MIRI JWST noise whose errors are scaled for N transits.","marker":"Batalha et al. (2017)"},{"why":"The POSEIDON retrieval code coupling the forward model to Bayesian parameter estimation and model comparison.","marker":"MacDonald & Madhusudhan (2017)"},{"why":"MultiNest nested sampling computes the Bayesian evidences whose ratios define the CO2, CH4, and H2O detection significances.","marker":"Feroz et al. (2009)"},{"why":"Establishes the POSEIDON terrestrial retrieval setup and the observing-time convention (twice the transit time plus the transit) reused here.","marker":"Lin et al. (2021)"},{"why":"Supplies the theory that unocculted cold spots create positive transmission-spectrum features and faculae create negative ones, motivating the contamination model.","marker":"Rackham et al. (2018)"},{"why":"Reports stellar contamination beyond 3 μm on TRAPPIST-1e, the observation the paper cites as its own 5.5 μm contamination cutoff may not capture.","marker":"Espinoza et al. (2025)"}],"fun_headline_variants":["CO2 on TRAPPIST-1f confirmed in just 10 JWST transits despite starspots","NIRSpec alone finds CO2 on TRAPPIST-1f in 10 transits, MIRI not needed","Starspot veiling can't hide CO2 on TRAPPIST-1f: 10 transits reveal it","Methane on TRAPPIST-1f needs 50 transits; water stays hidden even at 100"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the star's spots and faculae are identical on every transit, so a single set of six stellar parameters fits all visits; real TRAPPIST-1 observations show the contamination pattern changing between visits, and the paper itself labels the fixed-star situation a best-case baseline rather than a description of the real star.","fun_headline_variants_meta":{"raw":{"variants":["CO2 on TRAPPIST-1f confirmed in just 10 JWST transits despite starspots","NIRSpec alone finds CO2 on TRAPPIST-1f in 10 transits, MIRI not needed","Starspot veiling can't hide CO2 on TRAPPIST-1f: 10 transits reveal it","Methane on TRAPPIST-1f needs 50 transits; water stays hidden even at 100"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3452,"prompt_tokens":1172,"completion_tokens":2280,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":788,"completion_tokens_details":{"reasoning_tokens":2166}},"tokens_in":788,"tokens_out":2280,"duration_ms":14906,"temperature":1.0,"reasoning_tokens":2166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:33:10.426880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two observations would settle it. Monitoring the star across several visits — for instance with the back-to-back TRAPPIST-1b proxy strategy the paper cites — would show whether spot and facula parameters vary from transit to transit; if they vary at the level Lim et al. (2023) report, the fixed-star model breaks and the paper's own sensitivity test shows the atmosphere then becomes completely unconstrained. Separately, measuring the contamination spectrum beyond 5.5 μm, where this paper's model switches to a constant scaling: Espinoza et al. (2025) already report stellar contamination past 3 μm for TRAPPIST-1e, and contamination reaching the 4.3 μm CO2 band would directly invalidate the ten-transit strong-evidence result.","supporting_citations":[],"review_version":1}