{"id":"7c2bcbf2-1288-4ce6-acd3-c0bc603393f9","arxiv_id":"2501.09494","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Using 1000 simulated transit observations, the authors show that HRDS exoplanet detection significances are strongly pipeline-dependent and that Bayesian retrievals recover signals that CCF-based metrics miss.","lead":"This paper simulates 1000 noisy observations of a known water-vapor signature in the atmosphere of the hot Jupiter HD 189733 b and shows that the computed detection significance depends heavily on the analysis pipeline used. A smart generalist should read it because it quantifies how random noise alone can make the same true signal look like a non-detection or a strong detection, which matters for how exoplanet atmosphere discoveries are claimed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative instability percentages rest on white-noise-only simulations; correlated noise/telluric variability could shift them substantially.","rationale":"I agree with the Reader that the weakest assumption is the white, uncorrelated Gaussian noise model of Eq. 3. The paper is otherwise carefully calibrated: it reproduces the published HD 189733 b detection within 1-sigma, uses 1000 realizations, and explicitly demonstrates the oversampling mechanism behind Welch versus S/N differences. The independent support here is real, because the 3.2 km/s oversampling correction has a quantitative prediction that can be checked, and the validation against the real CARMENES detection anchors the simulator. The concern I raise is not that the simulations are wrong internally; it is that the specific numbers highlighted in the abstract and introduction are ensemble statistics of a white-noise ensemble. Real HRDS noise is correlated (telluric residuals, instrument systematics, stellar activity), and the paper itself concedes this. I do not think this overturns the central qualitative message: the significance metrics genuinely respond differently to noise, and noise realizations genuinely scatter significances. But the quantitative claims -- 15% non-detections, 15% high-significance outliers, 0.8% false positives -- are conditional on the noise model. The Bayesian-retrieval robustness claim is a second weakness because it is tested with the same generating model (closed loop), but it is less load-bearing for the title claim and would matter mostly for the 'revisit discarded datasets' recommendation. Adding a correlated-noise experiment is feasible and would either confirm the percentages or force them to be caveated. Since the Reader already assigned CONDITIONAL and flagged the noise model, no verdict change is needed.","tokens_in":27777,"tokens_out":3424,"duration_ms":33494,"concrete_test":"Add a correlated-noise model to EXoPLORE (e.g., a stationary Gaussian process with covariance C(t,t')=sigma^2 exp(-|t-t'|/tau) and/or common-mode per-exposure systematics, choosing tau comparable to telluric water-vapor variability and normalizing so that the per-pixel variance matches Eq. 3). Regenerate 1000 in-silico HD 189733 b nights, rerun BL19 and SYSREM pipelines, and compare the S/N and Welch distributions, the fraction of nights with S/N<4.5 and S/N>6.6, and the 0.8% false-positive rate at true KP-Vrest. If these quantities stay within their quoted 1-sigma widths, the concern is resolved; if the tails shift by more than about 1-2 percentage points, the published percentages should be reported with a correlated-noise caveat. A complementary real-data check is to inject the simulated H2O signal into out-of-transit CARMENES residuals and verify the same 1000-realization statistics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claims -- the 5.6/5.4 S/N and 7.9/8.0 Welch expectations, the 'about 30%' spread, ~15% weak and ~15% spuriously strong nights, and the 0.8% false-positive rate at the true KP-Vrest -- are all generated from noise matrices drawn independently from N(0, sigma_noise) with sigma_noise(lambda,t)=1/S/N(lambda,t) (Eqs. 3-4). The simulations contain no correlated noise and no time-variable telluric residuals. The authors themselves note in Sect. 3.1.2 that such structures 'play a key role in either hindering or inflating signals in real datasets.' This matters in two ways. First, every quoted percentage is a statement about this specific white-noise ensemble; an equivalent correlated-noise ensemble with the same per-pixel S/N could shift the tails that produce the 15%/15% split, since CCF significances are driven by how noise aligns with the template. Second, because BL19 and SYSREM leave different correlated residuals in real data, the paper's finding that preparation pipelines do not matter (Sect. 3.1.2) may be an artifact of the white-noise assumption rather than a robust conclusion. The qualitative claim that significances are metric-dependent is well supported, but the quantitative pipeline-contingency percentages need a correlated-noise test before they can be applied to real observations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EXoPLORE, a simulation framework for ground-based high-resolution Doppler spectroscopy (HRDS) observations of exoplanet atmospheres, and uses it to generate 1000 synthetic CARMENES-like transit datasets of HD 189733 b with a fixed injected H2O signal and random Gaussian noise matrices. The authors compare two preparation pipelines (BL19 and SYSREM) and three significance metrics (CCF S/N, Welch's t-test, and Bayesian retrievals), finding that significance values depend strongly on the noise realization and on the chosen metric, that Welch's values exceed S/N values by a factor of about 1.5 due to oversampling when the CCF step is below the instrumental resolution element, and that Bayesian retrievals recover the injected signal in datasets where CCF-based methods fail. The framework is validated by reproducing the real HD 189733 b detection within 1 sigma (Table 1). The paper concludes that detection significances in HRDS should be contextualized statistically and that Bayesian retrievals are the most robust approach.","tokens_in":1834,"tokens_out":1971,"duration_ms":41175,"significance":"If the quantitative claims survive, the paper makes a genuinely useful contribution to comparative exoplanet spectroscopy: it provides a public-style simulation tool, quantifies the spread of detection significances owed purely to noise draws, and gives a concrete warning that thresholds such as S/N > 5 are not transferable across metrics or CCF sampling choices. The validation against a real CARMENES detection (Table 1) and the dry-atmosphere control (Sect. 3.7, Fig. S5) are good checks that lend credibility to the framework. The explicit finding that Welch's t-test values are inflated by oversampling when the CCF step is smaller than the resolution element is a specific, falsifiable methodological result that should be of immediate use to the field.","major_comments":[{"comment":"All quantitative percentages in the central claim -- the ~30% fraction yielding S/N > 6.6 or S/N < 4.5, the ~15%/15% split, and the 0.8% false-positive rate at the truth velocities -- are computed for a noise ensemble in which every noise matrix element is drawn independently from N(0, 1/S/N(lambda,t)) with no correlated noise and no time-variable telluric residuals. The authors themselves note in Sect. 3.1.2 that such structures 'play a key role in either hindering or inflating signals in real datasets.' Because CCF significances are driven by how noise aligns with the template, the quoted percentages and the ranking of BL19 versus SYSREM are claims about this specific white-noise ensemble. An equivalent correlated-noise ensemble with the same per-pixel S/N could shift the tails that produce the 15%/15% split. The qualitative claim that significances are noise- and metric-dependent is robust, but the quantitative pipeline-contingency percentages need to be re-evaluated under correlated-noise or telluric-variability injections before they are applied to real observations.","section":"Sect. 3.1.2, Eqs. (3)-(4)"},{"comment":"The claim that Bayesian retrievals are the most robust technique is evaluated only on data generated with petitRADTRANS using the same radiative-transfer code, the same line lists, and the same prior ranges that are used as the retrieval forward model. Fig. S9 validates the retrieval only on noiseless data from the same generator. This does not make the comparison circular in the narrow sense, since the retrieval must still fit the noisy data without knowing the injected parameters, but it likely overstates the robustness advantage over CCF methods: the retrievals are tested inside the model manifold that the forward model can reproduce, whereas real data contain model-error and systematics components. A concrete test would be to inject signals with a different chemical/thermal model, or to add correlated residuals, and then compare retrieval recovery rates with CCF recovery rates; without such a test, 'most robust' should be softened to 'most robust under the model-generated ensemble.'","section":"Sect. 2.1, Sect. 2.3, Eq. (8), Fig. S9"},{"comment":"The robustness advantage of Bayesian retrievals is supported by only 100 retrievals (Sect. 3.5), and the decisive comparison in Fig. 5 is based on the single best and single worst dataset according to the S/N metric, rather than a systematic statistic over the 100 realizations. The statement in Sect. 3.6 that retrievals 'clearly recover signals in datasets for which common CCF-S/N strategies fail' rests on one extreme example. Reporting the fraction of the 100 retrievals that recover the injected H2O abundance or K_P/V_rest within their 1-sigma credible intervals, broken down by S/N quartile, would make the claim quantitative and more convincing.","section":"Sect. 3.5, Sect. 3.6"}],"minor_comments":[{"comment":"In Table 1, the Welch's t-test rows do not report K_P and V_rest values for this work, although the main text states that the Welch's tests found maximum signals away from the expected K_P-V_rest; a brief note with those values or a placeholder would be helpful for comparison with the original Alonso-Floriano et al. (2019) results.","section":"Table 1"},{"comment":"The sentence introducing the 30% statistic is a bit informal: 'detections with significances within the aforementioned expected-value intervals is between 68% and 70%' is then followed by 'About 30% of the in silico observations will yield either a very high-significance detection (S/N > 6.6) or, alternatively, weak signals...' -- the relationship between the 68-70% and the 30% should be stated as a complement and the thresholds connected explicitly to the expected-interval definition.","section":"Sect. 3.1.2"},{"comment":"The labels 'P .Pos.-Max.', 'P .Pos.-Diff.', 'Away-Max.', and 'Away-Diff.' in Fig. 8 are difficult to parse; spelling out the labels (e.g., 'Planet position, maximize recovered peak') or adding a legend with the full phrase would improve readability.","section":"Fig. 8 and Sect. 3.8"},{"comment":"The concluding sentence 'we find the latter is generally more robust for high exposure-quality' appears to refer to the Welch's t-test, but the evidence in Figs. 1-2 shows comparable spreads and the comparison is not discussed in a dedicated quantitative manner; either add the supporting statistic or soften the sentence to reflect the exploratory nature of that claim.","section":"Conclusion"},{"comment":"No code availability statement is provided for EXoPLORE despite the framework being a central product of the paper. Given the reproducibility goals of the manuscript, a link to a public repository or a clear statement of availability on request should be added.","section":"Global"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and addresses a timely question. The main quantitative claims are currently tied to a white-noise-only simulation ensemble, and the strong 'Bayesian retrievals are most robust' claim would benefit from a systematic comparison over the 100 retrievals rather than a single extreme-pair example. I would be happy to see a revised version that adds correlated-noise tests and reports the retrieval recovery statistics quantitatively. I also note the reference to Blain et al. (2024b) contains a placeholder page number ('X(X):X') that should be corrected during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's qualitative message is solid: HRDS detection significances are strongly pipeline- and noise-dependent. Earlier work (Cabot, Cheverall) made this point; this paper quantifies it. The authors built a self-contained simulation framework (EXoPLORE), validated it against the CARMENES HD 189733 b water detection within 1 sigma, and then ran 1000 noise realizations through two preparation pipelines and three significance metrics. The cleanest new result is the explanation of the Welch-versus-S/N discrepancy via CCF oversampling, with a correction factor close to sqrt(resolution/sampling). The noise-correlation maps, showing that a random noise matrix can correlate or anticorrelate with the H2O template around the true KP-vrest, are also genuinely informative.\n\nThe main soft spot is the noise model. Equations 3-4 draw independent Gaussian noise with sigma = 1/(S/N). No correlated noise, no time-variable telluric structure. The authors acknowledge this in 3.1.2, but the abstract and conclusion lean on the quantitative robustness percentages. I would demote the specific numbers (the 30% split, ~15%/15%, the 0.8% false-positive rate) to 'under white-noise conditions' in the text. The central claim survives; those percentages are ensemble properties of a Gaussian ensemble, not predictions for real observing runs. Second, the 'Bayesian retrievals are most robust' claim is tested in closed loop: the retrieval forward model is petitRADTRANS, the same code that generated the injected signal. That is a real limitation, though it does not much affect the CCF-focused conclusions. Third, no code release is mentioned. For a methods paper explicitly aimed at reproducibility, shipping EXoPLORE would materially raise its value.\n\nThe reader's conditional verdict is right. The paper deserves a serious referee: it is well-executed, honest about its limitations, and directly useful to anyone interpreting single-epoch HRDS detections or planning ELT-era surveys. I would accept it for peer review with revisions, primarily to scope the numeric claims to the white-noise assumption and, ideally, to release the code.","headline":"A solid simulation study that makes the qualitative reproducibility point convincingly, but whose headline percentages rely on white-noise-only simulations and should be treated as illustrative, not as expected rates on real nights.","tokens_in":28570,"tokens_out":1885,"would_cite":true,"duration_ms":21361,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The reported significance of an exoplanet atmospheric detection depends almost as much on the random noise of the night as on the signal itself.","keywords":["exoplanet atmospheres","high-resolution Doppler spectroscopy","cross-correlation function","detection significance","noise simulations","Bayesian retrievals","HD 189733 b","SYSREM"],"falsifier":"One concrete test: take many real transits of the same bright hot Jupiter with the same instrument over several nights, inject a known water-vapor signal of the strength used here, and measure the distribution of recovered S/N; if the scatter is substantially larger than $\\pm 1.0$ or the tails deviate from the simulated 15%/15% split, the paper's quantitative noise-only picture does not cover real observing conditions. Alternatively, compute CCF maps of real telluric and stellar residual matrices from cloud-free nights and count how often a peak above S/N 4.5 appears at the expected $K_P$--$v_{\\rm rest}$; a rate well above 0.8% would falsify the false-positive estimate.","tokens_in":27441,"feed_emoji":"🔭","tokens_out":12008,"duration_ms":107088,"temperature":0.7,"pith_summary":"This paper argues that the significance attached to an exoplanet atmospheric detection with ground-based high-resolution Doppler spectroscopy is not a fixed property of the signal: it depends on the preparation pipeline, the significance metric, the velocity sampling of the cross-correlation, and, critically, on the particular random noise realization of the night. Using 1000 simulated transits of the hot Jupiter HD 189733 b with an identical injected water-vapor signal, it shows the expected signal-to-noise ratio is about $5.6 \\pm 1.0$ on the BL19 pipeline and $5.4 \\pm 1.0$ on SYSREM, with Welch t-test values of $7.9 \\pm 1.1$ and $8.0 \\pm 1.1$, while roughly 30% of otherwise identical observations come out either as very high-significance detections (S/N > 6.6) or as weak, unclaimable hints (S/N < 4.5). It further finds that Bayesian retrievals recover the injected signal even in simulated nights where cross-correlation methods fail. If this is right, single-night significance numbers carry an intrinsic scatter that must be quoted, and noise-only simulations should become part of standard detection reporting.","feed_headline":"Noise alone swings exoplanet signal from miss to 6.6-sigma","feed_subtitle":"Simulations show one in three identical nights lands outside the expected significance range.","key_machinery":"The load-bearing machinery is EXoPLORE, a three-part simulator that builds a noiseless model of the telluric transmittance, stellar spectrum, and exo-atmospheric transmission, overlays a Gaussian noise matrix $\\epsilon(\\lambda,t)$ with $\\sigma_{\\rm noise}=1/(S/N(\\lambda,t))$, and then feeds the synthetic observations through the same two preparation pipelines (BL19 polynomial fitting and SYSREM principal-component analysis) and the same three significance assessments (CCF S/N, Welch's t-test, and Bayesian retrieval evidence). The mechanism that produces the paper's headline result is the template-noise correlation: a random noise draw can contain a component that resembles the Doppler-shifted H$_2$O template at the expected $K_P$--$v_{\\rm rest}$, inflating the signal, or an anticorrelated component that suppresses it. A second mechanism is CCF velocity oversampling: with a 1.3 km s$^{-1}$ step against a 3.2 km s$^{-1}$ resolution element, Welch's in-trail values are not independent, inflating its significances by roughly $\\sqrt{3.2/1.3}\\approx1.6$ relative to S/N.","core_discovery":"The paper's central claim is that in high-resolution Doppler spectroscopy of exoplanet atmospheres, the measured detection significance of a real injected signal is not a fixed quantity: it fluctuates substantially from night to night purely because of which random noise matrix happens to be drawn. For 1000 in-silico transits of HD 189733 b with an identical water-vapor signal, the expected CCF signal-to-noise ratio is $5.6 \\pm 1.0$ for the BL19 pipeline and $5.4 \\pm 1.0$ for SYSREM, with Welch's t-test values of $7.9 \\pm 1.1$ and $8.0 \\pm 1.1$; about 30% of otherwise identical nights fall outside the expected-value interval, with roughly 15% producing S/N above 6.6 and 15% below 4.5. The variability is traced to noise matrices that cross-correlate with the H$_2$O template: the extreme nights show noise-only CCF maps with a correlation or anticorrelation region at the planet's true velocities. The paper further claims that Bayesian retrievals are more robust than CCF-based metrics, recovering the injected signal even in the nights where S/N is lowest, at the cost of much larger computational expense.","pith_inferences":["Editorial inference: if the noise-driven S/N scatter found here is typical, then archival single-transit detections should be re-examined with the same pipeline run on noise-only realizations, because a reported 5-sigma peak could be the tail of a $5.5 \\pm 1.0$ distribution rather than a secure detection.","Editorial inference: because the simulations isolate noise as the only variable, the framework naturally extends to predicting where detection scatter will dominate for smaller planets or future instruments by varying exposure S/N and spectral resolution, a regime the paper's NSF analysis already begins to map.","Editorial inference: the paper's 0.8% false-positive rate at the true velocities implies that the location of the peak in the $K_P$--$v_{\\rm rest}$ plane is itself a powerful robustness test, so future studies should report whether the maximum-significance peak coincides with the expected velocities rather than only quoting the maximum S/N."],"forward_implications":["Welch's t-test values should be quoted with the CCF velocity step; using a step finer than the instrument resolution element inflates them by roughly the oversampling factor $\\sqrt{3.2/1.3}\\approx1.6$ relative to S/N.","A detection reported at the expected $K_P$--$v_{\\rm rest}$ is far more trustworthy than an equal S/N peak elsewhere in the CCF map; noise alone produces such peaks in only 0.8% of simulated dry-atmosphere nights at the true velocities.","Datasets previously classified as non-detections by CCF may still contain a real signal recoverable by Bayesian retrieval, because retrieval global evidence does not correlate with CCF S/N.","Order-by-order optimization of SYSREM by injecting a signal at the planet's velocity and maximizing the recovered CCF peak selects noise, so that approach should be avoided; injecting away from the planet's velocities is safer.","The noise-scaling analysis shows that the nominal data quality for HD 189733 b falls in the regime where significance metrics are most noise-driven, implying current hot-Jupiter observations are near the instrumental limit for this science case."],"supporting_citations":[{"why":"Provides the real CARMENES H2O detection dataset and the S/N=6.6 measurement that the simulations reproduce, anchoring the truth values of orbital velocity, rest-frame velocity, and water abundance.","marker":"[Alonso-Floriano et al., 2019]"},{"why":"Defines the BL19 polynomial-fitting preparation pipeline and the likelihood-based Bayesian retrieval formalism used throughout the paper.","marker":"[Brogi and Line, 2019]"},{"why":"Supplies the retrieval log-likelihood function, priors, MultiNest sampling settings, and the HD 189733 b model parameters used to generate the synthetic observations.","marker":"[Blain et al., 2024b]"},{"why":"Computes the petitRADTRANS transmission spectra of the planet that are Doppler-shifted and injected as the ground-truth exo-atmospheric signal.","marker":"[Molli\\`ere et al., 2019]"},{"why":"Generates the BATMAN transit light curve that modulates the ingress and egress of the simulated exoplanet signal.","marker":"[Kreidberg, 2015]"},{"why":"Introduces the SYSREM principal-component algorithm used as the second preparation pipeline in the comparison.","marker":"[Tamuz et al., 2005]"},{"why":"Provides the t-test statistic whose in-trail versus out-of-trail CCF comparison yields the Welch sigma values reported.","marker":"[Welch, 1947]"},{"why":"Reports earlier discrepancies between analysis techniques for the same high-resolution dataset, which this study turns into a controlled simulation experiment.","marker":"[Cabot et al., 2019]"},{"why":"Documents robustness issues and S/N dependence on velocity windows in molecular detections, informing the paper's significance-metric comparisons.","marker":"[Cheverall et al., 2023]"}],"fun_headline_variants":["Noise alone swings exoplanet signal from zero to 6.6-sigma","One in three identical nights fails exoplanet signal check","Exoplanet detection significance is a roll of the noise dice","Same signal, same pipeline, different verdict each night","Random noise reshuffles exoplanet atmosphere evidence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative claims assume that observational noise is purely random, bell-shaped, and independent from one pixel and exposure to the next, with no time-correlated telluric or stellar residual structures; real data contain such correlated structures, so the exact percentages could shift.","fun_headline_variants_meta":{"raw":{"variants":["Noise alone swings exoplanet signal from zero to 6.6-sigma","One in three identical nights fails exoplanet signal check","Exoplanet detection significance is a roll of the noise dice","Same signal, same pipeline, different verdict each night","Random noise reshuffles exoplanet atmosphere evidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1730,"prompt_tokens":1130,"completion_tokens":600,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":746,"completion_tokens_details":{"reasoning_tokens":514}},"tokens_in":746,"tokens_out":600,"duration_ms":6873,"temperature":1.0,"reasoning_tokens":514,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:58:28.878427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete test: take many real transits of the same bright hot Jupiter with the same instrument over several nights, inject a known water-vapor signal of the strength used here, and measure the distribution of recovered S/N; if the scatter is substantially larger than $\\pm 1.0$ or the tails deviate from the simulated 15%/15% split, the paper's quantitative noise-only picture does not cover real observing conditions. Alternatively, compute CCF maps of real telluric and stellar residual matrices from cloud-free nights and count how often a peak above S/N 4.5 appears at the expected $K_P$--$v_{\\rm rest}$; a rate well above 0.8% would falsify the false-positive estimate.","supporting_citations":[{"cited_title":"petitradtrans-a python radiative transfer package for exoplanet characterization and retrieval","cited_arxiv_id":null,"evidence_quote":"Computes the petitRADTRANS transmission spectra of the planet that are Doppler-shifted and injected as the ground-truth exo-atmospheric signal."},{"cited_title":"Correcting systematic effects in a large set of photometric light curves","cited_arxiv_id":null,"evidence_quote":"Introduces the SYSREM principal-component algorithm used as the second preparation pipeline in the comparison."},{"cited_title":"The generalization of ‘student's’problem when several different population varlances are involved","cited_arxiv_id":null,"evidence_quote":"Provides the t-test statistic whose in-trail versus out-of-trail CCF comparison yields the Welch sigma values reported."},{"cited_title":"On the robustness of analysis techniques for molecular detections using high-resolution exoplanet spectroscopy","cited_arxiv_id":null,"evidence_quote":"Reports earlier discrepancies between analysis techniques for the same high-resolution dataset, which this study turns into a controlled simulation experiment."}],"review_version":1}