{"id":"d76d4eaf-c0fe-4654-8ab0-ee917454084a","arxiv_id":"2608.10707","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Directly convolving a modeled transmission spectrum biases low-resolution exoplanet retrievals; the correct approach is to convolve the stellar flux spectra and then take the ratio, and HR and LR helium observations are complementary.","lead":"This paper shows that at low spectral resolution, astronomers must compute a model atmosphere the same way the data are made: smear the stellar light and then take the ratio, rather than smearing the absorption signal directly, or the answers are biased. It then compares three instruments for observing escaping helium from exoplanets and finds that high-resolution and low-resolution observations are complementary, not redundant.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central convolution-bias claim is robust, but the headline 16% bias and instrument detection maps are set by the fixed 0.7 Å FWHM and Gaussian LSF; Appendix B shows factor-of-5 variation, so the quantitative conclusions are not yet secure.","rationale":"The reader's weakest assumption already identified the fixed FWHM and Gaussian LSF as the main conditions on the quantitative claims, and I agree with that identification. My focus is narrower: the 16% bias figure and the instrument detection maps are the quantitative core of the paper's practical message, and both depend on the assumed 0.7 Å FWHM and Gaussian LSF. Appendix B explicitly concedes a factor-of-5 variation with FWHM, which means the numerical headline is not robust even though the qualitative recommendation to use Eq. 1 is. I do not see a flaw in the analytic derivation in Appendix A.1; the structure of the argument is correct and the conclusion that direct convolution is biased is model-independent. The absence of shipped code prevents end-to-end numerical verification, but that is a reproducibility condition rather than a demonstrated error. Because the central methodological recommendation remains valid regardless of the exact bias magnitude, the appropriate verdict stays conditional rather than moving to accept or reject.","tokens_in":23839,"tokens_out":9683,"duration_ms":161643,"concrete_test":"Reproduce Sect. 3.2 and Sect. 4.1 with the He I FWHM as a free parameter (e.g., a uniform prior over 0.4-2.0 Å) and with the actual non-Gaussian NIRISS/NIRSpec LSF from JWST calibration, marginalizing over FWHM. If the recovered WASP-69 amplitude shifts by more than about 30% from 2.52% and the NIRSpec/NIRISS detection-threshold curves in Fig. 4 cross, then the 16% headline and the instrument ranking are artifacts of the fixed-FWHM/Gaussian-LSF setup; if the bias and ranking remain stable, the conditional concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central identity behind Eq. 1 is not in question: for a narrow line embedded in a stellar line, direct convolution of the transmission spectrum cannot equal the ratio of convolved fluxes, so some bias must exist at low resolution. The load-bearing numerical part is different: Sect. 3.2 claims a 16% bias at NIRSpec resolution for WASP-69, and Sect. 4.1 uses S/N_EW maps to rank instruments. This part is conditioned on fixing the He I FWHM at 0.7 Å in the fit (Sect. 3.2) and in the detection maps (Sect. 4.1), while Appendix B itself shows the LR-to-HR amplitude conversion varies by up to a factor of 5 with FWHM (Fig. B.1). Real He I profiles are not Gaussians of known width: dynamical broadening, saturation, and extended outflows broaden the line. If the true width differs, the 16% figure and the relative NIRSpec/NIRISS/NIRPS thresholds can shift substantially. A Gaussian LSF is also assumed; NIRISS SOSS in particular has a structured, non-Gaussian spectral response, which can change how the Si I line blends with He I. Thus the paper establishes the existence and sign of the bias, but not yet a robust quantitative bias or instrument ranking.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses how synthetic transmission spectra should be compared to He I observations. Its central methodological claim is that high-resolution theoretical absorption spectra must not be convolved directly with the instrumental profile; because observed transmission spectra are formed as the ratio of convolved in-transit and out-of-transit stellar fluxes (Eq. 1), direct convolution produces a resolution-dependent bias. The claim is supported by two independent routes: mock EvE observations fitted with the standard p-winds approach (Sects. 3.1 and 3.2) and a closed-form Gaussian model for the bias factor κ (Appendix A.1). The paper then uses the recommended methodology to compare NIRPS, JWST/NIRSpec, and JWST/NIRISS for detection (S/N_EW maps, Fig. 4) and characterization (weighted-distance metric, Fig. 5), concluding that NIRPS and NIRISS have comparable detection thresholds, NIRSpec performs best for faint targets, and HR/LR observations are complementary for dynamics versus outflow extent and baseline reconstruction. It closes with recommendations to use Eq. 1 and to account for baseline contamination (Sect. 5, Eq. 8).","tokens_in":24030,"tokens_out":7416,"duration_ms":71290,"significance":"If the central claim is accepted, this is an important correction for He I escape studies and, more broadly, for low-resolution retrievals of features narrower than the instrumental kernel: models must be evaluated as ratios of convolved fluxes, not by convolving model absorption spectra. The paper deserves credit for demonstrating the bias with both numerical mock observations and an analytic κ expression that agree, for including realistic stellar grids with POLDs and error bars, and for providing the GJ3090 b robustness check in Appendix B. The main weakness is that the quantitative bias magnitude (16%) and the instrument ranking depend on assumptions (a fixed Gaussian He I width of 0.7 Å; a Gaussian LSF) that the manuscript acknowledges but does not quantitatively bound. This does not threaten the existence of the bias, whose origin is mathematical, but it does limit the quantitative conclusions as currently stated.","major_comments":[{"comment":"The headline 16% bias for WASP-69 (Sect. 3.2) and the detection-threshold maps (Sect. 4.1, Fig. 4) are computed with the He I FWHM fixed at 0.7 Å. Appendix B itself shows that the LR-to-HR amplitude conversion varies by up to a factor of 5 when the FWHM is treated as free (Fig. B.1), and that the maps change when the reference system is GJ3090 b with FWHM ~1.0 Å (Fig. B.2). The statement in Sect. 4.1 that \"the conclusions drawn here ... remain valid\" is an assertion rather than a demonstrated robustness result. Please provide a sensitivity run of the Sect. 3.2 recovery and of Fig. 4 over the FWHM range 0.5–1.5 Å, and ideally with non-Gaussian He I absorption profiles, reporting the resulting range on the 16% bias and on the relative NIRSpec/NIRISS/NIRPS thresholds.","section":"Sect. 3.2 and Sect. 4.1; Appendix B (Figs. B.1, B.2)"},{"comment":"All mock observations assume a Gaussian instrument kernel, but the NIRISS versus NIRSpec comparison in Sect. 4.1 is driven by how the stellar Si I line blends with He I at σ_R ≈ 7 Å versus ≈ 2 Å. The real NIRISS/SOSS line spread function is structured and wavelength-dependent rather than a single Gaussian. Because the bias factor κ in Appendix A.1 (Eq. A.5) depends on the kernel shape and on the relative positions of the lines, a non-Gaussian LSF can alter both the bias magnitude and the NIRISS detection thresholds. Please test at least one asymmetric or empirical NIRISS LSF profile, or substantially soften the quantitative instrument-ranking conclusions and state this limitation explicitly.","section":"Sect. 2.2 (Eq. 1) and Sect. 4.1"},{"comment":"The claim that the stronger LR retrieval bias \"cannot be explained by POLDs\" is supported only qualitatively. The comparison mixes two changes at once: the mock spectrum is generated with EvE's 3D stellar and geometric treatment and with Eq. 1, while the p-winds grid is fitted after direct convolution without the stellar spectrum, as the authors state. To isolate the convolution bias from residual POLD or 3D-model mismatch, I suggest repeating the LR fit with the same mock but with models computed from Eq. 1, as already done in Appendix B. This would make Fig. 1 a controlled test of the central claim.","section":"Sect. 3.1 (Fig. 1)"}],"minor_comments":[{"comment":"Equations (A.4) and (A.5) contain notation that appears dimensionally inconsistent: Eq. (A.4) writes exp[-σ_i^2(λ-μ_i)^2/2] while the figure captions quote σ in Å, and the convolution amplitude factors in Eq. (A.5) contain σ^2/√(σ^2+σ_t^2) where the standard Gaussian convolution requires σ/√(σ^2+σ_t^2). Please correct these expressions, which are likely typographical but make the analytic derivation hard to follow.","section":"Appendix A.1"},{"comment":"There is a typo in \"We estimate the errros on the number of photons\", and the instrument name is written inconsistently as both NIRSpec and NIRSPEC (including in Fig. 1 labels); please standardize.","section":"Sect. 3.1"},{"comment":"The black spectrum is labeled \"No stellar spectrum\"; this is the star-biased direct convolution discussed in the text, so consider renaming it \"Direct convolution (star-biased)\" for clarity.","section":"Fig. 2"},{"comment":"The error scaling in Eq. (3) should state explicitly whether the reference S/N for WASP-69 is per single transit or per the coadded three-transit dataset used by Allart et al. (2025a), since the derived J-magnitude scaling depends on this choice.","section":"Sect. 2.2 (Eq. 3)"},{"comment":"The derivation of the baseline bias assumes normalized continuum fluxes and a homogeneous stellar disk; please state these simplifications explicitly when introducing F⋆ = Σ F_i and Eq. (6).","section":"Sect. 5 (Eqs. 6–8)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well within A&A scope and the central methodological point is sound; the requested revisions concern bounding the quantitative claims, not redoing the core derivation. I see no concerns about citation practice or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is Eq. 1: when you build a transmission spectrum from flux ratios, convolving the theoretical transmission spectrum directly is not the same as taking the ratio of convolved fluxes, and for a narrow line like He I at JWST resolution the difference is not negligible. The authors show this two ways — with explicit mock observations and with the analytic Gaussian derivation in Appendix A.1 that gives the bias factor kappa — and the two lines agree. That central claim is robust and I expect it to stick. Prior work gestured at related issues, but the quantitative demonstration for He I with JWST instruments is new.\n\nThe instrument comparison is also solid. NIRPS, NIRSpec, and NIRISS are modeled with realistic S/N scaling, and the detection maps are a useful way to see where each instrument wins. The baseline-contamination discussion in Sect. 5 is worth reading: a contaminated stellar reference biases the whole absorption time series in a wavelength-dependent way, which is easy to overlook.\n\nWhere the paper gets soft is in the quantitative claims. The 16% bias for WASP-69 and the detection maps assume a fixed He I FWHM of 0.7 Å, and Appendix B shows the predicted HR amplitude can vary by up to a factor of 5 with FWHM. The authors acknowledge this in Sect. 3.2 and Appendix B, so it is not hidden, but it means the specific numbers should not be read as universal. The LSF is also assumed Gaussian; NIRISS SOSS in particular has a structured, non-Gaussian response that could matter for the Si I blend. No code or artifacts are shipped, so the EvE/p-winds pipeline cannot be checked end-to-end. And the paper directly contradicts Dos Santos et al. (2023) on whether JWST gives tighter constraints, but spends very little space reconciling that disagreement.\n\nThe molecular-feature test in Sect. 3.3 shows the bias is small for broadband features, so the generalization to atmospheric retrievals in general is less dramatic than the abstract hints; the bias really bites for narrow lines embedded in deep stellar lines.\n\nBottom line: the central argument holds, the authors are honest about the conditions, and the paper deserves a serious referee. I would ask for reproduction artifacts and a robustness study that varies FWHM and LSF shape rather than just acknowledging the dependence. I would take it to our reading group and would cite Eq. 1 in my own work.","headline":"The convolution-bias claim is solid and worth taking seriously, but the paper's quantitative results are conditional on assumptions about FWHM and the instrumental line profile that the authors themselves flag.","tokens_in":24671,"tokens_out":2458,"would_cite":true,"duration_ms":26251,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model transmission spectrum must be convolved as a ratio of stellar fluxes, not as an absorption profile; at JWST/NIRSpec resolution this changes a recovered He I signal from 3% to 2.52%.","keywords":["atmospheric escape","metastable helium triplet","transmission spectroscopy","instrumental convolution","low-resolution retrievals","high-resolution spectroscopy","JWST","NIRPS"],"falsifier":"Generate a NIRSpec-resolution mock observation from a known 3% Gaussian He I absorption using the WASP-69 stellar spectrum and Eq. 1, then fit it with the direct-convolution method: the paper predicts the recovered amplitude is $2.52\\pm0.09\\%$; a rerun with the same inputs that recovers 3% within errors would falsify the central claim. Alternatively, analyze one transit observed simultaneously with NIRPS and NIRSpec through Eq. 1 and check that the two instruments agree on the absorption amplitude while the direct-convolution NIRSpec fit is low by the predicted amount.","tokens_in":23508,"feed_emoji":"🪐","tokens_out":13122,"duration_ms":225730,"temperature":0.7,"pith_summary":"The paper's central claim is that the standard shortcut in low-resolution transmission spectroscopy—convolving a model absorption spectrum with the instrument profile and comparing that directly to observations—introduces a systematic bias that grows as spectral resolution drops. The observed transmission spectrum is by construction a ratio of in-transit to out-of-transit stellar fluxes, each already degraded by convolution and resampling (Eq. 1), and convolution and ratio-taking do not commute. For the WASP-69 stellar spectrum, a true 3% Gaussian He I absorption with 0.7 Å FWHM is recovered as only $2.52\\pm0.09\\%$ by the direct-convolution approach at NIRSpec resolution—a 16% relative underestimate—and the paper derives an analytic bias factor $\\kappa$ that predicts when the shortcut is safe. With the correct method, the paper then compares NIRPS, JWST/NIRSpec, and JWST/NIRISS, finding that NIRPS and NIRISS detect similar He I levels, NIRSpec wins for faint targets, and low-resolution space data cannot resolve outflow dynamics but are the right tool for measuring outflow extent and baseline coverage; if this is right, low-resolution helium pipelines must be rebuilt around Eq. 1.","feed_headline":"Direct convolution biases low-resolution helium signals by 16 percent","feed_subtitle":"Convolving the model absorption instead of the convolved flux ratio underreports He I depth and biases JWST retrievals.","key_machinery":"The load-bearing object is Eq. 1, the definition of the observed absorption spectrum as one minus the ratio of convolved in-transit to out-of-transit stellar fluxes, and the bias factor $\\kappa(\\lambda)$ derived in Appendix A.1. $\\kappa$ is the ratio between the correctly computed absorption and the naive direct convolution of the planetary absorption profile; for Gaussian lines it has a closed form (Eq. A.5) that depends only on the widths, depths, and relative positions of the stellar and planetary lines and on the instrumental kernel. In the high-resolution and broadband limits $\\kappa\\to 1$, so the shortcut is harmless; in the intermediate-resolution regime relevant to He I with JWST, $\\kappa$ can reach about 1.6 for a deep narrow stellar line, meaning the naive approach underestimates the true absorption and biases retrieved parameters. The rest of the paper uses this machinery to build instrument detection maps and to compare high- and low-resolution constraints.","core_discovery":"The paper establishes that the transmission spectrum is defined by $A(\\lambda,t)=1-\\mathcal{R}_s(\\mathcal{G}_R f_{\\rm in})/\\mathcal{R}_s(\\mathcal{G}_R f_\\star)$, so a model that instead smears a theoretical absorption profile and fits it to this ratio misestimates the signal. Assuming Gaussian stellar lines, planetary lines, and a Gaussian instrumental kernel, the deviation collapses to a single multiplicative factor $\\kappa(\\lambda)$ (Eq. A.5): at high resolution $\\kappa\\to 1$, for features much broader than the kernel $\\kappa\\to 1$, but in the intermediate regime of He I at NIRSpec resolution $\\kappa$ departs from 1 and deep, narrow stellar lines like those of WASP-69 push the inferred peak absorption below the true value. Numerically, fitting a mock NIRSpec observation built from a true 3% peak, 0.7 Å FWHM Gaussian He I absorption with the star-biased (direct-convolution) method returns $2.52\\pm0.09\\%$, a 16% relative bias, while the same test with the shallower WASP-121 stellar lines gives a much smaller bias. The same treatment shows broadband molecular features are only mildly affected, because their intrinsic width exceeds both the kernel and the stellar lines. These results are independent of the thermospheric model and hold for low-resolution retrievals in general.","pith_inferences":["Published low-resolution He I retrievals that directly convolved the absorption profile may be systematically low for stars with deep narrow lines; re-running them through Eq. 1 could revise reported amplitudes and mass-loss rates without new observations (editorial inference).","The analytic $\\kappa$ factor could serve as a fast archival correction: with a stellar model and line-spread function in hand, divide the naively convolved model by $\\kappa$ instead of redoing the full forward model (editorial inference).","The same ratio-versus-convolution logic should apply to other narrow atomic lines observed at low resolution, so the paper's recommendation likely generalizes beyond He I to future JWST observations of alkali and metal lines (editorial inference)."],"forward_implications":["Low-resolution He I retrievals should be computed from convolved flux spectra (Eq. 1) rather than by convolving the absorption profile; otherwise signals are underestimated for stars with deep narrow stellar lines.","The shortcut remains acceptable at high resolution (NIRPS-class) and for broadband molecular features, so the correction is needed mainly for narrow lines at JWST-class resolution.","NIRPS and NIRISS have nearly identical He I detection thresholds, while NIRSpec/G140H detects fainter targets; JWST does not broadly outperform ground-based HR for constraining temperature and mass-loss.","Low-resolution space-based data lose the velocity information needed to infer outflow dynamics near the planet, but their long, finely sampled time coverage makes them the primary tool for measuring outflow spatial extent and for establishing reliable baselines.","Choosing a baseline that already contains extended-outflow absorption biases the entire absorption time series; simultaneous space-based coverage is the practical remedy."],"supporting_citations":[{"why":"The EvE transit code the paper uses for synthetic spectra originates here and defines the flux-difference formalism behind Eq. 1.","marker":"Bourrier & Lecavelier des Etangs 2013"},{"why":"Supplies the 3D atmospheric and POLD treatment in EvE that separates stellar-line effects from the convolution bias at low resolution.","marker":"Dethier & Bourrier 2023"},{"why":"Previous work showing POLDs do not bias the helium band at low resolution, which isolates the convolution bias as the source of the LR retrieval offset.","marker":"Carteret et al. 2024"},{"why":"Supplies p-winds, the thermospheric model used to generate the He I absorption profiles and the fitting grid.","marker":"Dos Santos et al. 2022"},{"why":"Provides the Turbospectrum models of WASP-69 and WASP-121 whose stellar-line depths set the Appendix A.1 bias magnitude.","marker":"Plez 2012"},{"why":"Provides Pandexo, used to compute JWST noise in the detection and constraint maps.","marker":"Batalha et al. 2017"},{"why":"A literature example of direct convolution of the absorption spectrum, the star-biased approach the paper shows is biased at low resolution.","marker":"Fournier-Tondreau et al. 2024"},{"why":"Another direct-convolution low-resolution retrieval used to show the bias is embedded in current practice.","marker":"Piaulet-Ghorayeb et al. 2024"},{"why":"The JWST low-resolution He I outflow-dynamics measurement that motivates the HR-versus-LR comparison and is reinterpreted under Eq. 1.","marker":"Allart et al. 2025b"},{"why":"Establishes that stellar spectra must be sampled finely enough to resolve lines, a requirement Eq. 1 inherits.","marker":"Deming & Sheppard 2017"}],"fun_headline_variants":["Direct convolution underreports helium escape by 16 percent","JWST helium signals biased by convoluted model comparison","Low-res helium escape observations lose 16% to inference method","Convolving model spectra skews exoplanet helium detections","Helium escape bias: direct convolution vs proper ratio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported magnitude of the bias and the instrument rankings assume Turbospectrum stellar models for WASP-69 and WASP-121, a Gaussian instrument line-spread function, and a fixed 0.7 Å He I width, so those numbers would shift for other stars or instrumental profiles, even though the existence of the bias does not depend on them.","fun_headline_variants_meta":{"raw":{"variants":["Direct convolution underreports helium escape by 16 percent","JWST helium signals biased by convoluted model comparison","Low-res helium escape observations lose 16% to inference method","Convolving model spectra skews exoplanet helium detections","Helium escape bias: direct convolution vs proper ratio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000218,"raw_usage":{"total_tokens":1547,"prompt_tokens":1157,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":773,"completion_tokens_details":{"reasoning_tokens":309}},"tokens_in":773,"tokens_out":390,"duration_ms":5092,"temperature":1.0,"reasoning_tokens":309,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:55:24.324967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a NIRSpec-resolution mock observation from a known 3% Gaussian He I absorption using the WASP-69 stellar spectrum and Eq. 1, then fit it with the direct-convolution method: the paper predicts the recovered amplitude is $2.52\\pm0.09\\%$; a rerun with the same inputs that recovers 3% within errors would falsify the central claim. Alternatively, analyze one transit observed simultaneously with NIRPS and NIRSpec through Eq. 1 and check that the two instruments agree on the absorption amplitude while the direct-convolution NIRSpec fit is low by the predicted amount.","supporting_citations":[],"review_version":1}