{"id":"9b313069-69c4-4f64-8c59-52f9dc7c327e","arxiv_id":"2504.18815","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":16,"one_line_summary":"NEXOTRANS, a Bayesian plus machine-learning retrieval framework validated against eight existing codes, retrieves super-solar O/H, C/H, S/H, and C/O from combined JWST NIRISS, PRISM, and MIRI spectra of WASP-39 b.","lead":"This paper introduces NEXOTRANS, a new framework for analyzing exoplanet atmospheres that combines Bayesian statistics with machine learning. It applies the tool to JWST observations of WASP-39 b and reports super-solar oxygen, carbon, and sulfur abundances.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline abundances rest on a model-selection reversal: the paper's own Bayesian evidence favors equilibrium-offset chemistry (Delta lnZ ~ 14), while nearly identical reduced chi-squared values are used to crown hybrid equilibrium as best fit.","rationale":"The reader's verdict is CONDITIONAL, and my concern reinforces that conditionality rather than overturning it. The reader identified the cross-instrument offset correction as the weakest assumption; that is a serious issue, especially because offsets are derived from a preliminary retrieval and applied without propagating their uncertainties. However, I judge the model-selection contradiction to be more load-bearing for the specific headline claim: the abstract and conclusion quote super-solar abundances from the 'best-fit modified hybrid equilibrium model,' yet the paper's own Bayesian evidence strongly prefers equilibrium-offset chemistry. The reduced chi-squared values are 2.97 and 2.98, effectively indistinguishable, so they cannot justify reversing the evidence ranking. Since the framework's raison d'etre is Bayesian model comparison via nested sampling, selecting a model by reduced chi-squared when the evidence disagrees is internally inconsistent. This is not a criticism of the forward model: the MIRI benchmark in Section 2.4 and Appendix 6.3 shows agreement with eight independent codes for H2O, SO2, and cloud parameters, which is real independent support. The machine-learning uncertainty intervals are admittedly heuristic and uncalibrated, but the central abundance claim rests on the Bayesian retrievals. The concrete test I propose would settle the model-selection issue directly and would determine whether the abstract's numbers should be re-assigned or model-averaged. If the evidence still favors the offset model, the verdict should remain CONDITIONAL until the authors reconcile the selection criterion and re-report abundances accordingly.","tokens_in":34007,"tokens_out":3714,"duration_ms":41322,"concrete_test":"Re-run the hybrid and equilibrium-offset Bayesian retrievals with identical priors, likelihood, datasets, and aerosol treatment, computing ln(Z) with both UltraNest and PyMultiNest. If Delta lnZ = lnZ(offset) - lnZ(hybrid) remains greater than 5, the abstract's 'best-fit modified hybrid equilibrium' assignment is unsupported, and the headline abundances must be re-quoted for the evidence-favored model or model-averaged. If the evidence ranking flips, or if the two models' O/H, C/H, S/H, and C/O posteriors overlap within 1 sigma, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing soft spot is the model-selection step, not the forward model. Section 4.1.1 reports ln(Z) = 3148.86 for hybrid equilibrium and ln(Z) = 3163.09 for equilibrium offset, giving a decisive Bayes factor of roughly e^14.2 in favor of the offset model under the same data and priors. Yet the paper selects 'modified hybrid equilibrium' as the best-fit model using reduced chi-squared (2.97 vs 2.98), values that are nearly identical and do not discriminate between the models. The abstract's headline abundances—O/H = 14.12 x solar, C/H = 21.37 x solar, S/H = 5.37 x solar, C/O = 1.35 x solar—are quoted specifically for the hybrid model. The evidence-favored equilibrium-offset model gives different values (log O/H -1.72 vs -2.16 and C/O 0.89 vs 0.80; still super-solar, but shifted). Because NEXOTRANS is a Bayesian framework that explicitly computes evidence, using reduced chi-squared to overturn the evidence ranking is not justified. Reduced chi-squared near 3 for all models also indicates underfitting, so chi-squared is a weak arbiter even on its own terms. The MIRI benchmark against eight retrieval codes gives independent support for the forward model, but the specific WASP-39 b abundance values depend on a model-selection choice that the paper's own statistics contradict.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents NEXOTRANS, a new atmospheric retrieval framework that combines Bayesian nested sampling (PyMultiNest and UltraNest) with four machine-learning regressors, and applies it to JWST transmission spectroscopy of WASP-39 b using NIRISS (0.6-2.8 micron), NIRSpec PRISM (2.0-5.3 micron), and MIRI (5-12 micron). Four chemistry treatments are explored: free, equilibrium, modified hybrid equilibrium, and modified equilibrium-offset chemistry. The paper's central claims are that NEXOTRANS is validated as a hybrid Bayesian plus ML framework, that WASP-39 b's atmosphere is super-solar in O, C, and S, and that the best-fit model is the modified hybrid equilibrium chemistry, yielding O/H = 14.12 (+2.86/-1.82) x solar, C/H = 21.37 (+4.93/-3.18) x solar, S/H = 5.37 (+0.79/-0.65) x solar, and C/O = 1.35 (+0.05/-0.02) x solar. The authors also report constraints on SO2, Na, K, aerosols (ZnS, MgSiO3), and terminator cloud properties.","tokens_in":34464,"tokens_out":5315,"duration_ms":55905,"significance":"If the framework and the WASP-39 b results hold, NEXOTRANS would be a useful community tool: the MIRI benchmark in Table 5, comparing against eight independent retrieval codes across three pipeline reductions, provides strong external evidence that the forward model and sampler are sound, and the NEXOCHEM comparison against FastChem in Appendix 6.1 is a concrete, reproducible validation. The ML-Bayesian consistency checks are also a useful contribution. However, the headline elemental abundances are not reliable as presented because the model-selection step contradicts the paper's own Bayesian evidence, and because the instrument offsets used to align the datasets are derived from the same data and their uncertainties are not propagated. The scientific conclusions therefore require substantial reworking before the paper can be accepted.","major_comments":[{"comment":"The model-selection step is internally inconsistent with the paper's Bayesian framework. The text reports ln(Z) = 3148.86 for hybrid equilibrium and ln(Z) = 3163.09 for equilibrium offset under the same data and priors, a Bayes factor of roughly e^14 strongly favoring the equilibrium-offset model. The paper nevertheless selects hybrid equilibrium as the best statistical fit because its reduced chi-squared (2.97) is marginally lower than 2.98. These two reduced chi-squared values are statistically indistinguishable, and when Bayesian evidence has been computed, it is not justified to overturn the evidence ranking with an essentially flat chi-squared comparison. This choice directly affects the abstract's headline abundances, which are quoted for the hybrid model; the evidence-favored equilibrium-offset model gives different values, including log O/H of -1.72 versus -2.16 and C/O of 0.89 versus 0.80 (both still super-solar O/C/H/S, but shifted). The authors should either present the equilibrium-offset model as the primary result or provide a statistically justified argument, based on the computed evidences, for preferring the hybrid model.","section":"Section 4.1.1"},{"comment":"The alignment of the three instruments rests on constant multiplicative offsets (57 ppm for NIRISS and 311.13 ppm for MIRI, relative to NIRSpec PRISM) that were themselves retrieved from an initial free-chemistry retrieval on the same combined dataset; all subsequent retrievals are then run on the corrected data. This is circular in an important sense: the offsets absorb any absolute-scale or wavelength-independent systematics, but their uncertainties are not carried into the final posterior distributions, and any wavelength-dependent residual between instruments would be misattributed to chemistry. The authors should fit the relative offsets as nuisance parameters simultaneously in all models, or demonstrate explicitly that the retrieved mixing ratios, C/O, and elemental abundances are insensitive to the offset values and their uncertainties. As written, the quoted error bars on the headline abundances omit this systematic term.","section":"Section 3, offset correction paragraph"},{"comment":"The goodness-of-fit criterion itself is weakened by the fact that all combined-dataset reduced chi-squared values are near 3 (2.97-3.35), with per-instrument values of 2.67 (NIRISS), 3.19 (PRISM), and 2.14 (MIRI). These values indicate that the models underfit the data, yet the paper does not quote chi-squared probabilities, model the noise covariance, or rescale the error bars. Under these conditions, the small difference between reduced chi-squared values used for model ranking (2.97 versus 2.98) is not statistically meaningful. This reinforces Major Comment 1 and needs to be addressed by a proper noise model or by an explicit discussion of the underfitting and its consequences for parameter uncertainties.","section":"Section 4.2"}],"minor_comments":[{"comment":"The text states that the Stacking Regressor achieves the highest R^2 score of 0.855, but Table 4 reports R^2 = 0.76 for the Stacking Regressor; this numerical inconsistency should be corrected.","section":"Section 6.2.4 and Table 4"},{"comment":"The ZnS modal particle radius is printed as 'log(rc/um) = 1.29' with a positive sign, whereas all comparable values elsewhere in the paper are negative (e.g., -1.29 in the free-chemistry retrieval); this is likely a sign error and should be corrected.","section":"Section 3.1.3"},{"comment":"The planet name is written inconsistently as both 'WASP-39 b' and 'WASP-39b'; the manuscript should adopt one convention.","section":"Throughout"},{"comment":"The table header 'Redchi2' should be written as 'reduced chi-squared', and the caption should clarify which results correspond to NEXOTRANS's three cloud parametrizations for each reduction.","section":"Table 5 caption"},{"comment":"The text says the per-dataset reduced chi-squared values were obtained by performing retrievals with the global best-fit hybrid equilibrium model; it should be stated explicitly whether these values come from fixed global best-fit parameters re-evaluated on each dataset or from re-fits to each individual dataset, since the interpretation differs.","section":"Section 4.2 and Appendix 6.4"}],"recommendation":"major_revision","confidential_remarks":"The strongest asset of the paper is the independent MIRI benchmark against eight retrieval codes in Table 5, together with the NEXOCHEM versus FastChem validation; that portion is publishable quality. The blocking issues are the model-selection reversal in Section 4.1.1 and the unpropagated, data-derived instrument offsets in Section 3. Both are fixable within the manuscript's scope, so I do not recommend rejection, but the headline WASP-39 b abundances should not be published until these points are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know before you spend time on this one. The first is that the framework validation is the real contribution. On MIRI data, NEXOTRANS retrieves H2O and SO2 within 1 sigma of eight independent retrieval codes across three reductions, with cloud parameters consistent as well. NEXOCHEM tracks FastChem over the tested grid. That is independent, reproducible evidence that the forward model and sampling work, and it deserves credit. The second thing is that the WASP-39 b science headline does not survive contact with the paper's own statistics. The Bayesian evidence favors equilibrium-offset chemistry by Delta ln Z ~ 14, yet the paper selects hybrid equilibrium as best fit using reduced chi-squared values of 2.97 versus 2.98, which are statistically identical and both indicate underfitting. The super-solar O/H, C/H, S/H, and C/O quoted in the abstract are for the hybrid model. That is not a minor footnote; it is the paper's main scientific claim resting on a model-selection choice that its own evidence contradicts.\n\nWhat is genuinely new here is the specific assembly: nested sampling with four machine-learning regressors, the NEXOCHEM equilibrium grid, and the combined NIRISS+PRISM+MIRI retrieval including MIRI data. The ML results track the Bayesian medians reasonably well, though the uncertainty intervals come from a local +/-10% perturbation scheme and are not calibrated; the authors acknowledge this limitation.\n\nSoft spots, in proportion. The 57 ppm and 311 ppm instrument offsets are retrieved from the same data and then applied before all later retrievals, with no propagation of offset uncertainty. If those offsets are wavelength-dependent, the headline abundances shift. Reduced chi-squared near 3 across all models means the fits are underfitting, so chi-squared is a weak arbiter on its own terms. No code or data release limits independent verification.\n\nWho is this for: people building or comparing retrieval codes, and anyone using WASP-39 b as a benchmark target. The paper deserves a serious referee; the benchmark portion is strong enough to justify the referee time. Acceptance should be conditional on reconciling the evidence/chi-squared contradiction, propagating offset uncertainties, and releasing the code or detailed parameter files.","headline":"A credible retrieval-framework benchmark undercut by a model-selection choice that contradicts the paper's own Bayesian evidence; the WASP-39 b abundance headline should not be taken at face value.","tokens_in":35078,"tokens_out":3498,"would_cite":false,"duration_ms":37798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid Bayesian-plus-machine-learning retrieval framework reads three JWST instruments at once and finds that WASP-39 b's atmosphere is super-solar in carbon, oxygen, and sulfur.","keywords":["exoplanet atmospheres","transmission spectroscopy","atmospheric retrieval","Bayesian inference","machine learning","WASP-39 b","JWST","equilibrium chemistry"],"falsifier":"Re-run the hybrid equilibrium retrieval on the same combined dataset with the offsets left free or modelled as wavelength-dependent, and check whether the recovered $\\mathrm{O/H}$, $\\mathrm{C/H}$, $\\mathrm{S/H}$, and $\\mathrm{C/O}$ move outside the quoted 1$\\sigma$ intervals; if they do, the super-solar abundances are artifacts of the fixed-offset correction.","tokens_in":2272,"feed_emoji":"🪐","tokens_out":4193,"duration_ms":108722,"temperature":0.7,"pith_summary":"This paper claims that NEXOTRANS, a retrieval framework that couples Bayesian nested sampling with four machine-learning regressors, can extract reliable atmospheric abundances from the combined JWST NIRISS, NIRSpec PRISM, and MIRI transmission spectra of the hot Saturn WASP-39 b. Applied to the full 0.6–12 µm dataset, the framework's best-fit modified hybrid equilibrium model returns super-solar elemental abundances: $\\mathrm{O/H} = 14.12^{+2.86}_{-1.82}$ times solar, $\\mathrm{C/H} = 21.37^{+4.93}_{-3.18}$ times solar, $\\mathrm{S/H} = 5.37^{+0.79}_{-0.65}$ times solar, and $\\mathrm{C/O} = 1.35^{+0.05}_{-0.02}$ times solar. The paper also shows that pure equilibrium chemistry cannot reproduce the observed SO$_2$ feature, that disequilibrium chemistry plus high-altitude ZnS and MgSiO$_3$ aerosols are needed, and that the machine-learning retrievals agree with the Bayesian ones, which would make fast comparative exoplanetology practical in the JWST era.","feed_headline":"WASP-39 b holds 14× solar oxygen and 21× solar carbon","feed_subtitle":"A new hybrid retrieval of three JWST spectra finds a carbon- and oxygen-rich hot Saturn, with C/O at 1.35× solar.","key_machinery":"The load-bearing machinery is the NEXOTRANS pipeline itself. Its forward model converts layer-by-layer opacities into a transmission spectrum using the 1-D path distribution method of Robinson (2017), with absorption cross-sections from the POSEIDON opacity database and opacity contributions from Rayleigh scattering, collision-induced absorption, patchy grey clouds with hazes, and Mie-scattering aerosols (ZnS and MgSiO$_3$). Chemistry is handled by four prescriptions—free chemistry, NEXOCHEM equilibrium (a Gibbs free-energy minimizer benchmarked against FastChem), modified hybrid equilibrium, and modified equilibrium-offset—and the retrieval layer combines nested sampling (PyMultiNest/UltraNest) with a stacking regressor (Random Forest, Gradient Boosting, and k-Nearest Neighbor feeding a Ridge meta-model) trained on 60,000 generated spectra after wavelength-based feature reduction. The paper's central identity is the corrected combined 0.6–12 µm spectrum, which under the hybrid equilibrium model returns the reported super-solar elemental ratios.","core_discovery":"The central discovery is that a hybrid retrieval framework—Bayesian inference run alongside a stacking-regressor machine-learning surrogate—recovers a consistent chemical picture of WASP-39 b from three JWST instruments, and that the statistically preferred modified hybrid equilibrium model yields a super-solar and carbon-enhanced composition. For the best-fit model the paper reports volume mixing ratios for H$_2$O, CO$_2$, CO, H$_2$S, and SO$_2$, with SO$_2$ log VMR between $-6.25$ and $-5.73$ across all non-equilibrium models, and elemental ratios $\\mathrm{O/H} = 14.12^{+2.86}_{-1.82}\\times$ solar, $\\mathrm{C/H} = 21.37^{+4.93}_{-3.18}\\times$ solar, $\\mathrm{S/H} = 5.37^{+0.79}_{-0.65}\\times$ solar, and $\\mathrm{C/O} = 1.35^{+0.05}_{-0.02}\\times$ solar. The equilibrium-only model fails by underpredicting SO$_2$, which the authors interpret as evidence of photochemical disequilibrium. Benchmarked against eight other retrieval codes on MIRI data (Powell et al. 2024), NEXOTRANS returns consistent water and SO$_2$ abundances, and the ML predictions agree with the Bayesian posteriors, so the framework is put forward as a validated tool for JWST-era comparative exoplanetology.","pith_inferences":["If the constant-offset alignment is a fair approximation, the same hybrid Bayesian-plus-ML pipeline could be applied to other JWST targets with multi-instrument data, but the per-instrument chi-square disparity warns that cross-instrument systematics may dominate the error budget in such combined retrievals.","The paper's reported C/O ratio (0.80 absolute, 1.35 times solar) is consistent with a carbon-enriched but not carbon-dominated atmosphere; a direct testable extension would be to search for C$_2$H$_2$ or HCN features shortward of 5 µm, where the model predicts enhanced carbon chemistry at super-solar C/O.","Because the ML posterior is built from ±10% perturbations around the observed transit depths, the ML error bars likely underestimate true parameter uncertainty; a comparison with the Bayesian widths provides a cheap way to calibrate ML confidence intervals for future use.","The offset values (57 ppm for NIRISS, 311.13 ppm for MIRI) differ by an order of magnitude; if MIRI's offset were partly astrophysical (e.g., a different terminator or wavelength-dependent spot contamination), the retrieved S/H and SO$_2$ abundances from MIRI absorption features could be biased, a hypothesis testable by fitting the MIRI data alone with and without the offset."],"forward_implications":["If the hybrid equilibrium result holds, WASP-39 b's bulk elemental enrichment (O/H roughly 14 times solar, C/H roughly 21 times solar) is higher than earlier HST-era estimates, implying a more metal-rich formation environment for hot Saturns.","The failure of equilibrium chemistry to produce the observed SO$_2$ feature strengthens the case that photochemical and other disequilibrium processes are essential ingredients in JWST retrievals of hot giant atmospheres.","The consistency between the stacking-regressor ML retrievals and the Bayesian posteriors suggests that ML surrogates can replace or accelerate nested sampling for large JWST samples, provided their smaller error bars are interpreted carefully.","The per-instrument reduced chi-square values (2.67, 3.19, and 2.14 for NIRISS, PRISM, and MIRI) indicate that a single global model does not fit all three instruments equally well, so joint-instrument retrievals need per-instrument noise or model flexibility.","The constraints on ZnS and MgSiO$_3$ aerosols with modal sizes and terminator coverage fractions provide a path toward linking condensate chemistry to atmospheric dynamics in transmission spectra."],"supporting_citations":[{"why":"Supplies the MIRI observations and the eight-algorithm benchmark set used to validate NEXOTRANS.","marker":"Powell et al. 2024"},{"why":"Provides the 1-D path distribution radiative-transfer method that underlies the forward model's transit-depth computation.","marker":"Robinson 2017"},{"why":"Motivates the hybrid equilibrium and equilibrium-offset chemistry prescriptions and the Mie aerosol treatment that NEXOTRANS adapts.","marker":"Constantinou & Madhusudhan 2024"},{"why":"Provides the NIRSpec PRISM transmission spectrum used in the combined retrievals.","marker":"Rustamkulov et al. 2023"},{"why":"Provides the NIRISS SOSS transmission spectrum used in the combined retrievals.","marker":"Feinstein et al. 2023"},{"why":"Supplies prior equilibrium-chemistry and metallicity constraints that the paper compares against in Section 3.1.2.","marker":"Ahrer et al. 2023"},{"why":"Quantifies the equilibrium SO$_2$ abundance upper limits used to argue that disequilibrium chemistry is required to explain the SO$_2$ feature.","marker":"Tsai et al. 2023"},{"why":"FastChem, the reference code against which the NEXOCHEM equilibrium grids are benchmarked.","marker":"Joachim W. Stock & Sedlmayr 2018"},{"why":"Provides the Gibbs free-energy minimization method that NEXOCHEM implements.","marker":"White et al. 1958"},{"why":"Supplies the transit-depth expression used in the forward model.","marker":"MacDonald & Lewis 2022"}],"fun_headline_variants":["Hybrid retrieval finds C-rich WASP-39 b from JWST spectra","WASP-39 b: 14× solar O, 21× solar C, C/O 1.35×","ML + Bayesian retrieval reveals super-solar WASP-39 b","NEXOTRANS: JWST spectra yield super-solar C and O on WASP-39 b","SO2 hints at photochemistry on WASP-39 b in new retrieval"],"cache_read_input_tokens":36864,"weakest_assumption_plain":"The load-bearing premise is that the NIRISS and MIRI spectra can be aligned to NIRSpec PRISM by two constant multiplicative offsets (57 ppm and 311.13 ppm) derived from an initial free-chemistry retrieval, and that all later retrievals on the corrected data are unaffected by any error in those offsets.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid retrieval finds C-rich WASP-39 b from JWST spectra","WASP-39 b: 14× solar O, 21× solar C, C/O 1.35×","ML + Bayesian retrieval reveals super-solar WASP-39 b","NEXOTRANS: JWST spectra yield super-solar C and O on WASP-39 b","SO2 hints at photochemistry on WASP-39 b in new retrieval"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00097,"raw_usage":{"total_tokens":4310,"prompt_tokens":1316,"completion_tokens":2994,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":932,"completion_tokens_details":{"reasoning_tokens":2878}},"tokens_in":932,"tokens_out":2994,"duration_ms":19014,"temperature":1.0,"reasoning_tokens":2878,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:09:15.676519+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the hybrid equilibrium retrieval on the same combined dataset with the offsets left free or modelled as wavelength-dependent, and check whether the recovered $\\mathrm{O/H}$, $\\mathrm{C/H}$, $\\mathrm{S/H}$, and $\\mathrm{C/O}$ move outside the quoted 1$\\sigma$ intervals; if they do, the super-solar abundances are artifacts of the fixed-offset correction.","supporting_citations":[{"cited_title":"D., Lee, E","cited_arxiv_id":null,"evidence_quote":"Supplies the MIRI observations and the eight-algorithm benchmark set used to validate NEXOTRANS."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 1-D path distribution radiative-transfer method that underlies the forward model's transit-depth computation."},{"cited_title":"B., Johnson, S","cited_arxiv_id":null,"evidence_quote":"Provides the Gibbs free-energy minimization method that NEXOCHEM implements."},{"cited_title":"J., & Lewis, N","cited_arxiv_id":null,"evidence_quote":"Supplies the transit-depth expression used in the forward model."}],"review_version":1}