{"id":"281062bd-020c-45db-91a9-adcaded1e2f1","arxiv_id":"2505.18715","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"New flexible cloud modeling shows that JWST data alone cannot robustly pin down aerosol properties because its instruments lack simultaneous visible-to-mid-infrared coverage.","lead":"This paper presents new, open-source tools for modeling clouds and haze in exoplanet atmosphere retrievals, tested on Titan data and applied to JWST observations of four hot Jupiters. Its central finding is that JWST alone, without simultaneous visible and mid-infrared coverage, struggles to uniquely identify aerosol properties.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-retrieval limitation: the wide-coverage conclusion in §5–6 is demonstrated only inside the compact-sphere Mie model family, while the paper's own porous-particle test (§5.2) shows full coverage still biases recovered aerosols; hence the claim is conditional on aerosol microphysics, not…","rationale":"The reader's weakest assumption identified the same limitation: the aerosol forward model is assumed to span reality, whereas particles are modeled as compact homogeneous spheres with a limited menu of compositions and size distributions. I agree, and I frame it as the load-bearing issue because the paper's central message is fundamentally an information-content claim, and information-content claims are only as strong as the generative model. The authors are transparent about the restriction to compact spheres and about the known fractal nature of Titan's hazes, and they do test one deviation (porosity), which is commendable. However, that single deviation already shows that full wavelength coverage does not guarantee unbiased aerosol characterization: the spherical retrieval fits the porous truth well while biasing key parameters, including metallicity. This does not contradict the necessity claim that both visible scattering and mid-IR resonance features are needed, but it makes the practical conclusion conditional in a way the abstract and conclusion do not fully convey. A concrete out-of-family test with aggregates or more complex shapes would settle whether the wide-coverage prescription is robust. The real-data demonstrations, open-source plugin code, and multiple controlled cases are genuine supporting evidence; the concern is not that the retrievals are internally wrong, but that the headline guideline may be incomplete. Because the reader's verdict already conditions on this caveat, I do not recommend changing the verdict.","tokens_in":33976,"tokens_out":7305,"duration_ms":68720,"concrete_test":"Build an independent synthetic dataset with non-spherical aerosols (e.g., fractal aggregates of roughly 100–3000 monomers with Df ≈ 2, or DDA/T-matrix scattering) for the WASP-107b-like scenario of Section 5 (20x solar equilibrium chemistry, 10 ppm SO2, 10 Pa MgSiO3 layer, µr = 0.5 µm), and add JWST instrument noise from Batalha et al. (2017) for NIRISS-SOSS, NIRSpec-G395H, and MIRI-LRS. Then run the paper's TauREx-PyMieScatt spherical-particle retrievals on (a) NIRISS-only and (b) NIRISS+NIRSpec+MIRI, exactly as in Cases 1a/2. If the full-coverage retrieval (b) does not recover the input gas abundances and aerosol parameters markedly better than (a), the wavelength-coverage conclusion does not transfer to non-spherical aerosols; if it does, the concern is contained. Report posterior medians and biases for Z, log(SO2), µr, and χ in both setups.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 6 concludes that reliably constraining aerosol nature, size distribution, and abundances requires sensitivity to both the visible scattering slope and longer-wavelength resonance features, and that JWST's non-simultaneous coverage is the main obstacle. This practical conclusion is established by the controlled experiments of Section 5, but those experiments generate the truth with the same TauREx-PyMieScatt compact-sphere parameterization used in the retrievals (Appendix A.1, Eqs. A.3–A.4). Case 1a versus Case 2 therefore demonstrates that, within the assumed Mie-sphere, limited-species family, narrow coverage leaves degeneracies that full coverage resolves. It does not test whether the full-coverage remedy survives a more realistic aerosol truth. The one out-of-family test present is Case 1b/FM2 with 50% porous particles: spherical retrievals fit the full-coverage spectrum convincingly yet recover biased particle size, number density, and metallicity (Section 5.2, Figure 4). This is exactly the failure mode that would matter for real JWST data, since Titan's hazes are known fractal aggregates (Section 3; Rannou et al. 2022) and the authors explicitly omit aggregates. Thus the strongest version of the conclusion—that wavelength coverage is the key missing ingredient—has not been tested against aerosol shapes or compositions outside the retrieval family. The Titan benchmark is useful but validates flexibility rather than microphysical accuracy. The necessity claim may still be true, but the evidence for it is an in-family information-content result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a flexible aerosol parameterization implemented as TauREx plugins (TauREx-PyMieScatt, TauREx-MultiModel, TauREx-InstrumentSystematics) and applies it to atmospheric retrievals of Cassini Titan occultation data and JWST observations of HAT-P-18b, WASP-39b, WASP-96b, and WASP-107b. It also reports controlled simulation experiments with synthetic JWST data to study the information content of different wavelength coverages and aerosol model assumptions. The central claim is that robust characterization of cloud and haze properties requires sensitivity to both the visible scattering slope and longer-wavelength resonance features, and that JWST's non-simultaneous coverage from separate instruments makes such characterization difficult.","tokens_in":34350,"tokens_out":6728,"duration_ms":57077,"significance":"The paper makes a useful practical contribution by releasing open-source retrieval plugins, benchmarking a general-purpose retrieval code on an external solar-system dataset (Cassini/VIMS Titan occultations), and applying it to recent JWST data. The multiple independent demonstrations of aerosol degeneracies—identical posteriors for KCl, Na2S, and ZnS with NIRISS, the Na/K versus haze degeneracy for WASP-96b, and the reported NIRSpec/NIRCam spectral incompatibility for WASP-107b—support the qualitative conclusion that aerosol properties are hard to pin down with current JWST observations. The controlled simulations also provide useful parameterization guidance, e.g., weak sensitivity to the vertical aerosol profile and to the full particle size distribution. If the result holds, it is a valuable caution for the field, but the strength of the general conclusion is limited by the fact that the simulations are mostly self-retrievals within the same compact-sphere Mie framework.","major_comments":[{"comment":"The controlled experiments of Section 5 generate the true spectra with the same TauREx-PyMieScatt compact-sphere Mie parameterization used in the retrievals (Appendix A.1, Eqs. A.3–A.4). Case 1a versus Case 2 therefore demonstrates degeneracies and their resolution within that assumed model family, but does not test whether the full-coverage remedy survives a more realistic aerosol truth. The one out-of-family experiment present, Case 1b/FM2 with 50% porous particles, shows in Section 5.2 and Figure 4 that spherical-particle retrievals fit the full-coverage spectrum convincingly while returning biased particle size, number density, and metallicity. Since Titan's hazes are known fractal aggregates (Section 3, with references to Rannou et al. 2022) and aggregates are not included in the simulated truths, the conclusion in Section 6 that observations need both the visible scattering slope and longer-wavelength resonance features is not demonstrated to be sufficient; the paper should explicitly qualify this claim to the assumed spherical Mie family, or add an out-of-family aggregate forward-model test.","section":"Section 5 and Section 6, Appendix A.1"},{"comment":"The Titan benchmark is a genuine external validation and a strength of the paper, but it validates the flexibility of the parameterized retrieval and its ability to recover bulk chemistry, not the microphysical accuracy of the aerosol model. The retrieved haze radius changes by a factor of two depending on the opacity source (µ_tholins ~0.15 µm for HITRAN versus ~0.3 µm for ExoMol), and the authors note that fractal aggregates, which are crucial for Titan, are not considered. The claim in Section 3 that the Titan experiment offers guidance on which atmospheric properties can be reliably retrieved is therefore supported, but the benchmark should not be used as evidence that the adopted compact-sphere aerosol model is an accurate representation of real aerosol microphysics.","section":"Section 3 and Figure B1"},{"comment":"The WASP-107b analysis reports a significant incompatibility between the NIRSpec-G395H and NIRCam-F322W2 spectra, with retrieved inter-instrument offsets up to 250 ppm, and the paper uses this as evidence that combining non-simultaneous JWST datasets is problematic. However, the controlled full-coverage simulations in Section 5 assume idealized, offset-free combination of NIRISS, NIRSpec, and MIRI data. The information-content estimate from those simulations is therefore an upper bound that ignores the systematic combination errors demonstrated on real data. The paper should state this explicitly so that the Section 6 conclusion is not read as applying to real, imperfectly combined datasets without further caveats.","section":"Section 4.2 and Figure C4"}],"minor_comments":[{"comment":"The text lists C2H8 among the molecular species included in the Titan retrievals; this should presumably be C3H8, which is the species listed in Table 1 and discussed in Section 3.","section":"Section 2, Cassini retrievals"},{"comment":"There are several typographical errors: 'by a single instruments' in the abstract, 'think layer' for 'thick layer' in Section 4.1, and 'dis-equilibrium' for 'disequilibrium' in Section 4.2.","section":"Abstract and Section 4.1"},{"comment":"The statement that the WASP-96b data from Taylor et al. (2023) could not be fully compared because the reduced spectra in Radica et al. (2023) could not be found is vague; the authors should specify which data products were unavailable and how this affects the comparison.","section":"Section 4.1"},{"comment":"The phrase 'inverted corner plots' is unclear; the authors likely mean that the FRECKLL posteriors are shown in the upper-right triangle, but this should be stated more explicitly.","section":"Figure C3 caption"},{"comment":"Several URLs are broken across line breaks with inserted spaces (e.g., 'https://github.com/ucl- exopl anets/TauREx3'), which will make them unusable in the published version; the links should be formatted as proper hyperlinks.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a lot of useful material and the practical warning about JWST aerosol retrievals is well supported by multiple real-data experiments. My main concern is that the abstract and conclusion state the wavelength-coverage requirement more generally than the simulations justify, because the controlled experiments are largely self-retrievals within the compact-sphere Mie family and the paper's own porous-particle test shows that even full coverage can leave microphysical biases. This is fixable by qualification and possibly by an additional out-of-family test, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. First, the paper ships real, reusable code: three open-source TauREx plugins (PyMieScatt, MultiModel, InstrumentSystematics) that extend aerosol retrievals from the heuristic L13 cloud to Mie scattering with multiple species, size distributions, and porous-particle options. Second, its main practical conclusion—that JWST's non-simultaneous coverage makes aerosol characterization genuinely hard—is supported by several independent demonstrations: identical posteriors for KCl/Na2S/ZnS on NIRISS, the Na/K-haze degeneracy in WASP-96b, and the NIRSpec/NIRCam offset problem in WASP-107b. The Titan Cassini benchmark is a genuine plus: an external dataset with independent in-situ constraints, and the paper is candid about where it does not reproduce the full picture (fractal aggregates, for instance).\n\nWhat is new is mainly the toolkit and the information-content study, not the qualitative \"wide wavelength coverage helps\" message, which has been around since Lee et al. 2014 and others. The paper's version is more concrete and quantitative, and the open code and Zenodo opacity tables make it verifiable. The FRECKLL-on-JWST claim is a priority claim that should be softened or substantiated, but that is minor.\n\nThe real soft spot is the self-retrieval structure of Section 5. The controlled experiments generate the truth with the same compact-sphere Mie parameterization used in the retrievals, so Cases 1a/2 demonstrate identifiability within the assumed microphysical family, not robustness to aerosol shapes outside it. The one outside-family test the paper includes—the 50% porous particle case—shows spherical retrievals fit the full-coverage spectrum convincingly while returning biased particle size, number density, and metallicity. That is exactly the failure mode that matters for real data, and it means the strongest version of Section 6's conclusion (\"observations need both visible scattering slope and longer-wavelength resonances\") is conditional on the microphysical truth being close to compact spheres. The paper itself flags this in Section 5.2, but the conclusion does not carry the caveat.\n\nThis is not a fatal flaw. The necessity claim is plausible and consistent with earlier work; the Titan benchmark and real-data degeneracies provide independent support. But the information-content evidence is in-family, so the conclusion should be phrased as coverage plus microphysical prior, and the single-realization bias claims deserve either multiple noise draws or an explicit caveat.\n\nWho gets value: retrieval practitioners and JWST planning teams. It deserves a serious referee; I would send it out, with a request to tone down the Section 6 generality and to soften or evidence the FRECKLL priority claim.","headline":"A useful, honest toolkit paper whose main information-content conclusion is solid but slightly overgeneralized: the controlled simulations only test compact-sphere Mie aerosols, and the paper's own porous-particle test shows the bias risk.","tokens_in":34916,"tokens_out":2505,"would_cite":true,"duration_ms":18292,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper contends that cloud and haze properties in exoplanet atmospheres cannot be reliably retrieved from JWST spectra unless observations combine the visible scattering slope with mid-infrared resonance features, and it supports this…","keywords":["exoplanet atmospheres","atmospheric retrieval","clouds and hazes","aerosol parameterization","JWST","Mie scattering","Titan","hot Jupiters"],"falsifier":"Take a WASP-107 b-like simulated spectrum with fractal or porous aerosol particles, add JWST noise for NIRISS+NIRSpec+MIRI as in the paper's Case 1b, and run the spherical-particle retrieval: if the retrieved metallicity and cloud composition match the input, the full-coverage prescription survives; the paper's own FM2 experiment already shows metallicity bias, so a positive result would require additional model freedom such as porosity or aggregate parameters to be included in the retrieval.","tokens_in":33763,"feed_emoji":"☁️","tokens_out":10051,"duration_ms":76277,"temperature":0.7,"pith_summary":"This paper tries to establish that the physical properties of aerosol particles in exoplanet atmospheres—what they are made of, how big they are, and how many there are—cannot be reliably recovered from JWST spectra unless the observations cover both the visible scattering slope and mid-infrared resonance features at the same time. The authors build a flexible, Mie-theory-based aerosol parameterization inside the TauREx retrieval framework, validate it on Cassini occultation data of Titan, and apply it to JWST observations of HAT-P-18 b, WASP-39 b, WASP-96 b, and WASP-107 b. In parallel, controlled retrievals on simulated JWST data show that narrow wavelength coverage produces apparently good fits while retrieving wrong cloud composition, chemistry, and metallicity. The practical consequence is that JWST's single instruments cannot simultaneously provide the needed wavelength range, and stitching together observations from different visits introduces offsets and temporal-variability biases. If the claim holds, interpreting JWST clouds and hazes will require coordinated observations with other facilities or a relaxation of what can be claimed from JWST spectra alone.","feed_headline":"Clouds are unreadable unless JWST sees visible and mid-IR at once","feed_subtitle":"Single-instrument JWST spectra cannot fix cloud type or size, and that bias leaks into chemistry and metallicity.","key_machinery":"The carrier of the argument is TauREx-PyMieScatt, a plugin for the TauREx retrieval framework that computes Mie extinction cross-sections for spherical aerosol particles from their complex refractive index and a chosen particle size distribution (one-parameter gamma from Budaj et al. 2015, log-normal, or modified gamma), and places the particles in pressure-bounded layers with retrievable number density. It is complemented by TauREx-MultiModel, which mixes clear and cloudy atmospheric regions to model partial cloud coverage, and TauREx-InstrumentSystematics, which fits vertical offsets between combined datasets. The mechanism does the work by letting the same retrieval code treat everything from Titan's tholin hazes to hot-Jupiter silicate clouds, and by enabling controlled experiments in which the input aerosol truth is known and the retrieved parameters can be compared across wavelength subsets.","core_discovery":"The paper's central claim is stated in its conclusion: to minimally constrain aerosols, observations need to be sensitive to both the visible light scattering slope and longer wavelength resonance features, such as the 10µm Si-O stretch. Without that combined information, and in the absence of priors on aerosol composition, retrievals find spectra that fit the data but infer incorrect cloud species, abundances, and metallicities; in the simulated cases, metallicity is off by roughly 10σ and SO2 by about 4σ when only NIRISS+NIRSpec or MIRI data are used. JWST has no single instrument covering both regions simultaneously, and combining datasets from different visits is plagued by offset and shape incompatibilities, as seen between the NIRSpec and NIRCam data of WASP-107 b. The paper also finds that JWST is largely insensitive to the vertical aerosol distribution and to the full particle size distribution, so simple one-parameter size distributions suffice, while particle porosity and non-spherical shapes introduce biases that even full wavelength coverage does not fully remove.","pith_inferences":["If hot-Jupiter aerosols are as structurally complex as Titan's fractal hazes, the paper's information-content estimates are likely optimistic: its own porous-particle simulation shows biases persist even with full wavelength coverage, so coverage alone may not recover true sizes or metallicities.","The WASP-96 b degeneracy between alkali line wings and a scattering slope suggests that stellar activity or limb-darkening errors could masquerade as either clouds or enhanced Na/K; a direct test is to fit the same NIRISS spectrum with and without stellar heterogeneities.","The paper's conclusion that JWST is insensitive to vertical aerosol profiles sets a boundary for microphysical cloud models: they can predict observable spectra, but JWST retrievals cannot validate their vertical transport predictions.","If the NIRCam/NIRSpec incompatibility in WASP-107 b is astrophysical rather than instrumental, repeated NIRISS visits separated by days could serve as a direct probe of exoplanet aerosol weather."],"forward_implications":["JWST-only aerosol characterization cannot rely on a single instrument; combinations such as NIRISS plus MIRI are needed, and such combinations must be treated as potentially systematics-limited rather than cleanly constraining.","When no mid-infrared resonance feature is covered, simple phenomenological cloud models such as Lee et al. (2013) are sufficient and give results robust to the assumed cloud species, so complex microphysics is not needed for those datasets.","Detecting and identifying silicate clouds requires the 8–11 µm Si-O feature; the paper's WASP-107 b retrievals consistently favor SiO2, but the full feature shape needs a secondary component such as MgSiO3 or Mg2SiO4 and depends on particle-shape assumptions.","JWST data do not constrain the vertical profile of aerosol abundance or the shape of the particle size distribution, so simplified parameterizations capture the available information and should be used to avoid over-interpretation.","Breaking cloud–chemistry degeneracies will require simultaneous visible-to-mid-infrared coverage, which the paper argues can come from synergies with other observatories rather than from JWST alone."],"supporting_citations":[{"why":"Provides the TauREx3 retrieval framework into which the new aerosol plugins are integrated.","marker":"Al-Refaie et al. 2021"},{"why":"Supplies the PyMieScatt library used to compute Mie extinction efficiencies and cross-sections.","marker":"Sumlin et al. 2018"},{"why":"Defines the phenomenological cloud model L13 used as a generic, featureless aerosol baseline.","marker":"Lee et al. 2013"},{"why":"Provides the one-parameter gamma particle size distribution that defines most aerosol retrievals in this work.","marker":"Budaj et al. 2015"},{"why":"Consolidated Cassini/VIMS occultation spectra of Titan used as the benchmark retrieval example.","marker":"Robinson et al. 2014"},{"why":"Source of the JWST/NIRISS spectra for WASP-39 b and WASP-96 b analyzed in Section 4.","marker":"Holmberg & Madhusudhan 2023"},{"why":"Provides the HST/WFC3, JWST/NIRCam, and JWST/MIRI data for WASP-107 b and the comparison free retrieval.","marker":"Welbanks et al. 2024"},{"why":"Provides the JWST/NIRSpec G395H data for WASP-107 b.","marker":"Sing et al. 2024"},{"why":"Interpretation of the MIRI silicate feature for WASP-107 b, compared against the paper's SiO2/MgSiO3/Mg2SiO4 retrievals.","marker":"Dyrek et al. 2024"},{"why":"Defines the JWST instrument noise models used for the controlled simulated retrievals in Section 5.","marker":"Batalha et al. 2017"}],"fun_headline_variants":["Clouds need visible and mid-IR in one JWST spectrum","No single JWST instrument can read exoplanet clouds","Without visible light, JWST chemistry estimates slip 10 sigma","Combining JWST visits won't cure cloud-chemistry bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simple spherical-particle aerosol model used in the retrievals is close enough to reality that the wavelength-coverage requirements derived from it transfer to real exoplanet aerosols; if real particles are more complex, as Titan's fractal hazes and the paper's own porous-particle simulations suggest, the required coverage and achievable accuracy could differ.","fun_headline_variants_meta":{"raw":{"variants":["Clouds need visible and mid-IR in one JWST spectrum","No single JWST instrument can read exoplanet clouds","Without visible light, JWST chemistry estimates slip 10 sigma","Combining JWST visits won't cure cloud-chemistry bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000365,"raw_usage":{"total_tokens":2045,"prompt_tokens":1104,"completion_tokens":941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":871}},"tokens_in":720,"tokens_out":941,"duration_ms":8237,"temperature":1.0,"reasoning_tokens":871,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:26:48.415439+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a WASP-107 b-like simulated spectrum with fractal or porous aerosol particles, add JWST noise for NIRISS+NIRSpec+MIRI as in the paper's Case 1b, and run the spherical-particle retrieval: if the retrieved metallicity and cloud composition match the input, the full-coverage prescription survives; the paper's own FM2 experiment already shows metallicity bias, so a positive result would require additional model freedom such as porosity or aggregate parameters to be included in the retrieval.","supporting_citations":[{"cited_title":"2015, , 454, 2, 10.1093/mnras/stv1711","cited_arxiv_id":null,"evidence_quote":"Provides the one-parameter gamma particle size distribution that defines most aerosol retrievals in this work."}],"review_version":1}