{"id":"a7520835-5620-4feb-ac18-43ed988dbb09","arxiv_id":"2505.16876","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Kilonova ejecta morphology is distinguishable only when late-time JWST mid-infrared data are added to Rubin optical data, and AT2017gfo is best matched by the SuperNu TP2 (toroidal plus peanut wind) model.","lead":"This paper adds a new grid of neutron star merger light-curve simulations (SuperNu) to the NMMA inference pipeline and tests whether upcoming telescopes can tell apart four ejecta shapes. It finds that optical-only Rubin observations cannot, but adding late-time JWST mid-infrared points can, and it argues the 2017 kilonova AT2017gfo best matches a toroidal-plus-peanut morphology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 3.2 observing-strategy forecasts feed noiseless emulator-generated light curves into the same emulator-based inference, so the reported Bayes factors (2.4 to 7.4) may not survive realistic photometric noise or emulator error.","rationale":"The reader's weakest assumption identifies precisely the concern I consider most load-bearing: the observing-strategy forecasts in Section 3.2 use noise-free mock light curves drawn from the same emulators used for recovery, so the reported Bayes factors measure internal consistency rather than realistic discriminative power. This matters because the paper's headline claim is an observational-strategy recommendation: Rubin alone cannot determine morphology, while Rubin plus JWST MIR can. If noise or emulator bias weakens those Bayes factors, the recommendation loses its quantitative basis. The paper does provide useful validation that the four emulator classes are distinct under abundant, noise-free data (Figure 4), and the chi-squared validation against off-grid SuperNu light curves at sigma = 0.4 mag is a real check, but neither substitutes for noise-inclusive forecasting. The AT2017gfo conclusion has its own soft spots, notably the absence of reported Bayes factors for TP2 versus TS2 and the reliance on a disk-mass plausibility argument, but the observing-strategy claim is the more central and novel assertion, and it is the one directly threatened by the noiseless-mock issue. Since the reader already conditioned acceptance on addressing this and related uncertainties, my assessment does not change the verdict; it reinforces the condition. A single concrete test, injecting realistic photometric noise into the mocks and recomputing Bayes factors, would settle whether the JWST-necessity claim survives contact with real measurement scatter.","tokens_in":17192,"tokens_out":4007,"duration_ms":33883,"concrete_test":"Re-run the Section 3.2 strategy forecasts with the same 100 mock light curves but add per-point Gaussian photometric noise with standard deviations matching Rubin LSST filter and JWST MIRI exposure-time-calculator uncertainties, and optionally include the 0.4 mag emulator error as an additional covariance term in the likelihood; then recompute the Bayes factors for TP2 versus TP1, TS1, and TS2 using identical priors and sampler settings. If B(TP2|TS2) drops below 1, or if the Bayes factors no longer all favor TP2, the claimed ability to discriminate morphologies with the proposed Rubin+JWST strategy fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central forecast claim rests on mock observations produced directly by the trained emulators and analyzed with the same emulator family, with no described photometric noise added to the simulated Rubin or JWST data points (Section 3.2, Figures 5, 7, 8). This makes the exercise a closed-loop self-consistency check rather than a test against realistic measurements: with noise-free data and the true model inside the comparison set, likelihoods are artificially sharp and model-evidence differences are inflated. Real photometric uncertainties from Rubin and JWST MIRI will broaden each model's likelihood, and the reported Bayes factors are already modest, especially B = 2.4 for TP2 versus TS2, a pair whose spectra the paper itself notes are similar (Section 3.2, Figure 6). The emulator validation in Section 3.1 quotes a surrogate error of sigma = 0.4 mag but that uncertainty is never propagated into the mock-data likelihood, and the emulators' poor recovery of dynamical-ejecta parameters (Figure 2) means emulator bias could further shift the evidence ratios. If realistic noise or unmodeled systematic error shrinks the Bayes factors below the discrimination threshold, the conclusion that JWST MIR observations are required to distinguish morphologies would not be supported by these simulations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper incorporates a new grid of two-component SuperNu kilonova light curves, covering four axisymmetric ejecta morphologies (TP1, TP2, TS1, TS2), into the NMMA Bayesian inference pipeline. It validates neural-network emulators of this grid, uses them to simulate follow-up observations with Rubin and JWST, and evaluates whether ejecta morphology can be recovered. The main claims are that Rubin-only follow-up in the proposed ToO cadence cannot determine morphology, that adding late-time JWST MIRI photometry enables morphology discrimination (Bayes factors 2.4–7.4), and that AT2017gfo is best described by the TP2 morphology, with a slight preference over POSSIS-based models.","tokens_in":17340,"tokens_out":4900,"duration_ms":42377,"significance":"If the central forecast claim survives scrutiny, the paper would provide a concrete, observationally actionable conclusion: kilonova morphology determination requires late-time JWST MIR coverage, while Rubin-only optical follow-up is insufficient. The SuperNu grid and its integration into NMMA are a useful addition for kilonova parameter estimation, and the AT2017gfo comparison with POSSIS is valuable. The paper's strengths include the four-morphology simulation grid, the off-grid validation attempt, and the use of real AT2017gfo photometry. However, the forecast claim currently rests on closed-loop simulations with no described photometric noise and with unpropagated emulator error, and the AT2017gfo morphology preference is not supported by formal model comparison; the stated significance is therefore conditional on substantial revision.","major_comments":[{"comment":"The simulated Rubin and JWST data used to compute the reported Bayes factors are generated directly from the trained emulators and analyzed with the same emulator family, with no description of added photometric noise. In this closed-loop setting the likelihoods are artificially sharp, and the evidence differences between morphologies are inflated. This is load-bearing for the abstract claim that morphology discrimination is possible only with JWST MIR data. The authors should repeat the forecast with realistic photometric uncertainties for Rubin and JWST and propagate the emulator error quoted as sigma = 0.4 mag in Section 3.1 into the likelihood; the resulting Bayes factors should be reported with and without this noise.","section":"Section 3.2, Figures 5–8"},{"comment":"The emulator validation shows weak recovery of the dynamical ejecta parameters from noiseless training data (Figure 2), and Table 2 reports reduced chi-squared values as high as 12.6 (entries iv, vii, xi) even with the JWST filters included. The statement that 'most of the light curves result in a reasonable reduced chi-squared' is not consistent with the table, where five entries exceed 6.0. Because the emulator bias is evidently not negligible, it must be quantified and propagated into the Section 3.2 forecasts before the Bayes factors can be taken as representative of real observations.","section":"Section 3.1, Figure 2, Table 2"},{"comment":"There is an internal inconsistency in the idealized 10-day Rubin-only analysis. The text states that morphological differences can be discerned and reports Bayes factors of B = 33 and 48 favoring TP2 over TS1 and TP1, while the Figure 7 caption states 'the morphology cannot be distinguished.' These statements cannot both be correct. The authors should clarify what exactly is and is not discriminated: the analysis apparently separates the wind type (TP1/TS1 versus TP2/TS2) but not TP2 from TS2 (B = 1.4). This also bears on the abstract's claim that discrimination is possible 'only when late-time JWST observations in MIR are available.'","section":"Section 3.2, Figure 7, text"},{"comment":"The AT2017gfo conclusion that the event follows the TP2 morphology is based on excluding TP1 and TS1 because their posteriors 'rail' at prior boundaries, and on the physical plausibility of the recovered wind mass and velocity for TP2 relative to TS2 and the POSSIS model. No evidence ratios or Bayes factors are reported for the four models against the real AT2017gfo data, nor for SuperNu-TP2 against POSSIS. Since the paper's stated preference for TP2 and for SuperNu over POSSIS is a central claim, the authors should provide formal model comparison (e.g., nested-sampling evidences or leave-one-out cross-validation) for the real data.","section":"Section 3.3, Figures 9–10, Table 3"}],"minor_comments":[{"comment":"The phrase 'just mot as accurately' should read 'just not as accurately.'","section":"Section 3.1"},{"comment":"The caption states that three chi-squared values are shown, but the table only lists two columns (chi2_nu and chi2_jwst); the column for the model without the PS1 g-band is missing or the caption should be corrected.","section":"Table 2 caption"},{"comment":"The description of the simulated observations should state explicitly whether photometric noise was added and, if so, what error model was used; currently the text says only that light curves were 'generated from our emulators' and used as data.","section":"Section 3.2"},{"comment":"The text says an ideally trained model would have recovered parameter distributions centered on the 'diagonal,' but the diagonal is only defined visually in the figure; a formal definition of the recovery metric would improve reproducibility.","section":"Section 3.1"},{"comment":"The claim that 'each morphological model is distinct' is based on recovery with the TP2 emulator only; a cross-recovery matrix showing fits with each of the four emulators would more directly support this statement.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about closed-loop simulations is well founded and should be addressed before publication. The paper's abstract overstates the current support for the JWST-only discrimination claim, given the noiseless mock data and the unpropagated emulator error. The AT2017gfo TP2 preference is interesting but needs formal evidence comparison. The internal inconsistency about the 10-day Rubin-only run should be resolved in revision. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful integration of the SuperNu morphology grid into NMMA, and the AT2017gfo comparison with POSSIS is a genuine cross-check. But the strategy forecasts in Section 3.2 generate mock light curves from the trained emulators and do not add photometric noise, so the reported Bayes factors are self-consistency numbers, not realistic predictions. The qualitative message—Rubin-only optical follow-up won't classify morphology; late-time JWST MIR can—is plausible, but the quantitative support is weaker than it looks.\n\nWhat's new: first time the four SuperNu two-component morphologies are in NMMA with emulators; the validation is honest and shows the emulators struggle with dynamical ejecta parameters, which is a real limitation they don't hide. The forecast that JWST MIR data at 8-20 days can break the degeneracies is a useful planning input for the community. The SuperNu-vs-POSSIS analysis on real data is not circular and gives a reasonable result: both fit AT2017gfo photometry, with a slight preference for SuperNu based on wind mass/velocity being more consistent with disk simulations.\n\nSoft spots: the central forecast is a closed-loop test. The same emulators generate the mock data and do the inference, with no described photometric noise and no propagation of the sigma=0.4 mag surrogate error into the likelihood. That will artificially sharpen likelihoods and inflate Bayes factors. The reported values are already modest—2.4 for TP2 vs TS2 in the JWST case—so realistic noise could drop them below any discrimination threshold. The paper would be substantially strengthened by adding noise-inclusive simulations or at least a sensitivity check. Also, the AT2017gfo preference for TP2 over TS2 is not strong evidence; TP1/TS1 are discounted for prior railing, and TP2 wins over TS2 on a physical plausibility argument about disk mass, not on the data alone. The authors are reasonably careful about this, but the abstract's phrasing overstates the certainty.\n\nWho this is for: observers planning kilonova follow-up, and anyone doing NMMA-based parameter estimation. It deserves a serious referee. I'd send it to review with a request to address the noise-free mock-data issue, either by adding realistic uncertainties or by explicitly caveating the Bayes factors.","headline":"Useful NMMA integration and a plausible Rubin-vs-JWST forecast, but the strategy Bayes factors come from a closed emulator loop and should be treated as upper bounds, not predictions.","tokens_in":18025,"tokens_out":3795,"would_cite":true,"duration_ms":29006,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Kilonova ejecta morphology can be told apart only when JWST mid-infrared data are added to Rubin follow-up, and AT2017gfo is best matched by a toroidal ejecta with a peanut-shaped wind.","keywords":["kilonova","neutron star merger","ejecta morphology","Bayesian inference","radiative transfer","SuperNu","AT2017gfo","JWST"],"falsifier":"Re-run the Rubin+JWST injection study adding realistic photometric noise and propagating the emulator's 0.4-magnitude error into the likelihood; if the Bayes factors no longer favor the injected morphology over the others, the claim that MIR observations enable morphology discrimination fails.","tokens_in":16874,"feed_emoji":"🔭","tokens_out":9285,"duration_ms":65838,"temperature":0.7,"pith_summary":"The paper asks whether photometric follow-up of a neutron-star-merger kilonova can reveal the geometry of the ejected material, a question that matters because ejecta geometry encodes which channels produced the r-process elements. The authors add a new grid of 12,150 SuperNu radiative-transfer light curves, organized into four two-component ejecta morphologies, to a Bayesian inference pipeline. They conclude that the planned Vera Rubin Observatory cadence alone cannot distinguish these morphologies, but that supplementing it with three late-time JWST mid-infrared observations makes the distinction possible. Applied to the one well-observed kilonova, AT2017gfo, the analysis favors a toroidal dynamical ejecta with a peanut-shaped wind.","feed_headline":"Rubin alone can't map kilonova shape; JWST can","feed_subtitle":"Three late-time JWST mid-infrared points separate four ejecta geometries; AT2017gfo favors a peanut-shaped wind.","key_machinery":"The load-bearing machinery is a grid of 12,150 SuperNu Monte Carlo radiative-transfer light curves, organized into four axisymmetric two-component morphologies: a common toroidal lanthanide-rich 'dynamical ejecta' component, combined with a wind that is either spherical (TS1, TS2) or peanut-shaped (TP1, TP2) and has electron fraction 0.37 or 0.27. These light curves are compressed into neural-network emulators inside the NMMA Bayesian inference pipeline, which computes evidences and Bayes factors between models. The argument works by comparing parameter-recovery error distributions and Bayes factors for simulated follow-up campaigns, and by applying the same emulators to the AT2017gfo photometry.","core_discovery":"The paper's central claim is that kilonova ejecta morphology can be read from photometry only when late-time mid-infrared data are included. In injection tests, the four SuperNu morphology classes (TP1, TP2, TS1, TS2) are distinct when the data are dense, but under the proposed Rubin ToO strategy of six visits over four nights the recovered parameters overlap completely, so no morphology can be preferred. Adding three JWST MIRI observations at 8, 15, and 20 days in the f560w and f770w filters restores discrimination, with Bayes factors of 7.4 (TP2 over TP1), 4.1 (over TS1), and 2.4 (over TS2). For AT2017gfo, the TP2 model—toroidal lanthanide-rich dynamical ejecta plus a peanut-shaped, lanthanide-free wind with electron fraction 0.27—beats the spherical-wind TS2 model and the POSSIS Bu2019 model, because its recovered wind mass implies an accretion disk of 0.06–0.09 solar masses, consistent with numerical relativity expectations, while the POSSIS fit implies an implausibly large disk.","pith_inferences":["If a few future kilonovae are followed with MIR, the inferred morphologies could map the diversity of r-process ejection geometries and test whether GW170817's torus-plus-peanut structure is typical or rare.","The Bayes factors reported here are modest; combining several events in a hierarchical analysis would sharpen population-level morphology inferences without relying on any single event.","The same methodology could be applied to early spectra or polarimetry to test whether morphology can be recovered before eight days, which the paper's photometry-only results suggest is unlikely.","A sharp test of the TP2 assignment: the recovered inclination angle and wind geometry should be consistent with the gravitational-wave and gamma-ray-burst afterglow constraints for GW170817; a future joint analysis would confirm or overturn it."],"forward_implications":["The planned Rubin ToO cadence of four nights is insufficient for morphology; extending optical coverage beyond one week can separate wind type (spherical vs peanut) but not full geometry.","A campaign that adds three JWST MIRI epochs at 8, 15, and 20 days can classify kilonova ejecta morphology with Bayes factors between 2.4 and 7.4.","AT2017gfo's ejecta is best described as a toroidal lanthanide-rich component plus a peanut-shaped wind, making it a concrete template for interpreting future events.","SuperNu-based fits that return wind masses implying disk masses near 0.06–0.09 solar masses are physically preferred over POSSIS fits implying 0.2–0.3 solar masses."],"supporting_citations":[{"why":"Supplies the SuperNu simulation grid with the two-component composition and electron-fraction variants used throughout.","marker":"Wollaeger et al. 2021"},{"why":"Defines the four Cassini-oval axisymmetric morphologies (toroidal dynamical ejecta with spherical or peanut wind).","marker":"Korobkin et al. 2021"},{"why":"Provides the NMMA Bayesian inference framework and neural-network emulator approach the study builds on.","marker":"Pang et al. 2022"},{"why":"Supplies the POSSIS model grid used as the comparison for the AT2017gfo analysis.","marker":"Bulla 2019"},{"why":"Defines the Vera Rubin Observatory ToO observing strategy that the paper tests and finds insufficient.","marker":"Andreoni & Margutti 2024"},{"why":"Defines the JWST MIRI f560w and f770w filters and the observation plan that enables morphology discrimination.","marker":"Dicken et al. 2024"},{"why":"Provides the inclination-angle constraint for GW170817 used to compare the AT2017gfo fits.","marker":"Hotokezaka et al. 2018"},{"why":"Gives the numerical-simulation expectation for disk wind mass that motivates rejecting the TS2 and POSSIS fits.","marker":"Fernández et al. 2015"}],"fun_headline_variants":["Rubin alone can't map kilonova shape; JWST MIR can","Late JWST MIR points tease apart kilonova ejecta shapes","AT2017gfo hints peanut-shaped wind over spherical","Kilonova morphology hidden from Rubin, shown by JWST MIR","Four ejecta shapes, JWST MIR tells them apart, Rubin can't"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecasting analysis treats noise-free emulator light curves as if they were real Rubin and JWST measurements, without folding in the emulators' known errors (about 0.4 mag) or their difficulty recovering dynamical-ejecta parameters; if real noise or emulator bias is larger than assumed, the reported Bayes factors (2.4 to 7.4) could fall below the discrimination threshold.","fun_headline_variants_meta":{"raw":{"variants":["Rubin alone can't map kilonova shape; JWST MIR can","Late JWST MIR points tease apart kilonova ejecta shapes","AT2017gfo hints peanut-shaped wind over spherical","Kilonova morphology hidden from Rubin, shown by JWST MIR","Four ejecta shapes, JWST MIR tells them apart, Rubin can't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1755,"prompt_tokens":998,"completion_tokens":757,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":660}},"tokens_in":614,"tokens_out":757,"duration_ms":6624,"temperature":1.0,"reasoning_tokens":660,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:53:01.225218+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Rubin+JWST injection study adding realistic photometric noise and propagating the emulator's 0.4-magnitude error into the likelihood; if the Bayes factors no longer favor the injected morphology over the others, the claim that MIR observations enable morphology discrimination fails.","supporting_citations":[{"cited_title":"2024, Rubin ToO 2024: Envisioning the Vera C","cited_arxiv_id":null,"evidence_quote":"Defines the Vera Rubin Observatory ToO observing strategy that the paper tests and finds insufficient."}],"review_version":1}