{"id":"b5b219d9-115d-48a1-9fda-2a8f03e2e0a6","arxiv_id":"2507.11785","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A bidirectional LSTM forecasts simulated kilonova light curves from IGWN alert features with test MSE 0.19 (ZTF) and 0.22 (Rubin), improving to 0.1 when ejecta mass is added.","lead":"This paper trains a machine learning model to forecast kilonova light curves from gravitational wave alert data, reporting low prediction errors on simulated events for the ZTF and Rubin surveys. The tool is meant to help astronomers decide which gravitational wave events are worth following up with telescopes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported MSE is measured on simulations with deliberately restricted ejecta configurations, and real-event validation includes no confirmed kilonova, so the headline accuracy does not yet demonstrate forecasting skill for real kilonovae.","rationale":"The paper's internal test performance is credible: held-out simulated data, public code, and a reproducible training pipeline support the claim that the LSTM emulates the NMMA/POSSIS light curves it was trained on. The central claim about forecasting real kilonovae, however, rests on the representativeness of the training distribution. The custom ejecta setup in Section 2 is a deliberate restriction, and the model has no input feature that would let it adapt to ejecta configurations outside that slice. The O4 candidate comparisons are not a quantitative remedy because none of the candidates is a confirmed kilonova. I therefore identify simulation-to-real transfer, specifically the restricted ejecta parameter space, as the load-bearing concern. The recommended test, regenerating the evaluation set with full alpha/zeta/epsilon ranges, directly quantifies how much of the reported MSE is an artifact of that restriction. If the test passes and MSE remains stable, the headline claim is supported for a broader simulation family; if it fails, the claim should be narrowed. This is the same concern the reader identified, and the resulting CONDITIONAL verdict remains appropriate, so no change to the reader's verdict is proposed.","tokens_in":15119,"tokens_out":6643,"duration_ms":84075,"concrete_test":"Regenerate the NMMA/POSSIS test set for the same O4/O5 injection population with alpha, zeta, and epsilon sampled from the full physically motivated prior instead of fixed values, keeping the train/test split, input features, and architecture unchanged, and recompute the aggregate MSE for ZTF and Rubin filters. If the MSE rises by more than a factor of two relative to 0.19 and 0.22, or if per-event worst-case errors grow substantially, the reported accuracy is an artifact of the restricted ejecta setup and the claim must be narrowed to that subpopulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the model provides useful low-latency kilonova forecasts, supported by test MSE 0.19 for ZTF and 0.22 for Rubin. For that claim to hold, the training and test distributions must cover the light curves the model will actually be asked to predict. This condition is not secured. Section 2 fixes alpha = 0, zeta = 0.3, and epsilon = 0, turning off fallback and wind ejecta and allowing only 30% of the disk mass to be ejected. The POSSIS/NMMA labels therefore span only a restricted slice of the possible kilonova ejecta parameter space. The model inputs are distance, area(90), HasNS, HasRemnant, HasMassGap, and PAstro; none of these encode ejecta mass, velocity, composition, or viewing angle, so the LSTM can only reproduce the conditional average light curve of the restricted training population. An event outside that slice, such as one with significant wind or fallback ejecta, has no reason to be forecast accurately, and the reported MSE is silent on how large that error would be. The Section 4.4 validation against O4 candidates does not close the gap: S230627c is described as possibly a binary black hole with very small ejecta mass, and for 250206dm none of the candidates match the forecasts. No confirmed kilonova is used to quantify simulation-to-real transfer. The abstract's claim that the tool can help plan follow-up is therefore supported only for the restricted simulated subpopulation, not for real kilonovae generally.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a public machine-learning tool that forecasts kilonova light curves in ZTF and Rubin/LSST filters from low-latency IGWN alert features (distance, area(90), HasNS, HasRemnant, HasMassGap, PAstro). A bidirectional LSTM is trained on NMMA/POSSIS simulated light curves from 7,985 O4 BNS/NSBH events for ZTF and 4,223 O5 events for Rubin, reporting test MSE of 0.19 (ZTF, R2=0.82) and 0.22 (Rubin, R2=0.68). The authors also test a CNN-LSTM variant using skymaps, find no improvement, and show that adding ejecta mass as a feature improves MSE to about 0.10-0.11. They compare forecasts to O4a/O4b follow-up candidates, with mixed qualitative agreement. The paper emphasizes the operational goal of helping plan electromagnetic follow-up of gravitational-wave events.","tokens_in":15359,"tokens_out":3074,"duration_ms":37612,"significance":"If the central claim is properly scoped, the tool is a potentially useful, fast surrogate for NMMA/POSSIS light-curve generation from low-latency GW alert parameters, with clear practical value for target-of-opportunity follow-up planning. The paper's strengths include a public GitHub repository with pretrained models, reproducible training scripts, and use of a held-out simulated test set, which makes the internal MSE numbers checkable. The main scientific value is the demonstration that low-latency alert features carry enough information to reproduce the conditional average simulated kilonova light curve, and the analysis of which added features help. However, the applicability to real kilonovae is not established by the present evidence: the training population is deliberately restricted in ejecta configuration, and the O4 validation includes no confirmed kilonova. The reported accuracy is therefore best interpreted as a simulation-to-simulation emulation accuracy for a restricted model family, not as a measured forecast skill for real events.","major_comments":[{"comment":"The training set is generated with alpha = 0, zeta = 0.3, and epsilon = 0, turning off fallback and wind ejecta and allowing only 30% of the disk mass to be ejected. Since the six input features (distance, area(90), HasNS, HasRemnant, HasMassGap, PAstro) do not encode ejecta mass, velocity, composition, or viewing angle, the model can only learn the conditional average light curve of this restricted simulated population. The headline test MSE of 0.19 (ZTF) and 0.22 (Rubin) therefore measures agreement with a restricted simulation family, not expected error for a real kilonova with, for example, significant wind or fallback ejecta. The central claim in the abstract that the tool 'can help to add important information to help plan follow-up' needs to be scoped accordingly, or the training set needs to cover a broader ejecta parameter space.","section":"Section 2, Training Data"},{"comment":"The SNR-to-IFAR mapping log10(IFAR) = 2.357*SNR - 20.198 is a fitted relation from GWTC-3 BNS injections, and it is used together with assumed values of RCBC, the IFAR threshold, and the number of pipelines to synthesize PAstro. No uncertainty in the fitted slope and intercept is propagated, and the relation is applied to NSBH events and to O4/O5 sensitivity without validation. Because PAstro is one of the six input features, a biased mapping will directly bias the forecasted light curves. At minimum, the authors should compare their synthesized PAstro values against actual low-latency alert values for O4 events and quantify the sensitivity of the MSE to the assumed mapping parameters.","section":"Section 2, Eq. (1) and Eqs. (3)-(5)"},{"comment":"The O4 validation does not include a confirmed kilonova. For S230627c the text states the candidate may be a BBH merger with very small ejecta mass, and for 250206dm none of the candidates match the forecasts. The abstract's phrase 'We verify the performance of the model against merger events followed-up by ZTF' overstates what Section 4.4 can show; the section demonstrates only a qualitative comparison with candidate counterparts, not a validation of forecasting skill for real kilonovae. The manuscript should explicitly state this limitation and distinguish simulation-to-simulation accuracy from real-event transfer.","section":"Section 4.4, Comparison to kilonova candidates"},{"comment":"The improvement from adding ejecta mass (MSE 0.10-0.11) is close to tautological in this setup: the NMMA/POSSIS light curves are deterministic functions of the same ejecta properties that the added feature summarizes. The comparison is therefore not a fair test of whether an independent physical feature improves generalization. Moreover, ejecta mass is not currently available in low-latency GW alerts, as the authors acknowledge in the Discussion. The result should be framed as an upper bound on the possible gain from a future low-latency ejecta-mass product, not as a recommended input for the current tool.","section":"Section 4.3, CNN-LSTM and ejecta-mass feature"}],"minor_comments":[{"comment":"The section title contains a typo: 'Comparision' should be 'Comparison'.","section":"Section 4.4 title"},{"comment":"The text says 'We use MC dropout (with value 0.1) to estimate mean and uncertainity of forecasted light curves'; 'uncertainity' should be 'uncertainty', and the description would benefit from stating how many stochastic forward passes were used for the uncertainty estimates.","section":"Section 3, Model"},{"comment":"The first author name appears corrupted as 'Nataly aPletskov a'; this should be corrected in the final manuscript.","section":"Author affiliation line"},{"comment":"The per-filter rows for the Rubin model list g, r, i, u, y, z, which is not in wavelength order; reordering to u, g, r, i, z, y would make the trend in MSE easier to read.","section":"Table 3"},{"comment":"The figure caption mentions a black line for FAR = 22, but the main text does not explain the origin or rationale of this specific FAR threshold; a sentence of context would help.","section":"Figure 12"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful engineering contribution with public code, and the internal test-MSE numbers appear internally consistent for the restricted simulated population. The main risk is that the abstract and Discussion generalize beyond what the evidence supports, particularly regarding real-event forecasting skill. The restricted ejecta setup and the absence of a confirmed kilonova in the O4 validation are not fatal to the paper if the claims are scoped, but they need to be addressed prominently. I would encourage the editor to request a revised version that either broadens the training ejecta configurations or explicitly redefines the central claim as emulation of NMMA/POSSIS light curves for a restricted population."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a legitimate engineering contribution: the code and data are public, the MSE numbers are internally consistent on held-out simulations, and the authors are transparent about what they did and what did not work (the skymap null result, the restricted ejecta setup). Second, the abstract's claim that the tool \"predicts kilonova light curves\" is stronger than what is actually shown. The model is an emulator of one restricted family of NMMA/POSSIS simulations, and the reported accuracy is measured only against held-out light curves from that same family.\n\nWhat is genuinely new: mapping IGWN low-latency alert features (distance, area90, HasNS, HasRemnant, HasMassGap, PAstro) to multi-filter light curves for ZTF and Rubin is not in the prior ML-for-kilonova literature, which mostly classifies or infers parameters from observed light curves. The public GitHub and Zenodo data make the work reproducible, and the comparison with O4 candidates, while qualitative, is a good-faith attempt at real-event validation.\n\nThe soft spots are real but not fatal. Section 2 fixes alpha=0, zeta=0.3, epsilon=0, excluding fallback and wind ejecta and allowing only 30% of the disk mass to be ejected. The input features do not encode ejecta mass, velocity, composition, or viewing angle, so the LSTM can only reproduce the conditional average of this restricted population. An event with significant wind or fallback ejecta would be outside the training domain, and the reported MSE says nothing about that error. The SNR-to-IFAR mapping (Eq. 1) is a fitted relation with no propagated uncertainty, and PAstro synthesized from it inherits that. The real-event validation is weak: S230627c may be a BBH with very small ejecta mass, and for 250206dm none of the candidates matched. The ejecta-mass feature experiment improves MSE to 0.1, but ejecta mass is not available at low latency, and since NMMA light curves are deterministic functions of ejecta mass, that experiment is close to a tautology.\n\nNone of this invalidates the emulator result. The paper should be reframed as forecasting within the simulated parameter space covered by the training set, not general kilonova prediction. Readers building Rubin ToO tools or developing ML emulators for GW follow-up will find it useful; the discussion of model failure modes (e.g., underestimating brightness for rare bright events) is thoughtful.\n\nMy recommendation: send this to a serious referee. It deserves referee time, with the expectation of a revised abstract and a clear domain-of-validity statement, quantitative uncertainty on the SNR-IFAR mapping, and ideally an out-of-distribution test on a broader ejecta setup. As it stands, it is a solid, honestly-reported contribution with an overstated headline.","headline":"A reproducible emulator of one restricted family of simulated kilonova light curves, with honest reporting and a headline claim that runs ahead of the demonstrated domain.","tokens_in":16087,"tokens_out":2275,"would_cite":false,"duration_ms":29903,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that a bidirectional LSTM trained on simulated gravitational-wave alerts can forecast kilonova light curves in ZTF and Rubin filters, with a test MSE of 0.19 (ZTF) and 0.22 (Rubin), and that adding ejecta mass as an input…","keywords":["kilonova","gravitational waves","light curve forecasting","LSTM","neural networks","ZTF","Rubin Observatory","multi-messenger astronomy"],"falsifier":"Feed the low-latency alert parameters of GW170817 (or a future well-localized BNS with a detected kilonova) into the public model and compare its forecast to the observed AT2017gfo light curves in the same filters; a large mismatch would show that the model tracks the NMMA/POSSIS training distribution rather than real kilonova behavior.","tokens_in":14823,"feed_emoji":"🔭","tokens_out":4974,"duration_ms":51874,"temperature":0.7,"pith_summary":"This paper presents a public machine-learning tool that forecasts kilonova light curves from the low-latency alert data that gravitational-wave observatories issue minutes after a merger candidate is found. The authors train a bidirectional LSTM on simulated binary neutron star and neutron star-black hole mergers, using only alert-derived features such as distance, sky-localization area, and the probabilities HasNS, HasRemnant, HasMassGap, and PAstro, and show it predicts simulated light curves across ZTF g/r/i and Rubin u/g/r/i/y/z bands with test MSE 0.19 and 0.22 respectively. They argue this is accurate enough to help plan electromagnetic follow-up by estimating whether a kilonova will be visible, how bright it will peak, and how fast it will fade. They also find that adding ejecta mass as an input lowers the MSE to about 0.1, while including full skymap images does not help. A sympathetic reader would take the core claim to be that low-latency GW alert parameters contain enough information to support useful kilonova brightness forecasts before any detailed inference is available.","feed_headline":"Neural net forecasts kilonova light curves from GW alerts","feed_subtitle":"Public LSTM tool turns low-latency alert data into predicted brightness across ZTF and Rubin filters to plan follow-up.","key_machinery":"The mechanism is a bidirectional LSTM that maps a vector of six low-latency alert features — distance, 90% sky-localization area, HasNS, HasRemnant, HasMassGap, and PAstro — to a 90-unit output representing 30 time steps across three ZTF filters, or a 180-unit output for six Rubin filters. The training labels come from the NMMA framework with POSSIS radiative-transfer models, which convert binary masses, tidal deformability, and spins into ejecta masses and multiband light curves for each simulated merger. A hybrid CNN-LSTM variant additionally encodes Bayestar skymaps as 500x1000 three-channel images, while an augmented model adds ejecta mass as a scalar feature, and the feature-importance analysis shows that this single physically motivated input carries most of the predictive gain.","core_discovery":"The central claim is that a bidirectional long-short-term memory network, trained exclusively on simulated IGWN alert features, can serve as a low-latency emulator of kilonova light curves produced by the NMMA/POSSIS framework, and that this emulator is accurate enough to inform follow-up decisions. The model achieves a test mean squared error of 0.19 across ZTF g, r, and i filters and 0.22 across six Rubin filters, with more than half of the test light curves having individual MSE below 0.2. The authors verify the model against candidates followed up during O4a and O4b, finding reasonable agreement for S230627c and no match for 250206dm. They further show that a hybrid CNN-LSTM model using full skymap images does not improve over the LSTM-only model, whereas adding the physically motivated ejecta-mass feature improves the MSE to 0.1 on ZTF and 0.10 on Rubin filters, identifying a clear path for future alert products.","pith_inferences":["Because the MSE is measured against NMMA/POSSIS simulations, the reported numbers bound how well the model reproduces those simulations, not how well it predicts a real kilonova; the tool's value on the sky depends on how faithful the simulations are to real events.","The same architecture could be retrained on other synthetic light-curve libraries or on real kilonovae as they accumulate, and the feature-importance result suggests that any future alert product carrying an ejecta-mass estimate would yield the largest immediate accuracy gain.","A natural extension is to use the forecast not as a point prediction but as a prior for a matching stage that cross-correlates forecast light curves against alert-stream candidates, an idea the authors gesture toward with their planned integration into platforms like SkyPortal.","The model's tendency to under-predict brightness for faint, unusual light curves, if it persists on real data, would bias follow-up toward deeper searches, which is the safer direction for discovery."],"forward_implications":["Observers can run the public model on an incoming IGWN alert and receive a first estimate of kilonova brightness and evolution across ZTF or Rubin filters within seconds of the alert.","Follow-up planning can use the forecasts to decide whether an event is worth chasing and which filters and exposure times to use, including for high-FAR events that current selection cuts discard.","For Rubin/LSST, the model provides a preview in six bands, including near-infrared y and z where the simulated KNe are best predicted, which can inform Target of Opportunity program design.","If the LVK begins releasing coarse ejecta-mass estimates in public alerts, the model's expected error should drop to roughly half of its current level, based on the demonstrated improvement to MSE 0.1.","The skymap CNN experiment suggests that full sky-localization images add little beyond the scalar area(90), so simpler alert features suffice for light-curve forecasting."],"supporting_citations":[{"why":"Supplies the simulated O4/O5 IGWN alert dataset (17,009 BNS and 3,148 NSBH events) used for training and testing.","marker":"Weizmann Kiendrebeogo et al. 2023"},{"why":"The NMMA framework used to generate kilonova light curves for every simulated merger.","marker":"Pang et al. 2023"},{"why":"The POSSIS radiative-transfer models inside NMMA that translate ejecta properties into multiband light curves.","marker":"Bulla 2019"},{"why":"Provides the ejecta-mass predictions that set up the paper's disk-ejecta configuration and the feature whose addition lowers MSE to roughly 0.1.","marker":"Toivonen et al. 2024"},{"why":"Bayestar, used to generate the skymaps and extract the 90% localization area (area(90)) input feature.","marker":"Singer & Price 2016"},{"why":"Defines the EM-bright probabilities HasNS, HasRemnant, and HasMassGap used as input features.","marker":"Chatterjee et al. 2020"},{"why":"Supplies the assumed astrophysical merger rate and noise model used to derive PAstro and the SNR-to-IFAR mapping.","marker":"Abbott et al. 2023"},{"why":"Documents the S230627c follow-up candidates used to verify the model against real O4a observations.","marker":"Ahumada et al. 2024"}],"fun_headline_variants":["LSTM emulates kilonova light curves for ZTF and Rubin","Neural net forecasts kilonova brightness from GW alerts","Public tool predicts kilonova light curves for follow-up","Ejecta mass improves LSTM kilonova light-curve forecasts","Kilonova light-curve forecasts aid Rubin and ZTF searches"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training labels are simulated light curves from the NMMA/POSSIS framework, so the reported accuracy measures agreement with those simulations; if the simulations misrepresent real kilonova brightness or evolution, the forecast error on real events will be larger than the reported MSE.","fun_headline_variants_meta":{"raw":{"variants":["LSTM emulates kilonova light curves for ZTF and Rubin","Neural net forecasts kilonova brightness from GW alerts","Public tool predicts kilonova light curves for follow-up","Ejecta mass improves LSTM kilonova light-curve forecasts","Kilonova light-curve forecasts aid Rubin and ZTF searches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000664,"raw_usage":{"total_tokens":3082,"prompt_tokens":1048,"completion_tokens":2034,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":1945}},"tokens_in":664,"tokens_out":2034,"duration_ms":17469,"temperature":1.0,"reasoning_tokens":1945,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:01:55.599251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed the low-latency alert parameters of GW170817 (or a future well-localized BNS with a detected kilonova) into the public model and compare its forecast to the observed AT2017gfo light curves in the same filters; a large mismatch would show that the model tracks the NMMA/POSSIS training distribution rather than real kilonova behavior.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The NMMA framework used to generate kilonova light curves for every simulated merger."},{"cited_title":"R., et al","cited_arxiv_id":null,"evidence_quote":"Defines the EM-bright probabilities HasNS, HasRemnant, and HasMassGap used as input features."}],"review_version":1}