{"id":"da168789-89dc-4133-b0df-74a046c43906","arxiv_id":"2502.10021","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Analytic two-component fits to kilonova light curves systematically misassign physical ejecta components because post-merger ejecta powers both blue and red emission through reprocessing, though total ejecta mass estimates remain roughly robust.","lead":"This paper fits analytic kilonova models to simulated light curves with known ejecta properties and shows that the fitted blue and red component parameters do not match the actual dynamical and post-merger ejecta. The result warns observers against interpreting two-component analytic fits as physical ejecta components, while showing total ejecta mass remains recoverable to within a factor of a few.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground truth is the authors' own LTE radiative transfer with a single line list; if that opacity treatment is unrepresentative, the inferred red/blue hierarchy and reprocessing picture could be simulation-specific.","rationale":"The reader's weakest assumption identifies the reliance on the authors' own simulation suite as the main vulnerability; I agree that this is the core issue, but I would sharpen it to the radiative-transfer opacity treatment rather than the conventional dynamical/post-merger split. The split is admittedly conventional (Table 1), but the post-merger mass-variation experiment (DD2-135pm03/pm01, Section 3.2) shows that the inferred red component responds to changes in the post-merger ejecta regardless of exactly where the boundary is drawn, so the arbitrary split is not the most load-bearing point. The load-bearing point is the wavelength- and time-dependent opacity set by the LTE assumption and the Domoto et al. (2022) line list. That opacity determines whether blue photons from post-merger ejecta are absorbed and reprocessed to red by lanthanide-rich dynamical ejecta, which is the mechanism the paper claims analytic models misattribute. The paper contains no cross-check with an independent radiation transport code or an independent line list, so the demonstrated mismatch could be an artifact of the simulation's spectral calibration rather than a generic property of analytic models. The GW170817 fit in Appendix B is suggestive but does not break this circularity, because it uses the same analytic model and is interpreted with the same simulation-based expectations. I do not think this invalidates the conclusion; rather, it justifies the reader's conditional verdict. The proposed test—repeating the experiment with an independent radiative transfer code—is computationally expensive but well-defined and would settle whether the reprocessing picture is robust. Therefore I recommend no change to the reader's CONDITIONAL verdict.","tokens_in":22203,"tokens_out":10139,"duration_ms":96891,"concrete_test":"Repeat the controlled fitting experiment with mock light curves for the same (or closely matched) NR ejecta profiles computed by an independent radiative transfer code that does not assume LTE and/or uses a different line list, e.g., the code used by Shingles et al. (2023) or Just et al. (2023). Run the same MOSFiT-style two-component MCMC fits on those light curves. If the inferred blue/red masses and velocities still satisfy M_red > M_blue and v_red < v_blue, and reducing the post-merger mass still shifts both components, the concern is resolved. If the hierarchy flips or the total-mass recovery changes significantly, the headline claim is specific to the authors' simulation setup rather than a general limitation of analytic models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the mock light curves used as ground truth faithfully represent real kilonova emission. The radiative transfer code assumes LTE and uses the Domoto et al. (2022) line list (Section 2.1); these choices set the wavelength-dependent opacity that controls how much post-merger blue flux is absorbed and reprocessed to red by lanthanide-rich dynamical ejecta. The demonstration in Section 3.2 and the schematic in Figure 5 therefore rest entirely on this opacity treatment. Because the NR ejecta profiles and the radiative transfer code come from the same group, the paper contains no independent check that the simulated colors and light curve shapes are representative. If the line list overestimates NIR bound-bound opacity, or if LTE fails at kilonova densities, the inferred hierarchy M_red > M_blue and v_red < v_blue (Table 3) and the conclusion that post-merger ejecta contribute to red emission could be artifacts of this simulation suite rather than generic properties of analytic two-component modeling. The GW170817/AT2017gfo validation in Appendix B does not remove this circularity, because it uses the same analytic model and compares against the same simulation-based expectations.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether the ejecta parameters inferred from popular analytic two-component kilonova light curve models (mass, velocity, opacity for a 'blue' and a 'red' component) correspond to the physical dynamical and post-merger ejecta in neutron star mergers. The authors use mock light curves generated by multi-dimensional, wavelength-dependent radiative transfer simulations based on numerical relativity simulations (Kawaguchi et al. 2021, 2022, 2023) as ground truth, and fit them with an analytic model similar to Villar et al. (2017b) via MCMC. For a fiducial model (DD2-135, polar view) and for other merger models and viewing angles, they find that the inferred red component is more massive and slower than the blue component, opposite to the input hierarchy of the dynamical and post-merger ejecta. They demonstrate, by varying the post-merger ejecta mass in separate simulations, that the post-merger ejecta contributes to both blue and red emission, because blue emission from the post-merger ejecta is absorbed and reprocessed to red by lanthanide-rich dynamical ejecta. The paper additionally shows that the sum of the inferred blue and red masses recovers the total input ejecta mass to within a factor of about three, and that incomplete observational coverage, especially the lack of NIR data near peak, degrades the total mass estimate.","tokens_in":22433,"tokens_out":9591,"duration_ms":97731,"significance":"If the central claim holds, the paper provides a valuable caution against interpreting the parameters of the widely used two-component analytic kilonova models as physical masses and velocities of the dynamical and post-merger ejecta. The use of controlled mock data with known input ejecta properties is a strong aspect: it converts a conceptual worry about analytic model limitations into a quantitative demonstration. The paper also gives a practical, observationally actionable recommendation about the importance of multi-epoch NIR observations near peak for total mass estimation. The proposed reprocessing interpretation is physically plausible and is supported by the controlled variation of post-merger mass. The main caveat is that the mock ground truth comes from a single simulation suite with specific assumptions; however, the authors are transparent about these assumptions and include a comparison with GW170817/AT2017gfo that shows the simulated light curves resemble the observed ones.","major_comments":[{"comment":"The paper does not include a self-consistency or recovery test in which the analytic model is fit to light curves generated with the same analytic model for known input parameters. Such a test would establish whether the MCMC procedure and the analytic model can recover the true parameters when the model is correct, thereby isolating the effect of the radiative transfer physics (e.g., reprocessing) from any inherent bias or degeneracy of the fitting procedure itself. Without this test, the reported mismatch between input and inferred parameters could be partly due to the fitting procedure rather than the physical reprocessing that the paper emphasizes. I recommend adding a recovery test, at least for the fiducial parameter set.","section":"§2.2.2 and Appendix A"},{"comment":"The central conclusion that analytic parameters do not represent the actual ejecta configuration rests entirely on the fidelity of the radiative transfer simulations, which assume LTE and use the Domoto et al. (2022) line list, and on a conventional split between dynamical and post-merger ejecta. The paper acknowledges these limitations but does not discuss how the inferred hierarchy (M_red > M_blue, v_red < v_blue) might change if, for example, the dynamical ejecta were distributed more spherically around the post-merger ejecta, or if the line list were incomplete at NIR wavelengths. Since the claim is general rather than specific to the simulated models, I ask the authors to add an explicit discussion of the robustness of the main conclusion to these assumptions, and to state clearly which aspects of the result are expected to be generic.","section":"§2.1 and Table 1"},{"comment":"The fit to the DD2-125 model yields a variance parameter σ = 0.329 mag, which is about three times larger than the σ ≈ 0.1 mag obtained for the other models. This indicates that the analytic two-component model is a poor fit to the DD2-125 light curve. The paper nevertheless uses DD2-125 to support the cross-model statement that the inferred hierarchy M_red > M_blue and v_red < v_blue is common to all models. The authors should either discuss whether the inferred parameters are reliable for DD2-125 given the poor fit, or exclude it from the general trend and explicitly state the reason.","section":"§3.4 and Table 3"}],"minor_comments":[{"comment":"\"Despite of the challenges in the parameter estimation\" should read \"Despite the challenges in the parameter estimation\".","section":"Abstract"},{"comment":"\"The bolometric luminosity of the best-fit model (line) with that of GW170187/AT2017gfo\" contains a typo: \"GW170187\" should be \"GW170817\".","section":"Appendix B, Figure 12 caption"},{"comment":"\"we peform two additional radiative transfer simulation\" contains a typo: \"peform\" should be \"perform\".","section":"§3.2, first paragraph"},{"comment":"Some references are duplicated (e.g., Kasen et al. 2015 appears twice, Tanaka et al. 2013 appears twice). Please use a single entry for each work.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's conclusions are strongly tied to the authors' own suite of numerical relativity and radiative transfer simulations. While this is not disqualifying and the authors are transparent about the assumptions, the editors may wish to ensure that the framing of the central claim does not overreach beyond the specific simulation suite. The suggested recovery test and a discussion of the robustness of the hierarchy to the simulation assumptions would substantively strengthen the manuscript. The paper is otherwise well organized and likely to be of interest to the kilonova community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper makes a point the field has been circling for years, and it makes it properly. Using mock light curves from realistic radiative transfer simulations with known ejecta inputs, they show that the standard two-component analytic model (MOSFiT-style) recovers the wrong hierarchy—inferred red component more massive and slower than blue, opposite to the NR input—because post-merger ejecta contributes to both blue and red through absorption and reprocessing by lanthanide-rich dynamical ejecta. The same pattern appears in their fit to AT2017gfo, so this isn't just a simulation artifact.\n\nWhat's genuinely new: earlier papers noted the discrepancy between analytic fits and NR predictions, but this is the first systematic demonstration using controlled mock data with known ground truth, plus a clean causal test (varying only the post-merger mass and watching both blue and red components scale together). The viewing-angle dependence and the missing-data experiments are also useful. The total-mass robustness claim is honestly qualified by the heating-rate systematic, which they acknowledge.\n\nSoft spots, in order of severity. First, the ground truth is entirely the authors' own simulation chain: one NR code, one radiative transfer code with LTE and a single line list (Domoto et al. 2022). If that opacity treatment is unrepresentative of real kilonovae, the quantitative hierarchy could shift. This doesn't kill the central claim—the qualitative reprocessing picture is physical and the AT2017gfo validation supports it—but independent radiative transfer codes or merger models would strengthen it. Second, there's no data or code release, which makes independent replication harder; that's addressable and should be requested. Third, the surface covering factor f_Omega is a post-hoc explanatory device, not a predictive model; it's fine as a schematic but shouldn't be over-read. Minor: only four merger models, though they do span different EoS and mass ratios.\n\nWho this is for: anyone who fits kilonova light curves with analytic models, especially for GRB-associated events where only sparse NIR data exist. The caution against interpreting blue/red components as distinct physical ejecta components is worth taking seriously. The paper deserves a serious referee—it's methodologically sound, clearly written, and has a load-bearing result. I'd send it to review, with the request that the authors release mock light curves and consider an independent RT code check.","headline":"A convincing mock-data demonstration that two-component analytic kilonova fits mis-assign ejecta masses and velocities, with total mass still recoverable to a factor of a few; the main caveat is that the ground truth is the authors' own simulation suite.","tokens_in":22949,"tokens_out":1445,"would_cite":true,"duration_ms":17757,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fitting kilonova light curves with a standard two-component analytic model yields 'red' and 'blue' ejecta parameters that invert the true dynamical and post-merger ejecta configuration, because post-merger emission is absorbed and…","keywords":["kilonova","neutron star merger","light curve modeling","radiative transfer","r-process nucleosynthesis","ejecta reprocessing","GW170817","near-infrared observations"],"falsifier":"Run the identical fitting pipeline on mock light curves from a radiative transfer simulation in which the lanthanide-rich dynamical ejecta is placed near the poles instead of the equator; if the red component still comes out systematically more massive and slower than the blue component, the reprocessing mechanism proposed here is not the whole story.","tokens_in":22024,"feed_emoji":"🔭","tokens_out":12292,"duration_ms":104652,"temperature":0.7,"pith_summary":"Kilonovae are the optical and infrared flashes powered by the radioactive decay of neutron-rich material ejected in neutron star mergers, and observers routinely fit them with a simple two-component model labeled 'blue' and 'red'. This paper asks whether the masses and velocities those fits return correspond to anything physical. Using mock light curves from realistic radiative transfer simulations based on merger simulations with known ejecta configurations, the authors show that the fits do not: the fitted blue component is less massive and faster, the fitted red component is more massive and slower, the opposite of the true dynamical versus post-merger ejecta hierarchy. The reason is that light from the lanthanide-poor post-merger ejecta is partially absorbed by the lanthanide-rich dynamical ejecta and re-emitted at red wavelengths, so a single physical component feeds both fitted colors. The reliable output is the sum of the fitted masses, which recovers the true total ejecta mass within a factor of about three, provided near-infrared observations near the light-curve peak are included.","feed_headline":"Kilonova 'red' and 'blue' fits mislabel the real ejecta","feed_subtitle":"Post-merger ejecta feed both colors; total mass survives within a factor of three if NIR peak data are observed.","key_machinery":"The load-bearing object is the standard analytic two-component kilonova model: each component is a one-zone, homologously expanding radioactive-heated shell with mass $M$, velocity $v$, constant gray opacity $\\kappa$, and a temperature floor $T_c$, radiating as a blackbody, and the total flux is the simple sum of a blue and a red component. The mechanism that carries the argument is geometric reprocessing: the lanthanide-rich dynamical ejecta sit mostly around the equator, so they absorb blue photons emitted by the lanthanide-poor post-merger ejecta and re-emit them at redder wavelengths. The paper quantifies this with a surface-covering factor $f_\\Omega \\sim 0.6$--$0.7$ for the fiducial model, which makes the fitted blue mass roughly $(1-f_\\Omega)M_{\\rm pm}$ and the fitted red mass roughly $f_\\Omega M_{\\rm pm} + M_{\\rm dyn}$.","core_discovery":"For the fiducial DD2-135 model viewed from the pole, the input configuration is dynamical ejecta with $M_{\\rm dyn}=0.0015\\,M_\\odot$ and $v_{\\rm dyn}=0.19\\,c$ plus post-merger ejecta with $M_{\\rm pm}=0.080\\,M_\\odot$ and $v_{\\rm pm}=0.092\\,c$. The analytic two-component fit returns $M_{\\rm blue}=0.010\\,M_\\odot$, $v_{\\rm blue}=0.43\\,c$, $M_{\\rm red}=0.028\\,M_\\odot$, and $v_{\\rm red}=0.26\\,c$, so the inferred red component is more massive and slower while the blue component is less massive and faster. The same reversal appears for all four merger models and for the equatorial viewing angle, and it reproduces the pattern previously inferred for GW170817/AT2017gfo. By re-running the radiative transfer with 30% and 10% of the post-merger ejecta mass, the authors show that both fitted masses shrink together, proving that the post-merger ejecta contributes to both blue and red emission: part of its blue light is absorbed by the lanthanide-rich dynamical ejecta and reprocessed into the red. The paper concludes that the analytic blue and red components are not the physical post-merger and dynamical ejecta, and that only the total fitted mass is a trustworthy physical estimate.","pith_inferences":["A testable extension of the reprocessing picture is that the fitted blue velocity is set mainly by the diffusion timescale of the small uncovered fraction of post-merger ejecta, so the high blue velocities inferred for AT2017gfo need not imply a distinct fast ejecta layer; the same pipeline applied to a sample should show the blue/red hierarchy reversed even when the underlying ejecta hierarchy is","If the same mechanism operates in GRB-associated kilonova candidates, analytic fits to their sparse, late-time NIR data inherit the same mislabeling, leaving the total mass as the only parameter worth comparing across events.","Because the inferred total mass runs low when high-electron-fraction (high-$Y_e$) post-merger ejecta dominate, folding a composition-dependent heating rate into the analytic model could tighten the factor-of-three total-mass recovery into a more precise estimate."],"forward_implications":["The fitted red mass should not be read as dynamical ejecta mass, nor the fitted blue mass as post-merger ejecta mass; the paper shows these identifications fail in every tested merger model.","$M_{\\rm blue}+M_{\\rm red}$ recovers the true total ejecta mass within a factor of about three across equations of state, merger masses, and viewing angles, because total luminosity tracks total mass.","The inferred blue mass drops by roughly 60% for an equatorial observer in the fiducial model, so comparing fitted blue components across events without accounting for viewing angle is unreliable.","Missing near-infrared data near the peak can inflate the inferred total mass by up to a factor of two; multi-epoch NIR coverage near peak is therefore necessary for reliable mass estimates."],"supporting_citations":[{"why":"Provides the fiducial radiative-transfer mock light curves and the ejecta profiles used as ground truth for the fitting experiment.","marker":"Kawaguchi et al. 2021"},{"why":"Supplies the DD2-135 and DD2-125 models, their density/Ye structure, and the multiband light curves that constitute the main mock dataset.","marker":"Kawaguchi et al. 2022"},{"why":"Supplies the SFHo merger models with comparable dynamical and post-merger masses, used to test whether the inferred reversal depends on the model.","marker":"Kawaguchi et al. 2023"},{"why":"Gives the numerical-relativity ejecta masses and velocities for the DD2 models that define the input parameters the fits are compared against.","marker":"Fujibayashi et al. 2020"},{"why":"Gives the numerical-relativity ejecta masses and velocities for the SFHo models and the conventional dynamical/post-merger decomposition adopted in Table 1.","marker":"Fujibayashi et al. 2023"},{"why":"Supplies the analytic two-component model, the GW170817/AT2017gfo dataset, and the earlier fit whose inferred hierarchy this paper reproduces and explains.","marker":"Villar et al. 2017b"},{"why":"Supplies the analytic radioactive-heating and diffusion prescription on which the two-component fitting model is based.","marker":"Metzger 2017"},{"why":"Supplies the line list used in the radiative transfer code, setting the heavy-element opacities that shape the simulated blue and red colors.","marker":"Domoto et al. 2022"}],"fun_headline_variants":["Kilonova red-blue fits misrepresent true ejecta","Analytic kilonova fits invert red and blue ejecta","Post-merger ejecta drives both kilonova colors","Kilonova color fits swap ejecta masses and speeds","Total kilonova mass robust despite color-fit confusion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulation light curves are faithful stand-ins for real kilonovae: the radiative transfer code assumes local thermodynamic equilibrium with a particular set of heavy-element opacities, and the boundary between dynamical and post-merger ejecta is a modeling convention, so a mismatch found in these mock data could in principle be an artifact of those choices rather than a general property of analytic fitting.","fun_headline_variants_meta":{"raw":{"variants":["Kilonova red-blue fits misrepresent true ejecta","Analytic kilonova fits invert red and blue ejecta","Post-merger ejecta drives both kilonova colors","Kilonova color fits swap ejecta masses and speeds","Total kilonova mass robust despite color-fit confusion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1800,"prompt_tokens":1162,"completion_tokens":638,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":778,"completion_tokens_details":{"reasoning_tokens":557}},"tokens_in":778,"tokens_out":638,"duration_ms":5745,"temperature":1.0,"reasoning_tokens":557,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T19:40:25.931026+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical fitting pipeline on mock light curves from a radiative transfer simulation in which the lanthanide-rich dynamical ejecta is placed near the poles instead of the equator; if the red component still comes out systematically more massive and slower than the blue component, the reprocessing mechanism proposed here is not the whole story.","supporting_citations":[],"review_version":1}