{"id":"7b5d30ab-931c-4bde-aeea-27afc7185407","arxiv_id":"2411.14705","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Starkiller forward-models every catalog star in an IFU datacube using Gaia photometry and stellar atmosphere models, then subtracts the model cube to reveal extended sources, including trailed stars and satellite streaks.","lead":"This paper presents Starkiller, an open-source program that removes stars and satellite streaks from telescope data cubes by recreating them from a star catalog and subtracting the simulation. It lets astronomers recover observations of comets, asteroids, and nebulae that would otherwise be spoiled by dense star fields or satellite trails.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '>90% stellar flux removal' claim rests on in-sample model fits: spectral type, E(B-V), and the global flux correction are all derived from the same datacube being subtracted, so reported residuals do not yet bound worst-case model-mismatch errors.","rationale":"The reader's conditional verdict identifies the same fundamental risk: the subtraction quality depends on the finite stellar library plus a single extinction law reproducing the true stellar spectra. My stress-test sharpens that into a concrete evidentiary problem. The independent Gaia G-band scaling is a genuine external anchor for total flux, and the data-PSF construction is a real strength for trailed sources, so the concern is not that the method is circular in a fatal way. The issue is that the paper's headline success metric, 'removing more than 90% of the stellar flux in most cases,' is supported by in-sample demonstrations: the per-star spectral type and E(B-V) are chosen by maximizing correlation against the same datacube, and the global flux correction is literally derived from the median ratio of model to observed spectra for the same calibration sources. Therefore the reported residuals in Figs. 5, 10, and 16 are expected to be optimistically small, because the fitted parameters have already absorbed some of the discrepancy. The paper itself flags this weakness for extinction in Sec. 4.1, but does not quantify how much it affects the subtraction residuals. An injection-recovery test with known input spectra, including deliberately mismatched spectra and non-standard extinction laws, would settle whether the 90% claim holds outside the training population. Without that test, a conditional accept is appropriate: the package is useful and well-documented, but the headline quantitative claim should not be accepted at face value.","tokens_in":158,"tokens_out":3087,"duration_ms":96415,"concrete_test":"End-to-end injection-recovery on a real MUSE datacube: inject synthetic stars with a known PSF and spectra drawn (a) from the CK/MARCS/OB grids, (b) with Teff shifted by ±1000 K and log g shifted, and (c) with non-Fitzpatrick extinction (R_V = 2.5 and 4.0), at magnitudes spanning the Gaia range, using both sidereal and trailed PSFs; then run Starkiller and measure per-star residual flux. If the median residual exceeds about 10% for any injected population above S/N ≈ 10, or if recovered E(B-V) differs from injected values by more than 0.1 mag for grid members, the >90% claim requires qualification and the extinction diagnostic needs independent calibration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Sec. 4.4: 'effective at removing more than 90% of the stellar flux in most cases') is supported mainly by a few example stars in Figs. 5, 10, and 16, whose residuals are measured after the same datacube was used to select the spectral template, fit E(B-V) with a single Fitzpatrick R_V=3.1 law, and construct the global flux correction (Sec. 3.6). The paper itself warns in Sec. 4.1 that 'if the input stellar atmosphere library is insufficient, starkiller may be using extinction as a tool to reshape poorly matched spectra to improve the correlation,' and in Sec. 4.4 that 'the quality ... depends heavily on how well the spectral models fit.' Because only a single Gaia G-band magnitude anchors the absolute scale, any wavelength-dependent mismatch between the true star and the reddened library spectrum is not independently constrained. For trailed and satellite cases no per-source residual statistics are provided. Thus the headline 90% figure is not yet established for the population of stars Starkiller is designed to remove, especially faint, crowded, or non-library stars. This is not a critique of the method's concept but of the evidence supporting the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents starkiller, an open-source Python package that forward-models and subtracts catalogued stellar sources (and satellite streaks) from IFU datacubes. It uses Gaia DR3 positions and G-band magnitudes to scale stellar atmosphere model spectra, fits a trailed PSF (with a data-driven PSF option), matches spectral templates under a Fitzpatrick extinction grid, and constructs a synthetic scene that is subtracted from the datacube. The method is demonstrated on VLT/MUSE data of the interstellar comet 2I/Borisov, the Didymos system after the DART impact, the planetary nebula NGC 6563, and a satellite-contaminated observation of the blazar J0141-5427. The paper also claims that the package simultaneously provides stellar spectral classification, relative velocity, and line-of-sight extinction, and that it removes more than 90% of stellar flux in most cases.","tokens_in":22954,"tokens_out":5157,"duration_ms":51591,"significance":"The core idea is timely and useful: an IFU-specific synthetic difference-imaging tool that works for both sidereal and non-sidereal data, including trailed sources and satellite streaks, would make dense-field and satellite-affected IFU observations usable for science. The independent Gaia-based flux scaling is a genuine external anchor, and the data-PSF construction for trailed sources is an effective practical solution to seeing variability along streaks. The package is open source and the demonstrations use real, varied MUSE datasets. However, the headline quantitative claim of >90% flux removal is currently supported only by in-sample residuals, and the secondary products (relative velocity, extinction) are either explicitly not fully implemented or are shown to be in tension with external dust maps. The significance of the paper would be substantially higher if the subtraction claim were validated out-of-sample and the abstract were aligned with the actual capabilities.","major_comments":[{"comment":"The claim that starkiller is 'effective at removing more than 90% of the stellar flux in most cases' is not established by the evidence presented. The quoted residuals (5% for the two trailed stars in Fig. 5, 6% and 7% for the sidereal stars in Fig. 16, <10% for NGC 6563 in §3.7) are measured after the same datacube was used to select the per-source spectral template and E(B-V) via correlation (§3.5) and to construct the global flux correction spline from the same calibration sources (§3.6). This is an in-sample measure: a star whose true spectrum is not in the model grid, or whose extinction law differs from the single Fitzpatrick (1999) R_V=3.1 law, can be fit with a degraded model that the flux correction partially absorbs, and the reported residual will not reflect the true model-mismatch error. The paper itself warns in §4.1 that 'if the input stellar atmosphere library is insufficient, starkiller may be using extinction as a tool to reshape poorly matched spectra to improve the correlation,' and in §4.4 that subtraction quality 'depends heavily on how well the spectral models fit.' To support the population-level claim, the authors should provide out-of-sample validation: for example, withhold a subset of calibration sources from the flux correction and report their residuals, inject synthetic stars with spectra outside the library, or report residuals for sources not used in any fit step. Without such a test, the >90% figure cannot be taken as a statement about the general population of stars starkiller is intended to remove.","section":"§4.4, with §3.6 and Figs. 5, 10, 16"},{"comment":"The abstract and conclusion state that starkiller 'simultaneously provides stellar spectral classification, relative velocity, and line-of-sight extinction,' but §3.5.1 says the velocity routine is 'not currently fully implemented in starkiller' and is only effective for stars with prominent A–K features. No validation against known radial velocities or Gaia RVS is presented, and the example in Fig. 7 shows line-to-line scatter (e.g., Hβ: −61±23 km/s; Hα: −51±5 km/s; Na D: −32±4 km/s) whose error-weighted average is quoted without discussion of the inconsistencies. This claim should either be removed from the abstract and conclusion, or the routine should be completed and validated before publication.","section":"Abstract and §5, versus §3.5.1"},{"comment":"The paper's own analysis shows that the starkiller E(B-V) values for NGC 6563 are about twice the Schlafly & Finkbeiner (2011) dust map value (median ~0.5 mag versus 0.2271 mag), and the authors state that 'the reliability of extinction values generated by starkiller' requires further investigation. Since the abstract lists line-of-sight extinction as a primary product, the current manuscript does not support that claim. The abstract and conclusion should be qualified to state that the extinction estimates are preliminary and should not be used for scientific inference until validated, particularly given the known degeneracy between spectral type and reddening in template matching.","section":"§4.1 and Fig. 14"}],"minor_comments":[{"comment":"The phrase 'using the catalog for positions and fluxes to scale stellar models, independent of the datacube' is only partially true: the initial flux scaling is indeed independent, but the per-source E(B-V), the spectral template choice, and the global flux correction of §3.6 all depend on the datacube. This should be clarified to avoid overstating the independence.","section":"Abstract and §3.5"},{"comment":"Typo: 'spetcra' should be 'spectra'.","section":"§3.5, first paragraph"},{"comment":"The pipeline name is spelled inconsistently: 'PampelMUSE' in the main text and 'PempelMUSE' in the Appendix B title and text. Please use a single spelling.","section":"Appendix B and §3.7"},{"comment":"The caption states 'Total residual = X%' but does not define whether this is the sum of absolute residuals divided by total counts, or another metric, and no uncertainty is given. Please define the residual metric and, ideally, provide a measure of scatter across the spatial cutout or across wavelength.","section":"Figs. 5 and 16"},{"comment":"The satellite spectrum fit uses a solar spectrum with a multi-component atmospheric extinction model (aerosols, Rayleigh, ozone). The authors should note the degeneracies in this fit (e.g., aerosol parameter correlations) and that this is a single example, not a general validation of satellite spectral extraction.","section":"§4.2 and Fig. 15"}],"recommendation":"major_revision","confidential_remarks":"This is a useful software paper with a sound overall concept and credible demonstrations on real MUSE data. The main technical concern is that the headline subtraction-quality figure is in-sample: the flux correction and per-source extinction are fit to the same cubes used to measure residuals, so the reported 5–10% residuals do not bound model-mismatch errors for the general stellar population. The abstract also overclaims relative velocity and extinction capabilities that are either unimplemented or unvalidated. These issues are fixable within the paper's scope by adding out-of-sample validation (withheld sources, injected non-library spectra) and aligning the abstract/conclusion with the demonstrated capabilities. I would not reject; the method is likely to be adopted by the community. The authors' own caveats in §4.1 and §4.4 are candid and should be retained, but they should be reflected in the claims made in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nStarkiller is a real contribution: it forward-models catalog stars through stellar atmosphere models scaled to Gaia G-band photometry and subtracts them from IFU datacubes, with a data-PSF construction that handles strongly trailed, non-sidereal sources and a satellite-streak subtraction prototype. That is new relative to PampelMUSE (sidereal PSF fitting only) and TRIPPy (imagers). The package is on GitHub with a Zenodo DOI, the writing is clear, and the limitations section is unusually honest.\n\nThe strongest part is the data PSF. Averaging the trails of calibration stars captures the seeing variability along a streak in a way a Gaussian or Moffat cannot, and the residual comparisons in Fig. 5 and Fig. 16 support that: 5% for trailed examples, 6–7% for sidereal point sources. The external Gaia anchor for absolute flux scaling is a genuine plus—it means the model flux is not fit to the datacube's photometric scale. The demonstrations on 2I/Borisov, the DART debris trail, and the satellite-struck blazar all look like real recoveries of otherwise marginal or unusable exposures.\n\nNow the soft spots, in proportion. The “more than 90% of the stellar flux in most cases” claim (Sec. 4.4) rests on a handful of example stars. The per-source spectral type, E(B-V), and the global flux correction (Sec. 3.6) are all derived from the same datacube being subtracted. The Gaia G-band magnitude anchors the overall scale, but any wavelength-dependent mismatch between a star and the reddened library spectrum is not independently constrained. The paper itself warns in Sec. 4.1 that extinction can be used to reshape poorly matched spectra. That does not kill the method—the residuals shown are reasonable, and the comparison to PampelMUSE in Appendix B is a fair sanity check—but the population-level claim is not yet established, especially for faint, crowded, or non-library stars. A referee should ask for either a holdout test (fit on one set of stars, subtract another) or residual statistics over all catalog stars, not just the pretty examples.\n\nSecond, the abstract says the package “simultaneously provides stellar spectral classification, relative velocity, and line-of-sight extinction,” but Sec. 3.5.1 states velocity is “not currently fully implemented” and Sec. 4.1 says the extinction values are “yet to be tested.” That is an overclaim and should be toned down before publication.\n\nThird, the satellite removal is a prototype: demonstrated on one observation, with the satellite spectrum extracted via PSF photometry rather than forward-modeled. Given the paper's own framing (“prototype,” “attempt to remove”), that is acceptable, but the section should say more clearly that it is a demonstration of feasibility, not a validated capability.\n\nWho is this for? Anyone reducing MUSE IFU data of moving targets or fields with satellite contamination, and anyone building similar subtraction tools for other IFUs. It deserves a serious referee. My recommendation: send it to peer review, conditional on the authors qualifying the abstract and either adding population-level residual statistics or softening the 90% claim.","headline":"A genuinely useful open-source tool for forward-modeling subtractive cleaning of IFU datacubes, whose headline '90% removal' claim is stronger than the in-sample evidence supports.","tokens_in":23574,"tokens_out":3201,"would_cite":true,"duration_ms":32833,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Starkiller forward-models every catalog source in an integral-field spectrograph's field of view and subtracts the simulated scene, so single exposures of crowded, trailed, or satellite-struck datacubes become usable for science.","keywords":["integral field unit spectroscopy","forward modeling","synthetic difference imaging","Gaia DR3 catalog","point-spread function subtraction","satellite streaks","non-sidereal tracking","stellar atmosphere models"],"falsifier":"Take a MUSE datacube whose field contains stars with independently known spectra (for instance from the MUSE-specific stellar library or archival high-resolution surveys), run starkiller, and measure the residual flux at each star's position: the more-than-90% flux-removal claim fails if stars whose true spectra are absent from the model grid leave residuals that scale with brightness. A second, more immediate check is the extinction output, since the paper already notes that its NGC 6563 $E(B-V)$ values run about twice the Schlafly & Finkbeiner dust map for that region; comparing per-star starkiller extinctions against spectroscopic or multi-band dust measurements across several fields would directly test whether the extinction grid is compensating for model mismatch.","tokens_in":22494,"feed_emoji":"🔭","tokens_out":12268,"duration_ms":104491,"temperature":0.7,"pith_summary":"The paper presents starkiller, an open-source Python package that clears stars and other point-source contamination out of integral-field-unit (IFU) spectroscopic datacubes by forward-modeling every catalog source in the field. Instead of needing a reference observation, starkiller builds a synthetic copy of the stellar scene—positions and brightness from the Gaia DR3 catalog, spectra from stellar atmosphere libraries, shapes from a fitted (possibly trailed) PSF—and subtracts it from the data. The central claim is that this synthetic difference imaging removes more than 90% of stellar flux in most cases, whether the stars are round point sources, non-sidereally trailed streaks, or satellite trails. The authors demonstrate it on VLT/MUSE observations of the interstellar comet 2I/Borisov, the DART-impacted asteroid Didymos, the planetary nebula NGC 6563, and a satellite strike that crossed directly over a blazar, salvaging an exposure that would otherwise have been lost. If the claim holds, IFU observations of extended sources in dense stellar fields—including near the Galactic plane—become usable, which is exactly where moving-target and low-surface-brightness science has historically had to pause.","feed_headline":"Starkiller strips catalog stars from IFU datacubes, streaks included","feed_subtitle":"Simulating the whole stellar scene from Gaia recovers comets, nebulae, and satellite-hit exposures.","key_machinery":"The carrying mechanism is the synthetic scene: a simulated datacube built at ten times the data's spatial resolution by convolving a PSF with each catalog source's position, multiplying by that star's best-match model spectrum, rescaling the flux to the Gaia G-band magnitude (converted from Vega to AB with a +0.118 mag offset calibrated on DA white dwarfs), and reddening the result with an extinction grid. Three PSF models are offered—Moffat, Gaussian, and a data PSF formed by normalizing and averaging the trails of isolated calibration stars—and the data PSF is what captures the seeing variability recorded along a streak, which analytic profiles miss. Spectral matching uses the Pearson correlation coefficient between the treated observed spectrum and every reddened model, deliberately avoiding $\\chi^2$ so that flux normalization never enters the match. A final wavelength-dependent flux correction, built from the median model-to-observation flux ratio of well-correlated calibration sources, absorbs MUSE calibration trends. Satellite streaks get their own pathway: they are detected with Canny edge detection and a Hough line transform, then subtracted using the stellar PSF parameters stretched into a line, with the spectrum extracted by PSF photometry because satellite SEDs cannot be forward-modeled.","core_discovery":"On the paper's own terms, the discovery is that a stacked stellar scene constructed almost entirely from external information—Gaia DR3 positions and G-band magnitudes for every source in the field, model stellar spectra chosen by Pearson correlation against the datacube, a Fitzpatrick (1999) extinction law with $R_V=3.1$ applied over a grid of $E(B-V)$ values, and a PSF that can be trailed or built directly from the averaged streaks of calibration stars—can be subtracted from an IFU datacube to expose the extended emission underneath. The deliberate design choice is that the scene is scaled by the catalog, not by the datacube, so the stellar models are not biased by a foreground coma or nebula; the data are used only for shape comparison and PSF determination. The paper reports residuals below 10% for the nebula field, about 5% for trailed stars when the data PSF is used versus roughly 10% for analytic Moffat or Gaussian profiles, and a successful satellite subtraction whose corrected blazar spectrum matches an uncontaminated follow-up observation. As a by-product, starkiller returns for every catalog star a spectral classification, a relative velocity from Gaussian fits to H$\\beta$, H$\\alpha$, Na D, and the Ca II triplet, and a line-of-sight extinction value.","pith_inferences":["Because the scene is scaled by the catalog rather than by the data, running starkiller across many archival MUSE fields would expose any slow wavelength-dependent drift in the instrument's flux calibration, turning the flux-correction curve into a long-term diagnostic the paper only sketches.","The paper's own NGC 6563 result—fitted $E(B-V)$ values about twice the Schlafly & Finkbeiner dust-map estimate—suggests the extinction grid is absorbing model mismatches; a testable extension is to fit $R_V$ as a free parameter or to use Gaia BP/RP photometry to break the degeneracy between spectral type and reddening.","The data PSF idea—averaging the streaks of calibration stars as the PSF template—is portable beyond MUSE and beyond spectroscopy, and could sharpen trailed-source subtraction in any survey that tracks non-sidereal targets.","The satellite mode extracts clean reflectance spectra, and the finding that the satellite spectrum needs extra atmospheric extinction beyond the pipeline's correction suggests a practical route to measuring the line-of-sight airmass of satellites, which could feed brightness-prediction models for observatories."],"forward_implications":["Exposures previously rejected—crowded fields within about 10 degrees of the Galactic plane and datacubes with satellite strikes—can be recovered for science; of 51 MUSE exposures of 2I/Borisov, 27 needed starkiller and 4 were rescued from rejection.","IFU difference imaging becomes possible from a single exposure, with no reference cube, because the reference is synthesized from the catalog and the model libraries.","Moving-target science near the Galactic plane becomes feasible: the Solar System community can use IFUs to observe comets and asteroids in dense stellar fields that observation programs previously avoided.","The same machinery doubles as an instrument diagnostic, since the flux-correction curve independently checks the datacube's flux calibration against Gaia photometry, and the per-star $E(B-V)$ values give a rapid independent extinction estimate for cluster fields.","Residual stellar PSFs left after subtraction can flag variable sources, carrying standard difference-imaging practice into IFU data."],"supporting_citations":[{"why":"Supplies the DR3 catalog whose positions and G-band magnitudes anchor every source in the simulated scene.","marker":"Gaia Collaboration et al. 2023"},{"why":"Provides the default grid of stellar atmosphere models used as the primary spectral templates.","marker":"Castelli & Kurucz 2003"},{"why":"Defines the extinction law applied over an $E(B-V)$ grid to redden each model spectrum.","marker":"Fitzpatrick 1999"},{"why":"The TRIPPy PSF module that starkiller builds on to construct trailed point-spread functions.","marker":"Fraser et al. 2016"},{"why":"Defines VLT/MUSE, the instrument and datacube format used for all demonstrations.","marker":"Bacon et al. 2010"},{"why":"PampelMUSE, the sidereal PSF-subtraction pipeline starkiller is compared against on NGC 6563.","marker":"Kamann et al. 2013"},{"why":"The 2I/Borisov MUSE campaign whose crowded, trailed datacubes motivated the development.","marker":"Bannister et al. 2020"},{"why":"The Didymos DART-impact follow-up data used to show faint trailed stars subtracted across the debris tail.","marker":"Opitom et al. 2023"},{"why":"Provides the procedure for converting Gaia G Vega magnitudes to AB, the basis of flux scaling.","marker":"Axelrod et al. 2023"},{"why":"Supplies the well-calibrated DA white dwarf spectra used to measure the Gaia Vega-to-AB offset.","marker":"Narayan et al. 2019"}],"fun_headline_variants":["Gaia-based stellar scene subtraction cleans IFU datacubes","Starkiller's forward model lifts stars, streaks, and satellites off IFU data","Extract extended emission from IFU datacubes by modeling every star","Starkiller subtracts all catalog stars to uncover comets and nebulae","Model-the-scene method recovers faint extended sources in dense fields"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that some member of the finite grid of stellar atmosphere models, reddened by a single Fitzpatrick (1999) extinction law with $R_V=3.1$ and scaled to Gaia G-band photometry, can reproduce what each catalog star actually looks like across 4000–9300 Å; the paper itself states that subtraction quality depends heavily on how well the spectral models fit, and warns that if the library is insufficient, extinction may be used to reshape poorly matched spectra.","fun_headline_variants_meta":{"raw":{"variants":["Gaia-based stellar scene subtraction cleans IFU datacubes","Starkiller's forward model lifts stars, streaks, and satellites off IFU data","Extract extended emission from IFU datacubes by modeling every star","Starkiller subtracts all catalog stars to uncover comets and nebulae","Model-the-scene method recovers faint extended sources in dense fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1892,"prompt_tokens":1010,"completion_tokens":882,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":785}},"tokens_in":626,"tokens_out":882,"duration_ms":8132,"temperature":1.0,"reasoning_tokens":785,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:59:52.586595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a MUSE datacube whose field contains stars with independently known spectra (for instance from the MUSE-specific stellar library or archival high-resolution surveys), run starkiller, and measure the residual flux at each star's position: the more-than-90% flux-removal claim fails if stars whose true spectra are absent from the model grid leave residuals that scale with brightness. A second, more immediate check is the extinction output, since the paper already notes that its NGC 6563 $E(B-V)$ values run about twice the Schlafly & Finkbeiner dust map for that region; comparing per-star starkiller extinctions against spectroscopic or multi-band dust measurements across several fields would directly test whether the extinction grid is compensating for model mismatch.","supporting_citations":[],"review_version":1}