{"id":"6b0f4cc9-1c38-420f-879a-3b34d672a027","arxiv_id":"2505.00657","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A normalising flow trained on real Blip glitches provides a data-informed prior that, used jointly with the signal model in Bilby, removes glitches and reduces bias in gravitational-wave parameter estimates.","lead":"Researchers trained a machine-learning model on real LIGO noise 'glitches' and used it to clean gravitational-wave data while simultaneously measuring the astrophysical signal. The method, built into the standard Bilby pipeline, reduces parameter-estimation bias when a glitch overlaps a signal.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The glitch model could absorb signal power in clean data; Section IV B checks only model preference, not whether signal parameters stay unbiased when no glitch is present.","rationale":"The reader's weakest assumption, that the O1-trained glitch prior must represent target glitches, is real and supported by the Tomte degradation. I identify a different load-bearing gap: the glitch model's effect on clean signal parameter recovery is not directly checked. The paper explicitly says the glitch+signal model is 'only slightly disfavoured' on glitch-free data, which leaves open the possibility that the flexible glitch model absorbs part of the signal even when Bayes factors prefer signal-only. Since the central claim is about bias reduction in signal parameters, an untouched check of the signal posterior under the joint model on clean data is more directly load-bearing than the training-distribution mismatch, though both matter. The proposed test is simple and would settle the concern: compare standard accuracy with and without the glitch model on the same glitch-free injections already used in Section IV B. The reader's overall conditional verdict remains appropriate; the missing safety check and the reproducibility gaps reinforce the need for minor revisions rather than rejection.","tokens_in":19523,"tokens_out":10995,"duration_ms":132801,"concrete_test":"Repeat the Section IV B setup on the same 25 O3 glitch-free segments, but record the signal posterior parameters (q, chirp mass, inclination, luminosity distance, right ascension, declination) for both the signal-only and the glitch+signal models. Compute the standard accuracy (Eq. 4) for each model and compare the posterior shifts. If any parameter shifts by more than 0.5 sigma when the glitch model is added, the safety claim that the glitch model does not significantly affect signals fails; if all parameters stay within 0.5 sigma, this concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim is that the joint glitch+signal model removes glitches while reducing bias in signal parameter estimation. A necessary condition is that the glitch model does not harm the signal when no glitch is present. Section IV B only establishes, via Bayes factors, that the signal-only model is preferred over the glitch+signal model on glitch-free data. It does not report the recovered signal posterior from the glitch+signal model. The glitch model is flexible: 12 normalising-flow coefficients plus an amplitude scale A with prior 1e-3 to 1e3 and a 10 ms time shift. Even if the Occam penalty makes signal-only win the model comparison, the glitch+signal posterior could still be shifted or broadened because the glitch basis can partially fit the merger waveform, which is concentrated in the same 1/8 s window. The bias tests in Section IV D are run only on glitch-contaminated data, so they cannot detect this failure mode. If the glitch model absorbs signal power, the reported bias reduction could be partly achieved by removing signal as well, undermining the central claim that the improvement is due to glitch removal rather than signal distortion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a data-informed parametric glitch model for LIGO data. Blip glitches from the Gravity Spy O1 training set are decomposed with SVD into a 12-dimensional basis, and a normalizing flow is trained on the coefficients to form a prior. This prior is implemented in Bilby, together with an amplitude scale and a time-shift parameter, so that signal and glitch parameters are inferred jointly. The authors validate the model on real O1 and O3 glitches with injected signals, reporting Bayes-factor separation between glitch and noise data, model-selection tests for signal-only, glitch-only, and glitch+signal models, and standard-accuracy bias metrics before and after glitch removal. They find that the joint model is preferred for most Blip test cases and that the recovered signal parameters are less biased after glitch removal.","tokens_in":19700,"tokens_out":7368,"duration_ms":77324,"significance":"If the clean-data signal-absorption issue is resolved, this is a valuable contribution. The method sits between agnostic wavelet models such as BayesWave and fully physical glitch models, and it is integrated into the standard Bilby workflow, which lowers barriers to adoption. The public code and trained model, the use of real glitches from three observing runs, and the explicit transfer test to Tomte glitches are strengths, and the validation is more extensive than in many similar proof-of-concept papers. The paper is clearly written and the conclusions are appropriately cautious about the need to retrain for other glitch classes. However, the absence of a signal-parameter check on glitch-free data leaves a gap in the central claim that glitch removal, rather than partial signal absorption, is responsible for the reported bias reduction.","major_comments":[{"comment":"The signal-only validation reports only log Bayes factors, not the signal posterior recovered by the glitch+signal model on glitch-free data. Because the glitch model contains 12 SVD coefficients plus a scale parameter A with prior 1e-3 to 1e3 and a 10-ms time shift, it can partially absorb a compact-binary merger waveform, which is concentrated in the same 1/8-s window. If this occurs, the bias reduction in Section IV D could be partly caused by removing signal power rather than glitch power. Please add a glitch-free injection test that compares the posterior moments, or the standard accuracy, of the reported parameters between the signal-only and glitch+signal analyses, and report both the model evidence and the recovered posteriors.","section":"Section IV B, Fig. 7"},{"comment":"The results depend on several ad hoc preprocessing choices: the 20-400 Hz bandpass, the 97% SVD power cutoff, the 1/8-s window, the 10-ms time-shift prior, and the A prior range, but no sensitivity analysis is given. Since the paper's central claim is that the method effectively removes glitches and reduces bias, the reader cannot tell whether the reported Bayes-factor separation and standard-accuracy improvement are robust to reasonable variations of these choices or are tuned to the specific settings. Please add a sensitivity test varying at least the SVD cutoff, the window length, and the bandpass edges, and report the effect on the foreground/background separation and on the standard accuracy for a subset of O3 glitches.","section":"Sections III B and IV"},{"comment":"The Tomte analysis concludes that even an incorrect glitch model improves signal analysis, but the same section reports that the glitch+signal model is preferred in only 15% of the 20 injections and that the foreground false-alarm rate is 22%. These two observations need to be reconciled explicitly, for example by showing that the standard-accuracy improvement is not driven by a few events and by reporting per-event standard accuracy or a model-averaged posterior. Without this, the claim that an incorrect model still reduces bias is not fully supported, and the conclusion is more fragile than the Blip-only claims.","section":"Section IV E and Section V"}],"minor_comments":[{"comment":"The quantity xmaxL is described as the 'maximum likelihood posterior value'; this wording is ambiguous. Please state whether it is the maximum-likelihood sample, the maximum-a-posteriori value, or the maximum of the marginal posterior.","section":"Equation (4)"},{"comment":"The O1 foreground results in Fig. 5 use glitches that are in the training set, while the O3 results use held-out glitches. Please state this explicitly in the text, or separate the in-sample and out-of-sample O1 results, so the reader can gauge the potential optimism from training on the same glitches.","section":"Section IV A 1"},{"comment":"The y-axis is described as the 'ratio of log Bayes factors', but the text and the threshold at one indicate that the plotted quantity is the log Bayes factor between the glitch+signal and signal-only models. Please clarify the exact quantity plotted.","section":"Section IV C, Fig. 10"},{"comment":"The mean standard accuracy is reported without uncertainties or per-event scatter. Adding error bars or showing the distribution would make the comparisons more interpretable given the small sample sizes of 20-25 glitches.","section":"Section IV D, Figs. 12, 13, 16"},{"comment":"A direct comparison, or at least an explicit discussion, with BayesWave and gwsubtract on the same test glitches would help the reader judge the practical added value of the new model, since these are the standard methods in LVK analyses.","section":"Section V"},{"comment":"Please specify whether the 10-ms Gaussian time-shift prior is centered on the Omicron trigger time and whether that trigger time is the one reported in the Gravity Spy catalogue.","section":"Section III D"}],"recommendation":"major_revision","confidential_remarks":"The clean-data posterior check is the key requirement; if the authors add it and the glitch+signal posterior on glitch-free data is unbiased, I would support acceptance. The lack of comparison with BayesWave is not a blocker in my view, but it should be framed as future work. The O1 in-sample results are a minor concern because the O3 holdout results are provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nRead Malz & Veitch on joint signal+glitch inference with a normalising-flow glitch model. Verdict: worth engaging, and worth sending to a referee.\n\nWhat's new: they train a normalising flow on the SVD coordinates of real Blip glitches from Gravity Spy, then use that flow as a prior inside Bilby for simultaneous signal and glitch inference. That combination—data-driven glitch prior plus full Bayesian joint fit in the standard LVK stack—is not in the earlier literature. Merritt's Glitschen used a Gaussian PCA prior and no signal; Sun and Xiong used flows but not as a Bilby prior in this way. The method description is clear, and the 12-parameter SVD+flow compression is a sensible middle ground between an agnostic wavelet model and a fully physical one.\n\nThe validation work is solid. They test on real O1 and O3 Blips, show good foreground/background Bayes factor separation, and demonstrate that joint fitting reduces the standard-accuracy bias of injected signals. The Tomte cross-class experiment is honest: the Blip-trained model only partially works, and they say so. I agree with the reader that this is not circular—the flow is trained on one set of real glitches and evaluated on held-out times; the prior is an input, not a post-fit quantity.\n\nSoft spots, in approximate order of seriousness. First, there is no check of the recovered signal posterior when the data contain a signal but no glitch. Section IV B only reports that the signal-only model wins the Bayes factor comparison. The glitch model is flexible (12 flow coordinates plus an amplitude scale and a time shift), and it could partially absorb signal power even when the Occam penalty makes signal-only preferred. The paper's claim that the model does not 'significantly affect' signals needs a direct posterior comparison, not just a model selection statistic. That is a real gap, but a fixable one. Second, no direct comparison with BayesWave or gwsubtract on the same injections; the paper doesn't claim to beat them, but the reader can't calibrate how much of the bias reduction is specific to this method. Third, the O1 results are partly in-sample, though the O3 out-of-sample tests mitigate that. Fourth, reproducibility: the code reference in [64] has no URL or version/commit hash; a referee should ask for it.\n\nWho it's for: GW data analysts and method developers working on glitch mitigation. It's a proof-of-concept, not production-ready, but the path to training per-class models is clear and cheap. I'd bring it to a reading group. I'd accept it for peer review; the missing posterior check and baseline comparison should be requested in revision.","headline":"A solid, honest proof-of-concept for a data-informed glitch prior in Bilby; the core idea is new and works, but the authors should verify the glitch model is not absorbing signal power in clean data and should give readers a baseline comparison.","tokens_in":20308,"tokens_out":2639,"would_cite":true,"duration_ms":26971,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A normalising flow trained on real Blip glitches, embedded in a standard Bayesian inference code, allows gravitational-wave signal and glitch parameters to be fitted jointly; the paper reports that this removes the glitch and reduces…","keywords":["gravitational waves","glitch mitigation","normalising flows","Bayesian inference","singular value decomposition","LIGO","binary black hole parameter estimation","transient noise"],"falsifier":"Take a sample of Blip glitches from a later observing run not used in training, inject a signal with known parameters into each, and measure the standard accuracy of the recovered parameters with and without the joint glitch model: if the model does not bring the bias down to about 1, or if the signal-plus-glitch model is not preferred in most cases, the central claim fails. The paper's Tomte experiment already provides a partial version of this test, with a 22% false-alarm rate and joint-model preference in only 15% of cases when the prior does not match the glitch class.","tokens_in":19237,"feed_emoji":"🔭","tokens_out":11440,"duration_ms":114177,"temperature":0.7,"pith_summary":"Gravitational-wave detectors are frequently hit by short noise transients, glitches, that can masquerade as or distort astrophysical signals. This paper tries to establish that a glitch model whose prior is learned from real detector glitches—rather than from a generic wavelet basis—can be inserted into the standard Bayesian inference pipeline, so that signal and glitch parameters are estimated jointly and the glitch is removed from the data. The authors demonstrate this on real LIGO data from the first and third observing runs, injecting binary-black-hole signals on top of actual Blip glitches. They report that the joint signal-plus-glitch model is preferred in most test cases, that glitch-free data rarely trigger the model, and that recovered signal parameters are substantially less biased after the glitch is removed. If the result holds, it offers a practical, data-informed alternative to pre-subtraction or agnostic wavelet fitting for glitch mitigation in gravitational-wave astronomy.","feed_headline":"A glitch-trained flow model cleans LIGO data and cuts bias","feed_subtitle":"Joint signal-and-glitch fitting removes Blip glitches and sharpens recovered source parameters.","key_machinery":"The load-bearing mechanism is the pairing of a low-dimensional SVD subspace of observed glitches with a normalising-flow prior over the subspace coordinates. A normalising flow is a machine-learning density estimator built from invertible transformations, which can both evaluate the probability of a set of glitch amplitudes and generate new glitch realisations by sampling a latent Gaussian. The training pipeline bandpasses 20–400 Hz, whitens, cuts to 1/8 s, applies a Hann window, and keeps enough singular vectors to retain 97% of the power, yielding 12 amplitude parameters per glitch. The flow learns the distribution of these 12 parameters; at inference time the amplitudes are drawn from the flow, multiplied by the fixed basis, shifted in time under a 10-ms Gaussian prior, and scaled by an amplitude factor, and the subtracted glitch leaves a residual whose Gaussian likelihood scores the signal parameters. This construction is what lets the analysis move from an agnostic wavelet description of transients to one informed by the actual population of glitches.","core_discovery":"The central claim is that a parameterised glitch model with a data-learned prior can be jointly inferred with a gravitational-wave signal, and that doing so removes the glitch and reduces bias in the recovered source parameters. The paper states this directly: modelling Blip glitches using normalising flows can successfully remove glitches from the data, thus reducing bias in the signal parameter estimation. The demonstration uses real O1 and O3 glitches: a flow trained on O1 Blip glitches is applied to unseen glitches, separates glitch from Gaussian background in log Bayes factor, and when a known binary signal is injected over a glitch, the signal-plus-glitch model is preferred over the signal-only model for the majority of test cases. Quantitatively, the mean standard accuracy of recovered parameters moves toward about 1 when the glitch model is included, whereas the glitch-contaminated analysis shows biases several times larger for the most affected parameters.","pith_inferences":["If the SVD-plus-flow representation generalises as cleanly as the Blip results suggest, the same glitch-prior plugin could be used as a transient model for unmodelled astrophysical bursts, where no waveform template is available for joint inference.","The strong separation in log Bayes factors between glitch and no-glitch data indicates the glitch evidence itself could serve as a data-quality flag, not just a parameter-estimation fix, with thresholds set from the background distribution.","A mixture-of-flows prior spanning several glitch classes would be a natural extension; it would let the sampler choose which class is present rather than committing to one trained morphology, and would likely address the Tomte failure mode.","The per-run retraining requirement suggests a monitoring use: periodic retraining on fresh glitch classifications would keep the prior aligned as detector noise evolves, turning the method into a continuous glitch-calibration tool."],"forward_implications":["Existing compact-binary analyses could adopt the method without changing their inference software: the glitch prior plugs into the same likelihood and sampler used for standard parameter estimation.","For Blip-class contamination, glitch-removed posteriors should have standard accuracy near 1 across mass, distance, and sky-location parameters, rather than the multi-sigma biases seen when the glitch is ignored.","Training a new model is cheap, so the procedure can be re-run for each observing run or detector configuration to keep the prior matched to current glitch populations.","The technique is not restricted to Blip glitches; a quick retraining on Koi Fish and Power Line glitches also yields a model that removes those glitches.","Using a glitch model trained on a morphologically similar but different class still reduces parameter bias, but model selection becomes unreliable, so the paper recommends training one model per glitch class."],"supporting_citations":[{"why":"Provides the 1,785 O1 Blip glitch examples from which the SVD basis is built and on which the normalising flow is trained.","marker":"[53]"},{"why":"Supplies the Gravity Spy classifications that define unseen O2/O3 test glitches and the glitch-free times used as Bayes-factor background.","marker":"[23]"},{"why":"Gives the public LIGO strain data and the O1 amplitude spectral density used to whiten training data and build test sets.","marker":"[55]"},{"why":"The Bayesian inference library into which the glitch prior and joint signal-plus-glitch likelihood are implemented for parameter estimation and evidence evaluation.","marker":"[46]"},{"why":"The nested sampler used with Bilby to draw posteriors and estimate the Bayes factors that drive model selection.","marker":"[60]"},{"why":"Defines the neural spline flow architecture used as the normalising flow that learns the glitch amplitude prior.","marker":"[58]"},{"why":"The software library used to train the normalising flow on the reduced SVD amplitudes.","marker":"[59]"},{"why":"The agnostic wavelet glitch model against which the data-informed prior is positioned, framing the paper's motivation for an informative glitch model.","marker":"[24]"}],"fun_headline_variants":["Flow-trained glitch priors cut bias in gravitational wave inference","Joint signal-glitch fit with learned priors cleans LIGO data","Machine-learned glitch model reduces bias in GW parameter estimation","Data-informed glitch priors improve gravitational wave source recovery","Flow-based glitch model sharpens gravitational wave parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the Blip glitches in the data being analysed look like the O1 Blip glitches used for training—same morphology, same 20–400 Hz band, same short duration—so the learned SVD subspace and flow prior cover the target glitch, and the paper's Tomte results show what happens when they do not.","fun_headline_variants_meta":{"raw":{"variants":["Flow-trained glitch priors cut bias in gravitational wave inference","Joint signal-glitch fit with learned priors cleans LIGO data","Machine-learned glitch model reduces bias in GW parameter estimation","Data-informed glitch priors improve gravitational wave source recovery","Flow-based glitch model sharpens gravitational wave parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000981,"raw_usage":{"total_tokens":4131,"prompt_tokens":882,"completion_tokens":3249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":3164}},"tokens_in":498,"tokens_out":3249,"duration_ms":19709,"temperature":1.0,"reasoning_tokens":3164,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:36:42.307907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of Blip glitches from a later observing run not used in training, inject a signal with known parameters into each, and measure the standard accuracy of the recovered parameters with and without the joint glitch model: if the model does not bring the bias down to about 1, or if the signal-plus-glitch model is not preferred in most cases, the central claim fails. The paper's Tomte experiment already provides a partial version of this test, with a 22% false-alarm rate and joint-model preference in only 15% of cases when the prior does not match the glitch class.","supporting_citations":[{"cited_title":"Kobyzev, S","cited_arxiv_id":null,"evidence_quote":"Gives the public LIGO strain data and the O1 amplitude spectral density used to whiten training data and build test sets."},{"cited_title":"Cabero, A","cited_arxiv_id":null,"evidence_quote":"The nested sampler used with Bilby to draw posteriors and estimate the Bayes factors that drive model selection."},{"cited_title":"Thrane and C","cited_arxiv_id":null,"evidence_quote":"Defines the neural spline flow architecture used as the normalising flow that learns the glitch amplitude prior."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The software library used to train the normalising flow on the reduced SVD amplitudes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The agnostic wavelet glitch model against which the data-informed prior is positioned, framing the paper's motivation for an informative glitch model."}],"review_version":1}