{"id":"74c60e55-d398-47dd-aa26-cfbdecd4ecb5","arxiv_id":"2507.05209","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper introduces an iterative, ensemble-based hierarchical subtraction scheme powered by neural density estimators that recovers overlapping gravitational wave signals accurately and fast.","lead":"Overlapping gravitational wave signals, expected in future observatories, can be analyzed with repeated subtraction using fast neural samplers instead of expensive joint fits. The paper shows this iterative approach converges to accurate posteriors in simulations, offering a fast alternative for third-generation detectors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The greedy sampler is not reversible and the reported checks are marginal, so equivalence to full joint PE is not established; a joint-coverage comparison against bilby would settle the claim.","rationale":"I read the paper as a proof-of-concept that hierarchical subtraction plus NDE sampling and resampling can replace joint PE for overlapping BBHs. The Fisher-matrix analysis (Eqs. 4-5) is a useful consistency argument for point estimates, and the empirical comparison to bilby in Fig. 2 is encouraging. However, the step from 'the systematic bias of a point estimate shrinks under iteration' to 'the ensemble approximates the joint posterior' is not justified: the sampler is deliberately not reversible, and the validation is marginal. The paper's own limitations paragraph concedes that Nen=200 breaks sample correspondence and that importance sampling is left to future work, so the joint-posterior claim is the least secure load-bearing element. This partially agrees with the reader's weakest assumption: I agree that no convergence guarantee is the crux, but I locate the failure more specifically in the absent detailed balance and absent joint validation rather than in the NDE proposal accuracy. I would keep the conditional verdict but require a joint-coverage or energy-score test against bilby as a condition, and ideally code release.","tokens_in":13587,"tokens_out":9039,"duration_ms":117517,"concrete_test":"Re-run the 64-injection validation, or at minimum Case 4 with delta_tc = 5 ms, and compare the final hierarchical ensemble against bilby joint PE using joint credible-region coverage: for each pair, form the joint highest-posterior-density region over (theta_A, theta_B) and record whether the true injected pair falls inside at nominal 50%, 90%, and 99% levels. If the hierarchical ensemble's joint coverage deviates from nominal while its marginal P-P plot is uniform, the greedy update has converged to a distribution that is not the joint posterior, and the headline equivalence claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Sections IV and V) is that repeated NDE-based hierarchical subtraction converges, within about 15 iterations, to the same inference as full joint PE. The load-bearing condition is that the stationary distribution of the iterative update is the joint posterior of (theta_A, theta_B). That condition is not met by construction. Section IV states that the Metropolis step has no reverse-jump term and that a proper Gibbs or joint sampler is 'unfeasible' in the current GNPE framework; the algorithm is explicitly 'a greedy algorithm that approximates the joint parameter distribution.' The theoretical support (Eqs. 4-5, Fig. 1) is a linear systematic-bias recursion for maximum-likelihood point estimates, not a statement about the ensemble posterior, and it does not by itself guarantee posterior convergence. The empirical validation is also marginal: Fig. 2 shows four cases of one-dimensional posteriors, and the P-P plot in Fig. 3 is a per-parameter self-consistency check. Because the implementation deliberately keeps only Nen=200 representative parameters per element, cross-source joint structure can be missed even when marginals look reasonable, exactly in the 'confuse sources' regime the paper acknowledges. Thus the claim that the method is as accurate as joint PE is unsupported until the joint posterior is validated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes an iterative hierarchical-subtraction algorithm for Bayesian parameter estimation of overlapping compact-binary signals. Instead of subtracting a single point estimate, the method propagates an ensemble of Nen = 200 candidate parameter sets, subtracts each candidate waveform from the data, and re-infers the companion source with a pre-trained neural density estimator (DINGO), alternating between the two sources. A likelihood-based, temperature-annealed resampling step and a Metropolis-like accept/reject step stabilize the iteration, which is terminated by a Jensen-Shannon divergence criterion. Section II derives a Fisher-matrix linear recursion (Eqs. 4-5) for the systematic error of maximum-likelihood estimates under iterative subtraction and shows numerically that the bias decays with iterations (Fig. 1). Section V presents four case studies at 5 ms separation comparing marginal posteriors with bilby joint PE, plus a 64-pair P-P calibration study spanning 5-50 ms separations. The paper claims that hierarchical subtraction can be made as accurate as joint estimation at a fraction of the cost, offering a potential general solution for overlapping signals in the 3G era, while explicitly acknowledging that the sampler is a greedy approximation without a guaranteed stationary distribution.","tokens_in":13866,"tokens_out":15834,"duration_ms":180691,"significance":"If the central claim holds, this is a timely and practically important contribution: it would make overlapping-signal inference at 3G detector rates feasible with existing single-signal NDEs, avoiding overlap-specific retraining, and the reported cost contrast (a few hours on one CPU plus one GPU versus more than 20000 CPU-core hours for bilby joint runs) is compelling. The paper deserves credit for a genuine first-principles Fisher-matrix analysis of the subtraction dynamics (Section II), for testing the hardest regime of similar masses, similar sky locations, and 5 ms separation, for the 64-pair calibration study in Fig. 3, and for unusually candid statements of its own limitations (no reverse-jump term, Nen = 200, DINGO's imperfect posteriors, untested edge cases). The significance is conditional, however: the evidence establishes a well-calibrated (marginally) approximate sampler, not yet equivalence with joint PE, which is the load-bearing claim in the title and abstract.","major_comments":[{"comment":"The paper's own Section IV states that the Metropolis update has no reverse-jump term and that the algorithm 'should be thought of as a greedy algorithm that approximates the joint parameter distribution,' which is in direct tension with the abstract's claim that hierarchical subtraction 'can achieve accurate results with a sufficient number of iterations.' The theoretical support in Section II, Eqs. (4)-(5) and Fig. 1, is a linear recursion for the systematic bias of maximum-likelihood point estimates, not a statement about the stationary distribution of the stochastic ensemble, so it does not by itself establish posterior-level convergence. At early iterations in case 4 the systematic/statistical ratio reaches about 100, so the linearization behind Eqs. (4)-(5) is also questionable precisely where the recursion is most needed. Please either supply a convergence argument or a bias-correcting reweighting (e.g., an importance-sampling step over the final ensembles) for the sampler, or explicitly re-scope the central claim to empirically validated approximate marginal inference and adjust the abstract and title accordingly.","section":"Sec. II (Eqs. 4-5), Sec. IV"},{"comment":"Equivalence to joint PE is not yet demonstrated quantitatively. Fig. 2 is a visual comparison of one-dimensional marginals in only four cases, and the 64-pair P-P plot in Fig. 3 is a self-consistency test: it checks that the algorithm's own credible intervals cover the injected values at the nominal rate, not that the algorithm's posterior agrees with the bilby joint posterior. A posterior can pass marginal P-P while missing cross-source correlations, which is exactly the 'confuse sources' regime acknowledged in Section VI. I request a quantitative comparison with the bilby joint-PE posterior, for example per-parameter JSD or KL divergences between the hierarchical and bilby OS posteriors for the four case studies, and/or a joint (two-dimensional or higher) coverage test in which the reference distribution is the bilby joint posterior rather than the injected point values. Relatedly, since the method inherits the NDE's single-signal approximation error (the paper itself notes in Section V that DINGO 'does not provide perfectly accurate posterior samples'), the accuracy claim should be stated relative to the NDE's isolated-signal performance.","section":"Sec. V (Figs. 2-3)"},{"comment":"The convergence statistics behind the 'within 15 iterations' claim are not reported. Section IV says 'most cases converge within 15 iterations' and Section V reports 9-16 iterations for the four displayed cases, but the 64-pair study does not give the distribution of iteration counts, the number of pairs that reached the Nit = 30 cap, or the number of pairs for which DoMetropolis never activated. This last point matters because Algorithm 1 declares convergence only when DoMetropolis = 1, so runs that remain in the resampling mode cannot converge by the stated criterion. The JSD stopping rule is also underspecified: the manuscript does not state whether the JSD is computed per parameter or on the joint distribution, over which parameters, or with which density estimator. Please report the full iteration-count distribution and the JSD details.","section":"Sec. IV-V"},{"comment":"The ensemble design limits the method to marginal, not joint, inference. Section VI concedes that Nen = 200 with 80 samples per element breaks the one-to-one correspondence between samples of the two sources, so inter-source correlations are not tracked; this is precisely the information that distinguishes hierarchical subtraction from joint PE in the strongly correlated case 4. In addition, when the NDE 'samples both signals' in the initial draw (Section IV), no mechanism is described that enforces a consistent assignment of the ensemble elements to one source throughout the iterations, and label swaps between the two similar sources would silently corrupt the element-wise subtraction. Please specify and validate the label-assignment mechanism, or state explicitly that only marginal single-source inference is claimed and temper the 'general solution' phrasing accordingly.","section":"Sec. IV, Sec. VI"}],"minor_comments":[{"comment":"Section I contains several wording errors ('detecor's frequency-dependent response,' 'trained for signal signals'), and Fig. 1's caption 'arrivial time' should be 'arrival time'; the detection-rate estimate also differs between Section I ('approximately 10^5 CBCs per year') and Section VI ('10^5 - 10^6').","section":"Sec. I, Fig. 1"},{"comment":"The injection parameters for the four case studies are not tabulated; Fig. 2 shows only axis ranges with grey dotted true values. A table of injection parameters, network SNRs, and noise seeds would be needed to reproduce the results.","section":"Sec. V, Fig. 2"},{"comment":"The exponential temperature-annealing schedule in Section IV and Algorithm 1 is left unspecified (initial T = 10, final T = 1, but no per-iteration rule), and the phase-grid maximization over 100 values needs implementation details (grid bounds and alignment convention).","section":"Sec. IV, Algorithm 1"},{"comment":"Please specify the bilby configuration used as ground truth in Section V (waveform model, PSD, dynesty settings, and prior bounds) and confirm that it matches the DINGO O1 setup, so the comparison is fair.","section":"Sec. V"},{"comment":"In the version under review, the text in Fig. 3 is not legible due to font-encoding artifacts in the PDF; please verify the figure rendering and, if any parameter shows a KS p-value below 0.05, discuss those failures explicitly.","section":"Fig. 3"},{"comment":"The cost statement in Section IV ('it takes ~6s for the DINGO model to draw samples') should specify the batch size and GPU model, since the 20-minutes-per-iteration figure depends on both; releasing the implementation would also strengthen reproducibility beyond Algorithm 1.","section":"Sec. IV"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is more honest than most about its limitations, and the main gap is one of validation rather than of execution: the greedy sampler is acknowledged, yet the abstract and title assert a 'general solution.' I would ask the authors for a quantitative joint-posterior comparison in revision rather than accepting the P-P self-consistency plot as the evidence for equivalence. I also note that the paper builds on the author's previous work (Hu and Veitch 2022, 2023, 2024; Hu et al. 2025), which is cited appropriately, and that no code release is announced; a public implementation would materially help readers assess the method."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper has a real idea: instead of treating hierarchical subtraction as a one-shot process, iterate it with a fast NDE, keep an ensemble of possible signals, and use likelihood-based resampling to steer toward the joint posterior. That combination is new relative to Janquart et al. and Langendorff et al., and the Fisher-matrix recursion in Section II is a useful way to think about why errors might shrink with iterations. The four case studies against bilby joint PE are encouraging, and the author is unusually candid about the sampler being a greedy approximation and about DINGO's own limitations.\n\nThe soft spots are real, though. The central claim that the method recovers parameters 'as accurately as full joint PE' is not supported by the evidence. The theory in Eqs. (4)-(5) is about systematic bias in point estimates, not about convergence of the posterior ensemble. The algorithm is explicitly non-reversible—no reverse jump—so there is no stationary distribution argument that it converges to the joint posterior. The P-P plot in Fig. 3 is only a self-consistency check, not a test of accuracy against the true parameters; what is missing is a comparison of joint posterior coverage against bilby's joint PE, especially in the 'confuse sources' regime. The Nen=200 ensemble is small and the paper itself notes that this breaks one-to-one sample correspondence, so cross-source correlations may be missed. No code or data are released, which makes the empirical claims hard to verify. And the title's 'general solution' overshoots the tested scope: two BBHs with specific masses, O1 sensitivity, and a single NDE.\n\nNone of this is fatal. The idea is plausible, the experiments are a good first step, and the author's honest discussion of limitations is a positive signal. But the paper needs major revision before it can claim equivalence to joint PE. Concretely, I'd want to see a joint-coverage test against bilby, a discussion of when the greedy update can fail, and code release. As is, I'd send it to review if the editor expects the authors to address these points; I would not accept the current version.\n\nWho will get value: anyone working on 3G PE pipelines, ML for GW, or overlapping signal analysis. I'd bring it to a reading group.\n\nRecommendation: send to peer review with the expectation of major revisions. The core contribution is worth the referees' time.","headline":"A genuinely new iterative subtraction scheme for overlapping GW signals, with honest caveats, but the title overclaims and the validation does not yet establish equivalence to joint PE.","tokens_in":14377,"tokens_out":2460,"would_cite":true,"duration_ms":29056,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Overlapping gravitational wave signals can be recovered as accurately as joint estimation by iterating signal subtraction with a fast neural sampler.","keywords":["gravitational waves","overlapping signals","hierarchical subtraction","neural density estimator","conditional normalizing flows","parameter estimation","third-generation detectors","Bayesian inference"],"falsifier":"Apply the pipeline to many simulated overlapping binary black hole pairs with a fixed $5$ ms merger-time separation, compare each recovered per-source posterior to an independent fully joint stochastic sampling run on identical noise realizations, and compute the Jensen–Shannon divergence between the two posterior sets; the central claim fails if the divergence does not approach zero within 30 iterations or if a P–P plot over many realizations shows significant deviations for a substantial fraction of parameters.","tokens_in":13385,"feed_emoji":"🔭","tokens_out":17659,"duration_ms":165995,"temperature":0.7,"pith_summary":"The paper argues that the standard objection to hierarchical subtraction of overlapping gravitational wave signals—that errors from subtracting one imperfectly reconstructed source contaminate the next—does not hold when the subtraction is iterated enough times. The author derives a Fisher-matrix recursion showing the systematic error from the unsubtracted companion shrinks as the two sources are estimated alternately, and implements the loop with a fast neural network trained only to produce posterior samples for isolated signals. Because the neural sampler is cheap, the algorithm subtracts an ensemble of 200 possible versions of each signal rather than one best guess, and a likelihood-based resampling step accelerates convergence. On simulated binary black hole pairs separated by as little as 5 ms in merger time, the recovered posteriors match full joint estimation, which would make parameter estimation of the many overlapping signals expected in third-generation detectors computationally feasible.","feed_headline":"Subtraction loop cracks overlapping gravitational wave signals","feed_subtitle":"Iterating a fast neural-density-estimator sampler matches joint estimation for overlapping signals.","key_machinery":"Two mechanisms carry the argument. The first is the Fisher-matrix error recursion, Eqs. (4)–(5), which expresses the systematic error in one source's parameters after $m$ subtractions as a projection of the residual left by the other source onto the waveform gradient; the residual is either the companion signal itself (at $m=0$) or the linearized error of the companion's previous estimate. This recursion is what shows the bias can vanish with iteration. The second is the sampler: a conditional normalizing flow, trained on isolated signals, maps Gaussian draws to approximate posterior samples $q(\\theta|d)$ fast enough that the loop can draw an ensemble of 200 candidate parameter vectors at each step, subtract the corresponding waveforms, and draw the companion's posterior from the residual. Likelihood-based resampling with a temperature annealed from 10 to 1 prunes low-likelihood ensemble members, and Metropolis-type updates take over near convergence; together these steer the greedy ensemble toward the joint posterior.","core_discovery":"The central claim is that hierarchical subtraction is not inherently inferior to joint estimation for overlapping signals: if the alternating $A\\to B\\to A$ inference loop is run long enough, the systematic bias introduced by the overlapping source decays, so the posterior over each source's parameters converges to the joint posterior. The paper supports this with an iterative Fisher-matrix equation in which the residual left after subtracting one source acts as the driving term for the other source's systematic error, and shows numerically that the systematic-to-statistical error ratio falls below one after a handful of iterations even for nearly identical sources arriving $5$ ms apart. The implemented algorithm then uses a neural density estimator trained on single signals as the sampler, maintains an ensemble of candidate signals for subtraction to avoid committing to a biased point estimate, and alternates likelihood-based resampling with Metropolis-type updates until the Jensen–Shannon divergence between successive ensembles drops below $10^{-4}$ nats. In the paper's simulations convergence typically occurs within 15–20 iterations, a P–P plot over 64 overlapping binary black hole pairs is consistent with calibration, and the author notes that residual differences with the ground truth come from the neural sampler's own approximation rather than from the subtraction loop.","pith_inferences":["The same alternating-subtraction loop could in principle be run for more than two overlapping signals, one extra pass per added source; the paper's tests cover only pairs, so the behavior with three or more sources is an open extension.","A faster, parallelizable normalizing-flow sampler with an explicit flow likelihood—flagged in the paper as future work—would let the ensemble size match the full posterior sample count and turn the greedy loop into a proper Gibbs sampler with a convergence guarantee.","The Fisher-matrix recursion suggests convergence speed depends on how distinguishable the two sources are; quantifying how the required iteration count grows as masses, sky locations, and arrival times become nearly identical would be a natural stress test for third-generation source populations.","The resampling and Metropolis updates effectively use the neural network as a proposal distribution rather than a final answer, which points toward a fully importance-sampling or Markov-chain Monte Carlo-corrected version that removes residual neural density estimator bias."],"forward_implications":["Overlapping signals can be analyzed with a single-signal neural density estimator, so no retraining is needed for each overlap geometry, source type, or time separation.","Parameter estimation for overlapping sources becomes fast enough for third-generation detector data streams, with reported runs taking hours on one GPU rather than tens of thousands of CPU core-hours for joint estimation.","Rare and difficult overlaps—similar masses, similar sky locations, merger times 5 ms apart—are recovered, with inter-source correlations captured by the ensemble iteration.","Because the subtraction uses a resampled ensemble rather than a single best-fit waveform, early biases such as underestimated luminosity distance, wrong sky direction, and overestimated spin magnitudes diminish as the loop continues."],"supporting_citations":[{"why":"The previous simulation study showing hierarchical subtraction improves only 62% of events, the baseline failure mode this work claims to overturn.","marker":"[43]"},{"why":"Introduces the neural posterior estimation architecture that supplies fast approximate posterior samples for single signals in this paper.","marker":"[48]"},{"why":"Shows neural posterior samples can be refined by importance sampling and applied to real events, grounding the paper's reliance on neural density estimators as fast samplers.","marker":"[49]"},{"why":"Provides the maximum-likelihood systematic-error formula on which the paper's iterative Fisher-matrix error recursion is built.","marker":"[51]"},{"why":"Extends the systematic-error formalism that the recursion uses to describe how residual waveform error propagates between alternating estimates.","marker":"[52]"},{"why":"The concrete conditional normalizing flow model trained on isolated binary black hole signals that the paper runs as its proof-of-concept sampler.","marker":"[58]"},{"why":"A previous neural network approach that performs joint estimation for pairs within 50 ms, whose scenario-specific training cost motivates the general subtraction route.","marker":"[60]"},{"why":"The Bayesian inference code used to produce ground-truth posterior distributions, providing the accuracy benchmark the subtraction results are compared against.","marker":"[65]"},{"why":"The nested sampler used by that code, supplying the stochastic sampling that defines the ground-truth comparison.","marker":"[66]"}],"fun_headline_variants":["Iteration unlocks overlapping GW inference via subtraction","Neural density estimators tame overlapping GW signals","Subtraction loop converges, no joint estimation required","Fast overlapping GW analysis with iterated subtraction","General overlapping GW solution from neural density estimators"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method depends on the neural density estimator, trained on non-overlapping signals, still proposing sufficiently good parameter values for a source whose data contain another overlapping signal; if the proposals are too badly biased, the likelihood-based resampling and Metropolis updates cannot pull the ensemble back to the true posterior.","fun_headline_variants_meta":{"raw":{"variants":["Iteration unlocks overlapping GW inference via subtraction","Neural density estimators tame overlapping GW signals","Subtraction loop converges, no joint estimation required","Fast overlapping GW analysis with iterated subtraction","General overlapping GW solution from neural density estimators"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000346,"raw_usage":{"total_tokens":1887,"prompt_tokens":925,"completion_tokens":962,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":895}},"tokens_in":541,"tokens_out":962,"duration_ms":11526,"temperature":1.0,"reasoning_tokens":895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:29:57.468896+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the pipeline to many simulated overlapping binary black hole pairs with a fixed $5$ ms merger-time separation, compare each recovered per-source posterior to an independent fully joint stochastic sampling run on identical noise realizations, and compute the Jensen–Shannon divergence between the two posterior sets; the central claim fails if the divergence does not approach zero within 30 iterations or if a P–P plot over many realizations shows significant deviations for a substantial fraction of parameters.","supporting_citations":[{"cited_title":"+𝑛 𝑑\",$=𝑑−ℎ(𝜃′!,$)ℎ!-subtracteddata 𝑑","cited_arxiv_id":null,"evidence_quote":"The nested sampler used by that code, supplying the stochastic sampling that defines the ground-truth comparison."}],"review_version":1}