{"id":"a9f18b1b-7abb-4b0e-be4e-77336b9f9cc4","arxiv_id":"2412.18259","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A transformer-based model counts and separates up to five overlapping compact binary merger signals in simulated Cosmic Explorer noise, achieving 99.89% counting accuracy and high waveform overlap.","lead":"This paper presents UnMixFormer, a deep learning model that counts how many overlapping gravitational wave signals from merging black holes and neutron stars are present in detector data, then separates each one into its own waveform. It reports high accuracy on synthetic data simulating the future Cosmic Explorer detector, including cases with up to five simultaneous signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline accuracy is in-sample: same Gaussian CE-40km PSD is used for noise generation, whitening, and the overlap metric, so real-detector transfer is the load-bearing unsupported premise.","rationale":"The paper's strongest claim has two parts: (1) high counting and separation performance on the specific synthetic dataset, and (2) that this constitutes a working method for next-generation detectors. Part (1) appears supported by the confusion matrix and overlap histograms, though the mean-overlap headline should be precisely attributed. Part (2) is the load-bearing step: it assumes the synthetic Gaussian CE-40km noise is representative of real detector noise. This is the reader's weakest assumption, and I agree it is the most important unvalidated premise. The same PSD is used for noise generation, whitening, and the overlap metric, so the evaluation is matched to the training distribution; no test with different noise statistics exists. The precession and eccentricity generalization tests address waveform family shift but not noise distribution shift, and their sample sizes are not reported. The concrete test proposed, evaluating on a different PSD, non-stationary noise, and glitches, would directly probe whether the reported numbers survive domain shift. Because the reader's CONDITIONAL verdict already hinges on this assumption, my read does not change the verdict; I would add the out-of-distribution noise test as an explicit condition and request the sample sizes for Fig. 9.","tokens_in":13398,"tokens_out":8229,"duration_ms":76109,"concrete_test":"Re-run the trained model on a test set generated with a different noise PSD (e.g., CE-20km or ET-D), on a non-stationary PSD, and on noise with injected glitches, using the same signal waveforms and SNR range. If counting accuracy or mean overlap drops materially (e.g., below 95% or 0.95), the generalization claim is unsupported. Also state the number of samples for each boxplot in Fig. 9 and recompute the overall mean overlap across all 2-5 signal test samples to verify the abstract's wording.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II.A generates noise as Gaussian with the CE-40km PSD; the same PSD is then used for whitening and for the overlap metric (Eqs. 3-5). The model is therefore trained and evaluated under exactly the noise statistics for which the matched-filter overlap is optimal. The central claim of a 'working method' for next-generation detectors requires that this synthetic Gaussian background be representative of real detector noise, which is non-stationary, contains spectral lines and glitches, and has an imperfectly known PSD. The paper provides no validation on any out-of-distribution noise: all test samples in Figs. 3-7 share the training PSD, and the generalization experiments in Fig. 9 (precession, eccentricity) also use the same Gaussian CE noise while reporting no sample sizes. If the noise model changes, the counting head may mistake glitches for signals and the decoders may output spurious waveforms; thus the 99.89% counting accuracy and 0.9831 overlap are in-sample numbers, not demonstrated transferability. A secondary reporting issue is that the 0.9831 mean overlap is the average for the five-signal subset only, not the overall mean across all 2-5 signal test samples.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents UnMixFormer, an attention-based neural network that combines a counting head with multiple waveform decoders to count and separate 2-5 overlapping compact binary coalescence signals in simulated Cosmic Explorer (CE-40km) noise. The authors report 99.89% counting accuracy and a mean overlap of 0.9831 between separated and target waveforms on a held-out synthetic test set with SNR 10-50. They also show qualitative examples of separation for five overlapping BBH/BNS/NS-BH signals, including cases with spin precession, orbital eccentricity, and higher-order modes, and a single-signal denoising example.","tokens_in":13526,"tokens_out":9451,"duration_ms":81529,"significance":"The problem addressed here is important: next-generation detectors will see overlapping CBC signals, and current matched-filtering and deep-learning methods are largely limited to one or two concurrent sources. The architecture is coherent and the use of a permutation-invariant SI-SNR objective for multi-source separation is well matched to the task. The paper provides a held-out synthetic test set of 20,000 samples, confusion matrices, ROC curves, and per-SNR mismatch analysis, which are useful. However, the evaluation is entirely within a single synthetic Gaussian-noise setting, no baseline comparison is made, and the headline overlap number is reported in an ambiguous way. If the results are confirmed under more realistic noise and compared with existing methods, this would be a valuable contribution to GW data analysis.","major_comments":[{"comment":"The reported mean overlap of 0.9831 is the average of the five per-signal means in the five-signal case (Fig. 4d: 0.9965, 0.9940, 0.9896, 0.9783, 0.9573), not an aggregate over all 2-5 signal test samples. Weighting the per-signal means in Fig. 4(a)-(d) by the number of signals gives an overall mean of approximately 0.989, so the abstract and introduction should either report the correct aggregate or explicitly state that 0.9831 refers only to the five-signal subset.","section":"Abstract and Section III.B"},{"comment":"All experiments use Gaussian noise generated from the CE-40km PSD, and the same PSD is used for whitening and for the overlap metric through the inner product in Eq. (3). The evaluation is therefore carried out under exactly the noise statistics assumed by the matched-filter overlap. The paper provides no test with non-stationary noise, glitches, spectral lines, or a mismodeled PSD, so the conclusion that this is a working method for next-generation detectors is not supported by the experiments. Please add out-of-distribution noise tests and report counting accuracy and overlap under those conditions.","section":"Section II.A and Sections III.A-III.C"},{"comment":"The generalization experiments are not sufficiently specified. Fig. 5 is a single qualitative example; Fig. 9 reports boxplots of mismatch versus precession and eccentricity but does not state the number of test samples, the number of noise realizations, or whether these waveforms were included in training. The single-signal denoising test in Fig. 8 is also ambiguous because the training data only contains 2-5 signals and the counting head has never been trained on one-signal examples; it is not explained how decoder selection works in that test. Similarly, the inspiral-only result in Fig. 6 refers to a retrained model whose dataset composition is not described. Please specify the exact evaluation protocols, sample sizes, and whether the same trained weights are used.","section":"Section III.C, Figs. 5, 8, 9"},{"comment":"The paper claims a substantial advance over existing methods that 'can typically handle only one or two concurrent signals,' but no quantitative baseline is implemented or compared. Without at least one matched-filtering baseline or a previously published deep-learning separation method evaluated on the same test set, the reported counting accuracy and overlap values do not establish the claimed improvement. Please add a baseline comparison on the same data, or temper the claim accordingly.","section":"Section III (overall)"}],"minor_comments":[{"comment":"The phrase 'unknown number of concurrent signals' should be qualified as 'between 2 and 5'; the model is not designed for zero, one, or more than five signals, as stated later in Section II.B.c.","section":"Abstract and Introduction"},{"comment":"The declination δ is sampled uniformly from 0 to π, but physical declination is usually defined in [-π/2, π/2]; this should be clarified or corrected, as it affects the antenna-pattern coverage of the simulated sky.","section":"Section II.A, Table I"},{"comment":"The value of the loss trade-off parameter λ in Eq. (10) is not reported in the implementation details; please provide the value and any sensitivity analysis.","section":"Section II.B.d and Section II.C"},{"comment":"The description of the FAN layer in Eq. (11) is incomplete: the dimensions and roles of W_p, W_{\\bar p}, and B_{\\bar p} are not stated.","section":"Section II.B.b"},{"comment":"The histograms in Fig. 4 show per-signal overlaps sorted in descending order, but the text does not state how many test samples each panel contains or whether the same samples are used for all signal counts; adding this information would improve interpretability.","section":"Section III.B and Fig. 4"},{"comment":"Please correct typographical issues: 'burried' in the Fig. 5 caption, inconsistent spacing in 'F AN', and the use of 'CBSs' in the Introduction should be made consistent.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a gravitational-wave data-analysis journal and addresses a timely problem. The main scientific risk is not the architecture but the evaluation protocol: all results are in-sample with respect to the Gaussian CE noise model, and no baseline comparison is provided. I would encourage the editor to request the noise-robustness experiments and a baseline comparison before publication, as these are feasible within the manuscript's scope. The authors should also be asked to clarify the exact meaning of the headline overlap and to provide code or data for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look, but the headline numbers should be read as in-sample. UnMixFormer extends deep-learning separation of overlapping CBCs to a wider range than prior work: up to five concurrent signals, mixing BBH, BNS, and NS-BH, with a counting head and multiple decoders. The architecture is a reasonable synthesis of established components (dual-path transformer, FAN, multi-decoder), and the synthetic evaluation is internally consistent: confusion matrix, ROC, and overlap histograms all tell the same story. That is real, if narrow, progress.\n\nWhere it gets soft: first, the evaluation is entirely under Gaussian noise generated with the CE-40km PSD, and the same PSD is used for whitening and for the overlap metric. That means the numbers are best-case for exactly the noise statistics the metric is optimal for. There is no test on non-stationary noise, glitches, or an imperfectly known PSD, so the transfer claim to real detectors is not supported. Second, the paper reports a mean overlap of 0.9831 but that is the five-signal subset average, not the overall mean across 2-5 signals; the histograms show per-signal means ranging from 0.9973 down to 0.9573, so the headline is flattering. Third, there is no baseline comparison at all, despite citing several related methods; you cannot tell whether the architecture is actually better than a simpler transformer or even a well-tuned matched-filter bank. Fourth, the generalization sections for precession and eccentricity show only a handful of examples with no sample sizes or error bars. Fifth, no code, data, or weights are released, so exact reproduction is impossible.\n\nNone of this kills the paper. The core counting and separation results on synthetic data are credible, and the method is a legitimate step toward handling the overlap problem in next-generation detectors. But the conditions for accepting it as anything more than a proof-of-concept should be a baseline comparison, a clearer evaluation protocol (report the overall mean overlap, not the cherry-picked subset), and a release of the implementation and data.\n\nFor peer review: send it out. The problem matters, the method is new in scope, and the internal evidence is good enough to justify referee time. The outcome should be a major revision, not a desk reject.","headline":"Worth reading as a proof-of-concept, but the headline numbers are in-sample and the absence of baselines means the real advance is still unquantified.","tokens_in":14145,"tokens_out":2465,"would_cite":true,"duration_ms":21675,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"UnMixFormer achieves 99.89% counting accuracy and 0.9831 mean overlap in separating up to five overlapping compact binary coalescence signals.","keywords":["gravitational waves","compact binary coalescence","overlapping signals","signal separation","source counting","deep learning","transformer","Cosmic Explorer"],"falsifier":"Run the trained model on a segment of real or simulated non-Gaussian detector noise containing glitches and non-stationary transients with injected overlapping signals, and check whether counting accuracy and mean overlap remain close to 99.89% and 0.9831; a substantial drop would show the reported numbers depend on the Gaussian-noise assumption.","tokens_in":13092,"feed_emoji":"🔭","tokens_out":3950,"duration_ms":32639,"temperature":0.7,"pith_summary":"The paper introduces UnMixFormer, a neural network designed for next-generation gravitational-wave detectors where overlapping signals from multiple compact binary mergers will be common. The central claim is that one architecture can both count how many signals are present (between two and five) and separate the individual waveforms, reaching 99.89% counting accuracy and 0.9831 mean overlap on synthetic data embedded in Cosmic Explorer Gaussian noise. The authors argue this moves beyond existing methods that handle only one or two concurrent signals and limited waveform families. A sympathetic reader would care because overlapping signals will bias parameter estimation if treated as single events, so a working count-and-separate tool would directly improve astrophysical inference in the third-generation era.","feed_headline":"One network counts and unmixes up to five overlapping mergers","feed_subtitle":"UnMixFormer hits 99.89% counting accuracy and 0.9831 waveform overlap on synthetic Cosmic Explorer noise.","key_machinery":"The central mechanism is a dual-path transformer with intra- and inter-segment attention, augmented by Fourier Analysis Networks (FAN) in place of MLP feed-forward layers, paired with a multi-decoder selector. The counting head outputs a probability over signal multiplicity and activates the matching decoder; separation is trained with a permutation-invariant SI-SNR loss plus cross-entropy on the count, so the same model both estimates the number of sources and reconstructs each coherent waveform from the whitened one-second, 16,384-sample input.","core_discovery":"The discovery is that jointly optimizing a counting head and a bank of decoders within a dual-path attention architecture yields accurate source counting and waveform separation for up to five overlapping compact binary coalescence signals, across BBH, BNS, and NS-BH systems, at SNRs between 10 and 50. On held-out synthetic data, the model achieves 99.89% counting accuracy (with 100% accuracy within ±1 signal), AUC values above 0.9999 for counting each multiplicity, and a mean overlap of 0.9831 between reconstructed and template waveforms; mismatch degrades gracefully as the number of signals increases. The model also separates inspiral-only segments and generalizes to waveforms with spin precession, orbital eccentricity, and higher-order modes, which were not present in its training set.","pith_inferences":["Editors' inference: if the Gaussian-noise assumption holds, the same architecture could be adapted to multi-detector networks by treating each detector as a channel, which the authors note but do not implement; spatial diversity would likely improve localization as well as separation.","Editors' inference: the counting head is trained only for 2–5 signals, so on real data a single loud event would need to be handled separately or the model retrained with a '0 or 1' class; the denoising example hints at capability but does not test counting at multiplicity 1.","Editors' inference: a direct stress test would be to run the model on LIGO-Virgo O4 data with glitches and non-stationary noise; the resulting drop in counting accuracy would quantify how much of the reported performance depends on ideal Gaussian noise."],"forward_implications":["Third-generation detectors could use a single trained model to count and separate overlapping BBH, BNS, and NS-BH signals in about 1.5 ms per sample, making real-time analysis feasible.","Because the model separates waveforms rather than fitting one template to the mixture, it could reduce the parameter-estimation bias that overlapping events introduce when treated as single signals.","The approach extends to inspiral-only long-duration data, relevant to the longer signals expected in next-generation detectors.","Generalization to precessing, eccentric, and higher-mode waveforms suggests the representation learned on simpler aligned-spin templates captures enough structure to handle more complex physics without retraining."],"supporting_citations":[{"why":"Supplies the Cosmic Explorer 40 km PSD used to generate Gaussian noise for training and testing.","marker":"[5]"},{"why":"PyCBC library is used to simulate GW signals and embed them in detector noise.","marker":"[51]"},{"why":"Provides the SEOBNRv4 waveform template for BBH signals.","marker":"[52]"},{"why":"Provides the IMRPhenomT waveform template for NS-BH signals.","marker":"[53]"},{"why":"Provides the TaylorF2 waveform template for BNS signals.","marker":"[54]"},{"why":"The multi-decoder DPRNN architecture that the counting-and-separation design is inspired by.","marker":"[58]"},{"why":"Introduces Fourier Analysis Networks, the periodic-feature layer used inside the transformer blocks.","marker":"[59]"},{"why":"Defines the SI-SNR loss with permutation-invariant matching used to train the separators.","marker":"[60]"}],"fun_headline_variants":["UnMixFormer separates up to five overlapping gravitational-wave signals","Counts and unmixes five overlapping mergers with 99.89% accuracy","One dual-path network tackles five overlapping GW signals at once","99.89% accurate counting for five overlapping compact binary mergers","UnMixFormer: five-way GW separation with 0.9831 waveform overlap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Gaussian noise generated from the Cosmic Explorer 40 km power spectral density accurately represents real next-generation detector noise, including its non-stationarity and glitches, so that the reported counting and separation accuracy would transfer to actual observations.","fun_headline_variants_meta":{"raw":{"variants":["UnMixFormer separates up to five overlapping gravitational-wave signals","Counts and unmixes five overlapping mergers with 99.89% accuracy","One dual-path network tackles five overlapping GW signals at once","99.89% accurate counting for five overlapping compact binary mergers","UnMixFormer: five-way GW separation with 0.9831 waveform overlap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001022,"raw_usage":{"total_tokens":4308,"prompt_tokens":938,"completion_tokens":3370,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":3279}},"tokens_in":554,"tokens_out":3370,"duration_ms":21696,"temperature":1.0,"reasoning_tokens":3279,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:52:41.985518+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained model on a segment of real or simulated non-Gaussian detector noise containing glitches and non-stationary transients with injected overlapping signals, and check whether counting accuracy and mean overlap remain close to 99.89% and 0.9831; a substantial drop would show the reported numbers depend on the Gaussian-noise assumption.","supporting_citations":[{"cited_title":"Dal Canton, A","cited_arxiv_id":null,"evidence_quote":"PyCBC library is used to simulate GW signals and embed them in detector noise."},{"cited_title":"These parameters are then used to generate the corre- sponding waveforms for each source","cited_arxiv_id":null,"evidence_quote":"Provides the SEOBNRv4 waveform template for BBH signals."},{"cited_title":"Gravitational Wave Mixture Separation for Future Gravitational Wave Observatories Utilizing Deep Learning","cited_arxiv_id":"2407.13239","evidence_quote":"Provides the IMRPhenomT waveform template for NS-BH signals."},{"cited_title":"Boh´ e, L","cited_arxiv_id":null,"evidence_quote":"Provides the TaylorF2 waveform template for BNS signals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The multi-decoder DPRNN architecture that the counting-and-separation design is inspired by."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces Fourier Analysis Networks, the periodic-feature layer used inside the transformer blocks."},{"cited_title":"Multi-Decoder DPRNN: High Accuracy Source Counting and Separation","cited_arxiv_id":"2011.12022","evidence_quote":"Defines the SI-SNR loss with permutation-invariant matching used to train the separators."}],"review_version":1}