Pith. sign in

REVIEW 4 major objections 4 minor 74 references

Combining gravitational wave search pipelines to find subthreshold signals in GWTC-5.0

T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Combining gravitational-wave search pipelines with calibrated machine learning up-ranks sub-threshold mergers as genuinely signal-like.

desk verdict Solid extension of the Ashton et al. pipeline-combination work with new O4 catalogue results, but the real-data application is ambiguous about whether the model was retrained on the reduced feature set, and that ambiguity matters for the calibration claim. read the letter →

arxiv 2607.07272 v2 pith:SYK6VOLH submitted 2026-07-08 gr-qc astro-ph.HEastro-ph.IM

classification gr-qcastro-ph.HEastro-ph.IM
keywords gravitationalwavesconformalpredictionmachinelearningsearchpipelinecombinationsubthresholdcandidatesbinaryneutronstarsconditionalconfidencetransientcatalogues
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the noise in individual gravitational-wave search pipelines can be tamed by merging their outputs: a classifier trained on labelled mock data learns that genuine signals are ones seen coherently by several pipelines, and conformal prediction converts the classifier's raw output into a single calibrated confidence score for each candidate. The authors show this combination is stable across three classifier architectures and two mock training sets, and that it beats the standard 'take the most significant pipeline' rule on false-positive control. Applied to real candidates from the third and fourth observing runs, the method gives elevated confidence to several events that individual pipelines or the usual astrophysical-probability threshold ranked as marginal, most notably the binary neutron star candidate GW200311_103121. Follow-up checks—spectrograms, parameter estimation, and feature-attribution analysis—are consistent with these up-ranked events being astrophysical. If true, the framework would give astronomers a single, interpretable number for triage, including for low-latency alerts and electromagnetic follow-up.

What carries the argument

Conditional confidence from label-conditional conformal prediction applied to a supervised classifier. The classifier maps a candidate's multi-pipeline feature vector (inverse false-alarm rate, signal-to-noise ratio, chirp mass per pipeline, zero-filled for non-detections) to a signal probability; conformal prediction calibrates this output on a holdout set and returns, for each candidate, the largest error rate at which the signal label lies in the prediction set. That single number is the paper's significance measure, and the argument's load-bearing step is the exchangeability of the mock calibration data with the real candidates.

What would settle it

Take a fresh mock of the fourth observing run built with a realistic injection rate and noise extending below the one-hour false-alarm floor, train and calibrate the same classifiers, and check coverage: if the empirical rate at which true signals receive conditional confidence above 0.5 differs significantly from the roughly 50 percent the calibration implies, the real-candidate scores—including GW200311_103121—are not trustworthy. A complementary empirical falsifier would be targeted data-quality follow-up of the up-ranked candidates: if a substantial fraction turn out to have instrumental o

Watch

Extended reading notes

Core claim

At the centre of the paper is a claim that the pattern of which pipelines see a candidate, and how strongly, carries information that no single pipeline's significance captures. The authors train binary classifiers on mock data with known signal/noise labels, using per-candidate feature vectors built from each pipeline's inverse false-alarm rate, signal-to-noise ratio, and chirp mass (zeros for non-detections), then apply label-conditional conformal prediction to obtain a 'conditional confidence' in [0,1] — the largest error rate at which the signal label remains in the prediction set. They report that all classifiers outperform the maximum-inverse-false-alarm-rate approach in ROC terms and

Load-bearing premise

The calibration guarantee of conformal prediction, and therefore the meaning of every confidence score on real candidates, rests on the assumption that the labelled mock candidates used for training and calibration are exchangeable with the real candidates — an assumption the authors concede is violated by inflated injection rates, evolving pipelines, and a truncated noise distribution in their main mock dataset.

Editorial extensions

If this is right

  • One calibrated confidence number per candidate can replace the maximum-inverse-false-alarm-rate / astrophysical-probability combination for catalogues and alerts, with fewer false positives at the same operating threshold.
  • Sub-threshold multi-pipeline events, especially binary neutron star and neutron star-black hole candidates, may be astrophysical despite astrophysical probability below 0.5; confirming GW200311_103121 would make it only the third binary neutron star merger.
  • Real-time application of the framework could up-rank a marginal candidate quickly enough to trigger electromagnetic follow-up and multi-messenger science.
  • The method extends to new pipelines, new classifiers, additional features, and multi-class labels, provided labelled mock data exists for calibration.
  • Down-rankings by the most conservative classifier also flag high-significance single-pipeline events that later turn out instrumental, as with a known scattered-light event from the fourth observing run.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper is right, the field's current threshold-based cataloguing may systematically miss a population of distant, coherent multi-pipeline mergers whose individual signal-to-noise ratios are weak; re-running the framework on future observing runs would test whether the up-ranked candidates' distances extend the known horizon.
  • The authors' own exchangeability caveat suggests a decisive next experiment: repeat training and calibration on a mock dataset whose injection rate and false-alarm floor match real operations; persistence of the specific up-rankings (GW200311_103121 and the fourth-run candidates) would raise confidence substantially.
  • The two same-day pairs of up-ranked candidates from the later part of the fourth observing run are an oddity worth investigating: if either pair is confirmed, the coincidence rate itself would be a signal about the underlying astrophysical population or about correlated detector artifacts.
  • Because the method's up-rankings are driven by 'seen by several pipelines,' it may be particularly vulnerable to common-mode detector or analysis errors; a targeted study injecting correlated glitches across pipelines could bound this vulnerability.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper extends the conformal-prediction pipeline-combination framework of Ashton et al. to the full GWTC catalogue (O3, O4a, O4b). A binary classifier (LR, MLP, KNN, XGBoost) maps per-pipeline outputs to a signal probability; Mondrian conformal prediction converts this into a conditional confidence score. The framework is validated on the MDC mock-data challenge and the LLPIC dataset, including an 'unseen test' design in which MDC-trained models are evaluated on LLPIC. The authors find that the ML+CP classifiers outperform maximum-log10IFAR in false-positive control, and then apply the framework to real events, identifying several sub-threshold candidates with elevated confidence, most notably GW200311_103121, plus a set of O4 BBH candidates. They support these up-rankings with spectrograms, parameter estimation, and SHAP analysis, while acknowledging the exchangeability limitation.

Significance. If the central claim is correct, the paper provides a practical, calibrated significance measure for LVK candidates and a template for combining heterogeneous pipeline outputs. The validation is well designed: ROC comparisons across four classifiers, an independent LLPIC test set, and explicit false-positive counts. The paper is also unusually candid about the limits of the MDC and the meaning of CP under distribution shift. However, the headline 'well-calibrated' claim has not been established for the real-data feature set: the validation is performed on a 21-feature space while the application uses a 3-feature space, and the CP calibration guarantee does not transfer across feature spaces. The physical evidence for the up-ranked candidates is suggestive but not conclusive.

major comments (4)
  1. [Appendix A1, Section II (Figs. 2, 3, 7, 10)] The abstract and Section IV describe the confidence scores as 'well-calibrated', but the calibration is validated on a different feature space from the one used in the real-data application. Appendix A1 states that 'For mock data studies, we use all 21 features per pipeline, where available' and that 'When applying our method to the new candidates ..., we restrict the features to the log10(IFAR), SNR, and chirp mass.' The paper never states that the ML models and CP quantiles were retrained and recalibrated on the restricted feature set. If the full-feature model is fed real candidates with zero-filled missing features, the CP quantiles from the full-feature MDC calibration do not apply. If a restricted-feature model was used, the ROC/AUC/FP validation in Figs. 2–3 and Table I is for a different model from the one producing the O4 rankings. This is load-bearing for the calibration claim
  2. [Section IV, Introduction] The authors correctly concede in Section IV that CP exchangeability 'cannot be guaranteed' and that the MDC has 'known limitations' — inflated injection rate, pipelines under development, and a minimum IFAR threshold. This means the conditional confidence on real O4 candidates is not strictly calibrated. The LLPIC 'unseen test' (Section IIB2) is a valuable robustness check, but it is still a mock-mock transfer; it does not provide a calibration guarantee for real data. I recommend softening 'well-calibrated' in the abstract and Section IV to 'empirically robust' or similar, and explicitly labelling the real-data confidences as conditional on the exchangeability assumption.
  3. [Section IIIB, IV, Appendix B4] The claim that 'high-confidence predictions correspond to signal-like events' is only partially supported. The SHAP analysis of up-ranked real candidates uses the same XGBoost model that produced the rankings, so it is not an independent confirmation; it shows that the model's decision is internally coherent. The spectrograms (Figs. 8, 11) are subjective at SNR 7–8, and parameter estimation (Figs. 9, 12) cannot by itself distinguish a weak signal from a noise transient that masquerades as a CBC. Section IIIB correctly notes that PE cannot determine whether a candidate is real, but the cumulative wording in the discussion goes beyond this. Please either add an independent validation step (e.g., injections in real O4 noise at similar SNR) or temper the claim.
  4. [Section IIB2, Table I, Fig. 5] XGBoost, chosen as the preferred classifier, has the lowest TPR (53.8%) and the largest number of down-ranked high-IFAR signals under the LLPIC distribution shift (Table I, Fig. 5). The text says XGBoost's down-rankings 'warrant caution' but then uses XGBoost exclusively for the O4 claims. This asymmetry should be justified more explicitly, since a large fraction of catalogue events are assigned below-threshold confidence by XGBoost; if these are wrongly down-ranked, the framework's single-measure significance is not as reliable as the 'well-calibrated' language suggests.
minor comments (4)
  1. [Tables XI and XII captions] Typographical: 'parameters shows' should be 'parameters shown'; Table IX caption has 'maximump astro' with an erroneous space. These should be corrected in the final version.
  2. [Appendix B3] The threshold of 0.5 is used throughout; the two data-driven criteria give different thresholds (0.43–0.63) and orderings across classifiers. The choice is reasonable as a convention but should be presented as a pragmatic operating point, not a calibrated optimum.
  3. [Appendix B2, Fig. 15] The statement that the sensitivity curve is 'necessarily diagonal' is potentially confusing. It follows from the CP validity guarantee and the definition of conditional confidence, but a one-sentence explanation of why the diagonal is a validity check rather than a performance metric would help the reader.
  4. [General] KNN is excluded from the CP analysis because of discrete outputs, yet the paper later speaks of 'different classifier architectures' in the robustness check. The exclusion is reasonable, but it should be stated more prominently in the main text so that the comparison is not over-read as covering all four classifiers.

Circularity Check

1 steps flagged · score 3.0 of 10

Core ML+CP pipeline is not circular; the one circular leg is the SHAP-based validation of up-rankings, which reduces to the model's own output by Eq. (B3).

  1. self definitional [Section IV (Discussion and Conclusion), validation paragraph; Appendix B4, Eq. (B3)]
    "Second, SHAP analysis of the up-ranked events confirms they exhibit signal-like characteristics... SHAP decomposes the output into a baseline value E[f(x)] ... plus additive feature contributions: f(x)=E[f(x)]+Σ_i φ_i"

    Equation (B3) defines SHAP values as an additive decomposition that sums exactly to the model output f(x). For any event the model up-ranks, the SHAP waterfall must, by construction, show contributions pushing toward the signal label. Therefore the statement that SHAP 'confirms' the events are signal-like is a restatement of the model's own confidence, not an independent test. The paper cites this as one of two main consistency checks supporting the reliability of the up-rankings ('Taken together, these consistency checks provide evidence that the up-ranked events correspond to genuine signal-like candidates'), so one load-bearing validation leg is self-referential. The central derivation is not circular because the classifier is trained only on labeled mock data and frozen when applied to

full rationale

The central claim does not reduce to its inputs. A supervised classifier is trained and CP-calibrated on labelled mock data (MDC/LLPIC) and then applied unchanged to real GWTC candidates; no real-event labels or parameters are fitted, so the up-rankings are genuine out-of-sample outputs. The LLPIC-as-unseen-test experiment provides an external distribution-shift check, and the paper's own MDC/LLPIC ROC and confusion-matrix results are independent benchmark comparisons. Reliance on the authors' prior work [44] is substantial but this paper re-derives the method with new classifiers, new mock data, and new real-data applications, so it is not a self-citation chain that forces the result. The feature-set mismatch between mock validation (21 features) and real-data application (IFAR, SNR, chirp mass) is a real correctness/exchangeability risk, but it is not circularity: it does not make the prediction equal to the input by construction. The one identifiable circular step is the SHAP-based 'validation': SHAP values are defined to sum to the model's own output, so using them to confirm that up-ranked events are signal-like is tautological with the model's decision. This affects one supporting argument, not the derivation of the confidence scores themselves. Score 3 reflects one partial, non-central circular validation step alongside an otherwise self-contained analysis.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

No new physical entities are introduced. The paper's contribution is a statistical score (conditional confidence) built from existing pipeline outputs; its validity rests on exchangeability and on the fidelity of mock-data labels rather than on any new physical postulate.

free parameters (4)
  • Conditional confidence threshold = 0.5
    Used throughout to define up-ranked vs down-ranked candidates. Appendix B3 shows it is a compromise between precision-based and ROC-based thresholds, but the final value is hand-selected.
  • XGBoost n_estimators and max_depth = 96 trees, depth 6
    Chosen by cross-validated grid-search on MDC. These hyperparameters affect the model output and therefore the real candidate confidences.
  • MLP hidden-layer size = 100 neurons
    Selected after finding only minor improvements with deeper models; affects MLP confidence scores.
  • KNN k = 19
    Chosen by cross-validation; KNN is excluded from CP and real-data analysis but enters the ROC comparison.
assumptions (4)
  • domain assumption Training/calibration mock data are exchangeable with real GWTC data
    This is the central CP validity condition, acknowledged in Section IV as not strictly guaranteed. The paper probes it with LLPIC but cannot prove it.
  • domain assumption Injected-signal labels and injection-to-candidate matching in MDC/LLPIC are correct
    The supervised classifier and CP calibration depend entirely on the ground-truth labels assigned to mock candidates; incorrect matching would bias the learned mapping.
  • standard math Mondrian conformal prediction provides valid class-conditional coverage under exchangeability
    The coverage guarantee in Eq. (A10) is standard CP theory and is relied on for the calibrated confidence interpretation.
  • domain assumption Pipeline outputs (FAR/IFAR, SNR, chirp mass) are accurate and comparable across observing runs
    The feature vectors combine GstLAL, PyCBC, MBTA, and cWB outputs from different pipeline versions; inconsistencies would break the mapping learned from mock data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Combining gravitational wave search pipelines to find subthreshold signals in GWTC-5.0." pith.science (2026). https://pith.science/paper/SYK6VOLH

@misc{pith2026260707272,
  author       = {Pith},
  title        = {Pith review of: Combining gravitational wave search pipelines to find subthreshold signals in GWTC-5.0},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SYK6VOLH}},
  note         = {Machine review of arXiv:2607.07272}
}
read the original abstract

The detection of transient gravitational wave signals relies on independent search algorithms that analyse detector data and assign significance measures to candidate events. However, varying performance complicates their interpretation. We use supervised machine learning combined with conformal prediction, a framework to quantify uncertainties, to merge multi-pipeline information into well-calibrated confidence scores. We demonstrate that this approach is robust across different classifier architectures and remains stable when trained on different simulated datasets. When applied to events across the GWTC catalogue up to and including the second part of the fourth observing run, the framework identifies several subthreshold candidates with elevated confidence, including the binary neutron star candidate GW200311_103121. We examine the reliability of these up-rankings, finding evidence that high-confidence predictions correspond to signal-like events. This framework enables simplified systematic candidate assessment for gravitational wave catalogues and real-time alerts by providing a single, well-calibrated confidence measure per candidate.

Figures

Figures reproduced from arXiv: 2607.07272 by the authors.

Figure 1
Figure 1. FIG. 1. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. ROC curve comparing the performance of four different [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: FIG. 4. The conditional confidence in the signal label for the O3 events in the GWTC-2.1 and GWTC-3.0 catalogue as obtained [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5. The conditional confidence in the signal label for the [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8. Time-frequency spectrograms for some of the up-ranked O4a subthreshold candidates in GWTC-4.1. The dashed white [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 10
Figure 10. Figure 10: FIG. 10. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 9
Figure 9. Figure 9: FIG. 9. Violin plots of estimated posterior distributions for [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 11
Figure 11. Figure 11: FIG. 11. Time-frequency spectrograms for some of the up-ranked O4b subthreshold candidates in GWTC-5.0. The dashed white [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: FIG. 12. Violin plots of estimated posterior distributions [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: FIG. 13. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: FIG. 14. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p020_14.png]
Figure 14
Figure 14. Figure 14: FIG. 14. The conditional confidence in the signal label for [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]
Figure 15
Figure 15. Figure 15: FIG. 15. Sensitivity for the [PITH_FULL_IMAGE:figures/full_fig_p023_15.png]
Figure 16
Figure 16. Figure 16: FIG. 16. Precision curves for the [PITH_FULL_IMAGE:figures/full_fig_p024_16.png]
Figure 17
Figure 17. Figure 17: FIG. 17 [PITH_FULL_IMAGE:figures/full_fig_p024_17.png]
Figure 19
Figure 19. Figure 19: (a)), but this is rare in the MDC, partly because 41% of injections are BNS signals that cWB is not well suited to detect. We can also examine feature importance for individ￾ual events to understand the reasoning behind a specific prediction. Focusing on the up-ranked…
Figure 20
Figure 20. Figure 20: FIG. 20. Comparison of mean [PITH_FULL_IMAGE:figures/full_fig_p026_20.png]
Figure 21
Figure 21. Figure 21: FIG. 21 [PITH_FULL_IMAGE:figures/full_fig_p027_21.png]
Figure 22
Figure 22. Figure 22: FIG. 22 [PITH_FULL_IMAGE:figures/full_fig_p028_22.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

74 extracted references · 3 linked inside Pith

  1. [1]

    From the template- based pipelines, we use the reportedFAR, SNR, and χ2 (a matched-filter signal-consistency test [11, 25])

    Data preparation The pipeline outputs include both significance statis- tics and source parameter estimates. From the template- based pipelines, we use the reportedFAR, SNR, and χ2 (a matched-filter signal-consistency test [11, 25]). These pipelines also provide estimates of source parameters from the nearest waveform template. We record the detector- fra...

  2. [2]

    Here, we explore addi- tional classifiers and compare the performance of LR, MLP, K-nearest neighbours (KNN), and XGBoost, a tree-based method

    Machine learning classifiers In our previous work [44], we focused primarily on LR due to its interpretability. Here, we explore addi- tional classifiers and compare the performance of LR, MLP, K-nearest neighbours (KNN), and XGBoost, a tree-based method. We implement the first three using scikit-learn [82], and use theXGBoost package [66] for the latter....

  3. [3]

    Thus,CP requires no assumptions about the model or data dis- tributions

    Conformal prediction CP [45, 46] is a framework developed to provide sta- tistically rigorous and well-calibrated uncertainties for any point prediction.CP does not modify the underlying algorithm but instead uses its predictions on a labelled calibration dataset to learn the uncertainty. Thus,CP requires no assumptions about the model or data dis- tribut...

  4. [4]

    This allows us to move beyond black-box predictions and understand the physical reasoning behind each classification

    Feature importance Having established XGBoost as our preferred classi- fier, we now examine the decision-making process of all classifiers throughSHAP analysis [68], using XGBoost as the primary example.SHAP analysis quantifies the con- tribution of each input feature to individual predictions, revealing which pipeline outputs drove a given prediction by ...

  5. [35]

    Machine-learningpipelineforreal-time detection of gravitational waves from compact binary coalescences.Phys

    EthanMarxetal. Machine-learningpipelineforreal-time detection of gravitational waves from compact binary coalescences.Phys. Rev. D, 111(4):042010, 2025. doi: 10.1103/PhysRevD.111.042010

  6. [36]

    Vasileios Skliris, Michael R. K. Norman, and Patrick J. Sutton. Toward real-time detection of unmodeled grav- itational wave transients using convolutional neural networks.Phys. Rev. D, 110(10):104034, 2024. doi: 10.1103/PhysRevD.110.104034

  7. [37]

    Nitz, Sumit Kumar, Yi-Fan Wang, et al

    Alexander H. Nitz, Sumit Kumar, Yi-Fan Wang, et al. 4-OGC: Catalog of Gravitational Waves from Compact Binary Mergers.Astrophys. J., 946(2):59, 2023. doi: 10.3847/1538-4357/aca591

  8. [38]

    New binary black hole mergers in the LIGO-Virgo O3b data.Phys

    Ajit Kumar Mehta, Seth Olsen, Digvijay Wadekar, Javier Roulet, Tejaswi Venumadhav, Jonathan Mushkin, Barak Zackay, and Matias Zaldarriaga. New binary black hole mergers in the LIGO-Virgo O3b data.Phys. Rev. D, 111 (2):024049, 2025. doi:10.1103/PhysRevD.111.024049

Show all 74 references
  1. [39]

    Abbott et al

    R. Abbott et al. Gwtc-3: Compact binary coalescences observed by ligo and virgo during the second part of the third observing run.Phys. Rev. X, 13:041039, Dec

  2. [40]

    A. G. Abac et al. GWTC-4.0: Updating the Gravitational-Wave Transient Catalog with Observations from the First Part of the Fourth LIGO-Virgo-KAGRA Observing Run. 8 2025

  3. [41]

    Sharan Banagiri, Christopher P. L. Berry, Gareth S. Cabourn Davies, Leo Tsukada, and Zoheyr Doc- tor. Unified pastro for gravitational waves: Con- sistently combining information from multiple search pipelines.Phys. Rev. D, 108(8):083043, 2023. doi: 15 10.1103/PhysRevD.108.083043

  4. [42]

    Coughlin, Man Leong Chan, and Leo Singer

    Seiya Tsukamoto, Andrew Toivonen, Holton Griffin, Avyukt Raghuvanshi, Megan Averill, Frank Kerkow, Michael W. Coughlin, Man Leong Chan, and Leo Singer. Astrophysical or Terrestrial: Machine Learning Classifi- cation of Gravitational-wave Candidates Using Multiple- search Infor...

  5. [43]

    The best of all worlds: gravitational-wave transient detection with multiple pipelines

    Nikolas Moustakidis, Theofilo Moustakidis, Deep Chat- terjee, Anastasios Tefas, and Erik Katsavounidis. The best of all worlds: gravitational-wave transient detection with multiple pipelines. 2026. in preparation

  6. [44]

    Enhancing Gravitational-Wave Detection: A Machine Learning Pipeline Combination Approach with Robust Uncertainty Quantification.Phys

    Gregory Ashton, Ann-Kristin Malz, and Nicolo Colombo. Enhancing Gravitational-Wave Detection: A Machine Learning Pipeline Combination Approach with Robust Uncertainty Quantification.Phys. Rev. Lett., 136(1): 011402, 2026. doi:10.1103/yfb3-fgf2

  7. [45]

    Springer, 2005

    Vladimir Vovk, Alexander Gammerman, and Glenn Shafer.Algorithmic learning in a random world, vol- ume 29. Springer, 2005

  8. [46]

    A gentle introduction to conformal prediction and distribution-free uncertainty quantification.arXiv preprint arXiv:2107.07511, 2021

    Anastasios N Angelopoulos and Stephen Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification.arXiv preprint arXiv:2107.07511, 2021

  9. [47]

    Calibrating gravitational-wave search algorithms with conformal prediction.Phys

    Gregory Ashton, Nicolo Colombo, Ian Harry, and Surabhi Sachdev. Calibrating gravitational-wave search algorithms with conformal prediction.Phys. Rev. D, 109 (12):123027, 2024. doi:10.1103/PhysRevD.109.123027

  10. [48]

    Redshift Prediction with Images for Cosmology Using a Bayesian Convolutional Neural Network with Conformal Predictions.Astrophys

    Evan Jones, Tuan Do, Yun Qi Li, Kevin Alfaro, Jack Sin- gal, and Bernie Boscoe. Redshift Prediction with Images for Cosmology Using a Bayesian Convolutional Neural Network with Conformal Predictions.Astrophys. J., 974 (2):159, October 2024. doi:10.3847/1538-4357/ad6d5a

  11. [49]

    Improving Generalization and Uncertainty Quantification of Photometric Redshift Models.Astron

    Jonathan Soriano, Tuan Do, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Jack Singal, and Evan Jones. Improving Generalization and Uncertainty Quantification of Photometric Redshift Models.Astron. J., 171(2):114, 2026. doi:10.3847/1538-3881/ae2ffe

  12. [50]

    Uncertainty Estimation in Classification of Massive Stars Spectra using Conformal Predictions

    Raquel Pezoa, Luis Salinas, Claudio Torres, Felipe Ortiz, Michel Cure, and Ignacio Araya. Uncertainty Estimation in Classification of Massive Stars Spectra using Conformal Predictions. In2024 43rd Interna- tional Conference of the Chilean Computer Science Society (SCCC), page ...

  13. [51]

    Uncertainty quan- tification of the virial black hole mass with conformal prediction.Mon

    Suk Yee Yong and Cheng Soon Ong. Uncertainty quan- tification of the virial black hole mass with conformal prediction.Mon. Not. Roy. Astron. Soc., 524(2):3116– 3129, 2023. doi:10.1093/mnras/stad2080

  14. [52]

    Marlon M. S. Mendes, Roberta Duarte Pereira, Mariana Dutra da Rosa Louren, and César H. Lenzi. Certified Uncertainty for Surrogate Models of Neutron Star Equa- tions of State via Mondrian Conformal Prediction. 2 2026

  15. [53]

    Williams, and Sujit Ghosh

    Naomi Singer, Jonathan P. Williams, and Sujit Ghosh. Conformal prediction for astronomy data with measurement error.Monthly Notices Royal Astro- nomical Society, 539(2):1372–1380, May 2025. doi: 10.1093/mnras/staf515

  16. [54]

    Araz and Michael Spannowsky

    Jack Y. Araz and Michael Spannowsky. Another Fit Bites the Dust: Conformal Prediction as a Calibration Standard for Machine Learning in High-Energy Physics. 12 2025

  17. [55]

    Classification uncertainty for transient gravitational- wave noise artifacts with optimized conformal pre- diction.Phys

    Ann-Kristin Malz, Gregory Ashton, and Nicolo Colombo. Classification uncertainty for transient gravitational- wave noise artifacts with optimized conformal pre- diction.Phys. Rev. D, 111(8):084078, 2025. doi: 10.1103/PhysRevD.111.084078

  18. [56]

    Abbott et al

    R. Abbott et al. GWTC-2.1: Deep Extended Catalog of Compact Binary Coalescences Observed by LIGO and Virgo During the First Half of the Third Observing Run. 8 2021

  19. [58]

    Atutorialonconformal prediction.Journal of Machine Learning Research, 9: 371–421, 2008

    GlennShaferandVladimirVovk. Atutorialonconformal prediction.Journal of Machine Learning Research, 9: 371–421, 2008. doi:10.48550/arXiv.0706.3188

  20. [59]

    Gwtc-2.1: Deep extended catalog of compact binary coalescences observed by ligo and virgo during the first half of the third observing run - candidate data release, July 2021

    LIGO Scientific Collaboration and Virgo Collaboration. Gwtc-2.1: Deep extended catalog of compact binary coalescences observed by ligo and virgo during the first half of the third observing run - candidate data release, July 2021. URL https://doi.org/10.5281/ zenodo.5759108

  21. [60]

    Gwtc-3: Compact binary coa- lescences observed by ligo and virgo during the second part of the third observing run — candidate data re- lease, November 2021

    LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration. Gwtc-3: Compact binary coa- lescences observed by ligo and virgo during the second part of the third observing run — candidate data re- lease, November 2021. URLhttps://doi.org/10.5281/ zenodo.5546665

  22. [61]

    Gwtc-4.1: Candidate data release, May 2026

    LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration. Gwtc-4.1: Candidate data release, May 2026. URL https://doi.org/10.5281/ zenodo.20276095

  23. [62]

    Gwtc-5.0: Candidate data release, May 2026

    LIGO Scientific Collaboration, Virgo Collaboration, and KAGRA Collaboration. Gwtc-5.0: Candidate data release, May 2026. URL https://doi.org/10.5281/ zenodo.20348004

  24. [63]

    Low-latency gravi- tational wave alert products and their performance at the time of the fourth LIGO-Virgo-KAGRA observing run.Proc

    Sushant Sharma Chaudhary et al. Low-latency gravi- tational wave alert products and their performance at the time of the fourth LIGO-Virgo-KAGRA observing run.Proc. Nat. Acad. Sci., 121(18):e2316474121, 2024. doi:10.1073/pnas.2316474121

  25. [64]

    Low-latency gravita- tional wave alert infrastructure and new data products for multi-messenger astronomy during ligo-virgo-kagras fourth observing run

    Sushant Sharma Chaudhary et al. Low-latency gravita- tional wave alert infrastructure and new data products for multi-messenger astronomy during ligo-virgo-kagras fourth observing run. 2026. in preparation

  26. [65]

    Learning representations by back-propagating errors.nature, 323(6088):533–536, 1986

    David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. Learning representations by back-propagating errors.nature, 323(6088):533–536, 1986

  27. [66]

    XGBoost: A Scalable Tree Boosting System

    Tianqi Chen and Carlos Guestrin. XGBoost: A Scalable Tree Boosting System. 3 2016. doi: 10.1145/2939672.2939785

  28. [67]

    Conditional validity of inductive confor- mal predictors.Machine Learning, 92, 2013

    Vladimir Vovk. Conditional validity of inductive confor- mal predictors.Machine Learning, 92, 2013

  29. [68]

    A unified approach to interpreting model predictions

    Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish- wanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran As- sociates, In...

  30. [69]

    The elements of statistical learning, 2009

    Trevor Hastie, Robert Tibshirani, Jerome Friedman, et al. The elements of statistical learning, 2009

  31. [70]

    Soni et al

    S. Soni et al. LIGO Detector Characterization in the first half of the fourth Observing run.Class. Quant. Grav., 16 42(8):085016, 2025. doi:10.1088/1361-6382/adc4b6

  32. [71]

    Case studies with GPBilby of glitch- contaminated transient gravitational waves

    Mattia Emma, Ann-Kristin Malz, Adriana Dias, and Gregory Ashton. Case studies with GPBilby of glitch- contaminated transient gravitational waves. 4 2026

  33. [72]

    Measuring the rate of glitches in interferometric gravitational wave detectors with a hierarchical Bayesian model

    Gregory Ashton, Colm Talbot, Andrew Lundgren, Ann- Kristin Malz, and Joseph Areeda. Measuring the rate of glitches in interferometric gravitational wave detectors with a hierarchical Bayesian model. 4 2026

  34. [73]

    A. G. Abac et al. GWTC-4.0: Methods for Identifying and Characterizing Gravitational-wave Transients. 8 2025

  35. [74]

    Data release: Combining gravitational wave search pipelines to find subthreshold signals in gwtc-5.0, July 2026

    Ann-Kristin Malz, Samuel Russell, Gregory Ashton, and nicolo colombo. Data release: Combining gravitational wave search pipelines to find subthreshold signals in gwtc-5.0, July 2026. URLhttps://doi.org/10.5281/ zenodo.21258511

  36. [75]

    BILBY: A user-friendly Bayesian inference library for gravitational-wave astronomy.As- trophys

    Gregory Ashton et al. BILBY: A user-friendly Bayesian inference library for gravitational-wave astronomy.As- trophys. J. Suppl., 241(2):27, 2019. doi:10.3847/1538- 4365/ab06fc

  37. [76]

    Computationally efficient models for the dominant and subdominant harmonic modes of precessing binary black holes.Phys

    Geraint Pratten et al. Computationally efficient models for the dominant and subdominant harmonic modes of precessing binary black holes.Phys. Rev. D, 103(10): 104056, 2021. doi:10.1103/PhysRevD.103.104056

  38. [77]

    Ramis Vidal, Cecilio García- Quirós, Sarp Akçay, and Sayantani Bera

    Marta Colleoni, Felip A. Ramis Vidal, Cecilio García- Quirós, Sarp Akçay, and Sayantani Bera. Fast frequency- domain gravitational waveforms for precessing binaries with a new twist.Phys. Rev. D, 111(10):104019, 2025. doi:10.1103/PhysRevD.111.104019

  39. [78]

    Joshua S. Speagle. dynesty: a dynamic nested sampling package for estimating Bayesian posteriors and evidences. Mon. Not. Roy. Astron. Soc., 493(3):3132–3158, 2020. doi:10.1093/mnras/staa278

  40. [79]

    A. G. Abac et al. GWTC-4.0: An Introduction to Version 4.0 of the Gravitational-Wave Transient Catalog. Astrophys. J. Lett., 995(1):L18, 2025. doi:10.3847/2041- 8213/ae0c06

  41. [80]

    GW231109_235456: A Sub-threshold Binary Neutron Star Merger in the LIGO-Virgo-KAGRA O4a Observing Run? 9 2025

    Wanting Niu et al. GW231109_235456: A Sub-threshold Binary Neutron Star Merger in the LIGO-Virgo-KAGRA O4a Observing Run? 9 2025

  42. [81]

    A. G. Abac et al. GW231123: A Binary Black Hole Merger with Total Mass 190–265 M⊙.Astrophys. J. Lett., 993(1):L25, 2025. doi:10.3847/2041-8213/ae0c9c

  43. [82]

    API design for machine learning software: experiences from the scikit-learn project

    Lars Buitinck et al. API design for machine learning software: experiences from the scikit-learn project. 9 2013

  44. [83]

    Array programming with numpy.Nature, 585(7825):357–362, 2020

    Charles R Harris, K Jarrod Millman, Stéfan J van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. Array programming with numpy.Nature, 585(7825):357–362, 2020

  45. [84]

    pandas-dev/pandas: Pandas, February 2020

    The pandas development team. pandas-dev/pandas: Pandas, February 2020. URL https://doi.org/10. 5281/zenodo.3509134

  46. [85]

    John D. Hunter. Matplotlib: A 2d graphics environment. Computing In Science & Engineering, 9(3):90–95, May- Jun 2007

  47. [86]

    Klimenko, V

    Vaibhav Tiwari, S. Klimenko, V. Necula, and G. Mit- selmakher. Reconstruction of chirp mass in searches for gravitational wave transients.Class. Quant. Grav., 33 (1):01LT01, 2016. doi:10.1088/0264-9381/33/1/01LT01

  48. [87]

    The regression analysis of binary se- quences.Journal of the Royal Statistical Society Series B: Statistical Methodology, 20(2):215–232, 1958

    David R Cox. The regression analysis of binary se- quences.Journal of the Royal Statistical Society Series B: Statistical Methodology, 20(2):215–232, 1958

  49. [88]

    Byrd, Peihuang Lu, and Jorge Nocedal

    Ciyou Zhu, Richard H. Byrd, Peihuang Lu, and Jorge Nocedal. Algorithm 778: L-BFGS-B.ACM Trans. Math. Software, 23(4):550–560, 1997. doi: 10.1145/279232.279236

  50. [89]

    Rectified linear units improve restricted boltzmann machines

    Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. InProceedings of the 27th international conference on machine learning (ICML-10), pages 807–814, 2010

  51. [90]

    Deep Learning using Rectified Lin- ear Units (ReLU)

    Abien Fred Agarap. Deep Learning using Rectified Lin- ear Units (ReLU). 3 2018

  52. [91]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. 12 2014

  53. [92]

    Nearest neighbor pattern classification.IEEE transactions on information theory, 13(1):21–27, 1967

    Thomas Cover and Peter Hart. Nearest neighbor pattern classification.IEEE transactions on information theory, 13(1):21–27, 1967

  54. [93]

    Angelopoulos, Stephen Bates, Michael I

    Tiffany Ding, Anastasios N. Angelopoulos, Stephen Bates, Michael I. Jordan, and Ryan J. Tibshi- rani. Class-conditional conformal prediction with many classes.CoRR, abs/2306.09335, 2023. doi: 10.48550/arXiv.2306.09335. URL https://doi.org/10. 48550/arXiv.2306.09335

  55. [94]

    The use of hypermod- els to understand binary neutron star collisions.Nature Astron., 6(8):961–967, 2022

    Gregory Ashton and Tim Dietrich. The use of hypermod- els to understand binary neutron star collisions.Nature Astron., 6(8):961–967, 2022. doi:10.1038/s41550-022- 01707-x

  56. [95]

    Abbott et al

    R. Abbott et al. GWTC-2: Compact Binary Coales- cences Observed by LIGO and Virgo During the First Half of the Third Observing Run.Phys. Rev. X, 11: 021053, 2021. doi:10.1103/PhysRevX.11.021053

  57. [96]

    GW190426_152155: a merger of neutron star-black hole or low mass binary black holes? 12 2020

    Yin-Jie Li, Ming-Zhe Han, Shao-Peng Tang, Yuan-Zhu Wang, Yi-Ming Hu, Qiang Yuan, Yi-Zhong Fan, and Da-Ming Wei. GW190426_152155: a merger of neutron star-black hole or low mass binary black holes? 12 2020

  58. [97]

    Schäfer and Alexander H

    Marlin B. Schäfer and Alexander H. Nitz. From one to many: A deep learning coincident gravitational- wave search.Phys. Rev. D, 105(4):043003, 2022. doi: 10.1103/PhysRevD.105.043003

  59. [98]

    The relationship be- tween precision-recall and roc curves

    Jesse Davis and Mark Goadrich. The relationship be- tween precision-recall and roc curves. InProceedings of the 23rd international conference on Machine learning, pages 233–240, 2006

  60. [99]

    The precision-recall plot is more informative than the roc plot when evaluat- ing binary classifiers on imbalanced datasets.PloS one, 10(3):e0118432, 2015

    Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the roc plot when evaluat- ing binary classifiers on imbalanced datasets.PloS one, 10(3):e0118432, 2015

  61. [100]

    The in- consistency of “optimal” cutpoints obtained using two criteria based on the receiver operating characteristic curve.American journal of epidemiology, 163(7):670–675, 2006

    Neil J Perkins and Enrique F Schisterman. The in- consistency of “optimal” cutpoints obtained using two criteria based on the receiver operating characteristic curve.American journal of epidemiology, 163(7):670–675, 2006

  62. [101]

    From local explanations to global understanding with explainable ai for trees.Nature machine intelligence, 2 (1):56–67, 2020

    Scott M Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. From local explanations to global understanding with explainable ai for trees.Nature machine intelligence, 2 (1):56–67, 2020. 17...

  63. [102]

    analysing the event under the assumption that it is an astrophysical signal, its properties are highly consistent with that of a binary neutron star merger

    Model comparison The rawML classifier output lacks a statistically rig- orous significance measure, unlike theFARprovided by the maximum-log10IFAR approach, an essential compo- nent for assessing the significance of individual candidate events and motivating the use ofCP. We t...

  64. [103]

    Sensitivity and precision In theGW literature, the sensitive volume [23, 97] is a key metric used to compare pipeline performance on mock 23 data. It is estimated from injections as the fraction of simulated signals recovered above a givenFARthreshold (equivalently, the sensit...

  65. [104]

    If the main goal is catalogue purity, one approach to selecting a confidence threshold would be to choose the point at which precision reaches unity, ensuring no false positives

    Choosing the conditional confidence threshold For our analysis so far, we have used the arbitrary confidence threshold of0.5to distinguish between signal and noise; we now briefly explore alternative thresholds. If the main goal is catalogue purity, one approach to selecting a...

  66. [2023]

    URL https: //link.aps.org/doi/10.1103/PhysRevX.13.041039

    doi:10.1103/PhysRevX.13.041039. URL https: //link.aps.org/doi/10.1103/PhysRevX.13.041039

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.