REVIEW 3 major objections 3 minor 1 cited by
Using gravitational wave dark sirens to choose between host galaxy weighting models
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that Bayes factor comparison of dark-siren host-galaxy weighting models can pick the true luminosity weighting model, becoming decisive against uniform weighting once roughly 1000 detections accumulate.
desk verdict Sensible dark-siren forecasting abstract, but the supplied full text is an unrelated paper, so the claims are currently unverifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the Bayes factor—the ratio of evidence for one candidate model over another—computed across three host-galaxy weighting models: r-band luminosity, g-band luminosity, and uniform weighting. The galaxy catalog supplies candidate hosts and their luminosities; the dark-siren likelihood converts each event's sky-localization uncertainty into a posterior over possible host galaxies; and the Bayes factor accumulates evidence across all events. A small number of well-localized events carry most of the model-separating power, because they confine the possible host galaxies tightly enough to distinguish the weighting rules.
What would settle it
Remove the few best-localized events from the simulated 1000-detection dataset: if the decisive Bayes factor against uniform weighting collapses, the stated mechanism is confirmed; if it does not collapse, the mechanism is wrong. Alternatively, once a real observing run produces roughly 1000 dark sirens, recompute the same comparison—a non-decisive Bayes factor against uniform weighting would falsify the forecast.
Extended reading notes
Core claim
The paper's central claim is that model selection on dark-siren events can identify the correct host-galaxy luminosity weighting before the true weighting is known from theory. In the simulated dataset, the Bayes factor between the injected true model and the alternatives is decisive against uniform weighting once roughly 1000 detections accumulate, but only minor at roughly 200 detections. The abstract's decisive statement is specifically about preferring the true model over the uniform model; the comparison between r-band and g-band weighting is reported as yielding a minor preference at 200 detections, which means the two bands may not separate cleanly at that sample size. The analysis is
Load-bearing premise
The forecast stands on the assumption that the simulated observing scenario and mock galaxy catalog faithfully reproduce the real detectors' sensitivity, sky-localization accuracy, selection effects, and galaxy-catalog completeness, and that the true host weighting is one of the three models compared.
Editorial extensions
If this is right
- At detection counts expected in the next observing run, dark-siren data alone can begin to rule out uniform galaxy weighting.
- A few well-localized events are worth more than many poorly localized ones for this question, so improving localization on a handful of events pays off directly.
- The same Bayes-factor procedure can be extended to other weighting prescriptions, such as stellar mass or star formation rate, without changing the core machinery.
- The forecast gives a concrete benchmark: if the real run produces roughly 1000 dark sirens and the Bayes factor is not decisive against uniform weighting, the host-weighting models under test would be called into question.
Reading between the lines
- The forecast depends on the tail of the localization distribution; if real events localize worse than the mock, the number of detections needed for a decisive Bayes factor could be much larger than 1000.
- The decisive result is stated against uniform weighting, not between the two luminosity bands; a fair reading suggests distinguishing r-band from g-band is harder and may need more events or additional information.
- A direct testable extension is to rerun the simulated analysis with the best-localized events removed; since the paper says the decisive preference is driven by those events, removing them should collapse the Bayes factor.
- If the correct weighting is degenerate with cosmological parameters in dark-siren analyses, choosing it could shift inferred cosmological parameters; this is an implied consequence rather than a result the paper demonstrates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract of arXiv:2508.15574 reports a simulation-based forecast that gravitational-wave dark sirens from an O5-like LVK observing run, analyzed with gwcosmo and a mock galaxy catalogue from MICECATv2, can select among host-galaxy luminosity weighting models (r-band, g-band, uniform). It claims a minor Bayes-factor preference for the true model at ~200 detections and a decisive preference over the uniform model at ~1000 detections, driven by a few well-localized events. The full text supplied under the same arXiv ID, however, is a different manuscript, 'LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space,' which contains no gravitational-wave methodology, simulation details, or results. The central claim is therefore not supported by the manuscript as received.
Significance. If the forecast were substantiated, it would provide a concrete, data-driven path to distinguishing BBH host-galaxy weighting models with near-future O5 data, an important input for dark-siren cosmology and galaxy-formation modeling. The claimed result that a small number of well-localized events drives model selection is testable and interesting. However, because the manuscript as received contains none of the machinery needed to evaluate the forecast, no scientific significance can be assessed from the present text.
major comments (3)
- [Full text (entire manuscript)] The supplied full text is arXiv:2508.15577v2, a quantum-chemistry active-learning paper, not the gravitational-wave paper announced in the abstract. There is no description of the simulation, injection procedure, gwcosmo setup, MICECATv2 catalogue construction, likelihood/Bayes-factor computation, or event selection. Absent these, the abstract's claim is an unsupported assertion; the manuscript cannot be checked for correctness, selection effects, or robustness.
- [Abstract, '~200' and '~1000 detection cases'] The paper does not specify how the detection scenarios were generated, what detection threshold or selection function was used, how sky-localization errors were modeled, or how the mock spectroscopic catalogue's completeness was handled. The claim that the Bayes factor is 'strongly driven by a small number of well-localised events' cannot be verified without these definitions; it could reflect prior sensitivity or catalogue incompleteness rather than intrinsic model separability.
- [All results] No tables, figures, or numerical values of Bayes factors are presented in the full text. The only results reported are in the abstract. In a forecasting paper, the actual distributions, not a summary sentence, are the evidence; their absence is a load-bearing gap, not a stylistic issue.
minor comments (3)
- [Abstract] 'IGO-Virgo-KAGRA' appears to be a typo for 'LIGO-Virgo-KAGRA' (unless a nonstandard acronym is intended).
- [Full text] All equations and figures in the supplied full text belong to the unrelated LFaB paper; none support the gravitational-wave analysis described in the abstract.
- [General submission] The arXiv ID mismatch should be corrected; as submitted, the reader cannot distinguish a submission error from a placeholder.
Circularity Check
No circularity found; the supplied full text is a different, unrelated preprint, so the claim cannot be verified but is not circular.
full rationale
The abstract proposes a controlled forward simulation: a true host-galaxy weighting model is injected, mock LVK O5 dark sirens are generated with gwcosmo, a mock spectroscopic catalogue is built from MICECATv2, and Bayes factors among r-band, g-band, and uniform weighting are compared. This is ground-truth testing, not circular reasoning: the simulated data are generated under a known model and the analysis asks whether model comparison recovers that model. No fitted parameter is renamed as a prediction, no definition makes the Bayes factors equivalent to the injected weighting by construction, and no self-citation chain supplies the conclusion. The mention of gwcosmo, co-authored by one of the present authors (R. Gray), is a normal use of an established tool rather than a load-bearing self-citation. The full text supplied with this review, however, is arXiv:2508.15577, 'LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space', which is a different paper about active learning for quantum chemistry. It contains no gravitational-wave methods, no O5 simulation, no gwcosmo setup, no MICECATv2 catalogue construction, no sky-localization error model, no Bayes-factor computation, and no results supporting the abstract's quantitative claims. That is a manuscript-support/integrity problem, not a circularity problem: the derivation chain is absent, so no step can be shown to reduce to its own inputs. Per the rubric, missing evidence is not itself a circular step, and speculation about hidden circularity is disallowed. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- Detection scenario sizes (~200 and ~1000 events) =
~200, ~1000
- Injected BBH population properties (rate, mass distribution, redshift evolution) =
not stated in abstract
assumptions (4)
- standard math Bayes factors are a valid and calibrated model comparison statistic for this setting
- domain assumption The mock spectroscopic galaxy catalogue from MICECATv2 faithfully represents the real galaxy distribution and its completeness properties
- domain assumption Host galaxy weighting models are adequately captured by r-band luminosity, g-band luminosity, and uniform weighting
- domain assumption The gwcosmo likelihood correctly handles selection effects, completeness, and the redshift-distance relation in the simulation
Cite this review
Pith. "Pith review of Using gravitational wave dark sirens to choose between host galaxy weighting models." pith.science (2026). https://pith.science/paper/KDMNC2HS
@misc{pith2026250815574,
author = {Pith},
title = {Pith review of: Using gravitational wave dark sirens to choose between host galaxy weighting models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KDMNC2HS}},
note = {Machine review of arXiv:2508.15574}
}
read the original abstract
Binary black hole (BBH) mergers,, an important source of gravitational-waves(GWs), are assumed to be hosted in galaxies. The probability of a galaxy to host a BBH is related to its properties, for example stellar mass and star formation rate. These properties can be estimated from observables, such as the luminosity in certain observation bands. We refer to this description of host galaxy properties as host galaxy weighting models. However, the host galaxy weighting model for BBHs has yet to be accurately determined. Population synthesis has provided a variety of host galaxy weighting models. Here we investigate whether it is possible to distinguish different host galaxy weighting models using a data driven approach. We use the GW cosmology tool gwcosmo with a simulated IGO-Virgo-KAGRA (LVK) fifth observing run (O5)-like observing scenario. We also construct a mock spectroscopic galaxy catalogue from MICECATv2. Our analysis compares the Bayes factors among three simple luminosity weighting models, r-band, g-band, and uniform weighting. The Bayes factors among different host galaxy weighting models show a minor preference for the true model for a ~200-detection case, and a decisive preference for the true model over the uniform model for a ~1000-detection case, which is strongly driven by a small number of well-localised events.
Forward citations
Cited by 1 Pith paper
-
The impact of precession and higher-order multipoles for gravitational wave cosmological inference
For H0 inference via the black-hole mass-spectrum method, waveform models with spin precession and higher-order multipoles offer no significant advantage over the simplest quadrupole-only model, at up to six times low...
Reference graph
Works this paper leans on
-
[1]
LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space Vivin Vinod1* and Peter Zaspel1 1School of Mathematics and Natural Sciences, University of Wuppertal, Germany. *Corresponding author(s). E-mail(s): vinod@uni-wuppertal.de Abstract Active learning promises to provide an optimal training sample selection proce- dure in the co...
-
[4]
Here, the model using GPR variance to sample training data results in a MAE around 3 kcal/mol even after 2000 AL-iterations. For the same number of iterations the model built using the LFaB scheme results in an MAE of about 1.5 kcal/mol. That is, a model that is twice as accurate. An alternative way to study these curves is to fix a desired MAE from the M...
work page 2000
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.