Pith. sign in

REVIEW 3 major objections 3 minor 1 cited by

Using gravitational wave dark sirens to choose between host galaxy weighting models

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that Bayes factor comparison of dark-siren host-galaxy weighting models can pick the true luminosity weighting model, becoming decisive against uniform weighting once roughly 1000 detections accumulate.

desk verdict Sensible dark-siren forecasting abstract, but the supplied full text is an unrelated paper, so the claims are currently unverifiable. read the letter →

arxiv 2508.15574 v1 pith:KDMNC2HS submitted 2025-08-21 astro-ph.CO gr-qc

classification astro-ph.COgr-qc
keywords gravitationalwavesdarksirensbinaryblackholemergershostgalaxyweightingBayesfactorluminositycatalogpopulationsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Binary black hole mergers are assumed to occur in galaxies, but how a galaxy's observable properties set the probability that it hosts one is not yet established. This paper tests whether the gravitational waves themselves—specifically 'dark sirens,' mergers without an electromagnetic counterpart—can pick the correct host-galaxy weighting model through model comparison. The authors simulate a next-generation observing run, pair the detections with a mock catalog of candidate galaxies, and compare Bayes factors among three simple weighting rules: brightness in the r-band, brightness in the g-band, and uniform weighting. The result is a minor preference for the true model at around 200 detections and a decisive preference over uniform weighting at around 1000 detections, driven mainly by a small number of well-localized events. If the forecast holds, it gives a purely data-driven way to choose among the host-weighting models that population synthesis currently leaves uncertain.

What carries the argument

The central mechanism is the Bayes factor—the ratio of evidence for one candidate model over another—computed across three host-galaxy weighting models: r-band luminosity, g-band luminosity, and uniform weighting. The galaxy catalog supplies candidate hosts and their luminosities; the dark-siren likelihood converts each event's sky-localization uncertainty into a posterior over possible host galaxies; and the Bayes factor accumulates evidence across all events. A small number of well-localized events carry most of the model-separating power, because they confine the possible host galaxies tightly enough to distinguish the weighting rules.

What would settle it

Remove the few best-localized events from the simulated 1000-detection dataset: if the decisive Bayes factor against uniform weighting collapses, the stated mechanism is confirmed; if it does not collapse, the mechanism is wrong. Alternatively, once a real observing run produces roughly 1000 dark sirens, recompute the same comparison—a non-decisive Bayes factor against uniform weighting would falsify the forecast.

Watch

Extended reading notes

Core claim

The paper's central claim is that model selection on dark-siren events can identify the correct host-galaxy luminosity weighting before the true weighting is known from theory. In the simulated dataset, the Bayes factor between the injected true model and the alternatives is decisive against uniform weighting once roughly 1000 detections accumulate, but only minor at roughly 200 detections. The abstract's decisive statement is specifically about preferring the true model over the uniform model; the comparison between r-band and g-band weighting is reported as yielding a minor preference at 200 detections, which means the two bands may not separate cleanly at that sample size. The analysis is

Load-bearing premise

The forecast stands on the assumption that the simulated observing scenario and mock galaxy catalog faithfully reproduce the real detectors' sensitivity, sky-localization accuracy, selection effects, and galaxy-catalog completeness, and that the true host weighting is one of the three models compared.

Editorial extensions

If this is right

  • At detection counts expected in the next observing run, dark-siren data alone can begin to rule out uniform galaxy weighting.
  • A few well-localized events are worth more than many poorly localized ones for this question, so improving localization on a handful of events pays off directly.
  • The same Bayes-factor procedure can be extended to other weighting prescriptions, such as stellar mass or star formation rate, without changing the core machinery.
  • The forecast gives a concrete benchmark: if the real run produces roughly 1000 dark sirens and the Bayes factor is not decisive against uniform weighting, the host-weighting models under test would be called into question.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The forecast depends on the tail of the localization distribution; if real events localize worse than the mock, the number of detections needed for a decisive Bayes factor could be much larger than 1000.
  • The decisive result is stated against uniform weighting, not between the two luminosity bands; a fair reading suggests distinguishing r-band from g-band is harder and may need more events or additional information.
  • A direct testable extension is to rerun the simulated analysis with the best-localized events removed; since the paper says the decisive preference is driven by those events, removing them should collapse the Bayes factor.
  • If the correct weighting is degenerate with cosmological parameters in dark-siren analyses, choosing it could shift inferred cosmological parameters; this is an implied consequence rather than a result the paper demonstrates.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The abstract of arXiv:2508.15574 reports a simulation-based forecast that gravitational-wave dark sirens from an O5-like LVK observing run, analyzed with gwcosmo and a mock galaxy catalogue from MICECATv2, can select among host-galaxy luminosity weighting models (r-band, g-band, uniform). It claims a minor Bayes-factor preference for the true model at ~200 detections and a decisive preference over the uniform model at ~1000 detections, driven by a few well-localized events. The full text supplied under the same arXiv ID, however, is a different manuscript, 'LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space,' which contains no gravitational-wave methodology, simulation details, or results. The central claim is therefore not supported by the manuscript as received.

Significance. If the forecast were substantiated, it would provide a concrete, data-driven path to distinguishing BBH host-galaxy weighting models with near-future O5 data, an important input for dark-siren cosmology and galaxy-formation modeling. The claimed result that a small number of well-localized events drives model selection is testable and interesting. However, because the manuscript as received contains none of the machinery needed to evaluate the forecast, no scientific significance can be assessed from the present text.

major comments (3)
  1. [Full text (entire manuscript)] The supplied full text is arXiv:2508.15577v2, a quantum-chemistry active-learning paper, not the gravitational-wave paper announced in the abstract. There is no description of the simulation, injection procedure, gwcosmo setup, MICECATv2 catalogue construction, likelihood/Bayes-factor computation, or event selection. Absent these, the abstract's claim is an unsupported assertion; the manuscript cannot be checked for correctness, selection effects, or robustness.
  2. [Abstract, '~200' and '~1000 detection cases'] The paper does not specify how the detection scenarios were generated, what detection threshold or selection function was used, how sky-localization errors were modeled, or how the mock spectroscopic catalogue's completeness was handled. The claim that the Bayes factor is 'strongly driven by a small number of well-localised events' cannot be verified without these definitions; it could reflect prior sensitivity or catalogue incompleteness rather than intrinsic model separability.
  3. [All results] No tables, figures, or numerical values of Bayes factors are presented in the full text. The only results reported are in the abstract. In a forecasting paper, the actual distributions, not a summary sentence, are the evidence; their absence is a load-bearing gap, not a stylistic issue.
minor comments (3)
  1. [Abstract] 'IGO-Virgo-KAGRA' appears to be a typo for 'LIGO-Virgo-KAGRA' (unless a nonstandard acronym is intended).
  2. [Full text] All equations and figures in the supplied full text belong to the unrelated LFaB paper; none support the gravitational-wave analysis described in the abstract.
  3. [General submission] The arXiv ID mismatch should be corrected; as submitted, the reader cannot distinguish a submission error from a placeholder.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the supplied full text is a different, unrelated preprint, so the claim cannot be verified but is not circular.

full rationale

The abstract proposes a controlled forward simulation: a true host-galaxy weighting model is injected, mock LVK O5 dark sirens are generated with gwcosmo, a mock spectroscopic catalogue is built from MICECATv2, and Bayes factors among r-band, g-band, and uniform weighting are compared. This is ground-truth testing, not circular reasoning: the simulated data are generated under a known model and the analysis asks whether model comparison recovers that model. No fitted parameter is renamed as a prediction, no definition makes the Bayes factors equivalent to the injected weighting by construction, and no self-citation chain supplies the conclusion. The mention of gwcosmo, co-authored by one of the present authors (R. Gray), is a normal use of an established tool rather than a load-bearing self-citation. The full text supplied with this review, however, is arXiv:2508.15577, 'LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space', which is a different paper about active learning for quantum chemistry. It contains no gravitational-wave methods, no O5 simulation, no gwcosmo setup, no MICECATv2 catalogue construction, no sky-localization error model, no Bayes-factor computation, and no results supporting the abstract's quantitative claims. That is a manuscript-support/integrity problem, not a circularity problem: the derivation chain is absent, so no step can be shown to reduce to its own inputs. Per the rubric, missing evidence is not itself a circular step, and speculation about hidden circularity is disallowed. Therefore the circularity score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

Only the abstract was available for this assessment; the full text supplied is a different manuscript. No fitted numbers are stated in the abstract. The main assumed inputs are the simulated O5 detector response, the mock galaxy catalogue, and the restriction to three weighting models. No new entities are introduced.

free parameters (2)
  • Detection scenario sizes (~200 and ~1000 events) = ~200, ~1000
    Chosen O5-like scenarios at which the Bayes factor results are reported; the abstract does not explain how these event counts are derived from detector assumptions.
  • Injected BBH population properties (rate, mass distribution, redshift evolution) = not stated in abstract
    Any mock GW catalog requires population assumptions that determine event localization quality and hence the Bayes factor outcome; no values are given in the abstract.
assumptions (4)
  • standard math Bayes factors are a valid and calibrated model comparison statistic for this setting
    The abstract reports preferences as Bayes factors; interpreting them as 'minor' or 'decisive' assumes standard evidence thresholds and a correctly specified likelihood.
  • domain assumption The mock spectroscopic galaxy catalogue from MICECATv2 faithfully represents the real galaxy distribution and its completeness properties
    The dark siren likelihood depends on the galaxy catalog; if the mock catalog's selection function differs from the real one, the forecast could be optimistic or pessimistic.
  • domain assumption Host galaxy weighting models are adequately captured by r-band luminosity, g-band luminosity, and uniform weighting
    The abstract states that galaxy properties (stellar mass, star formation rate) are estimated from luminosities; restricting to three simple models assumes the true host-weighting process lies in this set.
  • domain assumption The gwcosmo likelihood correctly handles selection effects, completeness, and the redshift-distance relation in the simulation
    The analysis uses gwcosmo as a tool; the forecast's validity inherits the likelihood's assumptions, which are not described in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using gravitational wave dark sirens to choose between host galaxy weighting models." pith.science (2026). https://pith.science/paper/KDMNC2HS

@misc{pith2026250815574,
  author       = {Pith},
  title        = {Pith review of: Using gravitational wave dark sirens to choose between host galaxy weighting models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KDMNC2HS}},
  note         = {Machine review of arXiv:2508.15574}
}
read the original abstract

Binary black hole (BBH) mergers,, an important source of gravitational-waves(GWs), are assumed to be hosted in galaxies. The probability of a galaxy to host a BBH is related to its properties, for example stellar mass and star formation rate. These properties can be estimated from observables, such as the luminosity in certain observation bands. We refer to this description of host galaxy properties as host galaxy weighting models. However, the host galaxy weighting model for BBHs has yet to be accurately determined. Population synthesis has provided a variety of host galaxy weighting models. Here we investigate whether it is possible to distinguish different host galaxy weighting models using a data driven approach. We use the GW cosmology tool gwcosmo with a simulated IGO-Virgo-KAGRA (LVK) fifth observing run (O5)-like observing scenario. We also construct a mock spectroscopic galaxy catalogue from MICECATv2. Our analysis compares the Bayes factors among three simple luminosity weighting models, r-band, g-band, and uniform weighting. The Bayes factors among different host galaxy weighting models show a minor preference for the true model for a ~200-detection case, and a decisive preference for the true model over the uniform model for a ~1000-detection case, which is strongly driven by a small number of well-localised events.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The impact of precession and higher-order multipoles for gravitational wave cosmological inference

    astro-ph.CO 2025-11 conditional novelty 6.0 of 10

    For H0 inference via the black-hole mass-spectrum method, waveform models with spin precession and higher-order multipoles offer no significant advantage over the simplest quadrupole-only model, at up to six times low...

Reference graph

Works this paper leans on

2 extracted references · 2 canonical work pages · cited by 1 Pith paper

  1. [1]

    *Corresponding author(s)

    LFaB: Low fidelity as Bias for Active Learning in the chemical configuration space Vivin Vinod1* and Peter Zaspel1 1School of Mathematics and Natural Sciences, University of Wuppertal, Germany. *Corresponding author(s). E-mail(s): vinod@uni-wuppertal.de Abstract Active learning promises to provide an optimal training sample selection proce- dure in the co...

  2. [4]

    For the same number of iterations the model built using the LFaB scheme results in an MAE of about 1.5 kcal/mol

    Here, the model using GPR variance to sample training data results in a MAE around 3 kcal/mol even after 2000 AL-iterations. For the same number of iterations the model built using the LFaB scheme results in an MAE of about 1.5 kcal/mol. That is, a model that is twice as accurate. An alternative way to study these curves is to fix a desired MAE from the M...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.