REVIEW 2 major objections 3 minor 55 references
With a few thousand events, the Einstein Telescope can separate two neutron-star merger populations and trace their redshift evolution.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 18:59 UTC pith:QRJO7T7D
load-bearing objection Solid injection-recovery forecast for ET, but the redshift-separation claim rests on an uncalibrated statistic and is likely overconfident. the 2 major comments →
Probing the redshift evolution and sub-populations of binary neutron stars with the Einstein Telescope
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the two-population hypothesis can be tested with the Einstein Telescope's expected data volume. In the authors' model, heavy binaries merge after a fixed short delay of 30 Myr while light binaries follow a power-law delay, so the two subpopulations have different redshift evolutions even though they share the same cosmic star-formation history. Analysing mock catalogues of 100–5,000 events, they find that 500–1,000 detections suffice to establish that the total-mass distribution is bimodal, and that 1,000–5,000 events statistically separate the redshift distributions of the subpopulations for moderate delay indices. For steeper indices (α_L = –1.5) the diffe
What carries the argument
A hierarchical Bayesian mixture model over total mass, mass ratio, and redshift. Each component's redshift distribution is the convolution of the cosmic star formation rate with a time-delay distribution — the heavy component with a fixed short delay, the light component with a power-law delay. Bimodality is assessed with a standard dip test on posterior-predicted mass samples, and separability of redshift evolutions is quantified by comparing between-population and within-population cumulative-distribution statistics. This machinery converts the question 'are there two subpopulations?' into a statement about how many detected events are needed.
Load-bearing premise
The load-bearing premise is that the mock population — heavy/heavy binaries with a fixed 30 Myr delay and light/light binaries with a power-law delay — is a faithful stand-in for nature; if real systems include light–heavy pairings or a spread of heavy delays, the quoted event counts could be optimistic.
What would settle it
A concrete check: apply a dip test for unimodality to the total-mass distribution of the first ~1,000 ET detections; if in over 95% of posterior draws the distribution is classified as unimodal, the paper's bimodality claim fails. Equivalently, with 5,000 events, if the between-population cumulative-distribution statistic for light versus heavy redshift distributions does not exceed the 5% quantile of the within-population statistics, the claimed separability of the redshift evolutions would be contradicted.
If this is right
- With 500–1,000 Einstein Telescope detections, a bimodal total-mass distribution could be established independently of the assumed delay law.
- A few thousand events should distinguish the redshift evolutions of light and heavy populations when the light-population delay index is –0.5 or –1.
- Detection-rate estimates (roughly 2,400–78,000 mergers per year in the fiducial case) imply ET will accumulate enough events within its lifetime to run these tests.
- The reconstruction of redshift evolution is better constrained for the heavy population, whose louder signals give higher signal-to-noise ratios.
- If confirmed, a fast-merging heavy subpopulation would support unstable case BB mass transfer as a distinct formation channel.
Where Pith is reading between the lines
- If the same counting logic holds for steeper delay indices, future analyses may need non-parametric or machine-learning population methods, since hierarchical Bayesian sampling at tens of thousands of events could become computationally prohibitive.
- The method could be extended to other compact-object mixtures, such as neutron star–black hole binaries, where a short-delay subpopulation is suspected, since the bimodality and distribution-comparison statistics do not depend on the specific mass model.
- A null result — a unimodal total-mass distribution after 1,000 events — would not rule out two populations entirely, but would disfavour the specific heavy/heavy, short-delay construction assumed here and shift attention toward formation channels with overlapping masses or longer heavy delays.
- The paper's distinction between moderate and steep delay indices suggests that measuring the exact value of α_L for the light population is itself a prerequisite for designing optimal subpopulation searches.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a mock-injection study of the Einstein Telescope's capability to identify and characterize a hypothetical bimodal binary neutron star (BNS) mass distribution with two subpopulations: a 'light' population with long merger time delays and a 'heavy' population with short delays, motivated by GW190425 and case BB mass transfer. Using the latest LVK BNS merger-rate constraints and a Madau–Dickinson star-formation history, the authors generate mock ET catalogues of 100–5,000 events for three power-law indices of the light delay distribution (α_L = −0.5, −1, −1.5), then perform hierarchical Bayesian inference to recover the hyperparameters. They report that with a few hundred events the total mass distribution is robustly bimodal, and that a few thousand events suffice to distinguish the redshift evolutions of the two subpopulations for moderate α_L, with a steep α_L = −1.5 remaining challenging. The analysis includes selection effects, a simplified SNR-based parameter-estimation error model, and multiple realizations per catalogue size.
Significance. If the conclusions hold, the paper provides a timely, quantitative forecast for ET (and by extension next-generation detectors) to probe BNS formation channels through the redshift evolution of subpopulations. The work is clearly relevant to the ET science case and the ongoing discussion of a fast-merging massive BNS subpopulation. The authors are transparent about their modeling choices, use externally constrained inputs (LVK rate, Galaudage et al. 2021 population), and release a reproducible analysis pipeline (GWBENCH, Eryn, diptest). They also correctly flag the two-subpopulation limitation and the computational cap that prevents exploring α_L = −1.5. These strengths make the paper a useful contribution if the statistical measure used for the central claim is properly calibrated.
major comments (2)
- [§3, p_dist definition and Fig. 5] The p_dist statistic is not a calibrated hypothesis test. The 'between-population' KS statistic compares two posterior draws, one from each subpopulation, so its null distribution includes the posterior uncertainty of both subpopulations. Each 'within-population' KS statistic compares two draws from the same subpopulation only. When the light and heavy subpopulations have different measurement uncertainties (as the authors note in §4, the heavy population is better constrained), the between-population null distribution is not equal to either within-population distribution. The maximum-of-two-quantiles construction can therefore yield p_dist > 0.95 even when the true redshift distributions are identical, if one subpopulation is much better measured than the other. The thresholds quoted in Fig. 5 and the abstract ('a few thousand events are sufficient…') rest on this statistic. The authors
- [App. B, parameter-estimation error model] The measurement-error model assumes Gaussian uncertainties with variances scaling as (0.1ρ_th/ρ)^2, (0.15ρ_th/ρ)^2 etc., calibrated to GW170817 and GW190425. These nearby, low-redshift events are not representative of the bulk of ET detections at z ~ 1–3, where the luminosity-distance inference is governed by the redshift prior and cosmological degeneracies rather than simple SNR scaling. Since the central claim concerns the ability to separate redshift distributions, a more realistic treatment of distance errors, or at least an explicit stress test with altered error scalings, would strengthen the result. As written, the impact of this approximation on the required catalogue sizes is not quantified.
minor comments (3)
- [§4 and App. D, figure captions] The text and captions refer to 'α_L = 1' in the upper panel of Fig. 4 and in Fig. 1 of App. D. The model uses negative indices (Table 1); this should read 'α_L = −1'.
- [§3, p_bimodal terminology] The quantity p_bimodal is described as 'the probability that the mass distribution is bimodal with more than 95% confidence.' It is actually the fraction of hyperposterior samples for which Hartigan's dip test returns p < 0.05, and the rotation-and-minimum procedure alters the p-value's interpretation. The terminology should be softened or the calibration presented.
- [§3, p_dist interpretation] Even if the p_dist statistic is retained after calibration, the phrase 'probability that the two redshift distributions are different' is not a posterior probability. It should be described as a detection statistic with a threshold chosen by calibration.
Circularity Check
No significant circularity: this is an injection–recovery forecast built from external inputs, with limitations explicitly stated.
full rationale
This is a simulation/injection-recovery study. The authors construct mock ET catalogues from external inputs — the LVK local BNS merger rate (Abac et al. 2025b), the Galaudage et al. (2021) mass-population parameters, the Madau & Dickinson (2014) SFR, and the ET sensitivity curve — then run hierarchical Bayesian inference and assess recovery. No central claim is defined in terms of its own output. The bimodality claim is tested with Hartigan's dip test on posterior-predictive draws rather than being read off the mixture-model fit, and the redshift-distinction metric compares between- and within-population KS distributions; even if the p_dist statistic is statistically miscalibrated (a correctness risk, not a circularity), the forecast is not equivalent to its inputs by construction. Self-citations (Toubiana et al. 2021; Tenorio et al. 2025; Pellouin et al. 2025) appear only in caveats or future-work suggestions and are not load-bearing. The stated limitation to two subpopulations, and the inability to explore α_L = -1.5 due to computational resources, are acknowledged assumptions/constraints rather than disguised circular reductions. No quoted equation or fitted parameter reduces to its own input. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (17)
- λ (light mixing fraction) =
0.5
- α_L (light delay power-law index) =
-0.5, -1, -1.5
- ∆t_H (heavy fixed delay) =
30 Myr
- ∆t_L,min =
10 Myr
- ∆t_L,max =
13 Gyr
- µ_t,L =
2.63 M_sun
- σ_t,L =
0.092 M_sun
- µ_t,H =
3.27 M_sun
- σ_t,H =
0.21 M_sun
- µ_q,L =
1
- σ_q,L =
0.08
- µ_q,H =
0.82
- σ_q,H =
0.2
- Local BNS merger rate normalization =
129 Gpc^-3 yr^-1
- Mass-ratio uncertainty scaling =
0.15 ρ_th/ρ
- Total-mass uncertainty scaling =
0.1 ρ_th/ρ
- Θ uncertainty scaling =
0.21 ρ_th/ρ
axioms (7)
- domain assumption Mixture of two BNS subpopulations with only light-light and heavy-heavy pairings.
- domain assumption Redshift distribution of each subpopulation is convolution of SFR with time-delay distribution.
- domain assumption Heavy subpopulation has fixed short delay; light subpopulation power-law delays.
- domain assumption Cosmic SFR from Madau & Dickinson (2014).
- domain assumption ET sensitivity ET-10 and SNR distribution approximated by single-detector Gaussian.
- domain assumption Parameter-estimation errors scale as inverse SNR with calibrated constants.
- standard math Standard hierarchical Bayesian likelihood with selection function estimated by importance sampling.
read the original abstract
The formation channels of binary neutron stars (BNSs) currently remain uncertain, but important information can be gathered by observing their mergers with gravitational-wave detectors. The processes that lead to BNS coalescence are encoded in the time-delay distribution between stellar binary formation and BNS coalescence, and therefore in the BNS merger rate. Moreover, the detection of GW190425 by LIGO/Virgo/KAGRA (LVK) suggests a sub-population of massive BNSs, possibly formed through unstable 'case BB' mass transfer with short merger delays. We investigate whether next-generation detectors can constrain the time-delay distribution of BNSs and identify such sub-populations. Using the latest LVK constraints, we generated mock catalogues that contain a mixture of light and heavy sub-populations. We modelled the redshift distribution of each sub-population as the convolution of the cosmic star formation rate with a time-delay distribution. We first considered a scenario where the time-delay distribution is common to all BNSs. In the second scenario, heavy BNSs have fixed short delays, while light BNSs follow power-law delays. Hierarchical Bayesian analyses were then performed on catalogues of 100-5000 events. With thousands of events, we should be able to accurately characterise the time-delay distribution for moderate time-delay indices. We find that with hundreds of detections, we will be able to establish that the total mass distribution is bimodal. A few thousand events are sufficient to disentangle the redshift distributions of the two sub-populations for moderate time-delay indices. For steeper indices, the differences are more subtle and require larger catalogues, which was beyond what we could explore given our computational resources.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archiveprefix author booktitle chapter edition editor howpublished institution eprint journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 ...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in " " * FUNCTION format....
-
[3]
Abac, A. et al. 2025 a [ [arXiv] 2503.12263 ]
Pith/arXiv arXiv 2025
-
[4]
Abac, A. G. et al. 2025 b [ [arXiv] 2508.18083 ]
Pith/arXiv arXiv 2025
-
[5]
Abbott, B. P. et al. 2017, Phys. Rev. Lett., 119, 161101
2017
-
[6]
Abbott, B. P. et al. 2020, Astrophys. J. Lett., 892, L3
2020
-
[7]
O., & Berti, E
Alsing, J., Silva, H. O., & Berti, E. 2018, Mon. Not. Roy. Astron. Soc., 478, 1377
2018
-
[8]
Antoniadis, J., Tauris, T. M., Ozel, F., et al. 2016 [ [arXiv] 1605.01665 ]
Pith/arXiv arXiv 2016
-
[9]
2002, , 572, 407
Belczynski , K., Kalogera , V., & Bulik , T. 2002, , 572, 407
2002
-
[10]
Belczynski, K. et al. 2018, Astron. Astrophys., 615, A91
2018
-
[11]
2021, Class
Borhanian, S. 2021, Class. Quant. Grav., 38, 175014
2021
-
[12]
Branchesi, M. et al. 2023, JCAP, 07, 068
2023
-
[13]
2018, Mon
Chruslinska, M., Belczynski, K., Klencki, J., & Benacquista, M. 2018, Mon. Not. Roy. Astron. Soc., 474, 2937
2018
-
[14]
2025, Astrophys
Chu, Q., Lu, Y., & Yu, S. 2025, Astrophys. J., 980, 181
2025
-
[15]
& Zhang, T
Danilishin, S. & Zhang, T. 2023, ET sensitivity curves used for CoBA Science Study , Technical Report ET-0304B-22, Einstein Telescope Collaboration
2023
-
[16]
M., Rocha, L
de S \'a , L. M., Rocha, L. S., Bernardo, A., Bachega, R. R. A., & Horvath, J. E. 2024, Mon. Not. Roy. Astron. Soc., 535, 2041
2024
-
[17]
Dewi , J. D. M. & Pols , O. R. 2003, , 344, 629
2003
-
[18]
2023, Phys
Essick, R. 2023, Phys. Rev. D, 108, 043011
2023
-
[19]
Evans, M. et al. 2021 [ [arXiv] 2109.09882 ]
Pith/arXiv arXiv 2021
-
[20]
Evans, M. et al. 2023 [ [arXiv] 2306.13745 ]
Pith/arXiv arXiv 2023
-
[21]
Farah, A. 2022, GWMockCat : A lightweight code to generate mock catalogs of gravitational-wave sources, https://git.ligo.org/amanda.farah/GWMockCat, gitLab repository, release v1 (May 03, 2022)
2022
-
[22]
M., Edelman, B., Zevin, M., et al
Farah, A. M., Edelman, B., Zevin, M., et al. 2023, Astrophys. J., 955, 107
2023
-
[23]
Farr , W. M. & Chatziioannou , K. 2020, Research Notes of the American Astronomical Society, 4, 65
2020
-
[24]
2019, Astrophys
Farrow, N., Zhu, X.-J., & Thrane, E. 2019, Astrophys. J., 876, 18
2019
-
[25]
Finn, L. S. & Chernoff, D. F. 1993, Phys. Rev. D, 47, 2198
1993
-
[26]
M., & Holz, D
Fishbach, M., Farr, W. M., & Holz, D. E. 2020, Astrophys. J. Lett., 891, L31
2020
-
[27]
2021, Astrophys
Galaudage, S., Adamcewicz, C., Zhu, X.-J., Stevenson, S., & Thrane, E. 2021, Astrophys. J. Lett., 909, L19
2021
-
[28]
2020, Phys
Garc \' a-Quir \'o s, C., Colleoni, M., Husa, S., et al. 2020, Phys. Rev. D, 102, 064002
2020
-
[29]
& Bellotti, M
Gerosa, D. & Bellotti, M. 2024, Class. Quant. Grav., 41, 125002
2024
-
[30]
& Hartigan, P
Hartigan, J. & Hartigan, P. 1985, Annals of Statistics, 13, 70
1985
-
[31]
Iorio, G. et al. 2023, Mon. Not. Roy. Astron. Soc., 524, 426
2023
-
[32]
L., Korsakova, N., Gair, J
Karnesis, N., Katz, M. L., Korsakova, N., Gair, J. R., & Stergioulas, N. 2023, Mon. Not. Roy. Astron. Soc., 526, 4814
2023
-
[33]
& Safarzadeh, M
Korol, V. & Safarzadeh, M. 2021, Mon. Not. Roy. Astron. Soc., 502, 5576
2021
-
[34]
& Dickinson, M
Madau, P. & Dickinson, M. 2014, Ann. Rev. Astron. Astrophys., 52, 415
2014
-
[35]
M., & Gair, J
Mandel, I., Farr, W. M., & Gair, J. R. 2019, Mon. Not. R. Astron. Soc., 486, 1086
2019
-
[36]
& Nakar, E
Maoz, D. & Nakar, E. 2025, Astrophys. J., 982, 179
2025
- [37]
-
[38]
& Freire, P
\"O zel, F. & Freire, P. 2016, Ann. Rev. Astron. Astrophys., 54, 401
2016
-
[39]
Ozel, F., Psaltis, D., Narayan, R., & Villarreal, A. S. 2012, Astrophys. J., 757, 55
2012
-
[40]
2025, Astron
Pellouin, C., Dvorkin, I., & Lehoucq, L. 2025, Astron. Astrophys., 693, A283
2025
-
[41]
Qin, Y. et al. 2024, Astron. Astrophys., 691, A214
2024
-
[42]
Regimbau, T. et al. 2012, Phys. Rev. D, 86, 122001
2012
-
[43]
M., Farrow, N., Stevenson, S., Thrane, E., & Zhu, X.-J
Romero-Shaw, I. M., Farrow, N., Stevenson, S., Thrane, E., & Zhu, X.-J. 2020, Mon. Not. Roy. Astron. Soc., 496, L64
2020
-
[44]
2020, Astrophys
Safarzadeh, M., Ramirez-Ruiz, E., & Berger, E. 2020, Astrophys. J., 900, 13
2020
-
[45]
2020, Phys
Shao, D.-S., Tang, S.-P., Jiang, J.-L., & Fan, Y.-Z. 2020, Phys. Rev. D, 102, 063006
2020
-
[46]
& Golomb, J
Talbot, C. & Golomb, J. 2023, Mon. Not. R. Astron. Soc., 526, 3495
2023
-
[47]
Tauris, T. M. et al. 2017, Astrophys. J., 846, 170
2017
-
[48]
Tenorio, R., Toubiana, A., Bruel, T., Gerosa, D., & Gair, J. R. 2025, Astrophys. J. Lett., 994, L52
2025
-
[49]
Toubiana, A., Wong, K. W. K., Babak, S., et al. 2021, Phys. Rev. D, 104, 083027
2021
-
[50]
2025, diptest: Hartigan's dip test for unimodality (Python/C++), python package implementing Hartigan and Hartigan's dip test for unimodality
Urlus, R. 2025, diptest: Hartigan's dip test for unimodality (Python/C++), python package implementing Hartigan and Hartigan's dip test for unimodality
2025
-
[51]
J., Stevenson , S., et al
Vigna-G \'o mez , A., Neijssel , C. J., Stevenson , S., et al. 2018, , 481, 4009
2018
-
[52]
M., & Taylor, S
Vitale, S., Gerosa, D., Farr, W. M., & Taylor, S. R. 2020, in Handbook of Gravitational Wave Astronomy (Springer)
2020
-
[53]
S., Fong, W.-f., Kremer, K., et al
Ye, C. S., Fong, W.-f., Kremer, K., et al. 2020, Astrophys. J. Lett., 888, L10
2020
-
[54]
S., Berry, C
Zevin, M., Bavera, S. S., Berry, C. P. L., et al. 2021, Astrophys. J., 910, 152
2021
-
[55]
E., Adhikari, S., et al
Zevin, M., Nugent, A. E., Adhikari, S., et al. 2022, Astrophys. J. Lett., 940, L18
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.