REVIEW 4 major objections 5 minor 1 cited by
Analysis and simulations of binary black hole merger spins -- the question of spin-axis tossing at black hole formation
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Simulations with random spin-axis tossing at black hole formation reproduce the observed binary black hole spins, while simulations without it fail.
desk verdict A transparent, useful Monte Carlo comparison whose no-tossing result is robust under strict alignment assumptions, but whose tossing evidence is partly fitted to the same data — so the conclusion stays conditional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a Monte Carlo model of the final core collapse in a synthetic population of isolated binaries. The simulation draws nine parameters: the two black hole masses and spin magnitudes, the pre-supernova orbital separation, the kick velocity and two isotropic direction angles, and the tossing angle $\Phi_2$ of the second-born black hole's spin axis. The tilt assignments do the decisive work: accretion alignment gives $\Theta_1=\delta$, tidal locking gives $\Theta_2=\delta$ in the no-tossing case, and isotropic tossing replaces $\Theta_2$ with draws from $P(\Theta_2)=\frac12\sin\Theta_2$. Simulated $\chi_{\rm eff}$ distributions are compared with the 83 observed mergers using kernel density estimates, functional boxplots, and the Kolmogorov-Smirnov, Cramér-von Mises, and Anderson-Darling tests, with each observed value varied inside its 90% credibility interval.
What would settle it
Look for a single isolated-binary merger with a confidently negative effective spin and a small inferred natal kick: the no-tossing model predicts that such a system always has positive $\chi_{\rm eff}$ because the maximum tilt is only $6.42^\circ$, so one clean counterexample would falsify the no-tossing branch. Conversely, a large future sample of field mergers with only positive effective spins would remove the statistical need for tossing.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that the measured effective spin distribution of binary black hole mergers is a readable record of the second collapse. Because accretion torques align the first-born black hole's spin with the pre-supernova orbit, its tilt angle $\Theta_1$ is just the kick-induced misalignment $\delta$; in the widely used no-tossing picture, tidal locking gives the second-born black hole the same tilt, $\Theta_2=\delta$, so all merger spins stay aligned with the orbit and only positive $\chi_{\rm eff}$ values result (maximum tilt $6.42^\circ$). Against the 83 observed events this scenario is statistically excluded, with p-values below 0.001. If instead the second-born black hole's spin axis is drawn from an isotropic tossing distribution, $P(\Theta_2)=\frac12\sin\Theta_2$ on $[0,180^\circ]$, the simulated $\chi_{\rm eff}$ distribution matches the data with p-values as high as 0.882 (KS), 0.742 (CvM), and 0.250 (AD, capped). The authors therefore conclude that the observations support spin-axis tossing if isolated binaries dominate the merger channel.
Load-bearing premise
The argument stands on one premise: accretion fully aligns the first-born black hole's spin with the pre-supernova orbit, and tides lock the collapsing helium star to that same orbit, so without tossing both black hole tilts equal the kick angle.
Editorial extensions
If this is right
- Population synthesis of isolated binary black hole mergers should include spin-axis tossing of the second-formed black hole; the paper suggests a fully isotropic distribution as a simple default.
- A no-tossing interpretation is not dead, but it requires roughly 72±8% of detected mergers to come from dynamical channels with random spin directions.
- The best isolated-binary fits prefer mass reversal in about 30% of progenitor systems, aligning with independent estimates from spin data.
- The tossing scenario's high p-values mean it cannot be rejected with current data; future observing runs and third-generation detectors should sharpen the $\chi_{\rm eff}$ comparison and probe whether tossing is isotropic or depends on progenitor properties.
Reading between the lines
- If tossing is real, the toss angle should not be universal: the direction and magnitude of the toss likely depend on the pre-supernova structure, mass loss, and orbital period, so future data could search for a conditional, non-isotropic $\Phi_2$ distribution.
- The two surviving scenarios—tossing-dominated isolated binaries versus roughly 72% dynamical mergers—make different predictions for merger eccentricities, host environments, and rates, so combining $\chi_{\rm eff}$ with those observables could break the degeneracy.
- The ~30% mass-reversal preference implies that which black hole is spinning faster already encodes which one formed second; more precise individual-spin measurements could turn this into a direct formation-order test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes the LVK O1-O3 effective-spin (chi_eff) measurements of 83 binary black hole (BH+BH) mergers and compares the empirical distribution with Monte Carlo simulations of the second supernova in isolated binary evolution, with and without spin-axis tossing at BH formation. The authors use kernel density estimation, functional boxplots, and three two-sample tests (KS, CvM, AD). Their main finding is that isolated-binary simulations without spin-axis tossing always give p-values below 0.001, whereas simulations with tossing can match the data with p-values up to 0.882. They also infer a mass-reversal fraction near 29% and, in the no-tossing case, a dynamical-origin fraction of about 72±8%. The central claim is that the data support the spin-axis tossing hypothesis if isolated binaries dominate the BH+BH merger channel.
Significance. If the central contrast between tossing and no-tossing simulations is robust, the paper would provide population-level support for Tauris's (2022) spin-axis tossing hypothesis and a transparent way to constrain isolated versus dynamical formation fractions. The strengths of the paper are its explicit Monte Carlo setup, the use of three complementary statistical tests, the careful treatment of chi_eff measurement uncertainty through functional boxplots, and the fact that the predictions can be checked with O4+O5 and 3G detector data. The main risk is that the central no-tossing rejection depends on an alignment assumption that the paper itself flags as contested, and several model inputs are fitted to the same empirical data used for the final p-values.
major comments (4)
- [Section 4 (opening items i–ii), Section 4.4.1, Fig. 16] The statement that no-tossing isolated binaries cannot produce the observed negative chi_eff tail is load-bearing, and it depends on two assumptions stated in the text: the first-born BH spin is fully aligned with the pre-SN orbital angular momentum so that Θ1=δ, and tidal locking forces Θ2=δ in the no-tossing case. Footnote 2 cites Baibhav & Kalogera (2024) as challenging exactly this alignment story, but the manuscript does not quantify the sensitivity. If the first-born BH retains a misaligned spin component, or if tidal locking is incomplete, Θ1 or Θ2 can exceed 90° even with small kicks, producing negative chi_eff without any tossing. I request a quantitative test: rerun the no-tossing simulations with a misalignment distribution for the first-born BH (or a residual tilt for the helium star) and report the resulting p-values and the fraction of systems with chi_eff<0.
- [Sections 4.2.1–4.2.2 and Fig. 8] The chi1 and chi2 beta distributions are obtained by fitting the simulated chi_eff curve to the empirical functional-boxplot LSCV curve; these same empirical data are then used as the reference for all later p-values. This is a circular step that inflates the reported agreement: the model is scored against the same dataset used to set its spin inputs. The later use of the GWTC-3 chiA/chiB credibility intervals (Section 4.6) partially mitigates this, but the Section 4.2 spin inputs remain data-fitted. Please quantify how much of the tossing/no-tossing separation survives when chi1 and chi2 are instead drawn from independent population estimates or from the full prior range allowed by Fig. 18.
- [Section 4.4, Eq. (9)] The 'fitted tossing' distribution P(Φ2) = 1/2[β(1.36,2.51)+β(7.9,5.3)] is fitted to minimize RMSE to the same empirical LSCV curve, so the subsequent comparison (Fig. 15C and Section 4.6) uses a distribution that has already been tuned to the data. This makes the high p-values for the tossing scenario partly a measure of the flexibility of the two-beta mixture rather than evidence for isotropic tossing. The paper should state this explicitly and provide out-of-sample validation, for example by using only O1-O2 data for fitting and O3 for testing, or by forecasting the chi_eff distribution for O4.
- [Section 4.6, Fig. 18] The final high p-value (0.882) is obtained from panel A, which the text itself calls an extreme fit that maximizes chiA and minimizes chiB within the 90% credibility intervals, combined with the 29% mass-reversal fraction that was itself selected to maximize p-values in Section 4.5. As stated, the analysis selects the best case among several alternatives and then reports that best case as the headline. The choice of panel A needs to be presented as a selection effect: report the p-values for all panels (A-D) and, ideally, the distribution of p-values over randomly drawn spin-component curves inside the credibility region, so the reader can see how often the tossing/no-tossing contrast is achieved.
minor comments (5)
- [Eq. (1)] The definition q ≡ M2/M1 ≤ 1 is used in Eq. (1), but Section 4.5 explicitly discusses q > 1 after mass reversal; the definition should allow q > 1 or the mass-reversal text should refer to the re-labelled mass ratio.
- [Fig. 14 and Eq. (8)] The symbol Φ2 is used both for the tossing angle and for the post-SN spin tilt angle Θ2; please unify the notation to avoid confusion.
- [Table 1 and Section 4.5] The AD p-values are capped at 0.250, but the main text does not explain this cap until the table note; please state this where the p-values are first quoted.
- [Section 3.2.1] The detection-bias discussion dismisses the effect on the basis of Vitale et al. (2022), but a sentence explaining the sign and expected size of the bias would help the reader assess the sensitivity of the negative-tail inference.
- [Section 4.7] The phrase 'we cannot completely rule out this scenario' for Fdyn=1 is stronger than the preceding statistical discussion warrants; consider rewording to 'not strongly excluded' or similar.
Circularity Check
Partial circularity: the tossing model's high p-values are partly in-sample fits, because the spin magnitudes, mass-reversal fraction, and best spin-component panel are all fitted to the LVK χeff data; the no-tossing contrast is a genuine geometric argument but rests on an untested perfect-alignment assumption flagged in a footnote.
-
fitted input called prediction
[Section 4 bullet list; Sections 4.2.1–4.2.2 and 4.6]
"The simulation methodology and setup applied here is explained in detail in Sections 4.1–4.7 and similar to those presented in Tauris (2022), except for the following: BH spins (χ1, χ2) are chosen from χeff-fits to LVK data (or following LVK spin components, cf. Sections 4.6–4.7)."
The simulated χeff distribution is generated from Eq. (1) using spin magnitudes χ1 and χ2 that were themselves fitted to the empirical χeff distribution (Section 4.2.1: 'minimizing the root-mean-square error (RMSE) to a predetermined normalized curve' derived from the LVK functional boxplot). Reporting p-values up to 0.882 for the tossing simulations in Section 4.6 is therefore an in-sample goodness-of-fit measure, not an independent confirmation of the tossing hypothesis; the high p-value is partly guaranteed by the parameter fit. The no-tossing comparison is a separate geometric argument, but the positive support for tossing reduces in part to fitted spin inputs that were tuned to reproduce the target distribution.
-
fitted input called prediction
[Section 4.5, 'Mass reversal'; Section 4.6]
"Figure 17 and Table 1 show that choosing the fraction of binaries with mass reversal to about 30% results in significantly better agreement with the empirical data. Our result is thus very much in line with Mould et al. (2022)."
The 30% mass-reversal fraction is obtained by maximizing the p-value of the simulated χeff against the empirical LVK χeff sample (Fig. 17 and Table 1), and the same empirical sample is the comparison target for the 'final simulations' with 29% mass reversal in Section 4.6. Reporting 'a preference for mass reversal in ~30% of the progenitor binaries' is therefore a statement of the fitted optimum, not an out-of-sample prediction. The circularity is limited because the fraction independently agrees with Mould et al. (2022), and it is not the primary tossing claim.
1 more flagged steps
-
fitted input called prediction
[Section 4.6, 'Final simulations using LVK spin components']
"We notice that the simulation of χeff that produced by far the highest p-values from the two-sided tests is the one shown in panel A of Fig. 19 which is based on the individual BH spin components plotted in panel A of Fig. 18, where the difference in spins between the second-born BH (predominantly fast spinning) and the first-born BH (predominantly slow spinning) is the largest."
The headline p-value of 0.882 is the maximum over a set of spin-component distributions (panels A–D of Fig. 18) that were constructed to lie within the 90% credibility intervals of the LVK component-spin measurement (Abbott et al., 2023b). Selecting the best-fitting panel before reporting the p-value makes the quoted statistic a post-selection maximum rather than a predictive test of the tossing model. This does not undermine the independent no-tossing contrast, but it inflates the apparent evidential value of the tossing simulations and is another fitted input presented as simulation evidence.
full rationale
The central comparison has genuine independent content: in the no-tossing case the paper sets Θ2 = Θ1 = δ, finds δmax = 6.42° for its adopted kicks, and then χeff is positive by Eq. (1), so the p-values below 0.001 (Section 4.4, 4.6) are not an artifact of fitting. This geometric argument provides real evidence against the no-tossing scenario under the paper's stated assumptions. However, that conclusion depends on the load-bearing assumption that the first-born BH spin is fully aligned with the pre-SN orbital angular momentum and that the collapsing helium star is tidally locked (Section 4, intro items i–ii), and the paper's own footnote to Section 4.4 flags 'but see Baibhav and Kalogera, 2024' without testing this counter-scenario. That is a correctness/robustness caveat rather than a circular reduction. The circularity comes from the positive side of the claim: the spin magnitudes (χ1, χ2), the mass-reversal fraction (30%), and the spin-component panel A are all fitted to the LVK χeff or component-spin data, and the resulting simulated χeff distributions are then compared to the same LVK data to claim 'strong indications for spin-axis tossing' and p-values up to 0.882. The success of the tossing simulations is therefore partly constructed by the fit, although the no-tossing failure remains an independent and geometrically forced contrast within the paper's model space.
Assumptions & free parameters
free parameters (7)
- chi1 beta-distribution shape parameters =
alpha=1.01, beta=9.69
- chi2 beta-distribution shape parameters =
alpha=2.72, beta=2.34
- Spin-toss angle distribution P(Phi2) =
0.5[beta(1.36,2.51)+beta(7.9,5.3)]
- Mass-reversal fraction f =
approx 0.30 (29% in final runs)
- Natal kick speed w =
50 km/s
- Pre-SN orbital separation range =
[4, 40] R_sun
- Fast and slow spin component distributions (chiA, chiB) =
Panel A of Fig. 18
assumptions (8)
- domain assumption The spin axis of the first-born BH is aligned with the pre-SN orbital angular momentum vector through accretion, so its tilt equals the kick-induced misalignment angle delta.
- domain assumption Tidal torques align the collapsing helium star's spin with the orbital angular momentum, so without tossing Theta2 = delta as well.
- domain assumption Natal kicks of several hundred km/s are incompatible with observed Galactic BH binaries, so orbit-flipping kicks are excluded.
- domain assumption Pre-SN helium star mass relates to the second-born BH mass as M_He = M_BH,2 / 0.8.
- standard math Merger time of bound post-SN binaries follows Peters (1964).
- domain assumption Dynamical formation channels produce BH component spins with random isotropic orientations.
- domain assumption No significant chi_eff-dependent detection selection bias exists in LVK O1-O3.
- domain assumption Kick directions are isotropic; pre-SN orbital separations are uniform in [4,40] R_sun with constant kick w=50 km/s in the standard setup.
Cite this review
Pith. "Pith review of Analysis and simulations of binary black hole merger spins -- the question of spin-axis tossing at black hole formation." pith.science (2026). https://pith.science/paper/B5ON6HNY
@misc{pith2026250803809,
author = {Pith},
title = {Pith review of: Analysis and simulations of binary black hole merger spins -- the question of spin-axis tossing at black hole formation},
year = {2026},
howpublished = {\url{https://pith.science/paper/B5ON6HNY}},
note = {Machine review of arXiv:2508.03809}
}
read the original abstract
The origin of binary black hole (BH) mergers remains a topic of active debate, with effective spins (chi_eff) measured by the LIGO-Virgo-KAGRA (LVK) Collaboration providing crucial insights. In this study, our objective is to investigate the empirical chi_eff distribution (and constrain individual spin components) of binary BH mergers and compare them with extensive simulations, assuming that they originate purely from isolated binaries or a mixture of formation channels. We explore scenarios using BH kicks with and without the effect of spin-axis tossing during BH formation. We employ simple yet robust Monte Carlo simulations of the final core collapse forming the second-born BH, using minimal assumptions to ensure transparency and reproducibility. The synthetic chi_eff distribution is compared to the empirical data from LVK science runs O1-O3 using functional data analysis, kernel density estimations, and three different statistical tests, accounting for data uncertainties. We find strong indications for spin-axis tossing during BH formation if LVK sources are dominated by the isolated binary channel. Simulations with spin-axis tossing achieve high p-values (up to 0.882) using Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling tests, while without tossing, all p-values drop below 0.001 for isolated binaries. A statistically acceptable solution without tossing, however, emerges if ~72+/-8% of detected binary BH mergers result from dynamical interactions causing random BH spin directions. Finally, for an isolated binary origin, we find a preference for mass reversal in ~30% of the progenitor binaries. Predictions from this study can be tested with LVK O4+O5 data as well as the 3G detectors, Einstein Telescope and Cosmic Explorer, enabling improved constraints on formation channel ratios and the critical question of BH spin-axis tossing.
Forward citations
Cited by 1 Pith paper
-
Trails of clouds in binary black holes
Boson clouds around binary black holes generically deplete through orbital resonances, driving eccentricity and spin-orbit tilt toward fixed points—including off-equatorial ones—leaving observable gravitational-wave trails.
Reference graph
Works this paper leans on
-
[363]
Heinzel, J., Mould, M., Vitale, S., 2025
doi: 10.1086/429868, arXiv:astro-ph/0409422. Heinzel, J., Mould, M., Vitale, S., 2025. Nonparametric analysis of correla- tions in the binary black hole population with LIGO-Virgo-KAGRA data. Phys. Rev. D 111, L061305. doi:10.1103/PhysRevD.111.L061305. Heinzel, J., Vitale, S., Biscoveanu, S., 2024. Probing correlations in the bi- nary black hole populatio...
arXiv 2025
-
[2017]
Distinguishing spin-aligned and isotropic black hole populations with gravitational waves. Nature 548, 426–429. doi: 10.1038/nature23453, arXiv:1706.01385. Fragione, G., Loeb, A., Rasio, F.A., 2021. Impact of Natal Kicks on Merger Rates and Spin-Orbit Misalignments of Black Hole-Neutron Star Mergers. ApJ 918, L38. doi: 10.3847/2041-8213/ac225a, arXiv:2108...
arXiv 2021
-
[2018]
The spin of the second-born black hole in coalescing binary black holes. A&A 616, A28. doi: 10.1051/0004-6361/201832839, arXiv:1802.05738. Quessy, J.F., ´Ethier, F., 2012. Cram ´er–von mises and characteristic function tests for the two and k-sample problems with dependent data. Computational Statistics & Data Analysis 56, 2097–2111. Razali, N.M., Wah, Y ...
arXiv 2012
-
[2020]
Evolutionary roads leading to low e ffective spins, high black hole masses, and O1 /O2 rates for LIGO /Virgo binary black holes. A&A 636, A104. doi: 10.1051/0004-6361/201936528, arXiv:1706.07053. Biscoveanu, S., Isi, M., Vitale, S., Varma, V ., 2021. New Spin on LIGO- Virgo Binary Black Holes. Phys. Rev. Lett. 126, 171103. doi: 10.1103/ PhysRevLett.126.17...
-
[2022]
The cosmic evolution of binary black holes in young, globular, and nuclear star clusters: rates, masses, spins, and mixing fractions. MNRAS 511, 5797–5816. doi: 10.1093/mnras/stac422, arXiv:2109.06222. Marchant, P., Bodensteiner, J., 2024. The Evolution of Massive Binary Stars. ARA&A 62, 21–61. doi: 10.1146/annurev-astro-052722-105936, arXiv:2311.01865. M...
-
[2024]
Ultrasoft state of microquasar Cygnus X-3: X-ray polarimetry reveals the geometry of the astronomical puzzle. A&A 688, L27. doi: 10.1051/ 0004-6361/202451356, arXiv:2407.02655. Vigna-G´omez, A., Toonen, S., Ramirez-Ruiz, E., Leigh, N.W.C., Riley, J., Haster, C.J., 2021. Massive Stellar Triples Leading to Sequential Binary Black Hole Mergers in the Field. ...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.