REVIEW 3 major objections 6 minor 51 references
Recovering Pulsar Braking Index from a Population of Millisecond Pulsars
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that pulsar braking-index distributions are recoverable from a population, but only after 20 years of timing at 10^-5 ms noise.
desk verdict Novel pipeline, honest feasibility study, but the 20-year 'population recovery' is really a single pulsar. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing quantity is the braking index $n$, defined by $\dot{f} \propto -f^n$ and observable in principle as $n = f\ddot{f}/\dot{f}^2$; it labels the dominant spin-down mechanism. The argument runs on a two-stage statistical machine. First, each pulsar's arrival times are fit to produce a posterior on $f$, $\dot{f}$, and $n$ directly, with $n$ given a uniform prior on $[0,10]$. Second, those posteriors are stacked by a hierarchical Bayesian model that assumes the population's $n$ values are drawn from a Gaussian parent $N(n|\mu,\sigma)$ and approximates each marginal likelihood by the mean of that Gaussian evaluated over the pulsar's posterior samples: $$\int P(D_i|n)\,N(n|\mu,\$\sigma$)\,dn \approx \frac{1}{m_i}\sum_{j=1}^{m_i} N(n_{ij}|\mu,\$\sigma$).$$ The output is summarised by the odds ratio $\mathrm{OR}_{4-6}$, the probability that a random population $n$ lies between 4 and 6 divided by the probability it lies outside that range; $\mathrm{OR}=1$ means the signal is as likely as not.
What would settle it
Repeat the 47-pulsar analysis with red timing noise injected at the amplitude and spectral slope measured in the 12.5-year wideband sample, holding observation length at 20 years and RMS noise at $10^{-5}$ ms; the central claim fails if $\mathrm{OR}_{4-6}$ drops below 1.
Extended reading notes
Core claim
The central claim is that hierarchical Bayesian stacking can recover the distribution of braking indices of a millisecond-pulsar population even when no individual pulsar's $\ddot{f}$ is measurable, but only under conditions that current data do not meet. With all 47 simulated pulsars spinning down at $n=5$, the method returns the correct mean with odds ratio $\mathrm{OR}_{4-6} \approx 1.0$ at 20 years and $10^{-5}$ ms RMS noise—a coin flip between $n$ lying in $[4,6]$ and outside it. Extending the baseline to 50 years raises $\mathrm{OR}_{4-6}$ to 9.65, whereas reducing noise from $10^{-5}$ to $10^{-6}$ ms only reaches 1.76. The paper concludes that observation time is the parameter with the largest impact on feasibility, that current 12.5-year, $3\times10^{-4}$ ms data cannot support the measurement, and that the study is a proof of principle for future arrays rather than an immediate observational result.
Load-bearing premise
The feasibility thresholds assume the simulated arrival times contain only white noise—red timing noise is discussed but never simulated—and that every pulsar has the same injected braking index; if red noise matters over 20-to-50-year baselines, the required observation times and noise levels would be worse.
Editorial extensions
If this is right
- At the current mean RMS of $3\times10^{-4}$ ms and a 12.5-year baseline, the stacking method does not confidently recover the injected $n=5$ population; $\mathrm{OR}_{4-6}$ falls below 1.
- Observation length is the dominant lever: holding noise at $10^{-5}$ ms, $\mathrm{OR}_{4-6}$ rises from 1.00 at 20 years to 1.46 at 30 years and 9.65 at 50 years.
- Improving timing precision matters less: a factor-of-10 noise reduction from $10^{-5}$ to $10^{-6}$ ms only increases $\mathrm{OR}_{4-6}$ from 1.00 to 1.76.
- Growing the sample helps: increasing from 10 to 47 pulsars roughly doubles the odds ratio, and one exceptionally well-constrained pulsar contributes disproportionately, so targeted high-precision timing could accelerate the measurement.
- The method is not biased toward the injected value: separate runs with $n=3$, $n=5$, and $n=7$ are all recovered with similar confidence at 20 years and $10^{-6}$ ms, so the approach can in principle distinguish spin-down mechanisms.
Reading between the lines
- The simulations use a delta-function population in which all 47 pulsars share $n=5$, the most favorable case for a Gaussian parent model; a real population mixing $n=3$ and $n=5$ would blur the recovered mean and width, so the quoted thresholds are likely optimistic.
- Red timing noise is acknowledged as the main observational obstacle to measuring $\ddot{f}$ but is never injected; because such noise is correlated in time, longer baselines may accumulate it rather than average it away, so the 50-year projection could degrade faster than the white-noise curves suggest.
- A natural test of the method's reach is to inject a two-component mixture (for example 70% $n=3$ and 30% $n=5$) and ask whether the Gaussian parent model can detect the mixture; if not, a mixture or histogram parent distribution would be needed before applying the method to real data.
- The same hierarchical stacking recipe could be transported to other weakly measured pulsar population parameters, such as ellipticity or magnetic inclination, though the paper does not make that extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a simulation-based feasibility study of measuring the distribution of pulsar braking indices from a population of millisecond pulsars. The authors generate simulated TOAs for 47 MSPs drawn from the NANOGrav 12.5-year sample, fixing the second frequency derivative to correspond to an injected braking index of n=5, and add white noise only. Per-pulsar posteriors on n are obtained by sampling directly over the braking index with a uniform prior on [0,10], and the posteriors are combined with a hierarchical Gaussian parent model using the posteriorstacker package. The success metric is an odds-like ratio OR4-6, the posterior probability that a draw from the recovered parent distribution lies between n=4 and n=6 divided by the complementary probability. The main results are that at a mean RMS of 1e-5 ms, a 20-year baseline gives OR4-6=1.00, a 50-year baseline gives OR4-6=9.65, and increasing the number of pulsars improves the recovery. The paper concludes that population-level braking-index recovery is possible but requires observation times over 20 years and very low timing noise.
Significance. If the method were validated, it would offer a way to infer the dominant spin-down mechanism of MSP populations without precise individual measurements of the second frequency derivative, connecting observable timing data to the n=5 gravitational-wave spin-down scenario. The pipeline is coherent and uses appropriate machinery: sampling directly over n is a sensible reparameterization, equation (8) correctly exploits the uniform prior to treat posteriors as likelihoods, and the closed-box injection-recovery tests are the right calibration procedure. The authors also provide a useful technical contribution in the modified enterprise_warp code and document numerical-precision changes. However, the central quantitative claim is weakened by two structural issues: the default 20-year result is effectively driven by a single pulsar, and the simulations exclude red timing noise. The paper is an honest proof-of-principle, but the population-level interpretation at the headline thresholds is not yet demonstrated.
major comments (3)
- [§3.3, Table 3] The population-recovery claim for the default 20-year, 1e-5 ms run is not supported by the per-pulsar results in Table 3. For 46 of the 47 pulsars the posterior mean and standard deviation are close to the imposed Uniform(0,10) prior values (mean 5, std 2.89), while only B1937+21 is tightly constrained (n = 5.01 ± 0.06). Section 3.3 further states that this pulsar dominates and is force-included in every subset run. Consequently the reported OR4-6 = 1.00 for the 47-pulsar run is effectively a single-object measurement, and the improvement with pulsar number in Figure 4 does not by itself demonstrate that weak information is being accumulated across the population. Please report the recovered hyperparameters and OR4-6 with B1937+21 removed, and quantify per-pulsar information content (for example, the Kullback-Leibler divergence of each posterior from its prior) to show that the stacking step genuinely combines information from multiple pulsars.
- [§2.1, §3.1, Abstract] The headline feasibility thresholds (observation times over 20 years and RMS noise of order 1e-5 ms) are derived from simulated TOAs containing white noise only, as stated in Section 2.1. Red timing noise is cited as a major obstacle to measuring the second frequency derivative (Liu et al. 2019) but is never injected into the simulations, and Section 4 acknowledges it only qualitatively. Because red noise is correlated over the 20-50 year baselines considered here, it could mimic or contaminate the f-dotdot signal and bias the recovered braking index. The authors should either add red-noise injections (even a simple power-law component) to test the thresholds, or explicitly qualify the abstract's claims as conditional on red noise being negligible; as written, the quantitative feasibility statement is overstrong.
- [§2.3, Appendix C] The recovery tests use a population in which every pulsar has the same injected n=5, a delta-function parent distribution, and Appendix C tests n=3, 5, and 7 only in separate runs. A realistic MSP population is more likely to contain a mixture of spin-down mechanisms, and the Gaussian parent model's ability to identify the dominant mechanism should be tested on a mixed population (for example, a fraction with n=3 and the remainder with n=5). Without such a test, the claim that the method can 'allow the underlying energy loss process to be inferred' is not fully established, because the favorable delta-function case sidesteps the question of whether the hierarchy can separate a narrow signal from a broad or multi-component distribution.
minor comments (6)
- [Abstract, §3.2] The phrase 'observation times of over 20 years' is ambiguous: the 20-year run gives OR4-6 = 1.00, which is an equivocal result, while significant confidence appears only for the 50-year run (OR4-6 = 9.65). Please rephrase to make clear that a confident population-level measurement requires substantially more than 20 years.
- [§2.4, Eq. (13)] The quantity OR4-6 is called an 'odds ratio', but it is a ratio of posterior probabilities, not a Bayes factor or a classical odds ratio. Consider renaming it 'probability ratio' or 'odds' to avoid confusion with standard statistical terminology.
- [§2.3, Eq. (8)] The proportionality P(Di|n) ∝ P(n|Di) relies on the prior for n being uniform over the full support of the posterior; this is stated in the text but should also be stated immediately after the equation for clarity.
- [Table 3] The combined column heading 'Mean n * Std n*' is confusing; please use separate columns for the posterior mean and posterior standard deviation, and explain in the caption that these are for the default 20-year, 1e-5 ms run.
- [§2.1, Eq. (6)] The sqrt(N) cadence-to-noise equivalence is valid only for white noise; since the paper later discusses red-noise concerns, this assumption should be flagged at the point where equation (6) is introduced.
- [Figure 4] The OR4-6 values are given only in the caption and are easy to miss; consider annotating the curves or adding a small table with the numerical values.
Circularity Check
No circularity: the study is an injection-recovery calibration and its reported limitations are stated honestly.
full rationale
The paper is an injection-recovery calibration rather than a claimed first-principles prediction. The value n=5 is intentionally placed into simulated TOAs by fixing F2 via equation (4), and the enterprise/posteriorstacker pipeline is then tasked with recovering it; this is the standard and non-circular design of a sensitivity study. No output quantity in the derivation chain is defined in terms of an input quantity, and no fitted parameter is renamed as a prediction. The uniform prior on n between 0 and 10 does have midpoint 5, and Table 3 shows that at the headline 20-year, 10^-5 ms configuration most pulsars return posterior standard deviations near the prior value of about 2.887, with only B1937+21 well constrained (n = 5.01 ± 0.06). The paper itself states in Section 3.3 that this pulsar dominates the results, and it honestly reports OR4-6 = 1.00 at that threshold. That is a limitation of the population-level demonstration, but it is not a circular reduction: the paper does not claim the weakly informative pulsars independently constrain n, and Appendix C shows recovery of n=3 and n=7 under the same uniform prior, indicating the method is not simply forced to the prior midpoint. The acknowledged simplifications (white-noise-only TOAs, no red noise, delta-function injected population) are scope limitations and are flagged in Sections 2.1 and 4 as an optimistic proof-of-principle scenario. Self-citations with overlapping authorship (Woan et al. 2018; Goncharov et al. 2024) are motivational or software-related and are not load-bearing for the central simulation results. The derivation chain is self-contained with respect to the stated feasibility claim.
Assumptions & free parameters
free parameters (4)
- Injected braking index for all simulated TOAs =
5 (chosen injection)
- Target mean RMS noise (mu_desired) =
1e-6 to 2.9e-4 ms
- Hyper-priors on the parent Gaussian =
uniform on mu, log-uniform on sigma
- Cadence-to-noise equivalence factor sqrt(N) =
N from 4 to 30
assumptions (5)
- domain assumption White-noise-only TOAs adequately represent MSP timing for f-dotdot inference
- domain assumption Parent braking-index distribution is Gaussian
- domain assumption Measured F1 is intrinsic spin-down, uncontaminated by acceleration or Shklovskii effects
- ad hoc to paper Uniform prior on n in [0,10] justifies posterior-as-likelihood
- standard math Nested sampling (dynesty) and the Monte Carlo replacement of the integral in equation (9) converge
Cite this review
Pith. "Pith review of Recovering Pulsar Braking Index from a Population of Millisecond Pulsars." pith.science (2026). https://pith.science/paper/47X2OGEN
@misc{pith2026241210169,
author = {Pith},
title = {Pith review of: Recovering Pulsar Braking Index from a Population of Millisecond Pulsars},
year = {2026},
howpublished = {\url{https://pith.science/paper/47X2OGEN}},
note = {Machine review of arXiv:2412.10169}
}
abstract
The braking index, $n$, of a pulsar is a measure of its angular momentum loss and the value it takes corresponds to different spin-down mechanisms. For a pulsar spinning down due to gravitational wave emission from the principal mass quadrupole mode alone, the braking index would equal exactly 5. Unfortunately, for millisecond pulsars, it can be hard to measure observationally due to the extremely small second time derivative of the rotation frequency, $\ddot{f}$. This paper aims to examine whether it could be possible to extract the distribution of $n$ for a whole population of pulsars rather than measuring the values individually. We use simulated data with an injected $n=5$ signal for 47 millisecond pulsars and extract the distribution using hierarchical Bayesian inference methods. We find that while possible, observation times of over 20 years and RMS noise of the order of $10^{-5}$ ms are needed, which can be compared to the mean noise value of $3\times10^{-4}$ ms for the recent wideband 12.5-year NANOGrav sample, which provided the pulsar timing data used in this paper.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter doi edition editor eprint howpublished institution journal key month number organization pages publisher school series title misctitle type volume year version url label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts ...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION format.url url empty "" new.block "" url * "" * if FUNCTION format.eprint eprint empty "" archivePrefix empty "" archivePrefix "arXiv" = new.block " " eprint * " " * new.block " " eprint * " " * if if if FUNCTION format.doi doi empty "" " " doi * " " * if FUNCTION format.pid doi empty eprint empty ur...
-
[3]
6m F_)mKk> ۈQR0 FH) ܞ
thebibliography [1] 20pt to REFERENCES 6pt =0pt -12pt 10pt plus 3pt =0pt =0pt =1pt plus 1pt =0pt =0pt -12pt =13pt plus 1pt =20pt =13pt plus 1pt \@M =10000 =-1.0em =0pt =0pt 0pt =0pt =1.0em @enumiv\@empty 10000 10000 `\.\@m \@noitemerr \@latex@warning Empty `thebibliography' environment \@ifnextchar \@reference \@latexerr Missing key on reference command E...
2021
-
[4]
Abbott , R., Abbott , T. D., Abraham , S., et al. 2021 a , PhRvX, 11, 021053, 10.1103/PhysRevX.11.021053
-
[5]
2021 b , , 915, L5, 10.3847/2041-8213/ac082e
---. 2021 b , , 915, L5, 10.3847/2041-8213/ac082e
-
[6]
2021 c , , 922, 71, 10.3847/1538-4357/ac0d52
---. 2021 c , , 922, 71, 10.3847/1538-4357/ac0d52
-
[7]
Abbott , R., Abbott , T. D., Acernese , F., et al. 2023, Physical Review X, 13, 041039, 10.1103/PhysRevX.13.041039
-
[8]
Agazie , G., Anumarlapudi , A., Archibald , A. M., et al. 2023, , 951, L8, 10.3847/2041-8213/acdac6
Show all 51 references
-
[9]
F., Arzoumanian , Z., Baker , P
Alam , M. F., Arzoumanian , Z., Baker , P. T., et al. 2021 a , , 252, 5, 10.3847/1538-4365/abc6a1
2021 doi
-
[10]
2021 b , , 252, 4, 10.3847/1538-4365/abc6a0
---. 2021 b , , 252, 4, 10.3847/1538-4365/abc6a0
2021 doi
- [11]
-
[12]
2021, , 507, 2037, 10.1093/mnras/stab2236
Ashton , G., & Talbot , C. 2021, , 507, 2037, 10.1093/mnras/stab2236
2021 doi
-
[13]
D., et al
Ashton , G., H \"u bner , M., Lasky , P. D., et al. 2019, , 241, 27, 10.3847/1538-4365/ab06fc
2019 doi
-
[14]
2020, , 498, 5284, 10.1093/mnras/staa2684
Baronchelli , L., Nandra , K., & Buchner , J. 2020, , 498, 5284, 10.1093/mnras/staa2684
2020 doi
-
[15]
1996, , 312, 675
Bonazzola , S., & Gourgoulhon , E. 1996, , 312, 675
1996
-
[16]
Buchner , J. 2021. https://github.com/JohannesBuchner/PosteriorStacker
2021
-
[17]
N., Guo , Y
Chen , S., Caballero , R. N., Guo , Y. J., et al. 2021, , 508, 4970, 10.1093/mnras/stab2833
2021 doi
-
[18]
2002, , 66, 084025, 10.1103/PhysRevD.66.084025
Cutler , C. 2002, , 66, 084025, 10.1103/PhysRevD.66.084025
2002 doi
-
[19]
T., Hobbs , G
Edwards , R. T., Hobbs , G. B., & Manchester , R. N. 2006, , 372, 1549, 10.1111/j.1365-2966.2006.10870.x
2006
-
[20]
A., Vallisneri, M., Taylor, S
Ellis, J. A., Vallisneri, M., Taylor, S. R., & Baker, P. T. 2020, ENTERPRISE: Enhanced Numerical Toolbox Enabling a Robust PulsaR Inference SuitE, Zenodo, 10.5281/zenodo.4059815
2020 doi
-
[21]
, Arumugam, P
EPTA Collaboration and InPTA Collaboration , Antoniadis, J. , Arumugam, P. , et al. 2023, A&A, 678, A50, 10.1051/0004-6361/202346844
2023 doi
- [22]
-
[23]
2021, , 507, 116, 10.1093/mnras/stab2048
Gittins , F., & Andersson , N. 2021, , 507, 116, 10.1093/mnras/stab2048
2021 doi
-
[24]
Goncharov , B. 2021. https://github.com/bvgoncharov/enterprise_warp
2021
-
[25]
2024, mattpitkin/enterprise\_warp: Braking index sampling, braking\_index\_sampling, Zenodo, 10.5281/zenodo.13274448
Goncharov, B., Zic, A., Reardon, D., & Pitkin, M. 2024, mattpitkin/enterprise\_warp: Braking index sampling, braking\_index\_sampling, Zenodo, 10.5281/zenodo.13274448
2024 doi
-
[26]
2015, Physical Review D, 91, 10.1103/physrevd.91.063007
Hamil, O., Stone, J., Urbanec, M., & Urbancová, G. 2015, Physical Review D, 91, 10.1103/physrevd.91.063007
2015 doi
-
[27]
R., Millman , K
Harris , C. R., Millman , K. J., van der Walt , S. J., et al. 2020, , 585, 357, 10.1038/s41586-020-2649-2
2020 doi
-
[28]
2006, Chinese Journal of Astronomy and Astrophysics Supplement, 6, 189
Hobbs , G., Edwards , R., & Manchester , R. 2006, Chinese Journal of Astronomy and Astrophysics Supplement, 6, 189
2006
-
[29]
J., et al
Hobbs , G., Jenet , F., Lee , K. J., et al. 2009, , 394, 1945, 10.1111/j.1365-2966.2009.14391.x
2009
-
[30]
Hunter , J. D. 2007, CSE, 9, 90, 10.1109/MCSE.2007.55
2007 doi
-
[31]
Lam , M. T. 2018, , 868, 33, 10.3847/1538-4357/aae533
2018 doi
-
[32]
Liu , K., Verbiest , J. P. W., Kramer , M., et al. 2011, , 417, 2916, 10.1111/j.1365-2966.2011.19452.x
2011
-
[33]
J., Keith , M
Liu , X. J., Keith , M. J., Bassa , C. G., & Stappers , B. W. 2019, , 488, 2190, 10.1093/mnras/stz1801
2019 doi
-
[34]
E., Johnston , S., Dunn , L., et al
Lower , M. E., Johnston , S., Dunn , L., et al. 2021, , 508, 3251, 10.1093/mnras/stab2678
2021 doi
- [35]
-
[36]
2000, , 354, 163
Palomba , C. 2000, , 354, 163
2000
-
[37]
2005, , 359, 1150, 10.1111/j.1365-2966.2005.08975.x
---. 2005, , 359, 1150, 10.1111/j.1365-2966.2005.08975.x
2005
-
[38]
M., et al
Parthasarathy , A., Johnston , S., Shannon , R. M., et al. 2020, , 494, 2012, 10.1093/mnras/staa882
2020 doi
-
[39]
M., et al
Parthasarathy , A., Bailes , M., Shannon , R. M., et al. 2021, , 502, 407, 10.1093/mnras/stab037
2021 doi
-
[40]
2013, arXiv e-prints, arXiv:1311.3693
Perrodin , D., Jenet , F., Lommen , A., et al. 2013, arXiv e-prints, arXiv:1311.3693. 1311.3693
2013 arXiv
-
[41]
M., Sip o cz , B
Price-Whelan , A. M., Sip o cz , B. M., G \"u nther , H. M., et al. 2018, , 156, 123, 10.3847/1538-3881/aabc4f
2018 doi
-
[42]
J., Zic , A., Shannon , R
Reardon , D. J., Zic , A., Shannon , R. M., et al. 2023, , 951, L6, 10.3847/2041-8213/acdd02
2023 doi
-
[43]
P., Tollerud , E
Robitaille , T. P., Tollerud , E. J., Greenfield , P., et al. 2013, , 558, A33, 10.1051/0004-6361/201322068
2013 doi
-
[44]
M., Talbot , C., Biscoveanu , S., et al
Romero-Shaw , I. M., Talbot , C., Biscoveanu , S., et al. 2020, , 499, 3295, 10.1093/mnras/staa2850
2020 doi
-
[45]
Speagle , J. S. 2020, , 493, 3132, 10.1093/mnras/staa278
2020 doi
-
[46]
W., Keane , E
Stappers , B. W., Keane , E. F., Kramer , M., Possenti , A., & Stairs , I. H. 2018, Philosophical Transactions of the Royal Society of London Series A, 376, 20170293, 10.1098/rsta.2017.0293
2018
-
[47]
R., Baker, P
Taylor, S. R., Baker, P. T., Hazboun, J. S., Simon, J., & Vigeland, S. J. 2021, enterprise\_extensions. https://github.com/nanograv/enterprise_extensions
2021
-
[48]
R., Baker, P
Taylor, S. R., Baker, P. T., Hazboun, J. S., et al. 2024, mattpitkin/enterprise\_extensions: Keep float128 precision , float128\_precision, Zenodo, 10.5281/zenodo.13274450
2024 doi
-
[49]
2000, , 319, 902, 10.1046/j.1365-8711.2000.03938.x
Ushomirsky , G., Cutler , C., & Bildsten , L. 2000, , 319, 902, 10.1046/j.1365-8711.2000.03938.x
2000
-
[50]
D., Haskell , B., Jones , D
Woan , G., Pitkin , M. D., Haskell , B., Jones , D. I., & Lasky , P. D. 2018, , 863, L40, 10.3847/2041-8213/aad86a
2018 doi
-
[51]
2023, Research in Astronomy and Astrophysics, 23, 075024, 10.1088/1674-4527/acdfa5
Xu, H., Chen, S., Guo, Y., et al. 2023, Research in Astronomy and Astrophysics, 23, 075024, 10.1088/1674-4527/acdfa5
2023 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.