REVIEW 3 major objections 5 minor 1 cited by
This paper argues that the counterintuitive increase in gravitational-wave background significance when cutting EPTA DR2 from 25 to 10.3 years is a ~2σ random noise fluctuation driven by sparse frequency coverage, not an anomaly; it also pi
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 19:54 UTC pith:PUCNUCY4
load-bearing objection A careful simulation study that plausibly explains the EPTA DR2 significance flip as a ~2σ noise artifact and, more convincingly, identifies spectral leakage as the dominant bias in the short dataset; worth refereeing, though the headline rates are conditioned on a single injected SMBHB population. the 3 major comments →
Data span and frequency coverage requirements for robust detection and inference in PTAs: A case study with EPTA DR2
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the apparent paradox in the second data release of the European Pulsar Timing Array—where the 10.3-year subset shows stronger gravitational-wave background evidence (HD S/N ≈3.5) than the full 25-year set (HD S/N ≈1.3)—is a normal statistical outcome. In 100 realistic simulations with an injected SMBHB background, the first decade of observations contributes almost nothing to the significance because those measurements are mostly single-frequency, so dispersion-measure noise cannot be separated from the achromatic background. The HD S/N distributions of the full and short datasets therefore overlap substantially: the short dataset wins by chance in about 15%
What carries the argument
Three mechanisms carry the argument. (1) A set of 100 realistic DR2-like simulations, each injecting the same gravitational-wave background from a circular SMBHB population (amplitude 2.4×10⁻¹⁵, slope 13/3) with different noise realisations, cut at 25, 10 and 5 years. (2) The noise-marginalised optimal statistic HD S/N and the weighted-average Jensen–Shannon divergence degeneracy coefficient, which together quantify how well individual pulsar noise can be separated from the common background. (3) Spectral leakage: a finite observation window correlates Fourier frequency bins that the standard diagonal-covariance analysis neglects; the FFTInt method accounts for this and removes the parameter
Load-bearing premise
All simulations inject a single realisation of the SMBHB background with a fixed amplitude and slope, and adopt the noise parameters and frequency coverage of the real EPTA data; if the true background spectrum or noise budget differs, the quoted 15%/5% rates and leakage-induced bias could change.
What would settle it
Reanalyse the real 10.3-year dataset with a leakage-aware method and check whether the recovered (log₁₀ A, γ) moves from roughly (−13.9, 2.7) toward the steeper spectrum expected from a SMBHB population; if it does not, spectral leakage is not the dominant bias. Also repeat the 100-realisation study with a different injected GWB spectrum or noise budget and see whether the 15%/5% rates change by more than statistical uncertainty.
If this is right
- The observed DR2new versus DR2full discrepancy can be reproduced in controlled simulations as a ~2σ noise fluctuation, so no data or pipeline anomaly needs to be invoked.
- When spectral leakage is included in the model, the short 10.3-year dataset yields unbiased gravitational-wave background parameters, meaning short datasets can be trusted if properly modelled.
- Combining the full EPTA dataset with complementary long-baseline arrays and low-frequency radio observations roughly doubles the HD S/N (from about 2.0 to 3.8) and makes the significance grow monotonically with observation time.
- Very short baseline datasets (about 5 years) with low S/N bias amplitude estimates toward high values and constrain the GWB slope poorly; slope precision depends more strongly on observation time than amplitude precision.
- The orientation of the recovered (amplitude, slope) posterior is predictable from a Fisher matrix and depends on timespan, not S/N, explaining the different orientations seen in real datasets of different lengths.
Where Pith is reading between the lines
- The quoted 15% and 5% rates are computed for a single injected SMBHB background realisation and a single noise budget; with different true spectra or noise levels these probabilities would shift, so the 5% figure should not be read as a universal p-value.
- The spectral-leakage bias identified here is likely to affect other short-baseline pulsar timing array datasets, so any ~4-5 year dataset claiming gravitational-wave background parameters should be checked with a leakage-aware analysis.
- The timespan-dependent posterior orientation implies that when combining datasets of very different lengths, joint analyses should explicitly account for their different correlation structures rather than assuming a common covariance.
- A testable extension: if this explanation is right, reanalysing the real 10.3-year dataset with the leakage-aware method should move the recovered amplitude and slope toward the steeper spectrum expected from a SMBHB population.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses a counterintuitive feature of the EPTA second data release: the 10.3-year subset DR2new yields a higher GWB significance (HD S/N ~3.5) than the full 25-year DR2full (HD S/N ~1.3). Using 100 realistic simulations of EPTA DR2 that share a single injected SMBHB background (A=2.4e-15, gamma=13/3) and realistic per-pulsar noise, the authors show that the first ten years contribute little because of sparse frequency coverage, so the HD S/N distributions of DR2full and DR2new overlap substantially. They find that DR2new exceeds DR2full in 15% of noise realizations and that 5% of realizations match the real data within 90% C.I., which they interpret as a ~2 sigma noise fluctuation. They further show that DR2new parameter estimates are biased toward low gamma and high amplitude, that this bias is largely removed by modeling spectral leakage with the FFTInt method, and that adding IPTA-like frequency coverage from NANOGrav, PPTA, LOFAR, and NenuFAR restores monotonic S/N growth and unbiased parameters. Short-baseline (5 yr) datasets are shown to be strongly biased, consistent with the posterior orientation predicted by a Fisher-matrix calculation.
Significance. If correct, the paper provides a coherent explanation of the EPTA DR2new/DR2full significance flip without invoking data anomalies, and it underlines two practical lessons for PTA analyses: frequency coverage can matter as much as time span, and neglect of spectral leakage can dominate parameter biases. The simulation pipeline is grounded in the real DR2 setup (libstempo, SPNA noise maxima, actual TOA epochs and cadences), and the analysis uses standard tools with 100 realizations, P-P plots, and direct comparisons to the real data. The main quantitative conclusion—a 5% occurrence rate corresponding to roughly 2 sigma—is nevertheless conditional on the chosen fixed GWB realization and on a data-dependent selection criterion, which limits the strength of the claim.
major comments (3)
- [Sec. 3.1.1 and Sec. 3.1.2] The central probabilities (15% of realizations with DR2new>DR2full, 5% compatible with the real HD S/N values) are computed for a single fixed injection of the SMBHB population. The GWB is a stochastic process; different population realizations with the same power spectrum (A=2.4e-15, gamma=13/3) will change the source positions, phases, and individual binary contributions, and therefore the HD S/N distributions. Since the real data are one GWB realization, the '~2σ noise fluctuation' claim in the abstract and conclusions should either be marginalized over GWB cosmic variance (by generating several population realizations) or explicitly qualified as conditional on the one tested population. The note in Sec. 3.1.2 that 'we only tested one SMBHB population realisation' makes the limitation clear, but it is not reflected in the strength of the abstract and summary statements.
- [Sec. 3.1.1, Table 2] The '5% compatible with the real data' is not a calibrated p-value. It is the fraction of simulations that satisfy a criterion defined after examining the real data: DR2new yields higher HD S/N than DR2full and the median simulated HD S/N values are within the 90% C.I. of the real values. Because the event was not pre-specified and the match definition was chosen post hoc, interpreting this fraction as 'the probability of observing the EPTA scenario' is problematic. Please report the exact selection rule, quote the Poisson sampling uncertainty (5/100 leads to a 95% interval of roughly 2–11%), and either use a pre-defined statistic or clearly label this as a descriptive conditional frequency rather than a p-value.
- [Sec. 3.1.3, Fig. 4c] The P-P plot and the quoted bias significances (3σ for DR2full, 5.6σ for DR2new standard, 1.5σ with FFTInt) use γ=13/3 as the 'true' value. The injected GWB is the sum of individual SMBHBs (Fig. 1), whose spectrum is not a pure power law; the power-law recovery model is therefore misspecified by construction, as the authors acknowledge in Sec. 3.1.2. Consequently, the measured 'bias' and its reduction by FFTInt are not purely attributable to spectral leakage. To support the claim that leakage is 'the main source of bias', the comparison should be made against the actual injected spectrum (e.g., the effective spectral index of the simulated GWB), or a control injection with a strictly power-law GWB should be added.
minor comments (5)
- [Abstract / Conclusions] The '≈2σ' language should be softened to reflect the conditional nature of the 5% rate, e.g., 'within our controlled simulations', rather than appearing as a general statement about the real-data probability.
- [Fig. 1 and Sec. 2.2.3] The power-law approximation line in Fig. 1 is not defined; state how it is fit. Also, the caveat that only one SMBHB population realization is used would be better placed at the first statement of the 15%/5% numbers.
- [Sec. 3.2] The simulated IPTA backends are added with fixed TOA errors (2 µs for high-frequency, 7 µs for low-frequency) and fixed epochs. Since the real backends have heterogeneous errors and cadences, please justify these choices against the actual NANOGrav/PPTA/LOFAR/NenuFAR data or state how sensitive the conclusions are to them.
- [Appendix B, Eq. (B.7)] Equation (B.7) depends on constants a and b that are fit to the simulated data. This should be stated explicitly in the main text where the Fisher prediction is invoked (Sec. 3.3), and the posterior uncertainty on these constants should be reported.
- [Typos] Minor typographical issues: 'Weigthed' in the captions of Figs. 3 and 9; 'stastistic' in Sec. 3.1.1; 'prividing' in Sec. 4; 'ampliutde' in Table 2; and the inconsistent use of 'EPOCH' vs 'epoch' in Sec. 3.2.
Circularity Check
Simulation-based ~2sigma noise-fluctuation explanation is self-contained; only the Appendix B angle formula is a fitted-input prediction.
specific steps
-
fitted input called prediction
[Appendix B, Eq. B.7 and Figure B.1]
"Since for F^{-1}_{Aγ} and F^{-1}_{γγ} we do not know the exact value, but only an approximation, we can write the exact expression for θ by adding two unknown variables ... By fitting this function to the θ measured from the 100 realisation of DR2, we find that the relation is in very good agreement with the behaviour of the data, as can be seen from Figure B.1."
The constants a and b in Eq. B.7 are free parameters fit to the same 100 simulated orientation angles that are then used to claim agreement. The 'prediction' is therefore matched to the data by construction, so the agreement shown in Figure B.1 is not an independent verification of the Fisher-matrix formula. This is a side result used in Sec. 3.3 to support the posterior-orientation discussion; it does not drive the paper's central noise-fluctuation claim.
full rationale
The central claim—that the DR2new vs DR2full significance flip can be a ~2sigma noise fluctuation—is a genuine simulation-based statement. The 15% and 5% frequencies are counted from 100 noise realisations of a fixed injected SMBHB-population GWB (A=2.4e-15, gamma=13/3) and noise parameters taken from the EPTA DR2 SPNA; no parameter is adjusted to reproduce the observed significance flip. The 5% selection uses the real HD S/N values to define a compatibility box, which is a posterior predictive check rather than an equivalence-by-construction. The frequency-coverage explanation is supported by the DR2 vs DR2FC controlled comparison inside the paper, not only by the self-cited Ferranti et al. (2025) degeneracy metric, so that self-citation is not load-bearing. The spectral-leakage bias is diagnosed with an external FFTInt method (Crisostomi et al. 2025) and compared against the injected spectrum, so no circularity arises there. The only genuine circular step is in Appendix B, where free constants a and b are fit to the same 100 simulated angles that are then declared to be in good agreement with the fitted formula. That is a minor, side-result circularity and does not undermine the paper's main inference.
Axiom & Free-Parameter Ledger
free parameters (5)
- Injected GWB amplitude and slope =
A_GWB = 2.4e-15, γ = 13/3 (α = -2/3)
- Per-pulsar noise injection parameters (RN, DM, EFAC/EQUAD) =
Per-pulsar maximum-likelihood values from EPTA DR2 SPNA
- FFTInt low-frequency cutoff =
1/3 T_obs
- Fisher-formula constants a and b =
Fitted to 100 simulated principal-axis angles
- Simulated IPTA backend properties =
TOA errors 2µs and 7µs; start epochs 53000, 56500, 58300; 100 LOFAR and NenuFAR TOAs
axioms (6)
- domain assumption Gaussian stationary noise and timing-model-marginalized likelihood (van Haasteren & Vallisneri 2014)
- domain assumption Injected SMBHB population consists of circular binaries with a power-law spectrum
- standard math Standard analysis assumes a diagonal frequency-domain covariance matrix
- domain assumption Noise parameters from SPNA are representative of the true noise in all realizations
- domain assumption FFTInt/Discovery implementation correctly models spectral leakage
- ad hoc to paper The selection criterion for 'compatible with the real data' defines the reported 5% probability
read the original abstract
Pulsar Timing Arrays (PTAs) are approaching the sensitivity required for a $5\sigma$ detection of the nanohertz stochastic gravitational-wave background (GWB). This makes it crucial to deeply understand the behaviour of our analysis pipelines. A counterintuitive feature of the European Pulsar Timing Array (EPTA) second data release is that restricting the dataset to the last 10.3 years (DR2new) increases the inferred GWB significance from $\leq2\sigma$ for the full 25-year dataset (DR2full) to $\geq3.5\sigma$. We investigate whether this behaviour indicates an anomaly or is a possible outcome of the pipeline. Using realistic, DR2-like simulations with varying timespans, we find that the first 10 years contribute little to the GWB evidence due to their limited frequency coverage. This produces substantial overlap between the HD S/N distributions of DR2full and DR2new. Random noise fluctuations therefore yield a higher GWB evidence in DR2new than in DR2full in $15\%$ of cases. Furthermore, $5\%$ of simulations match the HD S/N of the real data, indicating that the observed behaviour is consistent with being a $\sim2\sigma$ outcome due to noise fluctuations. Regardless of significance, DR2new simulations introduce biases in the GWB parameter estimation due to spectral leakage effects that are ignored in standard analyses and which flatten the inferred spectrum. Including leakage removes these biases, demonstrating the reliability of DR2new when the signal is properly modelled. Furthermore, we demonstrate that combining EPTA DR2full with long-baseline data from NANOGrav and PPTA, as well as low-frequency data from LOFAR and NenuFAR, significantly enhances GWB evidence and parameter accuracy. Finally, we examine the impact of the observation timespan and find that short-baseline datasets introduce strong amplitude biases and are ineffective at constraining the GWB.
Figures
Forward citations
Cited by 1 Pith paper
-
The NANOGrav 15 yr Data Set: Customized Chromatic Noise Models
Customized chromatic noise models for 67 pulsars detect non-dispersive delays in 21 cases, alter achromatic noise inferences in 19, and enable solar wind density estimates over 1.5 cycles.
Reference graph
Works this paper leans on
-
[1]
2023, ApJ, 951, L11
Afzal, A., Agazie, G., Anumarlapudi, A., et al. 2023, ApJ, 951, L11
2023
-
[2]
Agazie, G., Anumarlapudi, A., Archibald, A. M., et al. 2023, arXiv e-prints, arXiv:2306.16221
Pith/arXiv arXiv 2023
-
[3]
M., et al
Agazie, G., Anumarlapudi, A., Archibald, A. M., et al. 2023, The Astrophysical Journal Letters, 951, L10
2023
-
[4]
2024, Phys
Babak, S., Falxa, M., Franciolini, G., & Pieroni, M. 2024, Phys. Rev. D, 110, 063022
2024
-
[5]
G., Janssen, G
Bassa, C. G., Janssen, G. H., Stappers, B. W., et al. 2016, MNRAS, 460, 2207
2016
-
[6]
& Figueroa, D
Caprini, C. & Figueroa, D. G. 2018, Classical and Quantum Gravity, 35, 163001
2018
-
[7]
Crisostomi, M., van Haasteren, R., Meyers, P. M., & Vallisneri, M. 2025, arXiv e-prints, arXiv:2506.13866
Pith/arXiv arXiv 2025
-
[8]
Y ., Verbiest, J
Donner, J. Y ., Verbiest, J. P. W., Tiburzi, C., et al. 2020, A&A, 644, A153
2020
-
[9]
T., Hobbs, G
Edwards, R. T., Hobbs, G. B., & Manchester, R. N. 2006, MNRAS, 372, 1549
2006
-
[10]
Ellis, J. A. 2013, Classical and Quantum Gravity, 30, 224004 EPTA and InPTA Collaboration, Antoniadis, J., Arumugam, P., et al. 2024, A&A, 685, A94 EPTA and InPTA Collaboration, Antoniadis, J., Arumugam, P., et al. 2023a, A&A, 678, A50 EPTA and InPTA Collaboration, Antoniadis, J., Arumugam, P., et al. 2023b, A&A, 678, A49 EPTA and InPTA Collaboration, Ant...
2013
-
[11]
2025, A&A, 694, A38
Ferranti, I., Falxa, M., Sesana, A., et al. 2025, A&A, 694, A38
2025
-
[12]
Foster, R. S. & Backer, D. C. 1990, ApJ, 361, 300
1990
-
[13]
& Sardana, S
Goncharov, B. & Sardana, S. 2025, Monthly Notices of the Royal Astronomical Society, 537, 3470–3479
2025
-
[14]
2025, Reading signatures of super- massive binary black holes in pulsar timing array observations
Goncharov, B., Sardana, S., Sesana, A., et al. 2025, Reading signatures of super- massive binary black holes in pulsar timing array observations
2025
-
[15]
Hellings, R. W. & Downs, G. S. 1983, ApJ, 265, L39
1983
-
[16]
2025, arXiv e-prints, arXiv:2510.04639 Article number, page 15 of 16 A&A proofs:manuscript no
Iraci, F., Chalumeau, A., Tiburzi, C., et al. 2025, arXiv e-prints, arXiv:2510.04639 Article number, page 15 of 16 A&A proofs:manuscript no. main
arXiv 2025
-
[17]
Jaffe, A. H. & Backer, D. C. 2003, ApJ, 583, 616
2003
-
[18]
Larsen, B., Mingarelli, C. M. F., Baker, P. T., et al. 2025, MNRAS, 542, 3028
2025
-
[19]
M., Coles, W
Lentati, L., Shannon, R. M., Coles, W. A., et al. 2016, MNRAS, 458, 2161
2016
-
[20]
Lorimer, D. R. & Kramer, M. 2004, Handbook of Pulsar Astronomy, V ol. 4
2004
-
[21]
Miles, M. T., Shannon, R. M., Reardon, D. J., et al. 2025, MNRAS, 536, 1489 Quelquejay Leclere, H., Li, K., V olonteri, M., et al. 2025, arXiv e-prints, arXiv:2510.14613
arXiv 2025
-
[22]
& Romani, R
Rajagopal, M. & Romani, R. W. 1995, ApJ, 446, 543
1995
-
[23]
J., Zic, A., Shannon, R
Reardon, D. J., Zic, A., Shannon, R. M., et al. 2023, ApJ, 951, L6
2023
-
[24]
A., Sesana, A., & Gair, J
Rosado, P. A., Sesana, A., & Gair, J. 2015, Monthly Notices of the Royal Astro- nomical Society, 451, 2417–2433
2015
-
[25]
2013, Classical and Quantum Gravity, 30, 244009
Sesana, A. 2013, Classical and Quantum Gravity, 30, 244009
2013
-
[26]
Sesana, A., Vecchio, A., & Colacino, C. N. 2008, MNRAS, 390, 192
2008
-
[27]
K., Falxa, M., et al
Speri, L., Porayko, N. K., Falxa, M., et al. 2023, MNRAS, 518, 1802
2023
-
[28]
C., Chalumeau, A., Tiburzi, C., et al
Susarla, S. C., Chalumeau, A., Tiburzi, C., et al. 2024, A&A, 692, A18
2024
-
[29]
2016, MNRAS, 455, 4339
Tiburzi, C., Hobbs, G., Kerr, M., et al. 2016, MNRAS, 455, 4339
2016
-
[30]
2020, libstempo: Python wrapper for Tempo2, Astrophysics Source Code Library, record ascl:2002.017
Vallisneri, M. 2020, libstempo: Python wrapper for Tempo2, Astrophysics Source Code Library, record ascl:2002.017
2020
-
[31]
2024, A&A, 683, A201 van Haasteren, R
Valtolina, S., Shaifullah, G., Samajdar, A., & Sesana, A. 2024, A&A, 683, A201 van Haasteren, R. 2024, ApJS, 273, 23 van Haasteren, R. 2025, MNRAS, 537, L1 van Haasteren, R. & Vallisneri, M. 2014, Phys. Rev. D, 90, 104012
2024
-
[32]
Verbiest, J. P. W., Lentati, L., Hobbs, G., et al. 2016, MNRAS, 458, 1267
2016
-
[33]
Verbiest, J. P. W. & Shaifullah, G. M. 2018, Classical and Quantum Gravity, 35, 133001
2018
-
[34]
J., Islo, K., Taylor, S
Vigeland, S. J., Islo, K., Taylor, S. R., & Ellis, J. A. 2018, Phys. Rev. D, 98, 044003
2018
-
[35]
Wyithe, J. S. B. & Loeb, A. 2003, ApJ, 590, 691
2003
-
[36]
2023, Research in Astronomy and Astrophysics, 23, 075024 Article number, page 16 of 16
Xu, H., Chen, S., Guo, Y ., et al. 2023, Research in Astronomy and Astrophysics, 23, 075024 Article number, page 16 of 16
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.