REVIEW 2 major objections 5 minor 4 cited by
A cross-correlation-only detection statistic, NPMV, that is as close as possible to the Neyman-Pearson optimal test raises pulsar-array GWB detection probability by about 47% at the 5-sigma threshold.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 23:29 UTC pith:EUHMRV6L
load-bearing objection The NPMV construction is clean and likely beats DFCC, but the headline 47% improvement rests on ROC curves whose computation is never described. the 2 major comments →
Optimal robust detection statistics for pulsar timing arrays
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For zero-mean complex Gaussian data, the NP-optimal quadratic filter is Q_NP = N^{-1} - C^{-1}. The standard DFCC filter instead removes the block-diagonal (autocorrelation) part from the deflection-optimal filter Q_DF, which maximizes signal-to-noise ratio, not detection probability. The paper proves that when the noise covariance is block-diagonal (the CURN null hypothesis), the quadratic filter that minimizes the null-hypothesis variance of its statistic's difference from the NP statistic, subject to using only cross-correlations, has the simple closed form Q_NPMV = Q_NP - Diag Q_NP. This statistic can be built directly from the whitened DF filter already used in PTA pipelines. ROC compar
What carries the argument
The central object is the quadratic detection statistic D(z|Q)=z†Qz and its filter Q. The NPMV filter is the solution to the constrained problem: minimize Var[D(z|Q)-D(z|Q_NP)] under the null hypothesis subject to Diag Q=0. For the block-diagonal noise covariance used in PTA analyses this solution collapses to Q_NPMV = Q_NP - Diag Q_NP, i.e., the NP-optimal filter with all autocorrelation blocks zeroed. The paper also derives the generalized chi-square CDF for indefinite filters via contour integration (Fourier representation of the step function and Cauchy residues), which is used to build ROC curves in the toy model and to compute p-values.
Load-bearing premise
The 47% improvement is read off ROC curves for the realistic NANOGrav-like model in Sec. 7.3, but the paper does not state how those curves were computed—the analytic generalized chi-square CDF cannot be used for realistic data, and no substitute method is specified—so the headline performance claim is only as reliable as that unspecified numerical procedure.
What would settle it
Recompute the Sec. 7.3 ROC curves using the published NANOGrav 15-year parameters (A_gw=2.1e-15, gamma=13/3, the quoted pulsar noise parameters) by Monte Carlo: draw many realizations under the null and signal hypotheses, evaluate the DFCC and NPMV statistics, and estimate the detection probabilities at FAP=2.9e-7. If the ratio is not close to 1.47, or the individual detection probabilities (≈2.7% vs ≈4.0%) are not reproduced, the central claim fails.
If this is right
- PTA collaborations can adopt NPMV as the detection statistic for FAPs and p-values, replacing DFCC, and obtain higher sensitivity at the same false-alarm rate without opening the analysis to pulsar-noise autocorrelation mismodeling.
- Because the NPMV filter is computed from the same whitened DF filter already in existing pipelines, retrofitting current analyses requires only minor code changes; the paper provides an implementation in the enterprise_extensions package.
- Reported detection significances for current datasets would change under NPMV: for a real GWB at the modeled amplitude, the same false-alarm threshold is crossed more often, so a claimed '5-sigma' threshold corresponds to a higher true detection rate.
- The NPCC statistic—the true constrained NP-optimal cross-correlation statistic—performs slightly better than NPMV, but it requires expensive numerical optimization; NPMV captures most of the available gain in closed form.
Where Pith is reading between the lines
- The minimum-variance projection idea is generic: any detection problem where the likelihood-ratio statistic is unusable because of nuisance-parameter sensitivity could define a 'practically optimal' statistic by projecting the NP filter onto the subspace of robust filters, and the closed-form solution would hold whenever the noise covariance is block-diagonal.
- The 47% gain is tied to the specific NANOGrav-like noise realization and GWB parameters; the improvement factor will likely vary with array sensitivity, number of pulsars, and signal amplitude, so the headline number should not be read as a universal sensitivity gain.
- The paper's Sec. 7.3 realistic-model ROC curves rely on an unspecified numerical method; until that method is documented, the 47% figure should be treated as a computational result to be verified by independent Monte Carlo.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a new quadratic detection statistic, NPMV, for pulsar timing array searches for a stochastic gravitational-wave background. The statistic is defined as the quadratic form that uses only cross-pulsar correlations and minimizes, under the null hypothesis, the variance of its difference from the Neyman-Pearson (NP) optimal quadratic statistic. For block-diagonal noise, the authors derive a simple closed form: the NPMV filter is the NP filter with its diagonal blocks removed. They also derive a generalized chi-square CDF for complex normal data with an indefinite quadratic form, and use it to compare DFCC, NP, NPCC, and NPMV in a toy model. They then present ROC curves for a more realistic NANOGrav-15-year-like model and report a 47% increase in detection probability for NPMV over DFCC at the 5-sigma FAP of 2.9e-7.
Significance. If correct, the NPMV statistic is a useful methodological contribution: it offers a cross-correlation-only statistic that lies closer to Neyman-Pearson optimality than the currently standard DFCC statistic, while retaining the practical robustness property of excluding autocorrelations from the quadratic form. The derivation of the NPMV filter is clean and is obtained from first principles with no free parameters fitted to the target result. The paper also ships code and uses external NANOGrav parameters as inputs, which are strengths. However, the headline quantitative claim rests on ROC curves whose numerical method is not described, and the central robustness motivation is not tested under noise mismodeling. These gaps need to be addressed before the performance claim can be considered established.
major comments (2)
- [§7.3, Fig. 3] The method used to compute the realistic-model ROC curves is not stated. Section 7.1 explicitly says that the analytic CDF (49)-(50) cannot be used for realistic data because the data are not described by a complex-valued random process, but Section 7.3 only says "We construct ROC curves" and gives no substitute: no Monte Carlo sample count, no Davies/Imhof algorithm, no convergence checks. The headline 47% improvement is read off these curves; the two quoted detection probabilities (2.7% vs 4.0%) differ by only 1.3 percentage points, so the result could be within the numerical error of an unvalidated tail approximation. Please specify the numerical procedure and attach uncertainties to the quoted improvement.
- [Abstract; §§1, 4.2] The paper motivates NPMV as "robust" to pulsar noise modeling errors, but no mismodeling test is performed. The filter is constructed from N and C via (20)-(21); although the quadratic form uses only off-diagonal data blocks, the filter itself depends on the assumed noise covariance N. The claim that NPMV is "less affected" by pulsar noise modeling errors than NP or DFCC is therefore not demonstrated. Please add a quantitative robustness study, e.g. ROC curves under perturbed N for NPMV vs DFCC vs NP, or state explicitly that robustness is only a structural property with no empirical validation.
minor comments (5)
- [§3] The sentence about "the autocorrelations can be reconstructed from the cross-correlations" for n>=3 is unsupported and seems to conflict with realistic pulsar-specific noise. It is not used later, but as written it is misleading; please clarify or remove.
- [§4.1] Typo: "modeling erorrs" should be "modeling errors". Also, after Eq. (4), "expectated value" should be "expected value".
- [§7.2, Fig. 2] The caption reports AUC=0.975 for both NPCC and NPMV, and equal detection probabilities at the 5-sigma FAP, while the text says NPCC curves lie above those of NPMV. Please state the precision to which the curves differ, or qualify the claim.
- [§6.2] The discussion of catastrophic cancellation is useful, but the paper does not say whether the toy-model CDF evaluations used arbitrary precision or standard IEEE doubles. Please specify, since the NPCC optimization relies on these evaluations.
- [References] The entry for Taylor et al. 2021 (enterprise_extensions) lacks a journal, arXiv identifier, or DOI; please complete the reference.
Circularity Check
No significant circularity: the NPMV filter is derived from a first-principles constrained optimization, and the reported 47% detection-probability improvement over DFCC is an evaluated output from externally published NANOGrav 15-year inputs, not a fitted or self-referential quantity.
full rationale
The paper's central derivation is self-contained and not circular. The NPMV statistic is defined by a new, well-posed objective (Eq. 11): minimize Var[D(z|Q) - D(z|Q_NP)] subject to Diag Q = 0. The solution is derived, not assumed, in Sec. 5 via Lagrange multipliers, leading to Eq. (12) for block-diagonal N; uniqueness is supported by an external mathematical source (Schur product theorem, Horn & Johnson). No parameter is fitted to the target result, and the detection-probability comparison is genuinely computed: the NANOGrav 15-year fixed-gamma parameters (A_gw, gamma, pulsar noise) enter as externally published inputs (Sec. 7.3, footnote 5), and the quoted detection probabilities (DFCC 2.7% vs NPMV 4.0% at FAP 2.9e-7) are outputs of ROC evaluation, not quantities tuned to produce a 47% gain. The self-citations (van Haasteren 2025 for the DF/NP filter formulas; Allen & Romano for HD background) are non-load-bearing: the formulas (7) and (10) are standard quadratic-form results, parameter-free, and re-derivable from the stated optimization objectives, so the citation does not reduce the argument to an unverified prior claim. The main caveat flagged by the manuscript is not circular: Sec. 7.1 states that the analytic generalized chi-square CDF (49)-(50) cannot be used for realistic data, yet Sec. 7.3 does not state the numerical method (Monte Carlo, Davies/Imhof, etc.) used to generate the realistic-model ROC curves in Fig. 3. This is a reproducibility and accuracy risk for the headline 47% number, but it is not a self-referential reduction of an output to an input; nothing in the derivation chain makes the reported improvement true by construction. The comparison uses a common external signal model, so the ordering of the statistics is an empirical computation, not a definitional equivalence.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Noise covariance N is block-diagonal (uncorrelated between pulsars).
- domain assumption Null and signal hypotheses are zero-mean multivariate Gaussian with known covariances (simple hypotheses).
- ad hoc to paper The minimum-variance-distance criterion defines the desired compromise statistic.
- domain assumption Autocorrelations are more vulnerable to pulsar noise modeling errors than cross-correlations.
- standard math Isserlis theorem, Schur product theorem, and contour-integration identities.
Cite this review
Pith. "Pith review of Optimal robust detection statistics for pulsar timing arrays." pith.science (2026). https://pith.science/paper/EUHMRV6L
@misc{pith2026250906489,
author = {Pith},
title = {Pith review of: Optimal robust detection statistics for pulsar timing arrays},
year = {2026},
howpublished = {\url{https://pith.science/paper/EUHMRV6L}},
note = {Machine review of arXiv:2509.06489}
}
read the original abstract
Pulsar timing arrays (PTAs) seek to detect a nano-Hz stochastic gravitational-wave background (GWB) by searching for the characteristic Hellings and Downs angular pattern of timing residual correlations. So far, the evidence remains below the conventional $5$-$\sigma$ threshold, as assessed using the literature-standard ``optimal cross-correlation detection statistic''. While this quadratic combination of cross-correlated data maximizes the {\em deflection} (signal-to-noise ratio), it does not maximize the detection probability at fixed false-alarm probability (FAP), and therefore is not Neyman-Pearson (NP) optimal for the assumed noise and signal models. The NP-optimal detection statistic is a different quadratic form, but is not used because it also incorporates autocorrelations, making it more susceptible to uncertainties in the modeling of pulsar timing noise. Here, we derive the best compromise: a quadratic detection statistic which is as close as possible to the NP-optimal detection statistic (minimizing the variance of its difference with the NP statistic) subject to the constraint that it only uses cross-correlations, so that it is less affected by pulsar noise modeling errors. We study the performance of this new $\NPMV$ statistic for a simulated PTA whose noise and (putative) signal match those of the NANOGrav 15-year data release: GWB amplitude $A_{\rm gw}=2.1\times 10^{-15}$ and spectral index $\gamma=13/3$. Compared to the literature-standard ``optimal" cross-correlation detection statistic, the $\NPMV$ statistic increases the detection probability by $47\%$ when operating at a $5$-$\sigma$ FAP of $\alpha = 2.9 \times 10^{-7}$.
Figures
Forward citations
Cited by 4 Pith papers
-
A comprehensive framework for phase-coherent mapping of the gravitational-wave sky with pulsar timing arrays
A new phase-coherent mapping framework for pulsar timing arrays that preserves the complete complex polarization state of the gravitational-wave sky in compact maps usable for multiple analyses.
-
Stochastic gravitational-wave background search using data from five pulsar timing arrays
Combined five-PTA dataset yields posterior on SGWB power-law amplitude and index consistent with nonzero signal but below 5-sigma significance, with reconstructed angular correlations matching the Hellings-Downs prediction.
-
The NANOGrav 15 yr Data Set: Impacts of Customized Chromatic Noise Models on Gravitational Wave Analyses
Customized chromatic noise models applied to NANOGrav 15 yr data raise the Bayes factor for Hellings-Downs GWB correlations by a factor of ~8, lower the amplitude to 2.1e-15, and increase the spectral index to 3.5.
-
Pulsar timing array analysis in a Legendre polynomial basis
Legendre polynomial basis for PTA analysis zeros first three modes from timing models and delivers closed-form expressions for the optimal quadratic Hellings-Downs estimator and variance under power-law GWB and noise spectra.
Reference graph
Works this paper leans on
-
[1]
2023, The Astrophysical Journal Letters, 951, L11, doi: 10.3847/2041-8213/acdc91
Afzal, A., Agazie, G., Anumarlapudi, A., et al. 2023, The Astrophysical Journal Letters, 951, L11, doi: 10.3847/2041-8213/acdc91
-
[2]
Agazie, G., Anumarlapudi, A., Archibald, A. M., et al. 2023a, The Astrophysical Journal Letters, 951, L8, doi: 10.3847/2041-8213/acdac6 —. 2023b, The Astrophysical Journal Letters, 951, L50, doi: 10.3847/2041-8213/ace18a —. 2024, The NANOGrav 15 Yr Data Set: Posterior Predictive Checks for Gravitational-Wave Detection with Pulsar Timing Arrays, arXiv, doi...
-
[3]
2023, Physical Review D, 107, 043018, doi: 10.1103/PhysRevD.107.043018
Allen, B. 2023, Physical Review D, 107, 043018, doi: 10.1103/PhysRevD.107.043018
-
[4]
Allen, B., Dhurandhar, S., Gupta, Y., et al. 2023, The International Pulsar Timing Array Checklist for the Detection of Nanohertz Gravitational Waves, arXiv, doi: 10.48550/arXiv.2304.04767
-
[5]
Allen, B., & Romano, J. D. 2023, Physical Review D, 108, 043026, doi: 10.1103/PhysRevD.108.043026 —. 2024, Optimal Reconstruction of the Hellings and Downs Correlation, arXiv, doi: 10.48550/arXiv.2407.10968
-
[6]
2009, Physical Review D, 79, 084030, doi: 10.1103/PhysRevD.79.084030
Siemens, X. 2009, Physical Review D, 79, 084030, doi: 10.1103/PhysRevD.79.084030
-
[7]
Antoniadis, J., Arumugam, P., Arumugam, S., et al. 2023, Astronomy & Astrophysics, 678, A49, doi: 10.1051/0004-6361/202346842 Optimal robust statistics for PTAs11 —. 2024a, Astronomy & Astrophysics, 685, A94, doi: 10.1051/0004-6361/202347433 —. 2024b, Astronomy & Astrophysics, 690, A118, doi: 10.1051/0004-6361/202348568
-
[8]
Chamberlin, S. J., Creighton, J. D. E., Siemens, X., et al. 2015, Physical Review D, 91, 044048, doi: 10.1103/PhysRevD.91.044048
-
[9]
Cox, D. 1962, Renewal Theory, 1st edn., Methuen’s Monographs on Applied Probability and Statistics (London: Methuen)
work page 1962
-
[10]
Das, A., & Geisler, W. S. 2021, Journal of Vision, 21, 1, doi: 10.1167/jov.21.10.1
-
[11]
A., Siemens, X., & van Haasteren, R
Ellis, J. A., Siemens, X., & van Haasteren, R. 2013, The Astrophysical Journal, 769, 63, doi: 10.1088/0004-637X/769/1/63
-
[12]
Ellis, J. A., Vallisneri, M., Taylor, S. R., & Baker, P. T. 2020, ENTERPRISE: Enhanced Numerical Toolbox Enabling a Robust PulsaR Inference SuitE, Zenodo, doi: 10.5281/zenodo.4059815
-
[13]
Archibald, A. M. 2023, Phys. Rev. D, 108, 104050, doi: 10.1103/PhysRevD.108.104050
-
[14]
Hellings, R. W., & Downs, G. S. 1983, The Astrophysical Journal, 265, L39, doi: 10.1086/183954
doi:10.1086/183954 1983
-
[15]
Horn, R. A., & Johnson, C. R. 2012, Matrix Analysis, 2nd edn. (Cambridge, UK: Cambridge University Press), doi: 10.1017/CBO9781139020411
-
[16]
1918, Biometrika, 12, 134, doi: 10.1093/biomet/12.1-2.134
Isserlis, L. 1918, Biometrika, 12, 134, doi: 10.1093/biomet/12.1-2.134
-
[17]
Powell, M. J. D. 2009, The BOBYQA algorithm for bound constrained optimization without derivatives, Technical Report DAMTP 2009/NA06, University of Cambridge, Cambridge. http: //www.damtp.cam.ac.uk/user/na/NA papers/NA2009 06.pdf
work page 2009
-
[18]
Rasmussen, C. E., & Williams, C. K. 2006, Gaussian Processes for Machine Learning (MIT Press)
work page 2006
-
[19]
Reardon, D. J., Zic, A., Shannon, R. M., et al. 2023, The Astrophysical Journal Letters, 951, L6, doi: 10.3847/2041-8213/acdd02
-
[20]
Romano, J. D., & Cornish, Neil. J. 2017, Living Reviews in Relativity, 20, 2, doi: 10.1007/s41114-017-0004-1
-
[21]
1998, IEEE Transactions on Aerospace and Electronic Systems, 34, 817, doi: 10.1109/7.705889
Spall, J. 1998, IEEE Transactions on Aerospace and Electronic Systems, 34, 817, doi: 10.1109/7.705889
-
[22]
Taylor, S. R., Baker, P. T., Hazboun, J. S., & Simon, J. 2021, Enterprise extensions
work page 2021
-
[23]
M., Chatziioannou, K., & Chua, A
Vallisneri, M., Meyers, P. M., Chatziioannou, K., & Chua, A. J. K. 2023, Physical Review D, 108, 123007, doi: 10.1103/PhysRevD.108.123007 van Haasteren, R. 2025, On the Calculation of P-Values for Quadratic Statistics in Pulsar Timing Arrays, arXiv, doi: 10.48550/arXiv.2506.10811 van Haasteren, R., & Vallisneri, M. 2014, Physical Review D, 90, 104012, doi...
-
[24]
Vigeland, S. J., Islo, K., Taylor, S. R., & Ellis, J. A. 2018, Physical Review D, 98, 044003, doi: 10.1103/PhysRevD.98.044003
-
[25]
2023, Research in Astronomy and Astrophysics, 23, 075024, doi: 10.1088/1674-4527/acdfa5
Xu, H., Chen, S., Guo, Y., et al. 2023, Research in Astronomy and Astrophysics, 23, 075024, doi: 10.1088/1674-4527/acdfa5
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.