{"id":"d7624866-509e-48dd-bb1a-b74c2e55b05f","arxiv_id":"2509.07090","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new pulsar-timing-array pipeline maps the gravitational-wave sky per frequency, folds cosmic variance into significance estimates, and detects a simulated loud source at p=0.01 versus 0.2 broadband.","lead":"Pulsar timing arrays have detected a gravitational wave background, and now scientists want to know where it comes from. This paper builds a faster statistical tool that can locate gravitational wave power on the sky at each frequency separately, which should help tell astrophysical sources from cosmological ones.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed p-value improvement may be an artifact of the power-law CURN mis-modeling the authors admit biases per-frequency PSD estimation.","rationale":"The reader's weakest_assumption is exactly the initial CURN power-law adequacy, and the paper's own Section VIII confirms that this assumption is violated and that biases propagate into the anisotropy pipeline. This is the most load-bearing concern because it directly undermines the calibration of the headline p-values, not just their interpretation. The multiple-comparison issue is real but secondary: the authors disclose it, and a reader can apply Bonferroni mentally. The simulation tuning (sources placed near pulsars) affects the generalization of the results but does not invalidate the internal comparison; the p-value improvement for S1 is still demonstrated under those conditions. The CURN mismatch, however, means that even the internal comparison may be comparing two biased quantities, so the claimed improvement could be spurious. The concrete test—using a more flexible CURN model—would settle whether the pipeline's frequency localization and p-value improvement survive when the initial spectral model is not grossly misspecified. I therefore agree with the reader's conditionality; the paper's honesty and open-source code are credits, but the central claim needs this validation before being taken at face value. Since the reader already set CONDITIONAL, my stress test does not change the verdict.","tokens_in":37273,"tokens_out":3443,"duration_ms":40516,"concrete_test":"Re-run the full pipeline on the same simulated dataset, but replace the single power-law CURN model in the initial Bayesian stage with a t-process PSD (Sardesai et al. 2024, ref [50]) that allows spectral excursions while constraining high-frequency noise. Then recompute the PFOS PSD estimates and the square-root spherical harmonic per-frequency p-values. If F3’s p-value remains below 0.01 after Bonferroni correction and the S1 signature does not leak into F4 (i.e., the F4 p-value stays near 0.5), the claim is robust. If the F3 p-value inflates above 0.05 or the leakage persists, the central claim is an artifact of the power-law mis-modeling.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—per-frequency search improves anisotropy detection from p~0.2 to p~0.01—requires the per-frequency p-values to be calibrated and unbiased. But Section VIII states: “the true PSD of our overall signal is a poor match to a power-law… the poor percentiles from the PFOS PSD estimation undoubtedly propagate biases further into the anisotropic stages of the pipeline.” This is a direct admission that the downstream anisotropy p-values are biased. Evidence of the bias: the PFOS PSD in bin 4 sits at the 0.3 percentile of the expected distribution (Fig. 3), and S1’s power appears in F4’s map (Fig. 7), i.e., spectral leakage from a mis-specified initial model. If the initial CURN model is wrong, the per-frequency pair correlations in Eq. (25) are constructed with the wrong Φ(f_n), so the null distributions in Eq. (36) and the resulting p-values are not representatives of the true noise+signal process. The headline improvement may therefore be an uncalibrated number, not a robust detection claim. The authors are transparent about this, but the abstract does not caveat it, and the stated p-value is also uncorrected for 12 frequency trials (Bonferroni would give ~0.07 for the F3 minimum).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a next-generation frequentist pipeline for pulsar timing array (PTA) gravitational-wave background (GWB) anisotropy searches, combining three previously developed ingredients: the pair-covariant optimal statistic (PCOS), the per-frequency optimal statistic (PFOS), and cosmic-variance-aware null distributions. The authors construct an IPTA-DR3-like simulated dataset containing an isotropic power-law GWB plus four injected continuous-wave (CW) sources in different frequency bins, and apply their pipeline in both broadband and per-frequency modes. On this dataset, the per-frequency square-root spherical-harmonic analysis recovers the loudest source (S1) with an uncorrected p-value of 6e-3 in frequency bin 3, compared to broadband p-values of 0.17 (OS) and 0.22 (PCOS). The paper claims that this demonstrates improved anisotropy detection prospects, while acknowledging several caveats, most notably that the initial power-law CURN spectral model is a poor match to the true spiky spectrum, leading to biased PFOS PSD estimates and spectral leakage.","tokens_in":37573,"tokens_out":2843,"duration_ms":32142,"significance":"If the central claim holds, the paper provides a practical, computationally efficient upgrade to the standard PTA anisotropy search: per-frequency localization of GWB anisotropies with calibrated null distributions that include cosmic variance. The methods are already implemented in community-available code (Enterprise, Defiant, MAPS) and are directly applicable to forthcoming IPTA datasets. The inclusion of cosmic variance in null distributions addresses a known false-detection problem (Konstandin et al. 2025), and the per-frequency decomposition is a natural step toward source-by-source characterization. However, the headline significance improvement is demonstrated on a single, deliberately optimized simulation, and the paper's own caveats about spectral mis-modeling and the use of uncorrected p-values substantially weaken the strength of the claim as stated in the abstract. The paper is transparent about these limitations, but the abstract overstates the robustness of the result.","major_comments":[{"comment":"The headline improvement from p~0.2 to ~0.01 is based on an uncorrected per-frequency p-value of 6e-3 for F3. As the authors note in Section III.B, a Bonferroni correction for N_freq=12 independent trials would raise this to ~0.07, which is not below the conventional 0.01 threshold. The abstract states '~0.01 in the per-frequency search' without mentioning this correction. Since the main claim is the significance improvement, quoting an uncorrected p-value as the headline number is misleading. The abstract should either quote the corrected value or explicitly state that the quoted p-value is raw.","section":"Section VI.B.2 and abstract"},{"comment":"The authors admit in Section VIII that 'the true PSD of our overall signal is a poor match to a power-law' and that 'the poor percentiles from the PFOS PSD estimation undoubtedly propagate biases further into the anisotropic stages of the pipeline.' This is not a minor caveat: the per-frequency pair correlations in Eq. (25) and the null distributions in Eq. (36) all use the PFOS PSD estimates S_n and Phi(f_n). If those estimates are biased (e.g., F4 at the 0.3 percentile in Fig. 3), then the per-frequency p-values are not calibrated under the true data-generating process. The visible spectral leakage of S1 into F4 (Fig. 7) is direct evidence. The abstract and Section VII do not carry this caveat. The authors should demonstrate that the headline p-value is robust to the initial spectral model choice, e.g., by repeating the analysis with a t-process or free-spectrum CURN model, or by expli","section":"Section VIII and Fig. 3"},{"comment":"The null distributions are generated using the pair-independent covariance matrix C_n,0 (Eq. 36), while the detection statistic itself is computed with the pair-covariant covariance matrix C_n (Eq. 28). The authors argue this avoids double-counting cosmic variance, but no calibration or coverage test is shown. Under the null hypothesis (statistically isotropic GWB), the p-values should be uniformly distributed. The paper does not verify this property, nor does it provide posterior predictive checks of the type recommended by Vallisneri et al. (2023). Without such validation, the quoted p-values cannot be interpreted as accurate frequentist significances. This is a load-bearing gap for the central claim.","section":"Section III.C, Eq. (36)"},{"comment":"The claimed improvement is demonstrated on a single realization of the noise and a single set of injected CW parameters, with S1 deliberately placed near pulsar sky locations and contributing ~90% of the total PSD in its bin. A p-value is a random variable; the paper provides no estimate of its variance or of the detection probability across realizations. The statement 'anisotropy detection prospects ... improve' is thus supported for one optimized dataset, not as a general property of the method. Multiple noise realizations (or at least a distribution of p-values under fixed injections, e.g., via bootstrapping the null draws) are needed to support the abstract's broader claim.","section":"Section IV"}],"minor_comments":[{"comment":"The definition of R_ab,k appears to have a typo: the second factor in each polarization term should be F^+_{b,k} and F^×_{b,k}, not F^+_{a,k} and F^×_{a,k}. The text just above defines R_ab,k as the response for a pair (a,b), so the current expression is inconsistent.","section":"Eq. (11)"},{"comment":"'at-process' should be 't-process' (referring to the t-process PSD model of Sardesai et al.).","section":"Section VIII"},{"comment":"The phrase 'challening' should be 'challenging'.","section":"Section VII"},{"comment":"The description of the null generation says '10^3 random draws ... and, for each draw, creating 10 realizations.' It would be clearer to state the total number of null realizations (10^4) and how they are combined into the final null distribution.","section":"Section V"},{"comment":"References [35] and [36] are the same paper (Vigeland et al. 2018) and should be merged.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong methodological contribution that fits the journal's scope, but the abstract overstates the robustness of the headline p-value improvement. The heavy reliance on the authors' own prior work (Gersbach et al. 2025, Konstandin et al. 2025) is appropriate given the pipeline is a direct extension, but the calibration gap is a substantive issue that needs to be addressed before publication. The uncorrected p-value and the admitted spectral mis-modeling are the two points I would emphasize to the authors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper does what it claims, mostly. It combines the pair-covariant optimal statistic, the per-frequency optimal statistic, and cosmic-variance-calibrated null distributions into one anisotropy pipeline, and demonstrates on a simulated IPTA-DR3-like array that a loudly injected binary source is detected far more clearly per-frequency (uncorrected p ~ 6e-3 in bin F3) than in a broadband search (p ~ 0.2). That is a useful, concrete result, and the frequency-resolved sky maps are a genuine step forward for distinguishing astrophysical from cosmological backgrounds.\n\nThe paper earns credit for being honest. The authors disclose in Section VIII that the power-law CURN prior is a poor fit to the true spiky spectrum and that the resulting PFOS PSD biases \"undoubtedly propagate\" into the anisotropic stages. They also flag the spectral leakage visible in F4 and the uncorrected nature of the p-values. Code is released. The heavy reliance on the same group's prior derivations (PCOS, PFOS, cosmic variance) is not a flaw; the new contribution is the integration and the demonstration, not the individual components.\n\nNow the soft spots, in proportion. The headline p-value is uncorrected; Bonferroni over 12 frequencies puts the F3 minimum near 0.07, which the paper itself notes. That alone changes the tone of the abstract's \"~0.01\" claim. More substantively, the stress-test concern lands: because the CURN model is wrong, the per-frequency pair correlations in Eq. (25) and the null distributions in Eq. (36) are constructed with a biased Phi(f_n). So the reported p-value is not a calibrated frequentist number. I don't think the improvement is entirely an artifact, since the injected signal in F3 is loud and spatially localized, but the true significance is probably weaker than stated, and the simulation was deliberately tuned (sources near pulsars, source PSD ~90% of the bin, 116 pulsars, 22 years) to give the method its best shot. The other three sources are not confidently recovered, and the extra hot-spots in F5 indicate unresolved interference.\n\nNone of this is disqualifying. The paper is a methods demonstration, not a detection claim, and the limitations mostly appear in the text. But the abstract should carry the caveat, and a revision should report corrected p-values and run at least one less-adversarial simulation (e.g., realistic binary population without a single dominant source) to show the pipeline behaves when the spectral model is not grossly misspecified.\n\nWho is this for? PTA methodologists and anyone planning anisotropy searches on future IPTA data. It deserves a serious referee and likely acceptance after revision. I'd bring it to reading group and cite it if I worked in this area.\n\nRecommendation: send to peer review, conditional on the authors reporting multiple-testing-corrected, calibration-checked p-values and softening the abstract's claim.","headline":"A competent integration paper whose headline p-value improvement is real but overstated: raw, uncorrected, and likely inflated by the authors' own admitted spectral mis-modeling; still worth refereeing.","tokens_in":38094,"tokens_out":1870,"would_cite":true,"duration_ms":25541,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Splitting the gravitational-wave background by frequency turns an inconclusive broadband anisotropy search into a per-frequency detection at p≈0.01.","keywords":["gravitational-wave background","pulsar timing arrays","anisotropy","per-frequency optimal statistic","cosmic variance","pair covariance","supermassive black hole binaries","frequentist detection"],"falsifier":"Construct a simulated dataset whose true gravitational-wave spectrum is a poor match to a power law but whose sky is statistically isotropic, run the pipeline, and count how often the per-frequency anisotropic p-value falls below 0.01; if false detections occur at a rate far above the calibrated level in bins adjacent to a spectral bump, the claimed frequency localization is partly spectral leakage rather than true anisotropy.","tokens_in":37144,"feed_emoji":"🔭","tokens_out":6696,"duration_ms":63450,"temperature":0.7,"pith_summary":"Pulsar timing arrays have detected a gravitational-wave background at nanohertz frequencies, but its origin is unsettled: a background from supermassive black hole binaries should be more anisotropic than a cosmological one. This paper upgrades the frequentist optimal-statistic search for that anisotropy with three changes: full covariance between pulsar-pair estimates, a per-frequency version of the statistic, and null distributions that include cosmic variance from a statistically isotropic background. On a simulated 116-pulsar sky containing an isotropic background plus four injected binary signals in different frequency bins, the per-frequency pipeline finds the loudest source in its own frequency bin with an uncorrected p-value of 6e-3, against 0.17-0.22 for the same broadband search. The point of the improvement is to give forthcoming datasets a computationally cheap way to map the gravitational-wave sky frequency by frequency, so that individual resolvable binaries can be distinguished from a smooth cosmological background.","feed_headline":"Per-frequency search lifts GWB anisotropy detection from p=0.2 to 0.01","feed_subtitle":"A pulsar-timing pipeline that splits the gravitational-wave background by frequency pinpoints loud sources a broadband search misses.","key_machinery":"The central object is the per-frequency optimal statistic (PFOS), a frequency-resolved version of the pulsar-pair cross-correlation estimator. For each frequency bin n it computes pairwise estimators of correlated power between pulsars a and b, their variances, and a full pair-covariance matrix that includes gravitational-wave self-noise rather than assuming the pairs are independent. These estimators enter a chi-squared fit whose model is the overlap reduction function—the expected correlation pattern between two pulsars—expanded in sky pixels or square-root spherical harmonics, producing a sky map and an anisotropic SNR per frequency. The significance of those SNRs is calibrated by drawing","core_discovery":"The paper's central claim is that anisotropy in the nanohertz gravitational-wave background is best searched for one frequency at a time. It assembles a pipeline from three existing pieces: the pair-covariant optimal statistic, which includes the covariance between pulsar-pair correlation estimates and therefore accounts for GWB self-noise; the per-frequency optimal statistic, which estimates correlated power and power-spectral density separately in each frequency bin; and null-hypothesis distributions built from explicit random realizations of a statistically isotropic background, so that cosmic variance is included in the significance calibration. In the authors' simulations, the loudest i","pith_inferences":["Because each frequency bin carries its own independent cosmic-variance realization, per-frequency searches lose the averaging that suppresses cosmic variance in broadband searches; this suggests a sensitivity floor that no increase in pulsar count alone can remove.","The leakage of S1 into frequency bin 4 predicts a concrete improvement: explicitly modeling covariances between adjacent frequency bins should remove the ghost source and further tighten the p-value in the true bin.","The authors' report of hot-spots offset from true sources in multi-source bins implies that map localization and detection p-value should be treated as separate diagnostics in realistic skies; follow-up deterministic searches should be triggered by the p-value, not by the map peak.","A posterior-predictive p-value calibration, which the paper notes is the more rigorous option but computationally expensive, would be a natural next test: if it preserves the per-frequency improvement, the frequentist shortcut is safe for real datasets."],"forward_implications":["A background from a finite population of supermassive black hole binaries, which is approximately monochromatic per source, will show up as excess power in specific frequency bins; per-frequency maps turn that spectral signature into a spatial detection.","Pair covariance should be standard in future optimal-statistic anisotropy searches, since neglecting it inflates significance in the strong-signal regime and can mimic anisotropy.","Per-frequency p-values must be corrected for the look-elsewhere effect, but the demonstrated improvement from about 0.2 to about 0.01 is large enough to remain meaningful after correction in the loudest bin.","The pipeline is massively parallelizable relative to full Bayesian sky-mapping analyses, so it can be deployed on next-generation pulsar-timing-array datasets without excessive computational cost.","Spectral leakage between adjacent frequency bins is the main practical limit: a loud binary leaves a ghost in the neighboring bin, so per-frequency localization should be complemented by frequency-covariance modeling."],"supporting_citations":[{"why":"Defines the pair-covariant and per-frequency optimal statistics, including the covariance expressions the pipeline uses.","marker":"[25]"},{"why":"Supplies the radiometer and square-root spherical harmonic anisotropy framework and the anisotropic SNR statistic.","marker":"[23]"},{"why":"Shows that cosmic variance must be included in anisotropy null distributions and gives the random-wave realization method used here.","marker":"[28]"},{"why":"Demonstrates that pulsar-pair independence is violated in current datasets, motivating the pair-covariant generalization.","marker":"[26]"},{"why":"Provides the noise-marginalized optimal statistic that supplies posterior draws for pulsar autocovariance estimates.","marker":"[35]"},{"why":"Derives the pulsar and cosmic variance decomposition used to construct the total variance in the estimators.","marker":"[24]"}],"fun_headline_variants":["Per-frequency GWB search sharpens anisotropy detection 20x","Frequency-split GWB analysis finds stronger anisotropy","Per-frequency PTA statistic localizes GWB anisotropies","Covariance-aware per-frequency search spots GWB anisotropy","Per-frequency method cuts GWB anisotropy p-value to 0.01"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The entire per-frequency chain inherits the initial power-law model for the background spectrum; if that model misrepresents the true spectrum, the per-frequency power estimates and the anisotropy maps and p-values built from them are biased, as the paper itself sees in frequency-bin 4.","fun_headline_variants_meta":{"raw":{"variants":["Per-frequency GWB search sharpens anisotropy detection 20x","Frequency-split GWB analysis finds stronger anisotropy","Per-frequency PTA statistic localizes GWB anisotropies","Covariance-aware per-frequency search spots GWB anisotropy","Per-frequency method cuts GWB anisotropy p-value to 0.01"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000945,"raw_usage":{"total_tokens":3915,"prompt_tokens":830,"completion_tokens":3085,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":3011}},"tokens_in":574,"tokens_out":3085,"duration_ms":23416,"temperature":1.0,"reasoning_tokens":3011,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:50:50.457562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a simulated dataset whose true gravitational-wave spectrum is a poor match to a power law but whose sky is statistically isotropic, run the pipeline, and count how often the per-frequency anisotropic p-value falls below 0.01; if false detections occur at a rate far above the calibrated level in bins adjacent to a spectral bump, the claimed frequency localization is partly spectral leakage rather than true anisotropy.","supporting_citations":[{"cited_title":"Disentangling Multiple Stochastic Gravitational Wave Background Sources in PTA Datasets","cited_arxiv_id":"2208.02307","evidence_quote":"Supplies the radiometer and square-root spherical harmonic anisotropy framework and the anisotropic SNR statistic."}],"review_version":1}