{"id":"fbb6bd36-1e0e-4663-882e-2262fba3cb92","arxiv_id":"2608.09929","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Shot-noise anisotropy in the nHz gravitational wave background has a broad, skewed distribution whose median is two to three times below the ensemble mean, and moment-based predictions overestimate it by orders of magnitude at high frequency.","lead":"This paper uses Monte Carlo simulations to map how much the shot-noise anisotropy of the gravitational wave background varies across random realizations of the supermassive black hole binary population. It finds that typical skies have two to three times less shot noise than the ensemble average, and that simple moment-based estimates fail at high frequencies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; central claims are internally consistent and robust to tested model variations.","rationale":"The reader's weakest assumption correctly identifies model dependence as the primary source of uncertainty. I agree that the empirical calibration of the SMBHB merger model, including the sharp high-mass and low-redshift cutoffs, is not externally validated at the level needed for precise quantitative predictions. However, the paper's central qualitative conclusions do not hinge on the exact model: the shot-noise PDF is broad and skewed in both LM24 and SZQ, the median lies below the mean in all tested variations, and the moment-based ratio overestimates the ensemble-averaged shot-noise regardless of the model, with the discrepancy growing at high frequency. The analytic bounds in Eqs. (12)–(15) provide independent support, and the Monte Carlo means for ĥ2_c and ĥ4_c agree with the analytic integrals, validating the sampling procedure. The quantitative headlines — factor ~3 at 0.1 yr^-1 and more than two orders of magnitude by 1 yr^-1 — are explicitly stated as fiducial-model results in Section V, so the paper does not overclaim universality. A robustness check on M_BH,max would strengthen the paper but is not required for the central argument to stand. Therefore the reader's ACCEPT verdict remains appropriate, with the same moderate confidence reflecting the stated model dependence.","tokens_in":16700,"tokens_out":21427,"duration_ms":204398,"concrete_test":"Rerun the fiducial Monte Carlo with M_BH,max = 10^10, 10^10.5, and 10^11 M_sun (and, separately, with the overall merger rate multiplied by 3) and compare the ensemble-mean shot-noise, the median-to-mean ratio, and the 95% width at f = 0.1 and 1 yr^-1; if these quantities shift by more than about 50% relative to the fiducial values, the abstract's quantitative numbers would need to be framed as model-specific rather than representative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central argument — that shot-noise realization variance is large, with a broad, positively skewed PDF whose median lies below the ensemble mean, and that the ensemble-averaged shot-noise differs from the moment-based ratio of the strain — is internally consistent. The Monte Carlo construction (Eqs. 8–11) correctly samples Poisson source counts, and the analytic bounding argument (Eqs. 12–15) is sound. The weakest point is model representativeness: the quantitative factors (e.g., ~3 at f=0.1 yr^-1, >2 orders by f=1 yr^-1, 95% width ~50) are computed for the fiducial LM24 model and may shift under variations in the high-mass cutoff M_BH,max or the overall merger-rate normalization. However, the authors test SZQ, z_min, and residence-time models, and the qualitative conclusions survive; the fiducial-specific numbers are explicitly qualified in Section V. This is a stated limitation, not an internal inconsistency, and it does not threaten the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses Monte Carlo realizations of Poisson-sampled supermassive black hole binary (SMBHB) merger populations, drawn from the empirically calibrated LM24 and SZQ models, to characterize the probability distribution of the shot-noise anisotropy amplitude C_hat/(4π) together with the strain moments h^2_c and h^4_c across sky realizations. The authors show that the ensemble mean of the shot-noise is not equal to the moment ratio <h^4_c>/<h^2_c>^2, that the shot-noise PDF is broad and positively skewed with median below the mean, and that conditioning on a measurement of h^2_c narrows the distribution. They test sensitivity to the merger model, to the low-redshift cutoff z_min, and to an alternative stellar-hardening residence-time model.","tokens_in":16970,"tokens_out":7043,"duration_ms":67581,"significance":"If correct, the result is important for interpreting pulsar timing array anisotropy measurements: it demonstrates that moment-based estimates can overpredict the shot-noise by orders of magnitude at high frequency and that full PDFs are needed for parameter inference. The Monte Carlo implementation is transparent and directly samples Poisson source counts, and the simulated means of h^2_c and h^4_c agree with the analytic integrals, providing a useful validation. The qualitative conclusion that the shot-noise distribution is broad, skewed, and with median below the mean appears robust to the tested model variations. The main caveats, acknowledged in the text, are the inclination/polarization-averaged strain and the ad hoc cutoffs z_min and M_BH,max; these affect quantitative amplitudes but do not appear to threaten the central qualitative claim.","major_comments":[],"minor_comments":[{"comment":"The orange dot-dashed curve in the shot-noise panels is labeled as an 'analytic prediction for the ensemble average value,' but it is actually the moment ratio <h^4_c>/<h^2_c>^2, which the text correctly argues is not the ensemble mean of C_hat/(4π). Please relabel this curve as 'moment-based estimate' in the caption and in the text of Sec. III A to avoid implying that the Monte Carlo mean should match it.","section":"Fig. 1 caption and Sec. III A"},{"comment":"The statement that the median h^4_c and C_shot/(4π) signals are insensitive to M_BH,max in SZQ is plausible but is not demonstrated with a numerical test; since the mean is known to be sensitive to this cutoff, a brief test varying M_BH,max would strengthen the robustness claim.","section":"Sec. IV A, near Fig. 4"},{"comment":"The paper acknowledges that inclination-angle and polarization averaging omits a source of variance, but the abstract and conclusions quote specific widths and median-to-mean ratios without this caveat. Please add a prominent qualifier that the quoted numbers are computed for averaged source orientations and will broaden when orientation scatter is included.","section":"Sec. II.A, after Eq. (6)"},{"comment":"The text refers to the 'current NANOGrav 95% upper-bound of 0.20' with reference [17] from 2023; please verify whether this is still the latest published limit at the time of submission and update the wording or reference accordingly.","section":"Sec. III A"},{"comment":"Reference [33] lists the DOI as '10.1103/ql8b-7q8v', which looks like a placeholder rather than a standard DOI; please verify and correct.","section":"Reference [33]"}],"recommendation":"minor_revision","confidential_remarks":"The paper is well within the scope of the journal and makes a useful, clearly presented contribution. The quantitative numbers are likely to be quoted by the community, so the authors should make the orientation-averaging caveat more prominent. No concerns about novelty or internal consistency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, useful paper. The new thing is not the statistical identity — everyone knows ⟨X/Y⟩≠⟨X⟩/⟨Y⟩ — but the demonstration that for PTA shot-noise the difference is quantitatively huge, and that the shot-noise PDF is broad and skewed with median 2–3 times below the mean. That changes how anisotropy upper limits should be interpreted. The Monte Carlo is straightforward and honest: it samples Poisson merger counts, reproduces the analytic means for h2_c and h4_c, and checks sensitivity to z_min, mass cutoff, and residence-time model. The proof that C_shot/4π is bounded by 1/N_eff and by unity is clean and worth having on record.\n\nSoft spots are the usual ones for population synthesis. The quantitative factors (roughly 3 at 0.1 yr^-1, more than two orders by 1 yr^-1, 95% width near 50) come from the LM24 model; SZQ shifts the medians by factors around 1.6–3.5. The inclination/polarization-averaged strain is a simplification and will add variance, as the authors say. And the differential quantities are not yet folded into finite PTA frequency bins; the paper flags that as future work. None of this threatens the central conclusion, which is model-robust: typical realizations have less shot noise than the moment-based estimate, and the full PDF is needed for interpretation.\n\nI checked the stress-test note and it matches my reading. There is no hidden circularity: no parameter is fit to the shot-noise target, and the self-citations are legitimate because the cited LM24 model and companion paper are the natural inputs, not padding.\n\nThis paper is for PTA analysts and anyone building parameter-inference pipelines for the nHz background. It deserves a serious referee: the methodology is transparent, the predicted median scaling f^{1.1–1.2} is falsifiable against future data, and the caveats are stated rather than buried. I would send it to review without hesitation.","headline":"A transparent Monte Carlo study showing that PTA shot-noise anisotropy is realization-dependent, with medians well below ensemble means, and that moment-based estimates fail badly; the central result is robust and deserves peer review.","tokens_in":17430,"tokens_out":1577,"would_cite":true,"duration_ms":15306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The shot-noise anisotropy of the nHz gravitational wave background has a broad, skewed probability distribution that the ensemble average misrepresents.","keywords":["gravitational wave background","pulsar timing arrays","shot noise","anisotropy","supermassive black hole binaries","Monte Carlo simulation","realization variance","nHz gravitational waves"],"falsifier":"A direct test would be a PTA measurement of the shot-noise anisotropy at multiple frequencies with the full PDF forward-modeled: if the observed shot-noise at $f \\approx 1\\,\\mathrm{yr}^{-1}$ consistently landed near the moment-based prediction (which exceeds the Monte Carlo mean by a factor of ~130) rather than near the median, or if the realization-to-realization scatter were measured to be far narrower than a factor of ~50, the proposed distribution would be ruled out.","tokens_in":16513,"feed_emoji":"📡","tokens_out":7587,"duration_ms":60205,"temperature":0.7,"pith_summary":"This paper argues that the shot-noise anisotropy of the nHz gravitational wave background—the spatial patchiness produced by discrete supermassive black hole binary sources—cannot be summarized by an ensemble average. Using Monte Carlo simulations that draw Poisson realizations of the merger population from empirically calibrated models, the authors find that the shot-noise amplitude has a broad, positively skewed probability distribution: the 95% range spans a factor of about 50 at fixed frequency, and the median and most probable values are factors of 2 to 3 below the mean. They further show that the standard analytic estimate based on strain moments, $\\langle h_c^4 \\rangle/\\langle h_c^2 \\rangle^2$, systematically overpredicts the ensemble-averaged shot-noise because the average of a ratio is not the ratio of the averages, with the discrepancy growing from a factor of ~3 at $f = 0.1\\,\\mathrm{yr}^{-1}$ to more than two orders of magnitude at $f = 1\\,\\mathrm{yr}^{-1}$. The paper concludes that interpreting pulsar timing array anisotropy measurements requires modeling the full shot-noise distribution, and that the median shot-noise, which scales roughly as $f^{1.1\\text{–}1.2}$, is within reach of current and future PTA observations.","feed_headline":"Shot noise skews gravitational wave sky maps","feed_subtitle":"Pulsar timing arrays should model the full distribution; the median sits 2–3x below the mean.","key_machinery":"The central object is the shot-noise anisotropy estimator $\\hat{C}_{\\rm shot}/(4\\pi) = \\left(\\sum_i N_i h_i^4\\right)/\\left(\\sum_i N_i h_i^2\\right)^2$, where $N_i$ are Poisson-random integer counts of SMBHB mergers in cells of black hole mass, mass ratio, and redshift, and $h_i^2$ is the inclination- and polarization-averaged strain per source. Writing the estimator as $\\sum_i N_i w_i^2$ with weights $w_i = h_i^2/\\sum_j N_j h_j^2$ yields the exact bounds $1/N_{\\rm eff} \\le \\hat{C}_{\\rm shot}/(4\\pi) \\le 1$, making explicit that shot-noise is the inverse effective source count. The Monte Carlo machinery samples these Poisson counts from factorized merger models (LM24 and SZQ) with GW-driven residence times, then evaluates the full probability distributions of the shot-noise, the characteristic strain $h_c^2$, and the fourth moment $h_c^4$, along with their conditional distributions given a noisy measurement of $h_c^2$.","core_discovery":"The central discovery is that the shot-noise anisotropy level in a single sky realization, $\\hat{C}_{\\rm shot}/(4\\pi)$, is bounded between $1/N_{\\rm eff}$ and 1, where $N_{\\rm eff}$ is the effective number of sources, and that its ensemble average is systematically below the naive moment-based estimate. Since the shot-noise is defined as the ratio of the fourth moment of the strain field to the square of the second moment, the ensemble average of this ratio is not equal to the ratio of the ensemble averages; the paper quantifies this difference using $7\\times10^5$ Monte Carlo realizations of the LM24 and SZQ merger models. In the fiducial model the mean shot-noise exceeds the median by a factor of 2.9 at $f = 0.1\\,\\mathrm{yr}^{-1}$ and by 2.0 at $f = 1\\,\\mathrm{yr}^{-1}$, while the moment-based prediction exceeds the Monte Carlo mean by factors of 2.8 and 130 at those same frequencies. The shot-noise PDF remains broad across the PTA band, with a 95% interval of about 1.7 dex, and the median scales as $f^{1.1\\text{–}1.2}$ rather than the $f^{8/3}$ scaling of the moment-based estimate. The authors also show that conditioning on a measurement of the characteristic strain $h_c^2$ narrows the shot-noise and $h_c^4$ distributions, and that the median shot-noise is insensitive to the low-redshift cutoff $z_{\\rm min}$.","pith_inferences":["A similar ratio-of-moments bias may affect other stochastic background searches, such as spectral index estimation at high frequencies where source counts are low.","Because the median shot-noise is insensitive to $z_{\\rm min}$ while the mean is not, the low-redshift cutoff uncertainty, though real, should have a smaller impact on detection forecasts than on ensemble-averaged amplitude predictions.","The exact bounds $1/N_{\\rm eff} \\le \\hat{C}_{\\rm shot}/(4\\pi) \\le 1$ provide a built-in consistency check: an anisotropy measurement exceeding unity would indicate a calibration error, a modeling error, or a contribution beyond the shot-noise assumption."],"forward_implications":["PTA anisotropy analyses must either measure many independent sky realizations or forward-model the full shot-noise PDF, since the ensemble mean is not a typical value.","The median shot-noise, reaching $\\hat{C}_{\\rm shot}/(4\\pi) \\approx 0.05$ at $f \\approx 30\\,\\mathrm{nHz}$, is close to the current NANOGrav 95% upper limit of 0.20, so upcoming data could plausibly detect it.","Conditioning on a precise measurement of the characteristic strain $h_c^2$ can narrow the shot-noise distribution by up to an order of magnitude, making combined $h_c^2$ plus anisotropy measurements more informative.","The moment-based estimate $\\langle h^4 \\rangle/\\langle h^2 \\rangle^2$ is unreliable in the sparse-source regime and should not be used for high-frequency predictions in GW-driven models."],"supporting_citations":[{"why":"Defines the shot-noise formalism and the moment-based estimate that this paper corrects; supplies the baseline analytic prediction.","marker":"[19]"},{"why":"Provides the current NANOGrav 95% upper limit on $\\hat{C}_{\\rm shot}/(4\\pi) \\le 0.20$ against which the median predictions are compared.","marker":"[17]"},{"why":"Supplies the fiducial LM24 SMBHB merger model (mass function, redshift and mass-ratio distributions) from which the Monte Carlo realizations are drawn.","marker":"[30]"},{"why":"Supplies the alternative SZQ merger model used for comparison in the paper's model-dependence study.","marker":"[31]"},{"why":"Provides the analytic shot-noise approach and the stellar-hardening residence time model adopted for the modified scenario.","marker":"[20]"},{"why":"Gives the inclination- and polarization-averaged strain per source formula and the general Monte Carlo approach to strain realizations.","marker":"[21]"}],"fun_headline_variants":["GWB shot noise median lies 2–3x below the mean","Shot noise anisotropies vary 50-fold across realizations","Typical GWB shot noise is 2–3x quieter than mean","Pulsar timing arrays face broad shot-noise scatter","GWB shot noise: mean overestimates typical skies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result depends on the assumed SMBHB merger model being representative of the real universe, especially the sharp low-redshift cutoff $z_{\\rm min}=0.05$ and the high-mass cutoff $M_{\\rm BH,max}=10^{10.5}M_\\odot$; rare nearby or very massive sources control the tails of the shot-noise distribution, so if those cutoffs are wrong the claimed PDF width and mean-to-median ratios would shift.","fun_headline_variants_meta":{"raw":{"variants":["GWB shot noise median lies 2–3x below the mean","Shot noise anisotropies vary 50-fold across realizations","Typical GWB shot noise is 2–3x quieter than mean","Pulsar timing arrays face broad shot-noise scatter","GWB shot noise: mean overestimates typical skies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001562,"raw_usage":{"total_tokens":6394,"prompt_tokens":1251,"completion_tokens":5143,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":867,"completion_tokens_details":{"reasoning_tokens":5055}},"tokens_in":867,"tokens_out":5143,"duration_ms":31111,"temperature":1.0,"reasoning_tokens":5055,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:12:51.010926+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be a PTA measurement of the shot-noise anisotropy at multiple frequencies with the full PDF forward-modeled: if the observed shot-noise at $f \\approx 1\\,\\mathrm{yr}^{-1}$ consistently landed near the moment-based prediction (which exceeds the Monte Carlo mean by a factor of ~130) rather than near the median, or if the realization-to-realization scatter were measured to be far narrower than a factor of ~50, the proposed distribution would be ruled out.","supporting_citations":[],"review_version":1}