{"id":"43406891-0acf-4d00-94d7-de37ab1bd9c1","arxiv_id":"2411.16948","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Covariance matrices from small-volume simulations can be rescaled to match large-volume ones at the 3% level using a new bin-centering correction, provided the large-scale power spectrum is known.","lead":"This paper tests methods for adding super-sample covariance to mock galaxy catalogues and shows that covariance matrices from small simulations can be rescaled to match large ones to 3% accuracy on scales relevant for current surveys. The new bin-centering correction helps, but it relies on knowing the large-box power spectrum, so the method is not fully predictive.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 4.5's bin-centering correction requires P^L(k), measured from the target large-box ensemble, so the 3% volume-scaling match is partly a calibration to the benchmark rather than an independent prediction.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing point: P^L enters Eq. 4.5 and is taken from the benchmark ensemble. I verified the algebra: the P^S factor cancels, so the correction amount is controlled entirely by P^L and the mode-volume ratio. The Gaussian diagonal covariance is then exactly re-centered on the large-box P^L, which explains why Figure 10 shows the low-k diagonal is recovered only with the correction. This does not invalidate the SSC-method comparison (Section 3), which is independently supported by prior work and is a solid contribution. Nor does it make the volume-scaling result useless: for the dark-matter power spectrum, accurate external P^L estimates exist. But the paper's 3% claim, as currently demonstrated, is a calibration to the target large-box ensemble; the missing step is testing with an external P^L and quoting the added uncertainty. The smallest-box failure (156.25 h^-1 Mpc) is honestly noted and is a separate mode-count limitation, not part of the 512x claim. A single re-analysis with an emulator-supplied P^L would settle whether the headline claim is a prediction or a consistency test. Hence I keep the reader's CONDITIONAL verdict.","tokens_in":19387,"tokens_out":4919,"duration_ms":47431,"concrete_test":"Re-run the volume-scaling test of Fig. 7 (especially the 512x case, L_small=312.5 h^-1 Mpc to L_large=2500 h^-1 Mpc) with P^L(k) in Eq. 4.5 replaced by an independent analytic power spectrum for the same cosmology (e.g., HMCode2020 or EuclidEmulator2), leaving everything else unchanged. If the C_scaled/C_large ratios stay within 3% on k < 1 h/Mpc, the method is predictive. Additionally, perturb P^L by the emulator's stated uncertainty and report the induced change in the covariance ratio to quantify the added error budget.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—a 3% covariance match for 512x volume scaling—rests on the bin-centering correction in Eq. 4.5. Algebraically, Pcorr,S(k) = (P^L/P^S) sqrt(Vk,S/Vk,L) P_S = P^L sqrt(Vk,S/Vk,L): the small-box realization scatter is renormalized so that the Gaussian part of the scaled covariance reproduces exactly the large-box Gaussian covariance built from P^L. The paper measures P^L from the same (2500 h^-1 Mpc)^3 ensemble that defines the benchmark covariance. Consequently the low-k diagonal agreement, which Figure 10 shows fails without the correction, is enforced by construction rather than predicted. In a real application the target survey's P^L is not known a priori; for the dark-matter power spectrum it could be supplied by an emulator or Halofit, and for galaxy power spectra a model is needed, but the paper neither tests this nor propagates the uncertainty in P^L into the 3% budget. The SSC and trispectrum parts are not affected by this circularity, and the comparison of SSC methods is independent and credible, but the volume-scaling headline as demonstrated is a calibration to the benchmark until an external P^L is used.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the estimation of dark matter power spectrum covariance matrices from ensembles of simulations, focusing on super-sample covariance (SSC) and on scaling covariance matrices between simulation volumes. It compares two separate-universe prescriptions (Sirko and spherical collapse) and two ways of incorporating SSC (addition and ensemble methods), finding that they recover the covariance measured from sub-boxes of a large simulation to within about 10%. The second half of the paper studies volume scaling: using an ensemble of smaller boxes to infer the covariance of a larger survey box, with a proposed bin-centering correction (Eq. 4.5) and an SSC term added via the addition method. The authors report a 3% match for the diagonal covariance elements when scaling the small-box volume up by a factor of 512 (from (312.5 h^-1 Mpc)^3 to (2500 h^-1 Mpc)^3), while noting that the smallest box tested, (156.25 h^-1 Mpc)^3, fails at low k because its lowest bins contain very few modes and become non-Gaussian. The paper argues this volume-scaling approach could substantially reduce the computational cost of covariance estimation for current and future surveys.","tokens_in":19651,"tokens_out":5153,"duration_ms":48044,"significance":"If the central claim holds in a predictive sense, the volume-scaling method would be a practically important shortcut: covariance matrices for large surveys could be estimated from much cheaper small-box simulations. The paper's strengths are its systematic comparison of SSC implementations, its identification of the bin-centering and discrete-mode-number effects, its use of a large sub-box benchmark to validate the SSC methods, and its public release of simulation parameter files and power spectra. The SSC method comparison is a solid, self-contained contribution that confirms and extends previous work. However, as demonstrated, the headline 3% volume-scaling result is partly a calibration rather than an independent prediction: Eq. (4.5) requires the target large-box power spectrum P^L(k) as an input, and the paper measures P^L from the same large-box ensemble that defines the benchmark. The paper therefore needs to show that the method works when P^L is supplied externally (e.g., from an emulator or analytic model) or to explicitly reframe the claim as a consistency test.","major_comments":[{"comment":"The bin-centering correction in Eq. (4.5) uses P^L(k), the ensemble-average power spectrum of the large-volume mocks, and the benchmark covariance is also computed from that same ensemble. Algebraically, after applying the correction the Gaussian part of the scaled small-box covariance equals the Gaussian covariance built from P^L(k) by construction: the corrected small-box power has mean P^L sqrt(V_{k,S}/V_{k,L}) and a Gaussian variance that, after volume scaling, reproduces the large-box Gaussian covariance. Consequently, the low-k diagonal agreement shown in Fig. 10 is enforced rather than predicted. The manuscript should either (a) demonstrate the method with P^L taken from an external model (e.g., Halofit or an emulator) and show that the 3% agreement persists, or (b) at minimum state clearly that Eq. (4.5) requires P^L as an external input and propagate the uncertainty in P^L into the quoted 3% budget. As written, the central volume-scaling claim is a calibration to the benchmark, not an independent validation.","section":"§4.1, Eq. (4.5)"},{"comment":"The abstract and conclusions claim a 3% match in the dark matter power spectrum covariance, but the evidence in Figs. 6-7 is for diagonal elements, while Fig. 9 shows that off-diagonal elements in the lowest k bin (k_j = 0.04 h Mpc^-1) deviate substantially from the large-box covariance after the bin-centering correction, becoming overestimated. The paper itself acknowledges that the correction is designed for the Gaussian piece and is inaccurate for the trispectrum/SSC-dominated off-diagonal terms. The claim should either be restricted explicitly to diagonal elements (with the off-diagonal caveat stated in the abstract) or the correction should be extended to handle the off-diagonal terms before claiming a 3% match for the full covariance matrix.","section":"§4.2, Figs. 9-10"},{"comment":"The paper notes in §3.4 that fast N-body codes such as L-PICOLA underestimate the full non-Gaussian covariance at k > 0.2 h Mpc^-1 compared to full N-body codes, as demonstrated by the comparison with L-Gadget2 results in Fig. 5. Because both the small-box ensembles and the benchmark sub-boxes are generated with L-PICOLA, the 3% volume-scaling agreement is only demonstrated for the approximate code and does not directly establish that the method recovers the true non-linear covariance. The manuscript should either qualify the main claim as L-PICOLA-specific or test the volume-scaling procedure against a full N-body sub-box benchmark in at least one configuration.","section":"§3.4 and §4.2"}],"minor_comments":[{"comment":"The caption contains a typo: it states 'the left panel shows the off-diagonal elements' twice; the second reference should be to the right panel.","section":"Figure 10 caption"},{"comment":"The notation in Eq. (4.5) is confusing: P^L(k) and P^S(k) are introduced as ensemble-average powers, while P_S(k) on the right appears to be the power spectrum of an individual small-box realization. Please use a distinct symbol (e.g., \\bar P or a hat) for the single-realization quantity to avoid ambiguity.","section":"Eq. (4.5)"},{"comment":"There is a missing space in 'using theSC addition method ensemble'; this should read 'using the SC addition method ensemble'.","section":"§4.2, first paragraph"},{"comment":"The sentence quoting 'a 3% match on scales k<1.0 h^-1 Mpc scaling the simulation volume by a factor 512' should explicitly say 'diagonal elements of the covariance' to be consistent with the presented figures, given the off-diagonal caveat in Fig. 9.","section":"§5, Conclusions"},{"comment":"The new contribution relative to Howlett and Percival (2017) [34] should be stated more explicitly. The paper extends that work with a wider range of volume ratios and a new bin-centering correction, but the reader has to infer this from context; a sentence in the introduction or in §4.1 clarifying the new elements would help.","section":"§1 and §4.1"}],"recommendation":"major_revision","confidential_remarks":"The SSC comparison part of the paper is solid and worth publishing. The volume-scaling claim, however, is not yet presented in a form that supports the abstract's predictive wording, because Eq. (4.5) requires the target power spectrum as input and the paper measures it from the benchmark ensemble. I would ask the authors to either supply an external P^L test (e.g., from Halofit at the same cosmology) or explicitly reframe the claim as a calibration/consistency check, and to consistently separate diagonal from off-diagonal statements. The L-PICOLA-only validation is also a concern; a single full N-body test of the 512x scaling would materially increase confidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Greg, quick take on 2411.16948. The SSC method comparison is the solid part: for the first time it directly compares Sirko and spherical-collapse parameter choices, and addition versus ensemble methods. The sub-box covariance match to within ~10% is credible and consistent with prior work. The bin-centering correction for discrete mode numbers is a genuinely new practical fix, and the paper is honest about where it works and where it doesn't.\n\nThe stress-test note is right about the circularity, and it matters. Eq. 4.5 multiplies the small-box power by P^L(k)/P^S(k), so the Gaussian part of the small-box diagonal covariance is rescaled to exactly reproduce the large-box Gaussian covariance built from P^L. Since they measure P^L from the same (2500 Mpc)^3 ensemble that defines the benchmark, the low-k diagonal agreement is enforced by construction, not predicted. That makes the '3% match' a calibration to the benchmark unless P^L is supplied externally. The paper never tests that route, and it doesn't propagate uncertainty in P^L into the 3% budget. For a real survey you'd need an emulator or model for P^L; the paper should say that clearly.\n\nBut not everything is circular. The off-diagonal and higher-k parts of the covariance are not enforced by Eq. 4.5, and the SSC comparison is independent and credible. The smallest-box failure is correctly attributed to non-Gaussianity, and the L-PICOLA underestimate of high-k covariance is acknowledged. They also ship the simulation parameter files and power spectra on Zenodo, which is a real plus.\n\nBottom line: the SSC section is publishable largely as is. The volume-scaling headline needs either an external P^L demonstration or a much more careful statement that the 3% is post-calibration. I would send this to peer review—a referee can push on the circularity—but I would not take the headline at face value. I'd cite it for the SSC comparison and the bin-centering idea, less for the volume-scaling promise.","headline":"Useful SSC method comparison, but the headline 512x volume-scaling claim is partly calibrated to the benchmark through Eq. 4.5, so treat it as a proof of concept until an external P^L is used.","tokens_in":20176,"tokens_out":1901,"would_cite":true,"duration_ms":18969,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Covariance matrices for large galaxy surveys can be estimated from simulations 512 times smaller, with only 3% error, by correctly handling super-sample covariance and correcting for discrete-mode bin centering.","keywords":["super-sample covariance","covariance matrix","volume scaling","separate universe simulations","power spectrum","mock catalogues","large-scale structure","DESI"],"falsifier":"Run the volume-scaling pipeline but supply the large-box power spectrum for Eq. (4.5) from an external analytic or emulated prediction instead of from the large-box ensemble itself; if the volume-scaled covariance then drifts beyond the claimed 3% agreement, the central claim fails as a prediction rather than a calibration.","tokens_in":19199,"feed_emoji":"🔭","tokens_out":4767,"duration_ms":42476,"temperature":0.7,"pith_summary":"The paper argues that the covariance matrix of the dark matter power spectrum, the statistic that sets error bars for galaxy surveys, can be estimated from ensembles of simulations much smaller than the survey volume, provided two effects are handled: super-sample covariance, the scatter induced by density fluctuations larger than the simulation box, and the bin-centering bias from the discrete set of Fourier modes a smaller box contains. It compares two ways of building the super-sample effect into mock ensembles, either by perturbing cosmological parameters in separate-universe simulations or by adding a response term to the covariance, and finds they agree. It then shows that, after rescaling a small-box covariance to a large-box volume and applying a new correction that re-centers the effective wavenumber of each bin, the scaled covariance matches the full-size simulation covariance to within 3% on scales of interest for current surveys, even at a volume ratio of 512. The payoff is large: covariance matrices are often the dominant computational cost of a survey analysis, so this would let them be produced from much cheaper simulations.","feed_headline":"Small mocks reach 3% match for big-survey covariance","feed_subtitle":"A bin-centering fix plus proper super-sample covariance lets 512x smaller boxes stand in for full-size survey simulations.","key_machinery":"The machinery is built on the separate-universe response: a long-wavelength density fluctuation $\\delta_b$ is reinterpreted as a change in the simulation's cosmological parameters, so the power-spectrum derivative $dP/d\\delta_b$ can be measured from pairs of simulations with perturbed background density. Super-sample covariance is then either generated inside an ensemble by drawing $\\delta_b$ for each mock from a Gaussian, or added analytically as $\\sigma_b^2 (dP/d\\delta_b)^2$ at the survey volume. The second piece is the bin-centering correction of Eq. (4.5), which rescales the small-box power spectrum by the ratio of $k$-space shell volumes $V_{k,S}/V_{k,L}$ and the ratio of the large- and small-box ensemble-average power spectra, correcting the fact that a discrete grid of modes measures $P(k)$ at a slightly different effective wavenumber in each box. Together these corrections make the covariance a function of volume that can be rescaled between box sizes.","core_discovery":"On its own terms, the paper establishes that volume scaling of covariance matrices works when super-sample covariance is added back at the survey volume and when the discrete-mode bin-centering effect is corrected. For five ensembles of L-PICOLA dark-matter simulations with volumes spanning a factor of 4096, the volume-scaled small-box covariance matches the large-box covariance to better than 3% on essentially all scales for volume ratios up to 512, with the exception of the smallest boxes whose low-k bins contain so few modes that the power spectrum distribution becomes non-Gaussian. The paper identifies the unresolved limitation: modes with wavelengths between the small and large box sizes couple to the measured modes, and their contribution cannot be reintroduced by the current method, which is why sub-percent agreement is out of reach. It also reports that the Sirko and spherical-collapse separate-universe parameterizations give nearly identical super-sample covariance, and that the additive method, which computes a power-spectrum response and adds the super-sample term separately, is the preferred route for volume scaling.","pith_inferences":["A testable extension would replace the measured large-box power spectrum in Eq. (4.5) with a theoretical power spectrum; if that preserves the 3% match, the method becomes a genuine prediction rather than a calibration to the target.","The off-diagonal failure at low k could plausibly be repaired by a higher-order correction that scales the trispectrum and super-sample terms the way Eq. (4.5) scales the Gaussian term, which the paper leaves for future work.","The missing intermediate-mode contribution might be captured with a tidal-field response in addition to the monopole background response, potentially pushing volume scaling below the 3% floor.","Real-survey complications such as window functions and redshift-space distortions likely degrade the 3% match; testing the same scaling on galaxy mocks with survey geometry would establish how much compute can actually be saved."],"forward_implications":["Surveys can estimate covariance matrices from ensembles of boxes hundreds of times smaller than the survey volume, cutting the dominant computational cost of mock-based error estimation.","The addition method for super-sample covariance should be preferred over the ensemble method when volume scaling is planned, since the response term is volume independent and can be rescaled to the survey volume.","Sub-percent covariance accuracy cannot be achieved by pure volume scaling, because modes with wavelengths between the two box sizes contribute to the covariance and are not included.","Very small boxes must be avoided for the lowest-k bins: with few modes per bin, the power spectrum becomes significantly non-Gaussian and a covariance matrix alone is no longer a complete description.","The 3% agreement is comparable to the agreement between semi-analytic and mock-based covariance estimates used for DESI Y1, suggesting it is adequate for current survey analyses."],"supporting_citations":[{"why":"Introduced the volume-scaling idea for covariance and previously demonstrated sub-box and separate-universe consistency that this work extends.","marker":"[34]"},{"why":"Defined super-sample covariance as the response of the power spectrum to a background density mode, giving the core additive formula.","marker":"[41]"},{"why":"Supplies the Sirko separate-universe parameter perturbation used as one of the two SSC prescriptions.","marker":"[68]"},{"why":"Provides the spherical-collapse separate-universe parameterization used for the SC method.","marker":"[48]"},{"why":"Established separate-universe simulation recovery of SSC and provides the comparison format used in this paper's Fig. 5.","marker":"[47]"},{"why":"Provides a recent independent validation of separate-universe SSC recovery against sub-boxes, used as a comparison.","marker":"[63]"},{"why":"Sets the DESI Y1 semi-analytic covariance target whose agreement level the 3% match is compared against.","marker":"[83]"},{"why":"Established that covariance error scales as 1/sqrt(N)/V, the basis of the computational-saving argument.","marker":"[45]"}],"fun_headline_variants":["Volume scaling with SSC hits 3% covariance match","512x smaller boxes: covariance within 3%","Bin-centering fix enables mock volume scaling","Sub-percent out of reach: modes between boxes","SSC plus binning: volume-scaled mocks to 3%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 3% volume-scaling match assumes the ensemble-average power spectrum of the large box is known in advance, because the bin-centering correction feeds it in; without an independent source for that quantity, the method calibrates to the very simulation it is meant to replace.","fun_headline_variants_meta":{"raw":{"variants":["Volume scaling with SSC hits 3% covariance match","512x smaller boxes: covariance within 3%","Bin-centering fix enables mock volume scaling","Sub-percent out of reach: modes between boxes","SSC plus binning: volume-scaled mocks to 3%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1549,"prompt_tokens":1048,"completion_tokens":501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":421}},"tokens_in":664,"tokens_out":501,"duration_ms":5328,"temperature":1.0,"reasoning_tokens":421,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:41:55.941202+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the volume-scaling pipeline but supply the large-box power spectrum for Eq. (4.5) from an external analytic or emulated prediction instead of from the large-box ensemble itself; if the volume-scaled covariance then drifts beyond the claimed 3% agreement, the central claim fails as a prediction rather than a calibration.","supporting_citations":[],"review_version":1}