{"id":"5b491232-6871-4778-a76e-4b864d7ad37a","arxiv_id":"2506.19506","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"PINOCCHIO, a fast semi-analytic halo simulator, reproduces N-body cosmic void size, shape, core density, and radial profile statistics to within roughly 2 sigma.","lead":"This paper checks whether a fast approximate simulation code, PINOCCHIO, can produce reliable cosmic void statistics compared with a full N-body simulation. The authors find that all four void statistics generally agree within about 2 sigma, which would make PINOCCHIO useful for large survey analyses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'better than 2σ' claim is not secured: the paired-simulation covariance is never computed, so the Maxerr estimator is not guaranteed to bound the true significance; the abstract overstates the body's 'mostly within 2σ'.","rationale":"I agree with the reader's CONDITIONAL verdict, but my most load-bearing concern is different. The reader's weakest assumption is the number-density matching in Section 2.3.3; that is a legitimate caveat, but the paper argues plausibly that matching tracer abundance is necessary for a fair void comparison and the residual differences are not erased by it. The more directly load-bearing issue is the significance quantification underlying the headline claim. The paper's central assertion is 'agreement for all void statistics at better than 2σ,' and this is supported by two ad hoc estimators plus Gaussian fits to aggregated residuals. The Maxerr estimator is asserted to be an upper limit on significance, but as written in Eqs. (12)-(13) this is not guaranteed: with positively correlated paired measurements, the true variance of the difference can be either smaller or larger than Var(G) depending on whether Cov exceeds Var(P)/2. Since the covariance is never computed, the claimed 2σ bound is not established. The body also contains statements that are weaker than the abstract, e.g., 'most values fall within ±2σ' and Appendix B's finding that the large-box CDF worsens at z>0. The paper has real strengths: identical initial conditions, matched halo samples, jackknife errors, a large-box robustness check, and physically sensible interpretations of residual trends. None of my concerns invalidate the conclusion that PINOCCHIO performs well for voids; they do mean the paper should either compute the covariance or soften the abstract and conclusions. This is consistent with the reader's CONDITIONAL verdict, so I recommend no change to the verdict.","tokens_in":17579,"tokens_out":7789,"duration_ms":85812,"concrete_test":"Compute the jackknife covariance between the PINOCCHIO and OpenGADGET3 void statistics using the same delete-one subvolumes, and evaluate z_bin = (P_bin - G_bin) / sqrt(Var_G + Var_P - 2Cov) for every VSF, VEF, CDF, and RDP bin. If any bin with Maxerr < 2 has z_bin > 2, the 'better than 2σ for all void statistics' claim fails; report the distribution of z_bin and compare it with the Gaussian fits shown in Figures 4, 6, 8, and 10.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.3.3 defines two significance estimators, Maxerr (Eq. 12) and Minerr (Eq. 13), and claims they bracket the true significance because the simulations share initial conditions. This bracketing is not rigorous. For a paired difference with variance Var(G)+Var(P)-2Cov, the denominator used in Maxerr, sigma_G, is smaller than the true error only if Cov > Var(P)/2. The paper does not demonstrate this inequality, and it can fail if the PINOCCHIO jackknife error is large relative to OpenGADGET3's or if the positive correlation is modest. In that regime Maxerr would underestimate the true significance, so the abstract's unconditional 'better than 2σ for all void statistics' is not established. The body itself repeatedly says 'most' bins fall within 2σ, and Appendix B reports that the large-box CDF worsens at z>0. The central claim therefore depends on an estimator that is not proven conservative and on a reading of the figures that is stronger than the text supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a validation study of the fast halo-catalog generator PINOCCHIO against the full N-body code OpenGADGET3 for the purpose of producing cosmic void statistics. The authors run two resolutions (512^3 and 1024^3 particles) in a 512 h^-1 Mpc box with identical initial conditions, identify voids with the VIDE watershed finder applied to halo catalogs, and compare four void statistics: the void size function, the void ellipticity function, the core density function, and stacked radial density profiles, over five redshifts. After matching halo number densities between the two codes, they find that most bins agree within 2 sigma, with deviations typically within 10%, and conclude that PINOCCHIO reliably reproduces void statistics at a fraction of the computational cost of N-body simulations. An appendix reports a larger-box test with a higher mass cut.","tokens_in":17740,"tokens_out":4197,"duration_ms":45298,"significance":"If the central claim holds, this work would establish PINOCCHIO as a computationally efficient tool for generating mock void catalogs for survey forecasts and covariance estimation, which is of practical importance for upcoming large-volume surveys such as Euclid and DESI. The study is carefully designed in several respects: it uses identical initial conditions, a fixed external N-body benchmark, matched halo number densities, jackknife errors, and a large-box robustness check. The comparison is measured, not fitted, and therefore has the character of a genuine validation. However, the abstract's unconditional 'better than 2 sigma for all void statistics' is stronger than the body's own repeated 'most bins fall within 2 sigma', and the statistical bracketing argument used to avoid computing the paired covariance is not just unproven but can fail in the very regime of strong positive correlation that shared initial conditions are expected to produce. The significance of the paper is real, but the quantitative headline claim needs to be either substantiated with a proper covariance treatment or restated in a more cautious form.","major_comments":[{"comment":"The abstract states 'agreement for all void statistics at better than 2sigma between PINOCCHIO and OpenGADGET3, with no systematic difference in redshift trends.' This is contradicted by the body's own qualified language: Section 3.1 says 'most values fall within +/-2sigma' for the VSF, Section 3.3 notes a systematic overestimation in the CDF at higher core densities, and Appendix B reports that the CDF agreement 'worsens, particularly at redshifts z>0' for the large box. The claim should be softened to 'mostly within 2sigma' or the specific exceptions (large-size VSF bins, high core-density CDF bins, low-resolution low-mass HMF bins) should be explicitly characterized as statistically consistent but with localized systematic deviations.","section":"Abstract and Sections 3.1-3.4"},{"comment":"The claim that the Maxerr and Minerr estimators bracket the true significance is not established. For paired simulations sharing initial conditions, the variance of the difference is Var(G)+Var(P)-2Cov. The Maxerr denominator, sigma_G, is smaller than the true error only if Cov < Var(P)/2; if the positive correlation is stronger (plausible for shared ICs, and likely when sigma_P is comparable to sigma_G), Maxerr can underestimate the significance of the difference. Similarly, Minerr is not a guaranteed lower bound for all covariance values. The paper either needs to compute the paired covariance (e.g., by paired jackknife resampling across both codes) or explicitly restrict its claims to the Maxerr statistic and avoid saying the true significance is bracketed. As written, the 'better than 2sigma' conclusion rests on an estimator whose conservativeness is assumed rather than derived.","section":"Section 2.3.3, Eqs. (12)-(13)"},{"comment":"The halo number-density matching procedure (selecting from PINOCCHIO the same number of halos as found in OpenGADGET3 above a fixed mass cut) removes the HMF normalization difference before voids are measured. The paper should state more explicitly that the validation is therefore conditional on a calibrated halo abundance: it tests whether PINOCCHIO's spatial clustering, given the same tracer number density as the N-body run, yields the same void population. This is a meaningful and standard comparison, but it does not validate PINOCCHIO's absolute prediction of void abundances, which would inherit any HMF systematics. A sentence clarifying this scope would prevent the 'reliably produce void statistics' claim from being over-interpreted.","section":"Section 2.3.3, Number Density Matching"}],"minor_comments":[{"comment":"There is a typo in 'enviroments' in the first paragraph of the introduction; also, 'V oid' and 'V oronoi' appear with stray spaces in several places (e.g., Section 2.1).","section":"Introduction"},{"comment":"The word 'resdshifts' appears instead of 'redshifts' in the first sentence after Figure 4.","section":"Section 3.1"},{"comment":"The second Heaviside step function is written as theta[-(r_j - delta r)^3]; the exponent 3 appears to be a typo, since the radial shell condition should involve theta[-(r_j - (r - delta r))] or an equivalent form without a cubic exponent.","section":"Eq. (7)"},{"comment":"The paper does not state how many jackknife resamplings were used or how the jackknife regions were defined; specifying this would improve reproducibility, since the error bars are central to the significance statements.","section":"Section 2.3.3 and Figures 1-10"},{"comment":"The large-box test uses a different halo mass cut (10^14 M_sun/h) and only three redshifts; this is appropriately acknowledged, but the caption or text should make clearer that the CDF discrepancy at z>0 is a resolution effect rather than a redshift effect, to avoid tension with the abstract's 'no systematic difference in redshift trends.'","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, useful code-validation study, and the authors are appropriately careful in most of the body text. The main issue is that the abstract and the significance-bracketing argument overstate the strength of the conclusion. The 'better than 2 sigma for all void statistics' claim is not supported by the displayed 'most bins' statements, and the Maxerr/Minerr bracketing is not rigorous for strongly correlated paired simulations. I believe these can be fixed by computing or approximating the paired covariance, or by rewording the claims to be explicitly limited to the Maxerr statistic and the observed 'mostly within 2 sigma' behavior. I also suggest the authors make the conditional nature of the number-density-matching comparison more prominent, since it affects how general the validation is. With those revisions, the paper would be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a careful validation of PINOCCHIO for void statistics, and it largely succeeds, but the abstract's 'agreement for all void statistics at better than 2σ' is stronger than what the analysis actually shows. The paper is the first to test PINOCCHIO on the VSF, VEF, CDF and stacked RDP against OpenGADGET3 with shared initial conditions, and the design is thoughtful: matched halo number densities, jackknife errors, two significance estimators, and a large-box robustness run. The authors are candid about weak spots—the CDF is the hardest statistic, and resolution dependence is visible in the VSF.\n\nThe real soft spot is the significance bracket. Maxerr divides by sigma_G only, Minerr by quadrature sum, and they claim these bound the true paired significance because the ICs are shared. That is not guaranteed. For Maxerr to be conservative, the covariance between the two codes would need to be large enough relative to PINOCCHIO's own variance; they never establish that inequality. The body itself hedges, saying 'most' bins fall within 2σ, and the large-box CDF worsens at z>0. So the abstract overstates the result. This is fixable: they could compute the paired covariance directly (they have the matched ICs, so it is just a Jackknife on the differences), or soften the claim.\n\nAlso, the number-density matching is a free parameter. The validation is therefore conditional on that calibration. That is standard and not a fatal flaw, but it means the good agreement reflects PINOCCHIO's ability to reproduce voids from a tracer set matched in abundance, not necessarily its raw predictive power from the same mass threshold.\n\nNone of this overturns the main result. The paper is a solid, honest validation that will be useful to anyone using fast simulations for void covariance estimation, forecasts, or parameter exploration. It deserves a serious referee. I would send it out, with the request that the authors either compute the covariance or carefully reword the 2σ claim.","headline":"A careful and useful first validation of PINOCCHIO for void statistics, but the blanket 'better than 2σ' claim in the abstract overstates what the error-bracketing argument actually supports.","tokens_in":18302,"tokens_out":3025,"would_cite":true,"duration_ms":30135,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fast simulations match N-body void statistics within 2σ.","keywords":["cosmic voids","large-scale structure","fast simulations","PINOCCHIO","Lagrangian perturbation theory","N-body simulations","void size function","radial density profiles"],"falsifier":"Repeat the same initial-condition comparison without the number-density matching step: if the void size and core density functions then differ by more than 2σ between PINOCCHIO and OpenGADGET3, the claimed agreement is an artifact of the calibration rather than an intrinsic property of PINOCCHIO's void modeling.","tokens_in":17352,"feed_emoji":"🌌","tokens_out":8721,"duration_ms":77591,"temperature":0.7,"pith_summary":"Cosmic voids are underdense regions whose size, shape, emptiness, and surrounding density profiles carry cosmological information. Generating them from full N-body simulations is expensive, so the paper asks whether PINOCCHIO, a fast halo-catalog generator based on Lagrangian perturbation theory, can reproduce the void population of a full N-body code. Using the same initial conditions and the same watershed void finder, it compares four void summary statistics across five redshifts and two resolutions. It finds agreement at better than 2σ for all statistics, with no systematic redshift trend, and concludes that PINOCCHIO can serve as a fast tool for void cosmology, at a computational cost more than a thousand times lower than N-body simulation.","feed_headline":"Fast code matches N-body void statistics within 2σ","feed_subtitle":"The 1,000x cheaper halo generator reproduces void sizes, shapes, cores, and profiles across five redshifts.","key_machinery":"The argument runs on four coupled pieces. PINOCCHIO is a fast halo-catalog code that combines ellipsoidal collapse, computed with third-order Lagrangian perturbation theory, with a fragmentation algorithm that builds halo merger histories; it converts an initial density field directly into a halo catalog without full gravitational evolution. OpenGADGET3 is the TreePM N-body code used as the benchmark, with halos found by SUBFIND and masses converted to a Friends-of-Friends definition via a low-particle-count mass correction. VIDE is a watershed void finder that tessellates the halo field with Voronoi cells, groups cells into basins around local density minima, and outputs each void's volume-weighted center, effective radius $R_{\\rm eff}$, inertia-tensor ellipticity $\\epsilon=1-(J_1/J_3)^{1/4}$, and core density $\\hat n_C$. The comparison machinery is the number-density matching step: after a mass cut on the OpenGADGET3 catalog, the same number of the most massive PINOCCHIO halos is selected, so both codes trace the density field with statistically comparable tracer sets before void statistics are measured.","core_discovery":"The paper's central claim is that when PINOCCHIO and OpenGADGET3 are run from the same initial conditions and their halo catalogs are matched in number density above a $10^{13}\\,M_\\odot/h$ mass cut, the void populations they produce are statistically indistinguishable. For the void size function, ellipticity function, core density function, and stacked radial density profiles, the residuals between the two codes stay within 10% for most scales and within 2σ across all bins and redshifts from 0.0 to 2.0, with no systematic redshift-dependent bias. The one statistic where the match is weakest is the core density function at high core densities, where PINOCCHIO overproduces voids; the paper attributes this to its fragmentation algorithm boosting low-mass halo counts, and notes that the agreement improves with resolution. The paper interprets the overall agreement as evidence that PINOCCHIO's quasi-linear treatment is adequate for the underdense regions where voids live, making it a candidate fast simulator for void surveys.","pith_inferences":["Inference: if PINOCCHIO also reproduces the covariance of void statistics, as it already does for halo statistics, then survey covariance matrices for void cosmology could be generated thousands of times cheaper; the paper explicitly leaves this validation to future work.","Inference: the 2σ agreement is partly downstream of the number-density matching calibration, so an uncalibrated run would isolate PINOCCHIO's intrinsic void accuracy; this is a direct, cheap test of the paper's load-bearing assumption.","Inference: the redshift-dependent core-density mismatch at large box and low mass resolution suggests a resolution floor for void-core modeling; interpolating between the tested resolutions would map where PINOCCHIO's void cores stop tracking N-body results.","Inference: because the paper's calibration is cosmology-independent within standard cosmology, the same matching procedure could be ported to modified-gravity simulations, but only after recalibration, which the paper flags as future work."],"forward_implications":["Void size functions, ellipticity functions, core density functions, and radial density profiles can be generated from PINOCCHIO instead of N-body simulations for survey forecasts and covariance estimates in the tested mass and resolution regime.","Because the agreement holds at every redshift without a systematic trend, PINOCCHIO can track the growth and tidal evolution of voids over cosmic time at a fraction of the cost.","The match improves with resolution, so applications should favor higher-resolution PINOCCHIO runs; the large-box test shows the core density function degrades when mass resolution drops.","Any residual difference between PINOCCHIO and N-body void statistics can be treated as a calibratable bias, to be folded into parameter inference rather than eliminated by expensive simulations.","The method is tuned to halo mass definitions and number-density matching; direct application to galaxy tracers such as luminous red galaxies requires the same calibration steps."],"supporting_citations":[{"why":"Supplies the PINOCCHIO code itself and the fragmentation and merger-tree construction that turns LPT displacements into halo catalogs.","marker":"Monaco et al. 2013"},{"why":"Introduces the PINOCCHIO method of predicting orbit crossing and halo formation from initial density fields.","marker":"Monaco et al. 2002"},{"why":"Calibrates PINOCCHIO's accuracy for the halo mass function and power spectrum, including the 3LPT and 2LPT setup used here.","marker":"Munari et al. 2017"},{"why":"Provides the VIDE void finder used to identify watershed voids and measure their properties.","marker":"Sutter et al. 2015"},{"why":"Describes the ZOBOV watershed algorithm that VIDE implements to define voids from tracer tessellation.","marker":"Neyrinck 2008"},{"why":"Documents the TreePM N-body solver underlying OpenGADGET3, the benchmark simulation code.","marker":"Springel 2005"},{"why":"Supplies the low-particle-count mass correction applied to OpenGADGET3 halo masses before the comparison.","marker":"Warren et al. 2006"},{"why":"Defines the Friends-of-Friends halo mass function to which PINOCCHIO is calibrated, justifying the FoF mass definition used for both catalogs.","marker":"Watson et al. 2013"},{"why":"Provides the analogous halo number-density matching procedure used to calibrate catalogs from different simulation codes.","marker":"Fumagalli et al. 2021"},{"why":"Supplies the radial density profile stacking method and the void-profile framework used in the RDP comparison.","marker":"Hamaus et al. 2014"}],"fun_headline_variants":["Fast sims match N-body void statistics within 2σ","Cheaper halo code reproduces void properties","Void sizes, shapes, cores match N-body in fast run","Fast PINOCCHIO voids agree with N-body within 2σ","Void statistics from fast code match N-body simulations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that matching the two halo catalogs to the same number density after a mass cut is a fair and sufficient calibration, so the measured void agreement reflects PINOCCHIO's intrinsic accuracy rather than the calibration itself.","fun_headline_variants_meta":{"raw":{"variants":["Fast sims match N-body void statistics within 2σ","Cheaper halo code reproduces void properties","Void sizes, shapes, cores match N-body in fast run","Fast PINOCCHIO voids agree with N-body within 2σ","Void statistics from fast code match N-body simulations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000295,"raw_usage":{"total_tokens":1763,"prompt_tokens":1041,"completion_tokens":722,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":639}},"tokens_in":657,"tokens_out":722,"duration_ms":7545,"temperature":1.0,"reasoning_tokens":639,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:31:38.035431+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same initial-condition comparison without the number-density matching step: if the void size and core density functions then differ by more than 2σ between PINOCCHIO and OpenGADGET3, the claimed agreement is an artifact of the calibration rather than an intrinsic property of PINOCCHIO's void modeling.","supporting_citations":[{"cited_title":"2013, MNRAS, 433, 2389","cited_arxiv_id":null,"evidence_quote":"Supplies the PINOCCHIO code itself and the fragmentation and merger-tree construction that turns LPT displacements into halo catalogs."},{"cited_title":"2017, MNRAS, 465, 4658","cited_arxiv_id":null,"evidence_quote":"Calibrates PINOCCHIO's accuracy for the halo mass function and power spectrum, including the 3LPT and 2LPT setup used here."},{"cited_title":"M., Lavaux, G., Hamaus, N., et al","cited_arxiv_id":null,"evidence_quote":"Provides the VIDE void finder used to identify watershed voids and measure their properties."},{"cited_title":"S., Abazajian, K., Holz, D","cited_arxiv_id":null,"evidence_quote":"Supplies the low-particle-count mass correction applied to OpenGADGET3 halo masses before the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the analogous halo number-density matching procedure used to calibrate catalogs from different simulation codes."},{"cited_title":"M., & Wandelt, B","cited_arxiv_id":null,"evidence_quote":"Supplies the radial density profile stacking method and the void-profile framework used in the RDP comparison."}],"review_version":1}