REVIEW 4 major objections 4 minor 39 references
Adding patch-level neural summaries to the power spectrum and bispectrum improves f_NL constraints by 30–45% at low k_max, capturing information beyond the 3-point function.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 09:53 UTC pith:ZQ5OXCND
load-bearing objection A credible Fisher study showing patch-level neural summaries add real f_NL information over P(k)+B(k), but the headline gains need error bars and a linearity check on the network derivative. the 4 major comments →
Hierarchical summaries for primordial non-Gaussianities
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a field-level neural summary computed on sub-volumes can be spliced onto standard large-scale statistics to recover nearly all available f_NL information. Concretely, the power spectrum and bispectrum of the full simulation box are measured up to a scale cut k_max, while PatchNet maps each 125^3 Mpc^3/h patch of the smoothed halo overdensity field to a predicted f_NL; the mean over 512 non-overlapping patches becomes a one-number summary. The Fisher information of the concatenated vector P(k)+B(k)+patches is computed with 5000 fiducial simulations for the covariance and finite-difference derivatives from f_NL = -50 and +50 simulation suites. The paper finds that the
What carries the argument
PatchNet: a convolutional neural network that takes a small cubic patch of the smoothed halo density field and outputs a single predicted f_NL value; averaging over all patches gives a field-level compression that is concatenated with the power spectrum and bispectrum. It extracts small-scale, higher-order information that the 2-point and 3-point functions miss, while the global n-point statistics retain the large-scale squeezed configurations that small patches cannot see. The hierarchy is motivated by the locality principle: late-time gravitational evolution produces local correlations, so the non-local primordial signal stays separable, making the combination of n-point statistics with a
Load-bearing premise
The forecast treats the neural network's output as a linear function of f_NL near zero, so the derivative measured from simulations at f_NL = +/- 50 is assumed to represent the true slope; this linearity is verified only for the power spectrum, not for PatchNet, and the quoted gains have no error bars.
What would settle it
Recompute the PatchNet Fisher derivative using several finite-difference offsets (e.g., f_NL = +/- 20, +/- 50, +/- 100) and verify that sigma_fNL and the 30–45% gain are stable; because the network is nonlinear, a strongly offset-dependent derivative would mean the reported gain is an artifact of the chosen delta f_NL.
If this is right
- If the Fisher gains hold, adding PatchNet to a survey analysis at k_max about 0.1 h/Mpc delivers the f_NL constraining power of standard statistics at k_max about 0.5 h/Mpc, avoiding much of the small-scale modeling cost.
- Because P(k)+patches beats P(k)+B(k), the hierarchical estimator captures information beyond the three-point function without explicitly computing higher-order correlation functions.
- The largest 30–45% gains occur exactly where standard statistics are easiest to model, so the method improves the science return of Stage IV surveys at modest added computational cost.
- The control experiment indicates that even when the halo mass function is fixed, the network extracts f_NL information from the density field, so the gain is not a trivial mass-function effect.
- The method is most effective for higher-mass, highly biased tracers at k_max >= 0.2 h/Mpc, suggesting particular value for quasar-like samples used in f_NL searches.
Where Pith is reading between the lines
- The paper's non-overlapping patch choice is deliberately conservative; if overlapping patches add 15–20% more information as the authors note, the realizable gains could exceed the reported 30–45%.
- The reported gains assume no marginalization over cosmological or nuisance parameters; a testable next step is whether the relative gain persists after marginalizing, since network summaries can break degeneracies shared by n-point statistics.
- The same hierarchical recipe could be retrained for equilateral or orthogonal PNG shapes, but the locality argument motivating the combination is shape-dependent, so whether the gains transfer is an open empirical question.
- A direct comparison with other small-scale summaries, such as wavelet scattering transforms or marked power spectra, would clarify how much of the gain is specific to the learned network versus the general idea of adding local field-level information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a hybrid estimator for constraining local primordial non-Gaussianity f_NL, combining the halo power spectrum and bispectrum with a convolutional-neural-network summary (PatchNet) computed on sub-volumes of the halo density field. Using Quijote-PNG halo catalogs, the authors perform a Fisher forecast, estimating covariance from 5000 fiducial simulations with Hartlap correction and derivatives from finite differences between f_NL = ±50 simulations. They report that adding the patch summary always improves the expected f_NL constraint, with gains of 30–45% at low kmax (~0.1 h/Mpc), and that the patches capture information beyond the bispectrum. The paper studies the gains as a function of redshift, halo mass cut, and scale cut, and includes a control removing halo-mass-function information.
Significance. If the result holds, this is a useful and computationally efficient step toward extracting more f_NL information from Stage IV surveys, particularly in the regime where standard perturbative modeling is most reliable. The paper’s strengths include the use of standard Fisher machinery with Hartlap correction, the exploration of several redshifts and mass cuts, and a control experiment for the halo mass function. The main uncertainty is the derivative of a learned neural summary, on which the entire Fisher gain rests; this derivative is not validated for linearity or statistical convergence. The claim that patch-based compression 'always enhances' constraints and captures information beyond the bispectrum is therefore not yet fully supported.
major comments (4)
- [§3.1, Eq. (3)] The Fisher derivative for the PatchNet component is computed as a central finite difference with δf_NL = 50. The mean output of a CNN trained with an MSE loss is a learned regression function and is not guaranteed to be linear over f_NL ∈ [−50, 50]. If the response has curvature, the finite-difference slope differs from the local derivative at f_NL = 0, directly biasing σ_fNL and the headline 30–45% gains. Footnote 4 checks derivative stability with simulation count only for the power spectrum, not for the network. Please provide a linearity check (e.g., derivative estimates at multiple δf_NL, or a plot of mean network output versus f_NL) and/or bootstrap uncertainties on the derivative.
- [§2, §3.3] The paper does not state whether the 5000 fiducial simulations and the 500+500 f_NL = ±50 simulations used for the Fisher analysis are disjoint from the PatchNet training set. If any of these simulations were used in training, the covariance and derivative estimates would be optimistically biased. Please explicitly confirm the separation, or describe any cross-validation procedure that prevents leakage.
- [§4] The reported gains of 30–45% are presented without error bars or significance tests. The Fisher quantities are estimated from a finite number of simulations (5000 for covariance, 500 per f_NL sign for derivatives), so the point estimates have sampling noise. Reporting the uncertainty on σ_fNL and on the gain, or at least demonstrating that the gain is robust to the number of simulations used for the PatchNet derivative, is necessary to support the 'always enhances' claim.
- [§4, HMF control] The halo-mass-function removal control reports Pearson correlation coefficients (r ~ 0.78 without HMF vs r ~ 0.99 with HMF) but does not translate these into Fisher constraints. The conclusion that 'the network extracts information beyond the halo mass function' is supported only qualitatively. To connect this control to the central information-gain claim, please provide a Fisher forecast for the fixed-number-density configuration, or explicitly state that the correlation test is not a Fisher-information statement.
minor comments (4)
- [References] The reference 'Maas, A. L. 2013, in' is incomplete; the proceedings or volume information is missing.
- [Fig. 2 caption] The mass cut is written as 'M >3.2×M ⊙ h−1'; the exponent 10^13 is missing.
- [§4] The phrase 'when k_patch_max < k_P,B_max point again in the direction' is awkward and should be rephrased.
- [§3.3] The description of the patch construction states that patches are non-overlapping and then notes that overlapping patches increase information by 15–20%. It would be clearer to state explicitly that the reported Fisher results use the non-overlapping choice, which is already implied but could be made explicit for reproducibility.
Circularity Check
No significant circularity: the Fisher forecast is self-contained and the reported gains are measured, not forced by construction.
full rationale
The paper's derivation chain is a standard Fisher forecast: it constructs three summary statistics (power spectrum, bispectrum, and the PatchNet patch-average), estimates the covariance from 5000 fiducial Quijote-PNG simulations, estimates the fNL derivative from finite differences at fNL = ±50, and combines them through Eq. (1). The central quantitative claims—30–45% gains at low kmax and information beyond the bispectrum—are empirical outputs of this pipeline, not identities. PatchNet is trained on a separate 1000-simulation Latin hypercube set and evaluated on held-out ±50 simulations, so its derivative is a measured property of a learned estimator rather than a parameter fitted to the forecast data. The claim that adding patch compression 'always enhances' constraints is partly a monotonicity property of Fisher information, but the magnitude of the gains and the beyond-bispectrum comparison are nontrivial. Self-citations to Bairagi & Wandelt (2025) supply the architecture and motivation, but the fNL forecast is independently evaluated on Quijote-PNG simulations; no equation reduces to its inputs by construction. The only noteworthy limitation is that derivative stability with simulation count is checked for the power spectrum only (footnote 4), not for the network summary, and no linearity test for the finite-difference step is reported. That is a robustness/correctness concern, not evidence of circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- PatchNet architecture & training hyperparameters =
8→32→128 conv filters, 512→256 dense, LeakyReLU slope 0.5, n_b=32, n_sb=64, 512 patches/box of 16^3 voxels; optimizer/lr
- finite-difference step δf_NL =
50
axioms (5)
- domain assumption Locality principle: late-time gravitational evolution produces only local, causal effects and cannot generate the long-range correlations produced by primordial non-Gaussianity (Baumann & Green 2022); hence P+B+local-patches approximates a map-level optimal analysis.
- domain assumption The neural summary's expectation responds approximately linearly to f_NL on [-50, 50], so the finite-difference derivative at ±50 equals the local derivative at f_NL=0.
- standard math Fisher information with Gaussian likelihood, covariance evaluated at the fiducial model (f_NL=0) with Hartlap correction, adequately represents the estimator's constraining power.
- domain assumption Quijote-PNG halo catalogs with all cosmological parameters fixed except f_NL faithfully represent the f_NL-dependence of the halo density field.
- domain assumption The patch density field (halos CIC-painted on a 128^3 grid, no mass weighting) retains the small-scale f_NL information the paper claims to extract.
read the original abstract
The advent of Stage IV galaxy redshift surveys such as DESI and Euclid marks the beginning of an era of precision cosmology, with one key objective being the detection of primordial non-Gaussianities (PNG), potential signatures of inflationary physics. In particular, constraining the amplitude of local-type PNG, parameterised by $f_{\rm NL}$, with $\sigma_{f_{\rm NL}} \sim 1$, would provide a critical test of single versus multi-field inflation scenarios. While current large-scale structure and cosmic microwave background analyses have achieved $\sigma_{f_{\rm NL}} \sim 5$-$9$, further improvements demand novel data compression strategies. We propose a hybrid estimator that hierarchically combines standard $2$-point and $3$-point statistics with a field-level neural summary, motivated by recent theoretical work that shows that such a combination is nearly optimal, disentangling primordial from late-time non-Gaussianity. We employ PatchNet, a convolutional neural network that extracts small-scale information from sub-volumes (patches) of the halo number density field while large-scale information is retained via the power spectrum and bispectrum. Using Quijote-PNG simulations, we evaluate the Fisher information of this combined estimator across various redshifts, halo mass cuts, and scale cuts. Our results demonstrate that the inclusion of patch-based field-level compression always enhances constraints on $f_{\rm NL}$, reaching gains of $30$-$45\%$ at low $k_{\rm max}$ ($\sim 0.1 \, h \, \text{Mpc}^{-1}$), and capturing information beyond the bispectrum. This approach offers a computationally efficient and scalable pathway to tighten the PNG constraints from forthcoming survey data.
Figures
Reference graph
Works this paper leans on
-
[1]
N., Adshead, P., Ahmed, Z., et al
Abazajian, K. N., Adshead, P., Ahmed, Z., et al. 2016, arXiv:1610.02743
Pith/arXiv arXiv 2016
-
[2]
G., Aguilar, J., Ahlen, S., et al
Adame, A. G., Aguilar, J., Ahlen, S., et al. 2025, JCAP, 2025, 028
2025
- [3]
-
[4]
Andrews, A., Jasche, J., Lavaux, G., et al. 2024, arXiv:2412.11945
arXiv 2024
-
[5]
2023, MNRAS, 520, 5746
Andrews, A., Jasche, J., Lavaux, G., & Schmidt, F. 2023, MNRAS, 520, 5746
2023
- [6]
-
[7]
Bairagi, A., Wandelt, B., & Villaescusa-Navarro, F. 2025, arXiv:2503.13755
arXiv 2025
-
[8]
2011, in Theoretical Advanced Study Institute in El- ementary Particle Physics: Physics of the Large and the Small, 523–686
Baumann, D. 2011, in Theoretical Advanced Study Institute in El- ementary Particle Physics: Physics of the Large and the Small, 523–686
2011
-
[9]
& Green, D
Baumann, D. & Green, D. 2022, JCAP, 2022, 061
2022
-
[10]
M., Lewandowski, M., Mirbabayi, M., & Si- monović, M
Cabass, G., Ivanov, M. M., Lewandowski, M., Mirbabayi, M., & Si- monović, M. 2023, Physics of the Dark Universe, 40, 101193
2023
-
[11]
M., Philcox, O
Cabass, G., Ivanov, M. M., Philcox, O. H. E., Simonović, M., & Zal- darriaga, M. 2022, Phys. Rev. D, 106, 043506
2022
-
[12]
2017, JCAP, 2017, 003
Cabass, G., Pajer, E., & Schmidt, F. 2017, JCAP, 2017, 003
2017
-
[13]
S., Barberi-Squarotti, M., Pardede, K., Castorina, E., & D’Amico, G
Cagliari, M. S., Barberi-Squarotti, M., Pardede, K., Castorina, E., & D’Amico, G. 2025, JCAP, 2025, 043
2025
-
[14]
S., Castorina, E., Bonici, M., & Bianchi, D
Cagliari, M. S., Castorina, E., Bonici, M., & Bianchi, D. 2024, JCAP, 2024, 036
2024
-
[15]
2019, JCAP, 2019, 010
Castorina, E., Hand, N., Seljak, U., et al. 2019, JCAP, 2019, 010
2019
-
[16]
2025, JCAP, 2025, 029
Chaussidon, E., Yèche, C., de Mattia, A., et al. 2025, JCAP, 2025, 029
2025
-
[17]
Chen, X., Padmanabhan, N., & Eisenstein, D. J. 2025, JCAP, 2025, 055
2025
-
[18]
R., Abel, T., & Banerjee, A
Coulton, W. R., Abel, T., & Banerjee, A. 2024, MNRAS, 534, 1621
2024
-
[19]
& Zaldarriaga, M
Creminelli, P. & Zaldarriaga, M. 2004, JCAP, 2004, 006
2004
-
[20]
P., Werner, M., Akeson, R., et al
Crill, B. P., Werner, M., Akeson, R., et al. 2020, in Society of Photo- Optical Instrumentation Engineers (SPIE) Conference Series, Vol. 11443, Space Telescopes and Instrumentation 2020: Optical, In- frared, and Millimeter Wave, ed. M. Lystrup & M. D. Perrin, 114430I
2020
-
[21]
2008, Phys
Dalal, N., Doré, O., Huterer, D., & Shirokov, A. 2008, Phys. Rev. D, 77, 123514 D’Amico, G., Lewandowski, M., Senatore, L., & Zhang, P. 2025, Phys. Rev. D, 111, 063514
2008
-
[22]
Davis, M., Efstathiou, G., Frenk, C. S., & White, S. D. M. 1985, ApJ, 292, 371 DESI Collaboration, Aghamousa, A., Aguilar, J., et al. 2016, arXiv e-prints, arXiv:1611.00036
Pith/arXiv arXiv 1985
-
[23]
& Seljak, U
Desjacques, V. & Seljak, U. 2010, Advances in Astronomy, 2010, 908640 Euclid Collaboration: Mellier, Y., Abdurro’uf, Acevedo Barroso, J. A., et al. 2025, A&A, 697, A1 Euclid Collaboration: Scaramella, R., Amiaux, J., Mellier, Y., et al. 2022, A&A, 662, A112
2010
-
[24]
Giri, U., Münchmeyer, M., & Smith, K. M. 2023, Phys. Rev. D, 107, L061301
2023
-
[25]
2020, JCAP, 2020, 040
Hahn, C., Villaescusa-Navarro, F., Castorina, E., & Scoccimarro, R. 2020, JCAP, 2020, 040
2020
-
[26]
2007, A&A, 464, 399 Article number, page 7 A&A proofs:manuscript no
Hartlap, J., Simon, P., & Schneider, P. 2007, A&A, 464, 399 Article number, page 7 A&A proofs:manuscript no. main
2007
-
[27]
2024, ApJ, 976, 109
Jung, G., Ravenni, A., Liguori, M., et al. 2024, ApJ, 976, 109
2024
-
[28]
2025, Phys
Kvasiuk, Y., Münchmeyer, M., & Smith, K. 2025, Phys. Rev. D, 112, 023540
2025
-
[29]
1989, Generalization and network design strategies, ed
LeCun, Y. 1989, Generalization and network design strategies, ed. R. Pfeifer, Z. Schreter, F. Fogelman, & L. Steels (Elsevier) LSST Science Collaboration, Abell, P. A., Allison, J., et al. 2009, arXiv e-prints, arXiv:0912.0201
Pith/arXiv arXiv 1989
-
[30]
2003, Journal of High Energy Physics, 2003, 013
Maldacena, J. 2003, Journal of High Energy Physics, 2003, 013
2003
-
[31]
2025, JCAP, 2025, 036
Marinucci, M., Jung, G., Liguori, M., et al. 2025, JCAP, 2025, 036
2025
-
[32]
Nagarajappa, C. G. & Ma, Y.-Z. 2024, MNRAS, 529, 3289
2024
-
[33]
2024, JCAP, 2024, 021 Planck Collaboration, Akrami, Y., Arroja, F., et al
Peron, M., Jung, G., Liguori, M., & Pietroni, M. 2024, JCAP, 2024, 021 Planck Collaboration, Akrami, Y., Arroja, F., et al. 2020, A&A, 641, A9
2024
-
[34]
2015, Phys
Scoccimarro, R. 2015, Phys. Rev. D, 92, 083532
2015
-
[35]
& Zaldarriaga, M
Senatore, L. & Zaldarriaga, M. 2012, Journal of High Energy Physics, 2012, 24
2012
-
[36]
2008, JCAP, 2008, 031
Slosar, A., Hirata, C., Seljak, U., Ho, S., & Padmanabhan, N. 2008, JCAP, 2008, 031
2008
-
[37]
2008, MNRAS, 391, 1685
Springel, V., Wang, J., Vogelsberger, M., et al. 2008, MNRAS, 391, 1685
2008
-
[38]
Springel, V., White, S. D. M., Jenkins, A., et al. 2005, Nature, 435, 629 Villaescusa-Navarro,F.2018,Pylians:Pythonlibrariesfortheanalysis of numerical simulations, Astrophysics Source Code Library, record ascl:1811.008
2005
-
[39]
2020, ApJS, 250, 2 Article number, page 8
Villaescusa-Navarro, F., Hahn, C., Massara, E., et al. 2020, ApJS, 250, 2 Article number, page 8
2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.