{"id":"56f2f198-399d-4d93-a7cd-9f25dd94178a","arxiv_id":"2506.05090","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On simulated weak lensing maps, the Neural Field Scattering Transform with trained filters improves constraints on sigma_8 and w by 6-11% and posterior density by about 17% over the standard Wavelet Scattering Transform.","lead":"A study tests a new type of summary statistic for weak lensing maps, the Neural Field Scattering Transform, which learns its filters from simulated data. It reports modest improvements over the standard Wavelet Scattering Transform in constraining dark matter and dark energy parameters from 500 simulated universes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"NFST-vs-WST improvements lack significance testing: the 10 repeats are reported without error bars or paired tests, and the text is ambiguous about whether NFST filter pre-training is included in the repeats.","rationale":"The reader's weakest assumption identifies precisely the gap that most threatens the central claim: without a test of whether the 6%/11%/17% gains exceed run-to-run noise, 'consistently outperforms' is not established. My independent reading confirms this is the load-bearing issue. The paper's method is coherent and the NFST is a reasonable extension of the WST, with a previously introduced architecture and an interpretability contribution that stands on its own. However, the quantitative support for the headline is thin exactly where it matters: Figure 1 shows individual runs with considerable spread and no accompanying uncertainty on the averages, and the text does not clarify whether the 10 repeats include the stochastic pre-training of the learned filters. This ambiguity matters because a single pre-trained NFST would underestimate total pipeline variance. The proposed check is inexpensive and decisive: rerun the full pipeline with fresh seeds and compute paired significance statistics. Because the reader already assigned a CONDITIONAL verdict, my stress test does not move the verdict; it strengthens the reason for that condition.","tokens_in":21810,"tokens_out":6077,"duration_ms":85506,"concrete_test":"Rerun the full pipeline, explicitly including NFST filter pre-training, with at least 20 fresh random seeds while keeping the data split fixed. For each seed, compute paired differences Delta_logprob = mean_logprob_NFST - mean_logprob_WST and Delta_RMSE for sigma_8 and w on the fixed 2000-object test set. Report a paired bootstrap 95% confidence interval for the mean Delta_logprob, and run a Wilcoxon signed-rank test across seeds. If the lower endpoint of the CI excludes zero and the paired differences are systematically positive, the 'consistently outperforms' claim is supported; if not, the reported 16--17% improvement should be downgraded to a directional, seed-dependent result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the NFST consistently outperforms the WST rests on three numbers: a 6% RMSE improvement for sigma_8, 11% for w, and a 16--17% increase in mean log probability (abstract says 16%, Section III.A says 17%). These are averages over 10 runs, but no confidence intervals, standard deviations, or significance tests are reported. The problem is amplified by the paper's own description of the repeats. Section II.C.1 describes the NFST and CNN pre-training as a single procedure ('we pre-train... then freeze'), while Section II.C.2 states 'we repeat each test 10 times' for direct parameter prediction and Section II.C.3 says 'Again, we report the results for 10 repeats' for posterior modeling. If those repeats cover only the downstream regression and NPE training, the run-to-run spread shown in Figure 1 does not include the stochasticity of learning the NFST filters themselves, which is exactly the novel component whose benefit is being claimed. Even if the full pipeline were repeated, the absence of paired-difference statistics means the reported improvements could easily be within the skewed and bimodal run-to-run scatter visible in Figure 1. The abstract/text inconsistency between 16% and 17% adds to the impression that the headline improvement is not being reported with the care needed to support the word 'consistently.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Neural Field Scattering Transform (NFST) as a learned summary statistic for weak lensing convergence maps and benchmarks it against the fixed Wavelet Scattering Transform (WST) and a CNN baseline. The NFST replaces Morlet wavelets with neural-field-parameterized filters, is pre-trained on the training cosmologies, then frozen; the resulting 64 coefficients are used both for direct MLP regression of Omega_m, sigma_8, and w and for Masked Autoregressive Flow posterior estimation. On 2000 held-out CosmoGrid cosmologies, averaged over 10 repeats, the authors report a 6% RMSE improvement for sigma_8 and 11% for w over WST, a 17% (abstract: 16%) higher average posterior density, and worse CNN performance. The paper also introduces a coefficient-maximization visualization to interpret learned filters and generated fields.","tokens_in":22129,"tokens_out":5401,"duration_ms":61185,"significance":"If the quantitative claims hold, the NFST offers a data-efficient middle ground between fixed scattering transforms and flexible CNNs, and the visualization method has independent utility as an interpretability tool for differentiable summary statistics. Strengths include a large disjoint test set (2000 cosmologies), repeated training runs, explicit comparison to an external WST benchmark, and physically motivated inductive biases. The main weaknesses are statistical: no error bars, confidence intervals, or paired significance tests are reported for the headline improvements, and the repeat protocol does not state whether the filter-learning step is repeated. These issues are fixable but currently prevent the word \"consistently\" from being justified.","major_comments":[{"comment":"The headline numbers (6% for sigma_8, 11% for w, and 17% posterior density increase) are averages over 10 runs, but the manuscript reports no measure of run-to-run variability. Figure 1 shows individual runs with large, skewed or bimodal scatter, and the text itself concedes this spread. Without paired-difference statistics, the reported improvements may be within stochastic noise. Please report per-run values, standard deviations or bootstrap confidence intervals, and a paired test (e.g., Wilcoxon signed-rank) for NFST versus WST and versus CNN.","section":"Section III.A, Figure 1"},{"comment":"It is unclear whether the 10 repeats apply to the full pipeline, including NFST and CNN pre-training and filter learning, or only to the downstream MLP and MAF training. The pre-training step in II.C.1 is described as a single procedure, while the repeats are mentioned only in II.C.2 and II.C.3. Since the learned filters are the novel component, this ambiguity determines what the scatter in Figure 1 represents. Please clarify the repeat protocol; if only the downstream training was repeated, repeat the full pipeline or explicitly state this limitation.","section":"Sections II.C.1-II.C.3"},{"comment":"The conclusion restates the central quantitative claims with numbers (15% and 17% RMSE improvement for sigma_8 and w, and a 33% posterior improvement) that do not appear in Section III.A, which reports 6% for sigma_8, 11% for w, and a 14% CNN deficit. If the conclusion's numbers are relative to the CNN rather than the WST, the baseline must be stated explicitly; as written, the two sections appear contradictory.","section":"Section V versus Section III.A"}],"minor_comments":[{"comment":"The abstract states a 16% increase in posterior density while Section III.A and the conclusion state 17%; please unify the number.","section":"Abstract and Section III.A"},{"comment":"There is a typo: \"over the the WST\" should read \"over the WST.\"","section":"Section III.A"},{"comment":"The notation N/2^j for Fourier-space truncation should clarify whether j is zero-indexed; the three scales are described later as approximately 8, 16, and 32 arcmin, so a concrete example would remove ambiguity.","section":"Section II.B.2"},{"comment":"The flat prior is stated on normalized parameters in [-5,5]; please state the normalization used for the physical parameters so the prior is reproducible.","section":"Section II.C.3"},{"comment":"Adding error bars or shaded regions to the averages, or at least a table of per-run values, would make the central comparison easier to assess; the current figure shows only point markers.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of astro-ph.CO and the methodology is reasonable. The central issue is statistical overclaiming rather than a fundamental flaw; with added error bars, significance testing, and clarification of the repeat protocol, it would be publishable. I would not reject on novelty grounds: the prior NFST paper [31] is appropriately cited, and the cosmology application plus the visualization method are new. One editorial concern is that the abstract's \"consistently outperforms\" should be softened unless the significance analysis supports it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First application of the NFST to weak lensing cosmology inference, evaluated honestly on an out-of-sample test set. The paper does several things right: a clean three-stage pipeline (pre-train frozen filters, direct regression, NPE), 2000 test cosmologies, and 10 repeats to probe stochasticity. The visualization technique for summary statistics is a legitimate extra, and it gives some physical intuition for why the trained filters behave differently from Morlet wavelets. The reported gains over the WST—6% on sigma8, 11% on w, and about 17% in mean log posterior density—are modest and directionally plausible.\n\nThe soft spots are real but not fatal. The headline 'consistently outperforms' is not backed by any significance test. The 10 repeats are averaged without error bars, and Figure 1 reportedly shows skewed or bimodal run distributions. The paper states repeats are meant to confirm systematic improvement, but without a paired-difference test that is an assertion, not a demonstration. More importantly, the repeat protocol as written covers the downstream regression and NPE training only; the NFST filter pre-training appears to be done once and frozen. If that is correct, the run-to-run scatter in Figure 1 does not include the stochasticity of the novel component whose benefit is being claimed. That is a substantive ambiguity and the authors should clarify it.\n\nThere is also a mismatch between the abstract's 16% and the main text's 17% for the posterior density improvement. Minor, but sloppy in a paper where that single number is the centerpiece. No code or data is provided, which is a drawback for a methods paper.\n\nOverall, the direction of the result is probably right, and the visualization contribution is useful. It is an incremental advance within the summary-statistics subfield, not a paradigm shift. The paper deserves a serious referee who can ask for significance testing, clarification of the repeat protocol, and ideally a release of code/data. I would not cite it in my own work within the next year, but I'd bring it to a reading group focused on simulation-based inference.","headline":"A plausible but under-quantified first application of the NFST to weak lensing inference, with a useful visualization tool and a central improvement claim that needs significance testing.","tokens_in":22627,"tokens_out":2988,"would_cite":false,"duration_ms":35384,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scattering transform with trainable neural-field filters beats the fixed Wavelet Scattering Transform on weak lensing maps, improving σ8 and w constraints and posterior density by 16–17%.","keywords":["neural field scattering transform","wavelet scattering transform","weak lensing","cosmological parameter inference","simulation-based inference","large-scale structure","summary statistics","CosmoGrid"],"falsifier":"Re-run the same pipeline on the same CosmoGrid split with many more training repeats (for example, 50) and compute a bootstrap or paired confidence interval on the NFST-minus-WST difference in mean log probability and in RMSE; if the interval includes zero for $\\sigma_8$ or $w$, the claimed consistent outperformance would be falsified.","tokens_in":21629,"feed_emoji":"🌌","tokens_out":12995,"duration_ms":133002,"temperature":0.7,"pith_summary":"This paper tries to establish that the Neural Field Scattering Transform (NFST), which replaces the fixed Morlet wavelets of the Wavelet Scattering Transform (WST) with trainable filters parameterized by a neural field, captures more cosmologically useful information from weak lensing convergence maps while preserving the scattering architecture's rotational and translational symmetries. Applied to 500 CosmoGrid simulations and evaluated on 2000 held-out maps, the NFST is reported to consistently beat the WST: direct prediction RMSE improves by 6% for $\\sigma_8$ and 11% for $w$, with no improvement for $\\Omega_m$, and the posterior model assigns higher density to the true parameters (mean log probability 0.47 vs 0.31, a 17% increase in the text's numbers; the abstract states 16%). The CNN baseline does worse than both, supporting the paper's claim that physics-motivated inductive biases beat generic flexibility in a limited-data regime. If true, the NFST offers a way to pull more non-Gaussian information from upcoming large-scale structure surveys without requiring huge simulation training sets.","feed_headline":"Neural field filters beat fixed wavelets on lensing maps","feed_subtitle":"A trainable scattering transform lifts posterior density by 16 percent and sharpens σ8 and w constraints.","key_machinery":"The central object is the Neural Field Scattering Transform: a two-layer wavelet scattering network whose convolution filters $\\hat{\\psi}_{jl}(\\vec{k}) = F_{jl}(\\vec{k}, j)$ are generated by a small neural network mapping Fourier-space coordinates and the scale index $j$ to filter values, with Fourier-space truncation at $N/2^j$ and trivial dilation and rotation operations to build the full filter bank. The scattering coefficients are first-order $S^1_{jl} = \\langle |\\delta \\star \\psi_{jl}| \\rangle$ and second-order $S^2_{j_1 l_1 j_2 l_2} = \\langle ||\\delta \\star \\psi_{j_1 l_1}| \\star \\psi_{j_2 l_2}| \\rangle$, with angular averaging over $\\Delta\\ell$ to preserve rotational symmetry. What carries the argument is the combination: the fixed WST architecture supplies symmetry inductive biases and data efficiency, while the neural-field parameterization supplies smooth, task-adapted filters; a dual-domain (real- and Fourier-space U-Net) field-maximization procedure makes the learned second-order coefficients interpretable in physical space.","core_discovery":"On the paper's own terms, the central discovery is that inserting a smooth, trainable neural field into the scattering transform's Fourier-space filters yields a summary statistic that is both flexible and robust enough to outperform a fixed WST when trained on only 500 simulations. The learned filters spontaneously widen their frequency pass bands compared with Morlet wavelets, indicating that for cosmology the information gain comes from averaging over more angular modes rather than from sharp directional resolution. The most cosmologically informative coefficient is the second-order term coupling the two smallest scales ($j_1=1$, $j_2=2$, $\\Delta\\ell=0$), i.e., sub-structure clustering deep in the non-linear regime; the paper's new coefficient-maximization visualization shows the NFST's version of this coefficient captures more filament-like structure than the WST's. The paper also finds no improvement over the WST for $\\Omega_m$, suggesting the WST already captures the matter-density information.","pith_inferences":["Because only $\\sigma_8$ and $w$ improve and $\\Omega_m$ does not, the NFST's advantage is concentrated in parameters most sensitive to non-linear structure formation; this predicts the margin widens for surveys probing smaller scales or higher redshifts.","The widened-filter result suggests a cheap design principle: a fixed scattering transform with hand-tuned wider band-pass filters might capture part of the NFST's gain at zero training cost, though likely not all of it.","The new coefficient-visualization method could be pointed at noise and systematics: maximizing a coefficient on fields with added survey masks or shear noise would reveal whether the statistic's information content is robust or contamination-driven.","If the NFST generalizes to 3D galaxy surveys as the authors propose, its symmetry-constrained flexibility may soften the simulation-volume bottleneck, but the noiseless, mask-free setup of this test means the reported margins are an upper bound on what observational data would deliver."],"forward_implications":["The NFST extracts additional non-Gaussian information from the small, non-linear scales of weak lensing fields, tightening $\\sigma_8$ and $w$ constraints without needing more training simulations than a fixed WST.","The learned filters' wider pass bands imply that wavelet-based summaries for cosmology should prioritize broad frequency coverage and angular averaging over precise directional resolution.","The CNN's underperformance relative to both scattering transforms supports the claim that symmetry-based inductive biases are a more effective regularizer than generic dropout in data-limited cosmological inference.","The new summary-statistic visualization method can be applied to any field statistic, providing a physical-space check of what the statistic measures and helping validate deep-learning-based summaries.","The two-stage freeze-then-train procedure (pre-train, freeze, then fit a lightweight regressor and posterior model) is presented as a compute-efficient alternative to end-to-end training that retains or improves performance."],"supporting_citations":[{"why":"Introduces the Neural Field Scattering Transform and supplies the implementation details used for the filters and scattering architecture.","marker":"[31]"},{"why":"Establishes the Wavelet Scattering Transform as a cosmology summary statistic and serves as the fixed-filter benchmark.","marker":"[27]"},{"why":"Provides the neural-compression CNN application to weak lensing that motivates the CNN baseline and the limited-data context.","marker":"[23]"},{"why":"Shows that learnable scattering parameters improve over fixed wavelets, motivating the NFST's trainable filters.","marker":"[29]"},{"why":"Supplies the neural posterior estimation framework used to turn predicted parameters into posteriors.","marker":"[36]"},{"why":"Provides the software package used to train the autoregressive density estimator for posterior modeling.","marker":"[37]"},{"why":"Defines the autoregressive density model used for posterior estimation.","marker":"[38]"},{"why":"Provides the summary-statistic field visualization framework that the new coefficient-maximization method extends.","marker":"[40]"},{"why":"Supplies the feature-attribution method used to identify the most cosmologically informative coefficients.","marker":"[44]"}],"fun_headline_variants":["Trainable neural fields sharpen lensing cosmology","Scattering transform with learned filters lifts lensing constraints","Neural field scattering beats fixed wavelets on lensing maps","Learned filters improve weak lensing parameter precision","NFST boosts lensing posterior density by 16 percent"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported gains (6% for $\\sigma_8$, 11% for $w$, and the 16–17% posterior-density improvement) are larger than the run-to-run training scatter, since the paper averages ten repeats without reporting error bars or significance tests on these percentages.","fun_headline_variants_meta":{"raw":{"variants":["Trainable neural fields sharpen lensing cosmology","Scattering transform with learned filters lifts lensing constraints","Neural field scattering beats fixed wavelets on lensing maps","Learned filters improve weak lensing parameter precision","NFST boosts lensing posterior density by 16 percent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000116,"raw_usage":{"total_tokens":1091,"prompt_tokens":978,"completion_tokens":113,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":37}},"tokens_in":594,"tokens_out":113,"duration_ms":2074,"temperature":1.0,"reasoning_tokens":37,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:24:46.188623+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same pipeline on the same CosmoGrid split with many more training repeats (for example, 50) and compute a bootstrap or paired confidence interval on the NFST-minus-WST difference in mean log probability and in RMSE; if the interval includes zero for $\\sigma_8$ or $w$, the claimed consistent outperformance would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the Neural Field Scattering Transform and supplies the implementation details used for the filters and scattering architecture."},{"cited_title":"Parametric Scattering Networks","cited_arxiv_id":"2107.09539","evidence_quote":"Provides the summary-statistic field visualization framework that the new coefficient-maximization method extends."}],"review_version":1}