{"id":"ce8d1f9d-ab62-42fa-8ce0-e0af0037db53","arxiv_id":"2411.08957","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Jointly training a graph neural network with a normalizing flow yields low-dimensional summary statistics from simulated galaxy catalogs that support likelihood-free inference of Omega_m, and can be interpreted via correlations with power spectrum and baryonic suppression.","lead":"Researchers train a graph neural network together with a neural density estimator to learn compact summary statistics from simulated galaxy catalogs, and use them to estimate the cosmological parameters Omega_m and sigma_8 without a likelihood. The learned summaries can be mapped onto known power spectrum statistics and used to compare different galaxy formation simulation models, offering a route toward near-optimal use of future galaxy survey data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Optimality claim is conditional on the chosen galaxy-graph representation; if that graph discards information relevant to Ωm or σ8, the learned summary is not optimal for the galaxy catalog but only for the reduced graph.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing concern: the optimality of the learned summary is conditional on the graph input representation. I agree with that assessment. The paper's theoretical argument (Section 3.3, based on Ref. [107]) shows that joint training of compression and inference maximizes mutual information between the summary and the cosmological parameters under full convergence and sufficient flexibility, but only over functions of the chosen input. The graph, as defined in Section 3.2, is a lossy representation of the galaxy catalog, and the authors themselves acknowledge this and list richer node features as future work. Therefore the headline 'optimal' is overstated relative to the actual guarantee. This is a correctness risk but not a fatal one: the method is sound, the inference results are competitive, and the caveat is identifiable in the text. A concrete test with an enriched graph would settle whether the omitted information is cosmologically relevant. The reader's CONDITIONAL verdict is appropriate; I do not see a reason to change it, hence UNCHANGED.","tokens_in":39997,"tokens_out":7011,"duration_ms":72063,"concrete_test":"Retrain the best-performing model on the same IllustrisTNG training set with an enriched graph that adds (i) stellar mass as a node feature, (ii) all three velocity components instead of only v_z, and (iii) a local density/count-in-cell feature; use the same hyperparameter budget, early stopping, and validation split. Compare the marginal posterior widths and χ²_red for Ωm and σ8 on the held-out validation boxes. If the enriched-graph model yields tighter posteriors or lower χ²_red beyond retraining scatter, the original graph is not information-sufficient, and the 'optimal summary' claim fails for the galaxy catalog.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 3.3 that a fully converged compression/inference joint network maximizes the mutual information between the summary statistic and the cosmological parameters is correct only for the fixed input representation. The graph in Section 3.2 encodes galaxy positions via edges (with linking length r), the z-component of peculiar velocity as the sole node feature, and the logarithm of galaxy count as the global feature; stellar mass, the full velocity vector, and environmental measures are omitted. The paper itself acknowledges the two lossy pre-network steps (catalog construction and graph assembly) and proposes 'to include astrophysical observables in the graph' as future work (Section 5), confirming the representation is a deliberate choice. If any omitted observable carries information about Ωm or σ8 — for example, mass-weighted clustering, anisotropic redshift-space distortions, or local density — then the learned summary cannot be information-optimal among summaries of the full galaxy catalog. The abstract and title's 'optimal' is therefore an overstatement; the actual guarantee is optimality among summaries of the chosen graph. This is not an internal inconsistency, but it is a correctness risk for the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a simulation-based inference pipeline in which a graph neural network (CosmoGraphNet) compresses galaxy catalogs from the CAMELS simulations into a low-dimensional summary vector, and a masked autoregressive flow estimates the posterior of the cosmological parameters Omega_m and sigma_8. The compression and inference networks are trained jointly using the expected KL divergence, so that, in the limit of perfect convergence, the summary should maximize the mutual information between the compressed representation and the parameters. The authors validate their posteriors with simulation-based calibration, compare performance with earlier moment-network work, interpret the summaries through principal-component correlations with cosmological and astrophysical parameters and with power spectra, and use Isomap and t-SNE embeddings to compare the three CAMELS feedback models. They also train an emulator to construct summaries marginalized over baryonic parameters.","tokens_in":40165,"tokens_out":4992,"duration_ms":49745,"significance":"If the optimality claim were established for full galaxy catalogs, the approach would be a valuable addition to simulation-based cosmological inference, particularly for non-Gaussian, non-linear scales where hand-crafted summaries are known to be insufficient. The paper has notable strengths: it uses publicly available simulations, holds out a validation set, applies SBC, performs systematic hyperparameter optimization, checks robustness across 30 trained models, and includes a useful n-body control (Figure 10e) for astrophysical correlations. However, the headline 'optimal' result is conditional on the specific graph representation, and the demonstrated inference is effectively one-dimensional: sigma_8 is not constrained beyond the prior. The paper is therefore best read as a proof-of-concept for joint GNN-compression and flow-based posterior estimation, with the optimality claim needing substantial qualification.","major_comments":[{"comment":"The optimality claim in Section 3.3 ('A fully converged compression/inference joint network therefore maximizes the mutual information between the summary statistic and the cosmological parameters') is valid only for the fixed input representation defined in Section 3.2. The graph encodes galaxy positions through edges (Eqs. 3.9-3.11), the z-component of peculiar velocity as the sole node feature, and log10 of the galaxy count as the global feature. The paper itself acknowledges in Section 3.2 that the construction of the galaxy catalog and the assembly of the graph are two lossy steps preceding the compression network. If information relevant to Omega_m or sigma_8 is contained in observables omitted from this graph (e.g., stellar mass, full velocity vector, or environmental measures), the learned summary cannot be information-optimal for the galaxy catalog, only for the chosen graph. Because the abstract and title claim 'optimal summary statistics' without this qualification, the central claim overstates the result. I recommend either revising the wording to 'optimal within the specified catalog and graph representation' or demonstrating empirically that adding such observables does not improve inference.","section":"Sections 3.2 and 3.3"},{"comment":"The pipeline does not constrain sigma_8: the marginal posteriors in Figure 4 span essentially the full prior range, and the text states that sigma_8 'shows a large degree of bias learning only the mean of the dataset.' The reported chi^2_red values close to 1 for sigma_8 in Table 3 are exactly what a posterior equal to the prior would produce, so they do not indicate successful inference. Since sigma_8 is one of the two parameters of interest, the paper does not demonstrate that the learned summaries are informative for half of the stated inference task. The authors attribute this to the small simulation box, which is plausible, but the consequence for the central claim should be made explicit: the learned summaries are not shown to be optimal summaries for sigma_8 on the simulated catalogs. Please report posterior contraction relative to the prior (e.g., the ratio of posterior to prior standard deviation) for both parameters, and either remove sigma_8 from the headline claims or restrict the optimality discussion to Omega_m.","section":"Section 4.1, Figures 2-4, and Table 3"},{"comment":"The simulation-based calibration test is used to support posterior validity, but the paper itself notes in Section 3.4 that a posterior estimate equal to the prior will pass the rank test. This is precisely the situation for sigma_8: if the learned posterior for sigma_8 is essentially the prior, the near-linear empirical CDF in Figure 6 for sigma_8 carries no evidence of informative calibration. The SBC result for sigma_8 should therefore be presented as a necessary but trivial consistency check, not as validation of a useful posterior. Additional diagnostics sensitive to informativeness, such as posterior contraction or coverage conditional on summary values, are needed before the sigma_8 posterior can be described as validated.","section":"Sections 3.4 and 4.1, Figure 6"},{"comment":"The correlations between the learned summary principal components and the stellar feedback parameters ASN1 and ASN2 are described as being 'implicitly learned' even though these parameters never enter the loss function. Because ASN1 and ASN2 are varied in the training simulations and directly affect the galaxy catalogs, it is expected that summaries sensitive to the catalog will correlate with them; this does not by itself show that the network has learned to encode these parameters. The n-body control in Figure 10e is a good check that the Latin-hypercube layout was not memorized, but it does not distinguish between (a) the summary containing physical information about feedback processes and (b) correlations arising through degeneracies with Omega_m or sigma_8 in the finite training set. Please rephrase the interpretation, or provide a direct test such as training a regressor on the summary to predict ASN1/ASN2 and comparing with the null distribution from the n-body runs.","section":"Section 4.2, Figures 7 and 10"}],"minor_comments":[{"comment":"In the bullet list, the sentence 'It is does not rely on either numerical derivatives...' contains a grammatical error and should read 'It does not rely on either numerical derivatives...'.","section":"Section 3.3"},{"comment":"The phrase 'baryon-robust summery' should be 'baryon-robust summary'.","section":"Section 4.4"},{"comment":"The phrase 'The sigma8 contour is not overcoming the prior' is awkward; consider rephrasing to 'The sigma8 posterior is not significantly narrower than the prior'.","section":"Figure 2 caption"},{"comment":"The definition of chi^2_red uses the index i both for the simulation index and implicitly for the parameter; please clarify that the sum runs over validation simulations and that theta_i, mu_i, and sigma_i refer to the scalar parameter value, posterior mean, and posterior standard deviation for that simulation.","section":"Equation (4.1)"},{"comment":"Reference [90] appears to duplicate Reference [78]; please consolidate or distinguish them.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a competent proof-of-concept for joint GNN compression and flow-based posterior estimation, and the authors are honest about the sigma_8 limitation and the lossy graph construction. The main issue is that the title and abstract claim 'optimal' summary statistics while the optimality is conditional on a specific, lossy graph representation, and the inference demonstration is effectively one-dimensional. This can be fixed with careful rewording and additional diagnostics, so I recommend major revision rather than rejection. The interpretability correlations with ASN1/ASN2 should also be framed more cautiously; they are expected consequences of the training distribution rather than independent discoveries."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent SBI paper with a genuinely useful new combination and an honest discussion of its failures. The title's 'optimal' is doing more work than the method can support, but that is a claim-refinement problem rather than a fatal flaw.\n\nWhat is new: they jointly train a GNN (CosmoGraphNet) as the compression network with a masked autoregressive flow for posterior estimation on galaxy catalogs. Previous joint-training work used CNNs, and GNN-based work used moment networks that only give means and variances. Here you get full posteriors from graph-structured input. They validate with simulation-based calibration, which is the right thing to do, and they report that Omega_m recovery works while sigma8 essentially does not (it collapses to the prior mean). They are explicit about that limitation and attribute it to the small box, which is plausible.\n\nThe interpretability section is the real contribution. PCA on the learned summaries shows that O(5) components capture most variance, correlations with the power spectrum localize relevant scales, and comparisons across the three CAMELS feedback models (IllustrisTNG, SIMBA, Astrid) give an interesting view of how the summary space varies. The emulator-based baryon marginalization is a nice trick, though the emulator architecture gets only a paragraph.\n\nSoft spots. The central 'optimal summary statistic' claim (Section 3.3) is stated as a consequence of the variational bound, but it is asymptotic and conditional on the input representation. The graph encodes positions via edges, a single velocity component as node feature, and galaxy count globally. Stellar mass, full velocity, and environment are left out. The authors themselves note the two lossy pre-network steps and list 'include astrophysical observables in the graph' as future work. So the title's 'optimal' should really be 'optimal among summaries of this particular graph.' That is not an internal contradiction, but it is a material qualification. The sigma8 result reinforces this: the network is not actually finding all information in the catalog, only what that cheap representation exposes.\n\nReproducibility is a real weakness: no code, no data, no trained models, and the emulator details are thin. For a methods paper in this space, that is a problem. The comparison to Ref. [117] is also not clean because the galaxy selection differs; they note this.\n\nBottom line: this is a solid methods paper that will be useful to people working on SBI for large-scale structure. Send it to review with a request to qualify the optimality claim and to release code and models. I would cite it.","headline":"A solid, honest SBI paper that overreaches on 'optimal' summary statistics but deserves a serious referee.","tokens_in":40739,"tokens_out":2245,"would_cite":true,"duration_ms":22762,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that jointly training a graph neural network and a neural density estimator produces summary statistics from galaxy catalogs that maximize the mutual information with the cosmological parameters $\\Omega_m$ and…","keywords":["simulation-based inference","summary statistics","graph neural networks","cosmological parameter inference","galaxy catalogs","baryonic feedback","neural posterior estimation","mutual information"],"falsifier":"Train the same joint network on galaxy catalogs that add the full three-dimensional velocity and stellar mass as node features while keeping everything else fixed, and compare the width of the $\\Omega_m$ and $\\sigma_8$ posteriors on the same validation boxes; noticeably tighter posteriors would show the original input representation threw away usable information, contradicting the optimality claim as stated.","tokens_in":39761,"feed_emoji":"🌌","tokens_out":8968,"duration_ms":76066,"temperature":0.7,"pith_summary":"This paper argues that a galaxy catalog can be compressed into a short vector of learned summary statistics that is information-optimal for inferring the cosmological parameters $\\Omega_m$ and $\\sigma_8$, without writing down a likelihood. The authors train a graph neural network and a neural density estimator jointly, so the loss is the expected Kullback-Leibler divergence; they claim a fully converged joint network maximizes the mutual information between summary and parameters. On hydrodynamical simulations with three different baryonic feedback implementations, the resulting summaries recover $\\Omega_m$ well, fail to constrain $\\sigma_8$ strongly (attributed to the small simulation box), remain interpretable via a handful of principal components, and let the authors identify the scales that matter. The point of the paper is that hand-crafted statistics can be replaced by learned, low-dimensional summaries that are also a diagnostic tool for baryonic physics.","feed_headline":"Learned galaxy summaries are information-optimal for cosmology","feed_subtitle":"Graph compression and density estimation maximize mutual information on Omega_m and sigma_8, and reveal the scales used.","key_machinery":"The machine is a joint network built from two parts: a graph neural network compression network $h_\\lambda$ that maps a galaxy graph to a vector $\\mathbf{t}$, and a masked autoregressive flow $q_\\phi(\\theta|\\mathbf{t})$ that estimates the posterior. The galaxy graph encodes galaxy positions in the edges (via separation and two rotation-invariant angle cosines), the $z$-component of peculiar velocity as the sole node feature, and $\\log_{10}$ of the galaxy number as the global feature. The flow's first affine block takes the compressed vector as a conditioning input, and backpropagation through the expected negative log posterior trains both networks together. That end-to-end KL loss is what carries the information-optimality argument: no separate compression loss, no covariance estimate, and no fiducial point in parameter space.","core_discovery":"The central claim is that automatic compression and inference can be solved as one optimization: instead of choosing a loss for the compression network, the graph network's output is fed directly into a masked autoregressive flow that estimates the posterior $p(\\theta|\\mathrm{summary})$, and the expected KL divergence is minimized end-to-end. A fully converged joint network therefore maximizes the mutual information between the summary statistic and the cosmological parameters over the prior volume, making the learned summaries optimal within the chosen data representation. The authors demonstrate this on galaxy catalogs from three hydrodynamical simulation suites: $\\Omega_m$ is inferred reliably, $\\sigma_8$ is not (they attribute this to the small $25\\,h^{-1}\\,\\mathrm{Mpc}$ box lacking large-scale modes), and the summaries encode baryonic feedback parameters even though those never enter the loss. They also show the summaries are low-dimensional in effect, with about five principal components explaining most variance, and that the physical scales the network uses include modes around $k = 5\\!-\\!30\\,h/\\mathrm{Mpc}$.","pith_inferences":["If the optimality claim holds, any hand-crafted summary (power spectrum, bispectrum, counts-in-cells) can be at most as informative for $\\Omega_m$ and $\\sigma_8$ within this graph representation; a direct comparison of posterior widths would make that concrete.","The same joint-training scheme should transfer to survey-like catalogs with selection effects, redshift-space distortions, and masks, provided the simulator reproduces them; the summaries would then be optimal for the survey rather than for an idealized box.","The sharply cut-off forbidden region in summary space, which the authors leave unexplained, is likely a prior boundary or a physical limit of the feedback model space; identifying it could sharpen the interpretation of the summaries.","The baryon-marginalized emulator could be inverted to calibrate subgrid feedback parameters against observations, in the spirit of simulation calibration but now in an information-optimal summary space."],"forward_implications":["The learned summaries can be used for likelihood-free cosmological parameter inference from galaxy catalogs, with performance at least competitive with existing machine-learning inference that predicts only means and standard deviations.","Because the summaries are low-dimensional and reproducible across training runs, they provide a candidate replacement for hand-crafted statistics in analyses where the likelihood is unknown.","The correlations between summary principal components and simulation parameters give a quantitative handle on which baryonic feedback processes matter for cosmological inference, and which do not.","Mapping simulations in summary space lets one compare feedback models directly; the observation that one suite occupies a larger volume explains why models trained on it generalize to the others.","Emulating the summary as a function of cosmology alone produces a baryon-marginalized summary that lies on a hypersurface in summary space, offering a route to baryon-robust inference."],"supporting_citations":[{"why":"Supplies the joint training scheme and the argument that a fully converged compression/inference network maximizes mutual information between summary and parameters.","marker":"[107]"},{"why":"Provides the galaxy graph input prescription (positions, one velocity component, mass threshold) and the baseline field-level inference results the paper compares its performance against.","marker":"[117]"},{"why":"Introduces the graph neural network architecture used as the compression network.","marker":"[114]"},{"why":"Provides masked autoregressive flows, the neural density estimator used for posterior estimation.","marker":"[39]"},{"why":"The simulation suite providing the galaxy catalogs and parameter ranges used for training and validation.","marker":"[19]"},{"why":"Demonstrates neural compression for likelihood-free inference with variational mutual information, the approach this work extends to graph networks.","marker":"[106]"},{"why":"Provides the simulation-based calibration test used to check the learned posteriors.","marker":"[160]"},{"why":"Supplies the result that mean squared error training yields the posterior mean, which grounds the baryon-marginalized emulator argument.","marker":"[101]"}],"fun_headline_variants":["SBI-trained galaxy summaries are info-optimal","No likelihood needed: optimal galaxy summaries via SBI","Joint compression and inference yields optimal summaries","Galaxy summaries learned to hit information bound"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The learned summaries are optimal only for the information already placed in the galaxy graph, which contains positions, the $z$-component of peculiar velocity, and galaxy count; if quantities omitted from this representation (such as the full velocity vector, stellar mass, or environment) carry information about $\\Omega_m$ or $\\sigma_8$, the summaries cannot be truly optimal.","fun_headline_variants_meta":{"raw":{"variants":["SBI-trained galaxy summaries are info-optimal","No likelihood needed: optimal galaxy summaries via SBI","Joint compression and inference yields optimal summaries","Galaxy summaries learned to hit information bound"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000413,"raw_usage":{"total_tokens":2142,"prompt_tokens":961,"completion_tokens":1181,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":1124}},"tokens_in":577,"tokens_out":1181,"duration_ms":11102,"temperature":1.0,"reasoning_tokens":1124,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:12:59.399143+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same joint network on galaxy catalogs that add the full three-dimensional velocity and stellar mass as node features while keeping everything else fixed, and compare the width of the $\\Omega_m$ and $\\sigma_8$ posteriors on the same validation boxes; noticeably tighter posteriors would show the original input representation threw away usable information, contradicting the optimality claim as stated.","supporting_citations":[],"review_version":1}