{"id":"eb1c4865-5ef2-4120-b6c3-deef45c530da","arxiv_id":"2507.21682","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A 3D CNN encoder plus a masked autoregressive flow recovers Ωm and σ8 from simulated HI tomographic data cubes at z=1 with R² ≥ 0.91 on test sets, with reduced but still useful accuracy out of distribution.","lead":"This paper trains a deep 3D convolutional encoder to compress simulated 21-cm hydrogen maps, then uses a neural density estimator to infer the cosmic matter density and density fluctuation amplitude, reporting R² ≥ 0.91 on held-out data. It also tests how well this approach works when the encoder sees data from a different simulation or a different redshift.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration claim rests on UCE, which is not a coverage test; 'most truths within 1σ' is unquantified, so posterior calibration is not yet established.","rationale":"The central claim is that the pipeline yields calibrated posterior uncertainties. The only quantitative evidence for calibration is UCE, which is a second-moment comparison, not a coverage check. Without an empirical coverage fraction, the phrase 'most of the ground truth fall within 1σ' is not backed by a statistic, and a visual impression of 150 points is insufficient. This is load-bearing because if coverage is poor, the posterior uncertainties cannot be used as credible intervals even though point predictions (R²) are high. The reader's weakest assumption (latent-code sufficiency) is a real and related issue, but it operates upstream: even if the codes were sufficient, the calibration evidence as reported would still not establish the claim. A coverage test is cheap and decisive. The same-distribution R² values (0.919–0.928) appear internally consistent with the held-out evaluation, and the paper's use of public CAMELS data and the SBI package are independent supports, so the appropriate action is to require the coverage check rather than reject the work. Thus the verdict remains CONDITIONAL, unchanged from the reader's assessment.","tokens_in":19254,"tokens_out":9052,"duration_ms":115247,"concrete_test":"Using the posterior samples (10,000 per test instance) or by re-running the pipeline, compute for each parameter the fraction of the 150 held-out truths that lie inside the 68% and 95% posterior credible intervals, and the analogous joint coverage using 2D highest-posterior-density regions. Compare with nominal 0.68 and 0.95, accounting for binomial sampling error (about ±0.04 and ±0.02 at 1σ for n=150). If marginal coverage is below ~0.60 at 1σ, or joint coverage is far below nominal while UCE remains low, the 'reasonably calibrated' claim is not supported. Report the coverage fractions explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's 'reasonably calibrated' conclusion (§4.3) is based on UCE (Eq. 6), which compares binned mean squared error ξ² to mean epistemic variance ζ². This is not a coverage metric: a predictor with a constant bias and a variance equal to the mean squared error can have UCE ≈ 0 while the 68% intervals contain almost none of the true values. The central claim 'most of the ground truth ... fall within 1σ' is supported only by visual inspection of Figures 3 and 4; no empirical coverage fraction or credible-interval containment statistic is reported, and no joint 2D coverage is given. The reported TNG Ωm UCE of 12.7% already signals imperfect calibration, but UCE alone cannot establish or refute the coverage claim. Since the paper's headline contribution is calibrated posterior inference, the missing coverage validation is the load-bearing weakness.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage pipeline for cosmological inference from 21-cm HI tomographic data cubes. In the first stage, a 3D convolutional encoder (Table 1) is trained with a supervised regression loss (Eq. 1) to predict Ωm and σ8 from 128^3 CAMELS HI boxes at z=1, yielding 1024-dimensional latent codes. In the second stage, a Masked Autoregressive Flow within the Automatic Posterior Transformation framework is trained on these latent codes to approximate the posterior p(Ωm, σ8 | data). The authors report R² ≥ 0.91 for same-distribution APT results (SIMBA→SIMBA, TNG→TNG), compare against GP/RF/linear baselines, and test robustness when the encoder is applied to out-of-distribution data (cross-simulation, and z=0/0.5 from the same simulation). The paper claims that the posterior uncertainties are reasonably calibrated, based on the Uncertainty Calibration Error (UCE) of Eq. (6), and that most test truths fall within the 1σ credible intervals.","tokens_in":19430,"tokens_out":4385,"duration_ms":54693,"significance":"If the claims hold, this is a useful proof-of-concept for field-level, likelihood-free inference from 3D 21-cm tomographic data, with potential applications to upcoming HI surveys. The use of two CAMELS simulations, held-out test sets with clean encoder/APT separation, a UMAP-based check of latent separability, and comparison against several regression baselines are concrete strengths. The main gaps are statistical: the calibration claim is based on UCE, which is not a coverage test, and the abstract's out-of-distribution performance claim (R² ≥ 0.80 in general) is contradicted by the paper's own z=0 SIMBA σ8 result. These issues are fixable within the manuscript's scope.","major_comments":[{"comment":"The central calibration claim is not established. UCE, as defined in Eq. (6), compares binned mean squared error ξ² to binned mean epistemic variance ζ². This is a variance-calibration diagnostic, not a coverage metric: a predictor with a constant bias and variance equal to its mean squared error can have UCE near zero while its 68% intervals contain almost none of the true values. The abstract and §4.2 state that 'most of the ground truth ... fall within 1σ uncertainty', but this is supported only by visual inspection of Figures 3 and 4; no empirical coverage fraction, prediction-interval coverage probability, or joint 2D coverage statistic is reported. Because calibrated posterior inference is a headline contribution, please add quantitative coverage tests, such as the fraction of test truths within the 68% and 95% marginal credible intervals and a 2D joint coverage metric, and report them for all four main setups.","section":"§4.2, §4.3, Eq. (6)"},{"comment":"The abstract's claim that out-of-distribution inference achieves 'R² ≥ 0.80 in general' is contradicted by the paper's own results: §4.6 states that 'in the worst case scenario at z=0, APT still boasts an R²~0.6 on σ8' for SIMBA, and the corresponding point is visible in Figure 9 (top-left panel). The claim should be qualified, for example by restricting it to z ≥ 0.5 or by explicitly noting the z=0 SIMBA σ8 exception. As written, the abstract overstates the demonstrated robustness of the method.","section":"Abstract, §4.6, Figure 9"},{"comment":"The sufficiency of the latent codes for posterior estimation is assumed, not tested. The encoder is trained with a regression loss (Eq. 1) that explicitly targets Ωm and σ8, so the latent codes are aligned with point prediction of these parameters. This is not full circularity because the APT is evaluated on held-out boxes the encoder never saw, but it does mean that the downstream posterior is conditioned on features already selected by the regression objective. The paper does not test whether these codes retain information needed for the tails of the posterior or for the reported calibration, for example by comparing against an unsupervised compression or by using an information-preserving summary. Please either add such a comparison or explicitly discuss this limitation when claiming that the latent codes are 'meaningful representations' in a general sense.","section":"§3.1, Eq. (1), §3.2"},{"comment":"The redshift robustness experiments are missing key experimental details. The text does not state how many z=0 and z=0.5 boxes are used, whether these are the same simulation boxes as the z=1 set, or how the train/validation/test split is performed for the APT, GP, RF, and LIN models at these redshifts. Since the z-dependent R² values are a central part of the robustness claim, these details are needed for reproducibility and for interpreting the comparison across redshifts.","section":"§4.6, Figure 9"}],"minor_comments":[{"comment":"The phrase 'most of the ground truth ... fall within 1σ' is unquantified; please report the actual fraction of test points inside the 68% and 95% intervals rather than relying on visual inspection.","section":"Abstract and §4.2"},{"comment":"A modified UCE is defined in Eq. (9), but only the UCE of Eq. (6) is reported in Figure 6. Please either report the modified values or clarify why they are omitted.","section":"§4.3, Eq. (9)"},{"comment":"R² values are point estimates computed on 150 test examples; no uncertainty intervals are provided. Given that several method comparisons differ by only ~0.01 (Table 2), bootstrap or other interval estimates would help assess whether the differences are meaningful.","section":"Tables and Figures"},{"comment":"There are typos in the conclusions, including 'theATP model' and 'esimator'; please proofread the final section and the abstract.","section":"§5"},{"comment":"The paper does not include a data or code availability statement. Since the CAMELS data are public and the SBI package is named, a brief statement (or link) would improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The method is promising and the same-distribution results appear to support the main technical pipeline. However, the calibration claim is not demonstrated by UCE, and the abstract overstates the out-of-distribution robustness. I see these as fixable with additional coverage analyses and a more careful abstract, rather than as fundamental flaws, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my take. The paper is a legitimate proof-of-concept: it's the first to apply a deep 3D CNN encoder plus APT/MAF density estimation to CAMELS HI tomographic cubes at z=1, and it includes cross-simulation and cross-redshift transfer tests. The same-distribution results (R²≥0.91) are credible: the encoder and APT are tested on held-out boxes, and the comparison against GP, RF, and a linear probe is useful. The paper is also honest about the supervised pretraining of the encoder — no attempt to hide that the latent codes are trained to predict the target parameters.\n\nThe soft spots are real but not fatal. The calibration claim is the weakest part. UCE (Eq. 6) is not a coverage metric; it compares binned squared error to binned epistemic variance. A predictor with constant bias and a variance equal to the mean squared error can have UCE ≈ 0 while its 68% intervals cover almost nothing. The paper's support for 'most truths within 1σ' is visual inspection of Figures 3 and 4, with no numerical containment fraction and no joint 2D coverage. Given TNG Ωm already has UCE = 12.7%, 'reasonably calibrated' overstates the evidence. This is the load-bearing issue for the paper's headline contribution.\n\nSecond, the abstract's 'R²≥0.80 in general' for out-of-distribution data is contradicted by the z=0 SIMBA σ8 case (R²≈0.6), which the paper itself reports in §4.6. The abstract should be revised to distinguish the different OOD regimes. Third, no error bars are attached to any R² values, despite test sets of only 150 boxes; bootstrap intervals would be straightforward. Fourth, no code or data are released, which limits reproducibility for a methods paper. None of these are fatal for a proof-of-concept, but they need to be addressed.\n\nThe paper deserves peer review. It is a serious application of existing methods to a new data modality, with honest reporting of results and clear limitations. I would request a coverage-based calibration analysis, numerical 1σ/2σ containment fractions, error bars on R², and a less sweeping abstract before acceptance.\n\nWho is this for? People working on SBI for 21-cm intensity mapping, and readers preparing analysis pipelines for HIRAX/CHIME. I'd bring it to a reading group, but I wouldn't cite it in my own work in the next year unless the calibration issue gets fixed.","headline":"A credible proof-of-concept for SBI on 3D HI tomographic data, but the calibration claim rests on a metric that is not a coverage test and the abstract overstates OOD robustness.","tokens_in":20035,"tokens_out":2518,"would_cite":false,"duration_ms":27577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage deep network — 3D convolutional encoder plus Masked Autoregressive Flow — recovers ($\\Omega_{\\rm m}$, $\\sigma_8$) from $128^3$ 21-cm boxes with $R^2 \\geq 0.91$ and calibrated uncertainties.","keywords":["21-cm intensity mapping","HI tomography","likelihood-free inference","normalizing flows","posterior estimation","cosmological parameters","neural compression","uncertainty calibration"],"falsifier":"On the held-out boxes, train the same MAF posterior directly on the raw $128^3$ cubes or on a summary known to be sufficient (e.g., the full power spectrum plus bispectrum) and compare its credible intervals with those obtained from the latent codes, or run an expected-coverage test across the full $(\\Omega_{\\rm m}, \\sigma_8)$ prior: if coverage deviates strongly from the nominal 68% in any region, or if the latent-code intervals are systematically wider than the field-level ones, the encoder's information sufficiency is falsified.","tokens_in":18981,"feed_emoji":"🌌","tokens_out":13543,"duration_ms":133605,"temperature":0.7,"pith_summary":"This paper shows that cosmological parameters can be extracted from three-dimensional 21-cm neutral-hydrogen maps at intermediate redshift without an analytic likelihood. A deep 3D convolutional encoder is first trained to compress each $128^3$ tomographic box into a 1024-dimensional latent vector by regressing the matter density $\\Omega_{\\rm m}$ and the fluctuation amplitude $\\sigma_8$. A Masked Autoregressive Flow, trained in an Automatic Posterior Transformation scheme on those latent vectors, then estimates the full posterior over ($\\Omega_{\\rm m}$, $\\sigma_8$) for a single box. On held-out test sets the posterior medians achieve $R^2 \\geq 0.91$ on both parameters, most ground-truth values lie inside the $1\\sigma$ credible interval, and the epistemic uncertainty is calibrated at the percent level. The same featurizer also performs reasonably on out-of-distribution data — a different simulation or a redshift offset up to $\\Delta z \\approx 0.5$ — which suggests the learned codes capture transferable information about matter clustering.","feed_headline":"Two-stage AI reads cosmology from 21-cm maps","feed_subtitle":"A 3D neural encoder plus density flow maps hydrogen cubes to cosmology with calibrated uncertainties.","key_machinery":"The load-bearing object is the encoder–posterior pair. The encoder is a 3D convolutional network with four stages of 3D convolution, batch normalization, ReLU, and max pooling, ending in adaptive average pooling, which maps each $128^3$ neutral-hydrogen box to a 1024-dimensional latent vector; it is pre-trained jointly with a two-layer regressor under the loss of Eq. (1), which balances the log-mean-squared prediction error against the log-squared error of the predicted per-parameter variance. The posterior estimator is a Masked Autoregressive Flow, a normalizing flow that factors a density into Gaussian conditionals and is invertible with an analytic Jacobian, used inside the Automatic Posterior Transformation framework, which trains the flow to map parameters and data (here, latent codes) to an approximation of $p(\\Omega_{\\rm m}, \\sigma_8 \\mid \\text{latent code})$. At test time, 10,000 samples are drawn from the fitted posterior; the median is the point prediction and the 68% and 95% percentiles give the epistemic credible intervals. The separability of the codes under UMAP is presented as evidence that the regression pre-training produces meaningful structure.","core_discovery":"In the paper's own terms, the claim is that the full 3D neutral-hydrogen field at $z \\approx 1$ carries enough information to infer $\\Omega_{\\rm m}$ and $\\sigma_8$ with a coefficient of determination $R^2 \\geq 0.91$ for both parameters, using only the 1024-dimensional codes produced by a regression-trained encoder. For an in-distribution test set, most ground-truth parameter values fall inside the 68% credible interval of the MAF-based posterior, and the Uncertainty Calibration Error ranges from about 2% to 13% depending on parameter and simulation, so the posterior uncertainty is described as reasonably calibrated. The paper also claims that the latent codes are separable in feature space, with points of similar parameter values clustering together, and that the featurizer generalizes: when the encoder built from one simulation compresses boxes from the other simulation, the posterior estimator still achieves $R^2 \\geq 0.84$, and when it compresses boxes from the same simulation at $z = 0.5$ (the encoder being built at $z = 1$), $R^2 \\geq 0.83$ holds. These results are put forward as a proof of concept for analyzing data from future 21-cm tomographic surveys.","pith_inferences":["The paper does not test whether the 1024-dimensional codes are sufficient summaries; a direct comparison with a field-level posterior or with a summary known to be sufficient would indicate whether the regression pre-training discards posterior-relevant information.","The cross-simulation robustness is demonstrated on two simulations of similar resolution and physics, so the transferability to real radio maps with noise and foregrounds remains open; the author explicitly flags domain adaptation as the next step.","The degradation at $z = 0$ suggests the useful range of a single featurizer is roughly $\\Delta z \\lesssim 0.5$; training on a multi-redshift set would be a natural way to extend the method, and the paper leaves that unexplored.","A survey-level application would combine the posteriors from many boxes; the paper does not address how to handle correlations between boxes when they are assembled into a single posterior."],"forward_implications":["If correct, the pipeline supplies a likelihood-free route to cosmological constraints from 21-cm tomographic data, usable wherever a simulation-based training set exists.","The cross-simulation result ($R^2 \\geq 0.84$) implies that a featurizer trained on one survey's simulations may compress another survey's data with only modest loss, as long as the voxel distributions are not too far apart.","The percent-level calibration means the $1\\sigma$ intervals can be read as physical error bars on a single inference without an additional recalibration step.","The linear test's $R^2 \\geq 0.80$ shows the codes are nearly linearly informative, so the encoder's compression does not destroy the cosmological content.","The worst transfer case ($R^2 \\approx 0.6$ for $\\sigma_8$ from $z = 0$ boxes encoded by a $z = 1$ featurizer) sets a practical redshift horizon for reusing one encoder."],"supporting_citations":[{"why":"Supplies the Automatic Posterior Transformation framework used to train the flow-based posterior estimator.","marker":"Greenberg et al. (2019)"},{"why":"Introduces Masked Autoregressive Flow, the density estimator modelling the posterior.","marker":"Papamakarios et al. (2017)"},{"why":"Provides the multifield 3D neutral-hydrogen grids used as training and test data.","marker":"Villaescusa-Navarro et al. (2022b)"},{"why":"Defines the simulation suite and parameter ranges from which the boxes and labels come.","marker":"Villaescusa-Navarro et al. (2021b)"},{"why":"Demonstrates 3D CNN plus likelihood-free inference on 21-cm lightcones, the approach extended here to intermediate redshift.","marker":"Zhao et al. (2022)"},{"why":"Provides the regression loss of Eq. (1) and the 2D HI-map baseline for comparison.","marker":"Andrianomena & Hassan (2023b)"},{"why":"Defines the Uncertainty Calibration Error used to assess posterior calibration.","marker":"Laves et al. (2020)"},{"why":"Supplies the software package used for the APT/MAF posterior estimation.","marker":"Tejero-Cantero et al. (2020b)"}],"fun_headline_variants":["Neural encoder plus flow reads cosmology from 21-cm data","AI maps neutral hydrogen to Ωm and σ8 with R²≥0.91","Probabilistic AI extracts cosmology from 21-cm tomographic cubes","AI reads cosmology from 21-cm maps even beyond training data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 1024-dimensional latent codes produced by an encoder trained only to predict $\\Omega_{\\rm m}$ and $\\sigma_8$ retain all the information the density estimator needs for a calibrated posterior; if the regression objective discards features relevant to the posterior's shape or tails, the reported calibration is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["Neural encoder plus flow reads cosmology from 21-cm data","AI maps neutral hydrogen to Ωm and σ8 with R²≥0.91","Probabilistic AI extracts cosmology from 21-cm tomographic cubes","AI reads cosmology from 21-cm maps even beyond training data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001162,"raw_usage":{"total_tokens":4911,"prompt_tokens":1146,"completion_tokens":3765,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":762,"completion_tokens_details":{"reasoning_tokens":3685}},"tokens_in":762,"tokens_out":3765,"duration_ms":30923,"temperature":1.0,"reasoning_tokens":3685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:29:09.574543+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the held-out boxes, train the same MAF posterior directly on the raw $128^3$ cubes or on a summary known to be sufficient (e.g., the full power spectrum plus bispectrum) and compare its credible intervals with those obtained from the latent codes, or run an expected-coverage test across the full $(\\Omega_{\\rm m}, \\sigma_8)$ prior: if coverage deviates strongly from the nominal 68% in any region, or if the latent-code intervals are systematically wider than the field-level ones, the encoder's information sufficiency is falsified.","supporting_citations":[{"cited_title":"pp 2404--2414","cited_arxiv_id":null,"evidence_quote":"Supplies the Automatic Posterior Transformation framework used to train the flow-based posterior estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces Masked Autoregressive Flow, the density estimator modelling the posterior."},{"cited_title":"D., 2022, The Astrophysical Journal, 926, 151","cited_arxiv_id":null,"evidence_quote":"Demonstrates 3D CNN plus likelihood-free inference on 21-cm lightcones, the approach extended here to intermediate redshift."},{"cited_title":"F., Kahrs L","cited_arxiv_id":null,"evidence_quote":"Defines the Uncertainty Calibration Error used to assess posterior calibration."}],"review_version":1}