{"id":"9262dcab-4055-4c3c-9276-e39f18d46e0a","arxiv_id":"2603.13011","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Simulation-based inference on Lyman-α P1D recovers unbiased Ωm and σ8 when training jointly on IllustrisTNG and SIMBA, while single-model training biases σ8 by ~10%.","lead":"This paper applies neural posterior estimation to the Lyman-alpha forest 1D power spectrum using CAMELS hydrodynamic simulations. Multi-domain training across galaxy-formation models recovers unbiased cosmological constraints when single-model training fails.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Multi-domain TNG+SIMBA training is only shown to de-bias within that closed pair; generalization to a third hydro model or real data is untested and load-bearing.","rationale":"The reader's weakest assumption is exactly the load-bearing gap: multi-domain training is validated only as an interpolator between TNG and SIMBA, not as a robust solution to baryonic model misspecification. The abstract's numerical claims (same-model accuracy, ~10% cross-model σ8 bias, multi-domain de-biasing) cannot be audited because the cached full manuscript is a different paper. No formal verification, public code check, or external validation is available. That leaves the CONDITIONAL verdict and low confidence intact; the concrete third-suite hold-out is the single check that would either secure or weaken the headline claim. No stronger internal inconsistency is visible from the abstract alone, and disagreement with consensus is not at issue—the concern is empirical generalization.","tokens_in":22719,"tokens_out":587,"duration_ms":13369,"concrete_test":"Hold out an entire third CAMELS-style hydro suite (e.g. Astrid or EAGLE if available) from all training. Train the normalizing flow only on the combined TNG+SIMBA set; evaluate median σ8 bias and posterior coverage on the third suite. If |bias| ≳ 5–10% reappears or coverage is miscalibrated, the multi-domain fix does not generalize beyond the two-model pair.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that joint training on IllustrisTNG and SIMBA recovers unbiased NPE posteriors on Ωm and σ8 for Lyman-α P1D(k), curing the ~10% positive σ8 bias seen in cross-model tests. That claim is only demonstrated inside the two-model closed set: train on both, test on held-out draws from the same two. Nothing in the abstract establishes that the union of TNG and SIMBA spans the baryonic systematics of a third independent hydro suite (or of nature). If both codes share common residual biases relative to a held-out model—e.g., in IGM thermal history, wind coupling, or AGN feedback not captured by the four SN/AGN parameters—multi-domain training can still produce biased σ8 while looking calibrated on TNG↔SIMBA. The abstract also notes astrophysical parameters remain unconstrained due to limited volume, so residual baryonic degeneracy is not ruled out. The supplied full text is the unrelated protein-ultrametricity manuscript (2603.13012), so no coverage diagnostics, third-model tests, or real-data validation can be checked.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The abstract claims the first full simulation-based inference (neural posterior estimation with a normalizing flow) on the Lyman-α forest 1D power spectrum P1D(k) at 2.0<z<3.5, using CAMELS hydrodynamic simulations with IllustrisTNG and SIMBA. Same-model train/test recovers Ωm and σ8 to better than ~10% precision with high accuracy, while four SN/AGN feedback parameters remain unconstrained (attributed to limited volume). Cross-model tests produce a ~10% positive bias on σ8; multi-domain training on the union of both models is reported to restore unbiased cosmological constraints. The supplied full manuscript body, however, is an unrelated cond-mat paper on ultrametricity of a toy protein energy landscape (arXiv:2603.13012), so none of the claimed methods, figures, coverage tests, or multi-domain results can be verified from the provided text.","tokens_in":22970,"tokens_out":982,"duration_ms":15031,"significance":"If the abstract’s results hold under proper scrutiny, the work would be a useful step for cosmology with the Lyman-α forest: it demonstrates that neural posterior estimation can extract competitive Ωm–σ8 constraints from P1D and that multi-domain training can mitigate known non-convergence between galaxy-formation models. That would be practically relevant for DESI/WEAVE-style analyses and for the broader SBI literature on baryonic systematics. The contribution cannot be credited as demonstrated until the correct manuscript, diagnostics, and external-validation tests are available.","major_comments":[{"comment":"The full text provided under this submission is not the Lyman-α / CAMELS paper: it is a complete, unrelated manuscript on ultrametricity of a disordered heteropolymer energy landscape (protein prototype, N=128, replica overlaps via Pearson correlation of pairwise energies). No sections, equations, figures, or tables corresponding to the abstract’s claims exist in the supplied body. A technical review of training protocol, architecture, posterior calibration, coverage, or multi-domain procedure is therefore impossible. The correct manuscript must be supplied before any scientific assessment can proceed.","section":null},{"comment":"From the abstract alone: the central claim that multi-domain (TNG+SIMBA) training recovers unbiased Ωm and σ8 is only stated for held-out draws from the same two models. There is no reported test on a third independent hydro suite or on real data. If TNG and SIMBA share residual IGM/feedback systematics not spanned by the four SN/AGN parameters, multi-domain training can remain biased while appearing calibrated inside the closed pair. This is load-bearing for the claim that multi-domain training is an “effective solution” to lack of convergence between galaxy-formation models.","section":null},{"comment":"The abstract states that the four astrophysical parameters are “generally unconstrained due to the limited probed volume.” That leaves open residual baryonic degeneracy with σ8 (and possibly Ωm). Without published coverage tests, TARP/SBC diagnostics, or explicit checks that the multi-domain posterior remains calibrated when feedback parameters are free, the reported ~10% de-biasing of σ8 cannot be taken as robust against unmodeled baryonic physics.","section":null}],"minor_comments":[{"comment":"Abstract phrasing “within 10% deviations in ≳75% and ≳90% of the cases for Ωm and σ8, and a precision better than 10% in both” should be clarified (fraction of test realizations vs. typical posterior width; absolute vs. relative error).","section":null},{"comment":"When the correct manuscript is provided, please ensure figure captions and tables report the exact train/test splits, number of CAMELS realizations per model, k-range and redshift binning of P1D, and the precise multi-domain mixing procedure.","section":null}],"recommendation":"uncertain","confidential_remarks":"The PDF/body mismatch (2603.13011 abstract vs. 2603.13012 protein-ultrametricity full text) looks like a packaging or arXiv-cache error rather than author misconduct, but it blocks review. Please request the correct manuscript and, if possible, code/data for the multi-domain NPE before reassigning. Scope-wise the intended paper is appropriate for astro-ph.CO; the protein paper is not. I would re-review promptly once the right text is available. The skeptic’s closed-set multi-domain concern is real and should be pressed even after the correct paper arrives."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is a methodological first—full neural posterior estimation on the Lyman-α forest 1D power spectrum from CAMELS (TNG and SIMBA)—and the abstract’s story is coherent. Same-model training recovers Ωm and σ8 to better than 10% precision with good accuracy; training on one hydro model and testing on the other produces a ~10% positive σ8 bias; joint multi-domain training on both models removes that bias. Astrophysical SN/AGN parameters stay unconstrained, which they attribute to limited volume. That is a useful, concrete demonstration of a real problem (baryonic model mismatch) and a practical fix inside the two-model set.\n\nWhat is new is the application and the quantitative cross-model result, not the tools. Normalizing flows and multi-domain training are established; putting them on P1D(k) at 2<z<3.5 with CAMELS and showing the bias pattern is the contribution. From the abstract alone the protocol looks standard (known ground truth, train/test splits), not circular.\n\nThe soft spot that matters is exactly the stress-test point: multi-domain training is only shown to de-bias within the TNG↔SIMBA closed pair. Nothing here tests a third independent hydro suite or real data. If both codes share residual systematics (IGM thermal history, wind/AGN details not spanned by the four feedback parameters), joint training can look calibrated while still biasing σ8 on nature or on a held-out model. The unconstrained astrophysics parameters make that residual degeneracy harder to dismiss. We also cannot check coverage, calibration plots, or code—the supplied full text is the unrelated protein-ultrametricity manuscript (2603.13012), so this is abstract-only.\n\nWho it is for: people building SBI pipelines for spectroscopic surveys and anyone wrestling with hydro-model systematics in the Lyman-α forest. It deserves a serious referee once the real manuscript is in hand; the claim is important enough and the abstract is sharp enough that desk rejection would be wrong. Engage if you care about Lyman-α cosmology or multi-fidelity SBI; treat the multi-domain “solution” as provisional until a third model or external validation appears.","headline":"Clean first SBI result on Lyman-α P1D: same-model NPE works, cross-model biases σ8 by ~10%, multi-domain TNG+SIMBA training removes that bias—but only inside that closed pair, and we only have the abstract.","tokens_in":23582,"tokens_out":588,"would_cite":false,"duration_ms":11553,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Training a neural posterior on both IllustrisTNG and SIMBA recovers unbiased Ωm and σ8 from the Lyman-α forest 1D power spectrum; training on one model alone biases σ8 by about 10% on the other.","keywords":["Lyman-alpha forest","simulation-based inference","normalizing flows","CAMELS","neural posterior estimation","baryonic feedback","Ωm","σ8"],"falsifier":"Hold out a third independent hydrodynamic suite (or real survey P1D measurements with known truth from a blinded cosmology), train only on the IllustrisTNG+SIMBA pool, and check whether the recovered σ8 still lies within the claimed unbiased interval or reappears with a systematic offset of order 10%.","tokens_in":23600,"feed_emoji":"🌌","tokens_out":946,"duration_ms":16212,"temperature":0.7,"pith_summary":"This paper shows that full simulation-based inference can be run on the Lyman-α forest one-dimensional power spectrum using CAMELS hydrodynamic simulations. A normalizing flow is trained to estimate the posterior on two cosmological parameters (Ωm and σ8) and four astrophysical feedback parameters from P1D(k) at redshifts 2.0–3.5. When the network is trained and tested on the same galaxy-formation model, cosmological posteriors recover the true values to better than 10% in most cases, while astrophysical parameters stay largely unconstrained because of the small simulated volume. Cross-model tests (train on IllustrisTNG, test on SIMBA, or the reverse) produce a roughly 10% positive bias on σ8. Combining both simulation suites in multi-domain training removes that bias. The result is offered as a practical route to cosmological constraints from the Lyman-α forest that does not assume a single baryonic-physics model is correct.","feed_headline":"Two simulation suites cancel a 10% σ8 bias in Lyman-α inference","feed_subtitle":"Multi-domain training on IllustrisTNG and SIMBA recovers unbiased cosmology from the forest power spectrum","key_machinery":"Neural posterior estimation with a normalizing flow trained on CAMELS P1D(k) spectra; multi-domain training that pools IllustrisTNG and SIMBA so the learned posterior is not locked to one baryon-physics implementation.","core_discovery":"When a normalizing flow is trained jointly on IllustrisTNG and SIMBA realizations of the Lyman-α forest P1D(k), neural posterior estimation returns unbiased constraints on Ωm and σ8; single-model training yields a ~10% positive bias on σ8 when the network is evaluated on the other model.","pith_inferences":["If a third, substantially different feedback implementation (e.g., a future CAMELS variant) reintroduces bias after multi-domain training, the method will need explicit domain-adaptation layers rather than simple data pooling.","The unconstrained astrophysical parameters suggest that P1D at these redshifts mainly encodes large-scale cosmology; adding higher-order or transverse statistics could break the remaining degeneracies.","Blind application to DESI or other survey P1D without an external truth test would still leave open the possibility that real baryonic physics lies outside the IllustrisTNG–SIMBA span."],"forward_implications":["Cosmological inference from Lyman-α P1D can proceed without committing to a single sub-grid feedback model if both major CAMELS suites are used for training.","Cross-model bias of ~10% on σ8 is a concrete, measurable diagnostic of baryonic-model mismatch for any future emulator or likelihood pipeline.","Astrophysical feedback parameters remain poorly constrained by P1D alone at the volumes simulated here, so joint probes will still be needed for SN/AGN physics.","The same multi-domain recipe can be applied to other summary statistics of the Lyman-α forest once they are available in CAMELS."],"fun_headline_variants":["Joint TNG+SIMBA training cancels 10% σ8 bias in Lyα P1D SBI","Multi-domain CAMELS flow recovers unbiased Ωm and σ8 from forest","Dual galaxy models erase model-dependent σ8 offset in Lyman-α inference","Combined IllustrisTNG and SIMBA training unbaises forest cosmology","Cross-model training removes ~10% σ8 bias in Lyman-α neural posteriors"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the two galaxy-formation models together cover enough of the real range of supernova and AGN feedback that the multi-domain posterior will stay unbiased on real data or on any third, unseen simulation suite.","fun_headline_variants_meta":{"raw":{"variants":["Joint TNG+SIMBA training cancels 10% σ8 bias in Lyα P1D SBI","Multi-domain CAMELS flow recovers unbiased Ωm and σ8 from forest","Dual galaxy models erase model-dependent σ8 offset in Lyman-α inference","Combined IllustrisTNG and SIMBA training unbaises forest cosmology","Cross-model training removes ~10% σ8 bias in Lyman-α neural posteriors"]},"model":"grok-4.5","effort":"low","cost_usd":0.004456,"raw_usage":{"total_tokens":1394,"prompt_tokens":881,"num_sources_used":0,"completion_tokens":116,"cost_in_usd_ticks":44560000,"prompt_tokens_details":{"text_tokens":881,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":397,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":881,"tokens_out":116,"duration_ms":3861,"temperature":1.0,"reasoning_tokens":397,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T21:55:48.502630+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold out a third independent hydrodynamic suite (or real survey P1D measurements with known truth from a blinded cosmology), train only on the IllustrisTNG+SIMBA pool, and check whether the recovered σ8 still lies within the claimed unbiased interval or reappears with a systematic offset of order 10%.","supporting_citations":[],"review_version":1}