{"id":"d431fc63-b16e-4a3d-9309-59ce92b582ee","arxiv_id":"2607.17942","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"On Euclid-like Gaussian mocks, likelihood-free inference stays well-calibrated under emulator error and non-Gaussian summaries, while explicit Gaussian-likelihood inference becomes miscalibrated and can disagree with a CNN probe by factors up to two.","lead":"This paper compares two ways of inferring cosmological parameters from simulated weak-lensing data: explicit-likelihood inference with a Gaussian approximation, and simulation-based inference that learns the likelihood from mock maps. On Euclid-like Gaussian mocks it finds the explicit approach can become miscalibrated when the emulator or Gaussian assumption fails, while the simulation-based approach stays calibrated and robust to how the data are compressed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Null test for CNN-vs-2PCFs is not clean: finite 8-bin compression and mildly non-Gaussian forward model can make the CNN genuinely more informative, so the factor-of-two ELI discrepancy is not purely a framework artifact.","rationale":"The reader's weakest_assumption identifies the Gaussian-sufficiency premise as the key load-bearing assumption, and I agree. The paper's main methodological finding — that ELI becomes miscalibrated when its emulator or Gaussian-likelihood assumptions fail, while LFI on shear-2PCFs stays calibrated — is well supported by the TARP/PIT/coverage diagnostics and the synthetic-dataset consistency test. That part of the conclusion should stand. However, the Section 8 extension to the CNN comparison inherits an idealized premise: that ξ± with 8 angular bins is information-equivalent to the full map for Gaussian fields. The forward model is not exactly Gaussian at the level of observed ellipticities (reduced shear, non-Gaussian shape-noise draw), and the binning is finitely lossy. The LFI panel of Fig. 10 already shows a ~30% residual difference and a trade-off in constraining power between Ωm and S8, which is consistent with genuine information loss or extra CNN information rather than a pure framework effect. Because the factor-of-two claim and the 'LFI bounds the genuine information difference' statement both rely on this null test, the interpretation is conditional. This does not change the reader's CONDITIONAL verdict: the methodological core is sound, but the CNN comparison needs either a finer-binning control or a linear-Gaussian forward-model control to make the null test quantitative. The proposed test is a direct, feasible check of whether the residual discrepancy is a binning/non-Gaussianity artifact or a true information difference.","tokens_in":31212,"tokens_out":13990,"duration_ms":131145,"concrete_test":"Repeat the LFI and ELI analyses for shear-2PCFs on the same GLASS mocks and the same mock observation, but with substantially finer angular binning (e.g., 16 or 32 log-spaced bins, or a full map-level pseudo-Cℓ estimator). If the LFI shear-2PCF posterior tightens by ~30% and moves toward the CNN posterior, the residual discrepancy is partly a binning artifact and the null test needs revision. In addition, run a control simulation with a strictly linear-Gaussian forward model (eobs = eint + γ, with Gaussian amplitude noise); if the CNN-vs-2PCFs residual disappears in that control, the non-Gaussianity of the forward model is the cause.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central cross-probe claim in §3 and §8 rests on the premise that, since the simulations are Gaussian random fields, shear-2PCFs capture all available information, so any CNN advantage must be an ELI artifact. This premise is not exact. (1) The observed ellipticity maps enter through the reduced-shear composition of Eq. (4), which is nonlinear in the Gaussian shear and the shape noise; the adopted phase-only noise with fixed amplitude is also non-Gaussian per component. (2) The shear-2PCFs are binned into only 8 logarithmic angular bins, which is a lossy compression of the full two-point information; the CNN sees the full 64×64 map. The paper acknowledges binning loss in passing but does not quantify it. (3) The paper's own LFI results show a ~30% residual difference between NN-compressed 2PCFs and the CNN, with the two probes trading constraining power between Ωm and S8. That is exactly what a small genuine information difference or asymmetric binning loss would look like, so it cannot be assumed to bound only the 'genuine information difference' in the way §9 claims. Consequently, the factor-of-two ELI discrepancy could be partly due to the CNN summary containing real extra information rather than solely to ELI's emulator/likelihood failures, weakening the clean null-test interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using GLASS Gaussian random-field weak-lensing mocks tailored to Euclid DR3, the paper compares explicit likelihood inference (ELI: GP emulator + Gaussian likelihood + simulation covariance) with likelihood-free inference (LFI: NDEs) on shear 2PCFs under MOPED/NN compression and on a map-level CNN summary, focusing on Ωm and S8. It reports that ELI and LFI agree when ELI's assumptions are met; that ELI is miscalibrated when emulation is inaccurate or the compressed likelihood is non-Gaussian; that compression choice affects ELI but not LFI; and that in the Gaussian-field null test the CNN-vs-2PCF discrepancy is up to a factor of two for ELI but only ≈30% for LFI. It concludes that LFI is more robust and advocates calibration diagnostics for ELI.","tokens_in":31369,"tokens_out":9966,"duration_ms":90174,"significance":"The paper is a carefully controlled empirical benchmark with real methodological value. It shows good train/validation/test hygiene (GP on 250 nodes, NDE on 3,687 nodes, diagnostics on 692 held-out nodes, mock observation reserved), and it extends calibration diagnostics (TARP, PIT, coverage) to ELI. The synthetic-LFI test and bin-0-only control are convincing internal-consistency checks. If the central claim survives the null-test caveat below, it would be a useful reference for future Stage-IV WL analyses choosing between ELI and LFI.","major_comments":[{"comment":"The null-test interpretation is not clean. The paper's central claim—that the factor-of-two CNN-vs-2PCF discrepancy under ELI is an ELI artifact—rests on the assertion that ξ± captures all information in Gaussian fields. However, the simulated observable is not Gaussian: the observed ellipticity is formed via the reduced-shear composition of Eq. (4), which is nonlinear in the Gaussian shear and in a phase-only noise field whose fixed amplitude makes the two noise components non-Gaussian; and ξ± is binned into 8 logarithmic angular bins, which is lossy, while the CNN sees the full 64×64 map. Each of these effects can make the CNN genuinely more informative. The LFI comparison itself shows a ~30% residual difference with asymmetric Ωm/S8 constraining power, exactly what genuine extra information in the CNN would look like. I recommend (i) repeating the comparison on maps constructed to be","section":"§3, §8, Eq. (4), §3.1"},{"comment":"The ELI and LFI frameworks are trained with very different amounts of data: the GP emulator uses 250 nodes, the NDE uses 3,687 nodes. The ELI-LFI discrepancy under MOPED is traced to poor emulation of bin 1 (Appendix A: slope a=0.710, WARN). The paper states in §5.1 that increasing beyond 250 nodes yields consistent posterior constraints, but no evidence is shown; if emulator accuracy improves with training-set size, the 'emulation inaccuracy' failure may be an artifact of the 250-node cap rather than an intrinsic ELI limitation. Please provide an emulator accuracy vs training-size test (e.g., slope/intercept for bin 1 at 500/1000/3687 nodes) or otherwise justify that the 250-node choice is not responsible for the ELI-LFI gap.","section":"§5.1 vs §6.1, Appendix A"},{"comment":"The abstract and conclusions state that ELI becomes 'strongly miscalibrated' under emulation inaccuracies or likelihood non-Gaussianity. However, the headline TARP curves in Fig. 4 show only a mild departure from the diagonal, mostly at low credibility levels; the larger departures appear in the marginal coverage and PIT diagnostics (Appendix C). Please either temper the wording or quantify the miscalibration (e.g., maximum ECP deviation, fraction of nodes outside the bootstrap band, or a combined statistic) so the strength of the claim is commensurate with the evidence. As written, the summary overstates the calibration difference shown in the main diagnostic.","section":"Abstract, §9, Fig. 4"}],"minor_comments":[{"comment":"The statement that 'Both frameworks well recover (Ωm^fid,S8^fid) within 1σ' is not supported for LFI S8: the reported median is S8=0.801+0.015−0.016, while the fiducial is 0.823, an offset of ~1.4σ given the lower error. Please correct or qualify this claim.","section":"§7.1"},{"comment":"The text first says shape noise is 'independent Gaussian noise' and then introduces the phase-only fixed-amplitude draw. Clarify that the nominal noise model is not Gaussian per component, since this is relevant to the Gaussianity discussion in §5.2.","section":"§2.3"},{"comment":"The flag label 'W ARN' appears to be a typo for 'WARN'. Also, the 'cbar: Omega_m' / 'cbar: sigma8' labels in the panels are cryptic; use 'colored by Ωm' / 'colored by σ8' in the caption.","section":"Appendix A, Fig. A.1"},{"comment":"There are duplicate reference entries for LeCun et al. 1998 ('Lecun' and 'LeCun'). Unify the spelling and merge the entries.","section":"References"},{"comment":"The 102 flat-sky patches are described as 'quasi-independent'. A brief quantitative statement (e.g., typical correlation between adjacent patches or effective number of independent samples) would help assess the covariance estimation and the KS-test critical values.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: The manuscript is a useful benchmark with a strong internal-consistency design. My main reservations are the clean interpretation of the CNN-vs-2PCF null test and the asymmetric training-set sizes between ELI and LFI; both are addressable with additional experiments or quantified diagnostics. I would be comfortable with acceptance after a major revision that resolves these two points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the most careful side-by-side ELI vs LFI comparison I've seen on matched Gaussian mocks, and the main conclusion — ELI diverges from LFI when its emulator or Gaussian-likelihood assumptions fail, and converges once they're satisfied — holds up. The porting of TARP/PIT/marginal coverage to the explicit-likelihood pipeline is a genuinely useful contribution; Euclid/Rubin analysts should be doing this.\n\nWhat I like: the synthetic-LFI test (training an MDN on GP-predicted means plus Gaussian noise) is a neat way to isolate framework differences; the bin-0-only control isolates emulator degradation; the 20-realization scatter test shows the realization-to-realization trends are shared. The compression result — NN compression hurts ELI but not LFI — is convincingly tied to the non-Gaussianity of the compressed summaries via the KS and chi-square diagnostics.\n\nSoft spots, in ascending order of importance. The 'both frameworks well recover within 1σ' statement in §7.1 is simply contradicted by their own numbers: LFI S8 intervals (0.801±0.015 and 0.803±0.013) exclude the fiducial 0.823, and the CNN LFI S8 is ~2.3σ low. Fixable wording, but it undermines trust in the abstract. The 20-realization bias test subtracts the median bias and doesn't include the CNN, so absolute CNN calibration is unvalidated — worth adding. TARP capped at 50 test nodes is a computational constraint; they verified consistency with 200 for MOPED, so minor.\n\nThe bigger issue is the 'null test' for CNN-vs-2PCFs. The claim that shear-2PCFs should be information-optimal because the fields are Gaussian is approximately true, but not exactly: the reduced-shear composition (Eq. 4) is nonlinear in shear and noise, the phase-only noise is not Gaussian per component, and the 8-bin compression is lossy next to a 64×64 map. The paper itself notes binning loss 'in passing' but never quantifies it. So the factor-of-two ELI discrepancy with the CNN cannot be attributed purely to ELI's failures — some of it may be real extra information in the CNN. That weakens the clean null-test interpretation in §8 and the 'robustness of the deep-learning probe' headline. The 30% residual LFI discrepancy, with the two probes trading constraining power between Ωm and S8, is exactly what a small genuine information difference would look like.\n\nNone of this overturns the central comparative claim about ELI vs LFI calibration, which is well supported by the synthetic test and the bin-0 control. But the paper should be revised to (a) fix the 1σ overstatement, (b) quantify the Gaussianity of the actual observed ellipticity maps, and (c) either drop the strong null-test claim or validate the CNN information content on truly Gaussian fields with lossless summary statistics.\n\nVerdict: send to a serious referee. The experimental design is reproducible in principle, but shipping code/data would materially increase confidence. Worth reading for anyone building Euclid/Rubin inference pipelines.","headline":"Careful matched ELI-vs-LFI benchmark with a useful calibration protocol; the main conclusion holds, but the CNN null test is cleaner in the text than in the actual forward model, and the 'within 1σ' phrasing is contradicted by their own numbers.","tokens_in":32081,"tokens_out":2513,"would_cite":true,"duration_ms":23932,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On Gaussian weak-lensing mocks, explicit-likelihood inference goes miscalibrated when its assumptions fail; likelihood-free inference stays calibrated, making the CNN's factor-of-two edge an artifact.","keywords":["weak lensing","cosmic shear","simulation-based inference","likelihood-free inference","posterior calibration","TARP","neural density estimators","Gaussian random fields"],"falsifier":"Replace the Gaussian random-field simulations with N-body mocks at matched resolution and noise, keeping everything else identical. If under LFI the CNN then beats shear-2PCFs by substantially more than ~30%, the 'CNN-versus-2PCFs discrepancy is an ELI artifact' reading fails; if under ELI the factor-of-two persists even after replacing the GP emulator with the true simulation mean and using a non-Gaussian likelihood, then the attribution to emulator and likelihood assumptions is wrong.","tokens_in":30878,"feed_emoji":"🔭","tokens_out":7124,"duration_ms":64200,"temperature":0.7,"pith_summary":"The paper sets up a controlled side-by-side test of two ways to infer cosmology from weak-lensing shear maps: explicit likelihood inference (ELI), which assumes a Gaussian likelihood fed by an emulator and a covariance matrix, and likelihood-free inference (LFI), which learns the likelihood from simulations with neural density estimators. On Gaussian random-field mocks tailored to a Stage-IV-style survey, it finds that ELI becomes strongly miscalibrated whenever its emulator is inaccurate or the compressed data vector is non-Gaussian, while LFI stays well calibrated. These ELI-specific failures produce an apparent factor-of-two difference in Ωm constraints between shear two-point functions and a map-level CNN; once the assumptions are repaired, ELI and LFI agree, and the remaining ~30% gap in LFI is the genuine information difference. The practical message is that posterior calibration tests such as TARP should be run on ELI pipelines too, and that comparing summary statistics without validating the inference framework can create false discovery.","feed_headline":"Explicit likelihood mis-calibrates; simulation-based inference holds","feed_subtitle":"On Gaussian cosmic-shear mocks, the apparent CNN information gain disappears once the inference framework is calibrated.","key_machinery":"The decisive mechanism is the matched-assumption test: feeding the GP emulator's predicted means and the Gaussian covariance into LFI reproduces the ELI posterior exactly, isolating emulator inaccuracy and likelihood non-Gaussianity as the sole drivers of the discrepancy. Alongside this, the Test of Accuracy with Random Points (TARP)—a joint posterior calibration diagnostic that compares expected coverage against credibility level—is applied to ELI as well as LFI, exposing miscalibration that standard per-bin KS and χ² Gaussianity tests miss. These two tools together convert a seemingly large cosmological discrepancy into a clean statement about which inference assumptions are actually viola","core_discovery":"The central claim is that the observed disagreements between ELI and LFI—up to a factor of two on Ωm and S8 constraints from shear-2PCFs, and between shear-2PCFs and a CNN—are not intrinsic to the algorithms but follow from ELI's two working assumptions: an accurate Gaussian-process emulator and a Gaussian likelihood for the compressed summaries. Using LFI retrained on a synthetic dataset built from the emulator's predictions plus Gaussian covariance noise, the authors reproduce the ELI posterior, and restricting to the single accurately emulated, near-Gaussian bin brings the two frameworks into agreement. LFI, which learns the likelihood shape directly, is unaffected by non-linear NN compre","pith_inferences":["If the simulated fields carry even mild non-Gaussianity from the power-spectrum emulator, noise injection, or the Kaiser–Squires inversion, the CNN-versus-two-point null test is not perfectly clean, so part of the residual ~30% LFI discrepancy could be genuine higher-order information captured by the CNN.","The same calibration audit could be applied to ELI analyses on real survey data as a sanity check: a failing TARP curve would flag unreliable error bars even when marginal Gaussianity tests pass.","The paper's controlled setup suggests a testable ordering for non-Gaussian mocks: ELI's miscalibration should grow as the likelihood becomes more non-Gaussian, while LFI should remain calibrated; checking this on N-body mocks would extend the result."],"forward_implications":["When ELI's emulator is inaccurate or the compressed likelihood is non-Gaussian, ELI posteriors become miscalibrated and can disagree with a map-level CNN by up to a factor of two on Ωm, even on Gaussian fields where two-point statistics should be optimal.","Repairing both assumptions—by using an accurately emulated bin or by feeding the emulator mapping and Gaussian noise into LFI—restores agreement between ELI and LFI, showing the discrepancy is an ELI artifact.","LFI remains well calibrated under both linear (MOPED) and non-linear (NN) compression, and NN compression modestly improves LFI's Ωm precision, whereas for ELI the same non-linear compression broadens Ωm errors by ≈1.3× and S8 errors by ≈2.2×.","Posterior calibration diagnostics developed for LFI (TARP, marginal coverage, PIT) applied to ELI reveal miscalibration that standard KS and χ² Gaussianity tests miss.","For future Stage-IV analyses the paper recommends running ELI and LFI in parallel, using these calibration diagnostics, and investing in non-Gaussian likelihood models for ELI."],"fun_headline_variants":["Cosmic shear: simulation-based inference out-calibrates explicit likelihood","Explicit likelihood miscalibration doubles weak-lensing errors","SBI robust, ELI off—calibration is the culprit in weak lensing","Weak-lensing inference discrepancies vanish after calibration"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The null-test conclusion rests on the premise that the simulated shear fields are exactly Gaussian random fields, so the two-point correlation functions carry all the information; if residual non-Gaussianity sneaks in through the power-spectrum model, noise injection, or map making, the CNN-versus-two-point comparison is no longer a clean test of the inference framework.","fun_headline_variants_meta":{"raw":{"variants":["Cosmic shear: simulation-based inference out-calibrates explicit likelihood","Explicit likelihood miscalibration doubles weak-lensing errors","SBI robust, ELI off—calibration is the culprit in weak lensing","Weak-lensing inference discrepancies vanish after calibration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001006,"raw_usage":{"total_tokens":4149,"prompt_tokens":862,"completion_tokens":3287,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":3215}},"tokens_in":606,"tokens_out":3287,"duration_ms":23687,"temperature":1.0,"reasoning_tokens":3215,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:34:50.010762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the Gaussian random-field simulations with N-body mocks at matched resolution and noise, keeping everything else identical. If under LFI the CNN then beats shear-2PCFs by substantially more than ~30%, the 'CNN-versus-2PCFs discrepancy is an ELI artifact' reading fails; if under ELI the factor-of-two persists even after replacing the GP emulator with the true simulation mean and using a non-Gaussian likelihood, then the attribution to emulator and likelihood assumptions is wrong.","supporting_citations":[],"review_version":1}