{"id":"efecc5b4-3c9d-44a9-b2e3-4c4c6bd01501","arxiv_id":"2412.04443","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The inferred initial matter power spectrum from the SELFI algorithm reveals misspecified galaxy bias, selection, mask, redshift, and gravity models, exposing a >2σ cosmological bias before parameter inference.","lead":"This paper tests a way to catch mistakes in the computer models that turn galaxy survey data into cosmological measurements. It shows that a subtly wrong model shifts the reconstructed early-universe power spectrum, which can be spotted before the final cosmological parameters are calculated.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detection is uncalibrated: dM(γB) is compared to prior draws, not to the null distribution of posterior means under a well-specified model, so the \"unambiguous\" detection claim lacks statistical support.","rationale":"The paper convincingly demonstrates on mocks that a misspecified forward model shifts the SELFI posterior and biases the ABC parameter posterior, and the code and validation are credible. The load-bearing step for the advertised practical use is the detection rule in Section IVA1. The statistic dM(γB,θ0|S) is compared to a histogram of prior samples, but a posterior mean is not a prior sample; Eq. (20) shows γ is an optimally shrunk estimator, so its null distribution is determined by the data noise, the prior width, and the gradient/covariance estimates. Since the paper does not provide this null distribution, the word \"unambiguously\" in the abstract is not supported. This is not a criticism of the physics but a missing statistical calibration that the authors can supply with their own linearized model at negligible cost. If the calibration shows dM(γB) is a rare event under Model A, the central claim is strengthened. The reader's CONDITIONAL verdict already anticipates this; our independent stress test therefore does not move the verdict.","tokens_in":38593,"tokens_out":8878,"duration_ms":177126,"concrete_test":"Compute the null distribution of the detection statistic under the well-specified Model A. Using the SELFI Gaussian approximation, draw N=10^4 synthetic observed summaries ΦO from N(f(θgt), C0) (with θgt from Table VIII), propagate each through Eq. (20) to obtain γ, and record dM(γ, θ0|S). Report the 95th and 99th percentiles and compare dM(γB)=2.827. If dM(γB) lies below the 99th percentile, the claimed unambiguous detection fails and a calibrated threshold must be supplied; if above, the central claim is supported. A complementary check with a handful of full N-body realizations (varying phase ψ only) would test the Gaussian-assumption dependence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the SELFI posterior mean γ reliably flags model misspecification before parameter inference. Section IVA1 (Eq. 24, Fig. 8) supports this by comparing dM(γB, θ0|S)=2.827 with the distribution of dM(θn, θ0|S) for 5000 prior draws (mean 2.431). That reference distribution is not the null distribution of the detection statistic. γ is a posterior mean, not a draw from P(θ); under a well-specified model it is shrunk toward the prior mean by Eq. (20), so its Mahalanobis distance has a different, narrower distribution. The paper reports γA=1.816 but gives no distribution of γA over noise/phase realizations, no threshold, and no false-positive rate. A user therefore cannot know whether dM=2.827 is an unambiguous misspecification signal or a tail fluctuation of the well-specified model. The abstract's \"unambiguously detected and avoided\" is thus not yet established; the detection rule needs calibration before the method is used as a gate for Stage-IV analyses.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a two-step framework for diagnosing model misspecification in field-based implicit-likelihood cosmological inference. In the first step, the SELFI algorithm is used to infer the initial matter power spectrum from a hidden-box forward model of a spectroscopic galaxy survey, with cosmological parameters treated as nuisance parameters inside the simulator. The inferred latent spectrum is then proposed as a diagnostic: systematic effects that bias the cosmological parameter posterior are expected to leave an imprint on the inferred initial power spectrum. In the second step, the simulations used for SELFI are recycled to build a score compressor, and ABC-PMC is used to infer cosmological parameters from the compressed summaries. The method is demonstrated on a mock survey with a well-specified model (Model A) and a subtly misspecified model (Model B), differing at the percent level in galaxy bias, selection-function amplitudes, and extinction. The authors show that Model A yields an unbiased SELFI posterior and an unbiased cosmological posterior, while Model B yields a SELFI posterior that departs from the prior and a >2σ bias in the (Ωm, σ8) plane. They also provide a practical guide to diagnosing individual systematics: galaxy bias, extinction, selection functions, masks, redshift errors, and gravity-solver choices.","tokens_in":38690,"tokens_out":7640,"duration_ms":84820,"significance":"If the central claim is established, this would be a practically valuable tool for upcoming surveys such as DESI, Euclid, and LSST, because field-based implicit likelihood pipelines currently have little protection against model misspecification. The paper has clear strengths: the method is validated on mocks with a known ground truth, a wiggle-less prior test recovers BAOs, ten ground-truth draws are reported as consistency checks, and the code and data are publicly available. Reusing a single set of N-body simulations for both the SELFI step and the score compressor is an attractive feature. However, the headline claim that misspecification can be 'unambiguously detected and avoided' is not yet supported by a calibrated decision rule, and the quantitative evidence for the >2σ bias rests on a single synthetic realization. These issues are fixable and do not undermine the overall idea, but they are load-bearing for the strongest claim in the abstract.","major_comments":[{"comment":"The quantitative detection criterion is not calibrated. Equation (24) defines d_M(γ, θ0|S) for the posterior mean γ, but the reference distribution in Fig. 8 is the distribution of d_M(θ_n, θ0|S) for 5,000 draws θ_n = T(ω_n) from the prior. Under the well-specified model, γ is not a draw from P(θ); using the Gaussian effective likelihood, the posterior covariance is Γ = [(∇f0)^T C0^{-1} ∇f0 + S^{-1}]^{-1} (Eq. 21), which is generally tighter than S. The null distribution of d_M(γ_A, θ0|S) over noise and phase realizations is therefore narrower than the prior-draw histogram shown. The paper reports d_M(γ_A)=1.816 and d_M(γ_B)=2.827, but with no threshold and no false-positive rate derived from the distribution of γ under a well-specified model, one cannot determine whether 2.827 is an unambiguous misspecification signal or a tail fluctuation. The abstract's 'unambiguously detected' claim requires either a posterior predictive check or a decision rule calibrated on the null distribution of γ, not on prior draws.","section":"IVA1 (Eq. 24, Fig. 8)"},{"comment":"The 'avoided before parameter inference' step is not demonstrated through a decision procedure. Section IVB shows that running ABC-PMC with the misspecified Model B produces a >2σ bias in the (Ωm, σ8) plane, but this posterior is only obtained by performing the very inference step that the framework is supposed to avoid. The paper never specifies the rule by which a user, seeing only the SELFI diagnostics of Section IVA, would select Model A over Model B, nor does it quantify the false-positive and false-negative rates of such a rule. Without this, the workflow demonstration is incomplete: the reader sees that Model B is detectable in hindsight but not how it would be excluded prospectively.","section":"IVB (Fig. 12)"},{"comment":"The 'bias exceeding 2σ in the (Ωm, σ8) plane' is established from a single synthetic observation. Under a well-specified model, a 95% credible region excludes the true parameter 5% of the time, so one realization cannot by itself distinguish a systematic bias from a rare statistical fluctuation. The paper reports that the Model A posterior is unbiased, but no repeated-realization or coverage test is shown for the ABC-PMC stage. To support the bias claim, the authors should either show the distribution of ABC posterior means over multiple ΦO realizations for Models A and B, or demonstrate empirically that the Model A posterior has correct frequentist coverage in this setup.","section":"IVB and Fig. 12"}],"minor_comments":[{"comment":"The ten-universe unbiasedness check is described in a single sentence; please include a figure or table showing the ten SELFI posteriors or at least their means and credible intervals, so the 'visually and quantitatively' claim can be inspected.","section":"Appendix C1"},{"comment":"The text says d_M measures deviation from 'the prior distribution', but the formula is evaluated with respect to θ0, the expansion point, not the prior mean θ̂ω of Eq. (8). Please clarify this distinction and report Δ = θ̂ω - θ0, since a nonzero shift enters d_M directly.","section":"IIB3 (Eq. 24)"},{"comment":"The scope limitation that some systematics may not imprint on the initial matter power spectrum, while still biasing cosmological inference, is important and currently appears only near the end of the paper; consider stating it explicitly in the abstract or introduction to prevent over-generalization.","section":"V"},{"comment":"The caption is very long, and the color-scale definitions for each sub-panel are only given in the caption; labeling the color bars inside each panel or adding a legend would greatly improve readability.","section":"Fig. 9"},{"comment":"The reported d_M values are point estimates that inherit noise from the finite numbers of simulations used to estimate f0, C0, and ∇f0; a bootstrap over the N0=500 expansion-point simulations would quantify the uncertainty on d_M and make the comparison more informative.","section":"IVA1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a promising methods demonstration with real strengths: it builds on a principled hierarchical framework, ships public code and data, and validates the SELFI recovery on mocks. My main concern is that the headline claim of 'unambiguous' detection goes beyond the evidence: the detection statistic is uncalibrated against the null distribution of posterior means, and the bias claim relies on a single realization. Both issues are fixable with additional simulation-based calibration and repeated-realization tests. I therefore recommend major revision rather than rejection; the central idea is sound and within the scope of the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Florent and Tristan have written a careful methods paper that does something genuinely new: it turns the SELFI posterior on the initial power spectrum into a diagnostic for model misspecification in field-based implicit likelihood inference. The formal generalization of the SELFI posterior to the case where the prior mean differs from the expansion point (Eq. 20, Appendix B) is a real addition, and the treatment of cosmological parameters as nuisance parameters inside the hidden-box model closes a gap in the original SELFI derivation. The mock validation is also solid: Model A recovers the truth, a wiggle-less prior still finds the BAOs, ten ground-truth draws show no bias, and the catalogue of systematics (bias, extinction, selection, masks, redshifts, gravity solver) is a practical contribution. Credit where due: this is reproducible, with code and data promised publicly.\n\nWhere I part company with the abstract's stronger claim: the 'unambiguous detection' is not statistically calibrated. The diagnostic in Fig. 8 compares dM(γB) with the distribution of dM for prior draws θn ~ P(θ). But γ is a posterior mean, shrunk toward the prior; under a well-specified model, its Mahalanobis distance has a different, narrower distribution. The paper reports dM(γA)=1.816 but gives no distribution of γA over noise/phase realizations, no threshold, and no false-positive rate. So a user cannot yet tell whether dM=2.827 is signal or tail. This doesn't sink the method—the posterior plots are visually telling, and the bias-shift in the (Ωm, σ8) plane is real—but it does mean the detection rule needs to be formalized before this becomes a gate for Stage-IV analyses. The proxy assumption (systematics that don't imprint on the initial spectrum are out of scope) is explicitly acknowledged in Section V; that's honest, but it's a limitation worth restating in the abstract.\n\nThe other soft spots are minor: the SELFI linearization is validated only on these mocks, and the public code has no commit hash, which makes reproducibility claims slightly weaker than they should be. The self-citation pattern is not a problem—SELFI is the right prior work, and the extension is clear.\n\nBottom line: this is a serious, useful paper that deserves a real referee. It should be sent to peer review, with a request that the authors calibrate the detection statistic (even a simple simulation-based null distribution of γ under Model A would do).","headline":"A genuinely useful SELFI-based misspecification diagnostic with an uncalibrated detection rule; deserves peer review, but the 'unambiguous' claim needs a null distribution.","tokens_in":39327,"tokens_out":2147,"would_cite":true,"duration_ms":21710,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The inferred initial power spectrum can expose hidden survey biases before cosmological parameters are fitted.","keywords":["model misspecification","implicit likelihood inference","initial matter power spectrum","SELFI","galaxy surveys","systematic effects","ABC-PMC","score compression"],"falsifier":"Build a forward model whose only misspecification adds a component to the galaxy power spectra that lies in the null space of $(\\nabla_\\theta f_0)^\\top C_0^{-1}$, so the inferred initial power spectrum posterior is unchanged, and show with the paper's ABC-PMC pipeline that $(\\Omega_\\mathrm{m}, \\sigma_8)$ still shifts by more than $2\\sigma$; that would refute the diagnostic premise.","tokens_in":38272,"feed_emoji":"🔭","tokens_out":6084,"duration_ms":56143,"temperature":0.7,"pith_summary":"The paper claims that the posterior on the initial matter power spectrum, obtained with the SELFI algorithm, can serve as a diagnostic for model misspecification in field-based implicit-likelihood cosmological inference. Using a spectroscopic galaxy survey forward model with non-linear gravitational evolution and several engineered systematic errors, the authors show that percent-level misspecifications of galaxy bias, selection functions, masks, redshifts, and the gravity solver leave recognizable imprints on the inferred initial power spectrum. They further show that one subtly misspecified model shifts the final cosmological posterior by more than $2\\sigma$ in the $(\\Omega_\\mathrm{m}, \\sigma_8)$ plane, and that the shift appears in the initial-power-spectrum diagnostic before the cosmological parameters are inferred. The practical stake is that this gives a pre-inference alarm for hidden-box forward models of surveys such as DESI, Euclid, and LSST.","feed_headline":"Initial power spectrum catches hidden survey biases","feed_subtitle":"Percent-level modeling errors show up in the inferred early-Universe spectrum before they derail cosmological parameters.","key_machinery":"The carrying object is the SELFI effective posterior for the initial matter power spectrum, a Gaussian with mean $\\gamma = \\theta_0 + \\Gamma (\\nabla_\\theta f_0)^\\top C_0^{-1} (\\Phi_O - f_0) + \\Gamma S^{-1}\\Delta$ and covariance $\\Gamma = [(\\nabla_\\theta f_0)^\\top C_0^{-1} \\nabla_\\theta f_0 + S^{-1}]^{-1}$, where $f_0$ and $C_0$ are the empirical mean and covariance of the hidden-box model at the expansion point, $\\nabla_\\theta f_0$ is the finite-difference gradient, and $S$ is the prior covariance on the spectrum obtained by sampling cosmological parameters through a Boltzmann solver. A Mahalanobis distance between the posterior mean and the prior, $d_M(\\gamma,\\theta_0|S)$, is the quantitative misspecification check. The same linearisation also defines the score compressor $\\tilde{\\omega} = \\omega_0 + F_0^{-1}(\\nabla_\\omega f_0)^\\top C_0^{-1}(\\Phi - f_0)$ for the second inference step. This machinery turns one set of $N$-body simulations into both a systematic-effect scan and a data compressor, which is what makes the diagnostic affordable.","core_discovery":"The central claim is that the SELFI posterior of the initial matter power spectrum can flag systematic effects that would otherwise bias a field-based implicit likelihood inference. In the two-step framework, the latent initial power spectrum $\\theta$ is inferred first via a Gaussian effective likelihood built from a first-order Taylor expansion of the forward model around a fiducial spectrum; the simulations used for that step also yield a score compressor for the second step, in which ABC-PMC produces the cosmological parameter posterior. The well-specified model recovers an unbiased initial power spectrum, while the misspecified model produces an implausible posterior with roughly $2\\sigma$ excess power at large scales and a matching deficit at small scales. The paper's headline demonstration is that the same misspecification induces a bias exceeding $2\\sigma$ in the $(\\Omega_\\mathrm{m}, \\sigma_8)$ plane, so the latent-spectrum posterior catches the problem before cosmological parameter inference.","pith_inferences":["Because the diagnostic sees only misspecifications that move summary statistics along directions captured by $\\nabla_\\theta f_0$, a systematic effect whose main influence is orthogonal to that gradient would be invisible here even if it biased cosmological parameters.","A direct test of how far the diagnostic reaches would be to inject a known systematic into real or realistic survey data and measure the smallest bias in $(\\Omega_\\mathrm{m}, \\sigma_8)$ that still triggers a Mahalanobis-distance alarm calibrated on the prior.","For surveys like DESI, Euclid, and LSST, the framework's practical bottleneck is the $O(10^5)$-simulation ABC step, so the largest gain comes from letting the $O(10^3)$-simulation diagnostic screen model choices before the expensive parameter step is launched.","The same two-step latent-function logic could transfer to other theoretically predictable summaries, such as the bispectrum or wavelet coefficients, if one wants a diagnostic sensitive to systematics that miss the power spectrum."],"forward_implications":["The same simulation set used for the diagnostic can be recycled for score compression, so the pre-inference check adds no separate simulation campaign.","Individual systematic effects can be separated by their scale-dependent signatures in the inferred spectrum, even when the direct galaxy power spectra of the well- and misspecified models look nearly identical.","Gravity-solver approximations that are invisible at the percent level in galaxy power spectra can produce percent-level errors in the inferred initial power spectrum, forcing a stricter accuracy target for forward models.","Using 2LPT instead of full N-body evolution over the scales studied rejects the ground truth by almost $2\\sigma$, so non-linear gravitational evolution is required for the inferred initial spectrum to be trusted."],"supporting_citations":[{"why":"Supplies the SELFI algorithm and the Gaussian effective likelihood that produces the initial matter power spectrum posterior.","marker":"Leclercq et al. (2019)"},{"why":"Introduces the two-step Bayesian hierarchical model procedure with latent function and score compression that this paper applies to cosmology.","marker":"Leclercq (2022)"},{"why":"Provides the score compression that turns the gradient of the effective log-likelihood into compressed summaries.","marker":"Alsing & Wandelt (2018)"},{"why":"Gives the inverse-covariance correction factor used in the effective likelihood and posterior.","marker":"Hartlap, Simon & Schneider (2007)"},{"why":"Supplies the adaptive ABC-PMC variant used to obtain the cosmological parameter posterior.","marker":"Simola et al. (2021)"},{"why":"Provides the adaptive population Monte Carlo ABC foundation used in the sampler.","marker":"Beaumont et al. (2009)"},{"why":"Motivates the problem by showing that model misspecification can bias or over-concentrate ABC posteriors.","marker":"Frazier, Robert & Rousseau (2017)"},{"why":"Provides the fiducial and prior cosmological parameters used to build the prior on the initial matter power spectrum.","marker":"Planck Collaboration (2020)"}],"fun_headline_variants":["Initial spectrum flags hidden survey biases before they hit cosmology","Hidden survey biases show up in initial spectrum before they skew parameters","Spectrum catches survey errors before they contaminate cosmological inference","Inferred primordial spectrum exposes survey flaws before they bias cosmology"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that any systematic effect strong enough to bias the cosmological parameter posterior will also leave a detectable imprint on the inferred initial matter power spectrum; the paper demonstrates this for the effects it studies and explicitly notes that systematics bypassing the initial power spectrum are outside its scope.","fun_headline_variants_meta":{"raw":{"variants":["Initial spectrum flags hidden survey biases before they hit cosmology","Hidden survey biases show up in initial spectrum before they skew parameters","Spectrum catches survey errors before they contaminate cosmological inference","Inferred primordial spectrum exposes survey flaws before they bias cosmology"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001173,"raw_usage":{"total_tokens":4889,"prompt_tokens":1022,"completion_tokens":3867,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":3799}},"tokens_in":638,"tokens_out":3867,"duration_ms":21899,"temperature":1.0,"reasoning_tokens":3799,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:24:28.477282+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a forward model whose only misspecification adds a component to the galaxy power spectra that lies in the null space of $(\\nabla_\\theta f_0)^\\top C_0^{-1}$, so the inferred initial power spectrum posterior is unchanged, and show with the paper's ABC-PMC pipeline that $(\\Omega_\\mathrm{m}, \\sigma_8)$ still shifts by more than $2\\sigma$; that would refute the diagnostic premise.","supporting_citations":[],"review_version":1}