{"id":"f3ef5ca0-0881-423e-8697-834adadb2538","arxiv_id":"2607.21721","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"partial","parameter_count":3,"one_line_summary":"Training a learned prior on legacy reconstructions equals one EM step of the old regularizer, freezing its belief on the operator's blind subspace and yielding overconfident, measurement-undetectable uncertainty.","lead":"Imaging models trained on old reconstructions instead of real images inherit the old method's blind spots: on directions a sensor cannot see, the uncertainty they report is the old assumption wearing a data-driven mask, and no check on the measured data can expose it. This paper proves the mechanism, gives closed-form coverage shortfalls, and shows the effect on seismic and groundwater imaging.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the acknowledged population idealization is the weakest point, but it is explicitly scoped and empirically supported.","rationale":"The reader's weakest_assumption identifies the same point I would: the posterior-sample/single-best archive idealization and the population-limit refit are the least secure premise. However, I do not regard this as a load-bearing objection that should change the verdict. The paper explicitly labels the identity a population idealization (Section 4.1), provides a separate, more robust freeze argument that requires less of the sampler, and validates the practical conclusion with capacity-gated experiments on two operators and two prior families. The single-best collapse, which is the more striking and more practically relevant branch, is proved under a very weak variational-estimator condition and demonstrated in the groundwater experiment where the regularizer is correctly specified, ruling out misspecification as the cause. The undetectability claim is carefully qualified to genuine kernel directions in Section 3.1 and Theorem A.3, and the near-null crossover is given in closed form. I found no internal inconsistency, no unacknowledged circularity, and no unsupported step that would invert the ACCEPT verdict. The concrete test above would further quantify the finite-archive approximation, but its absence is a limitation, not a fatal flaw.","tokens_in":25215,"tokens_out":16775,"duration_ms":173704,"concrete_test":"Run the released linear-Gaussian synthetic pipeline with finite posterior-sample archives of size n=1e2, 1e3, 1e4 and a flexible normalizing-flow prior; estimate the trained prior's blind-conditional precision N^T Sigma_q^{-1} N and compare it with the theoretical freeze value N^T Sigma_rho^{-1} N. If deviations decrease with n and with capacity, the population idealization is not load-bearing for the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theorems are conditional on the archive being an i.i.d. sample from the legacy posterior (or a deterministic MAP map) and on a nonparametric refit in the population limit. If real archives are approximate, Theorem 4.1's identification could fail and the freeze of Theorem 4.2 would hold only approximately. The paper states this in Section 4.1 and gates the experiments on model capacity; the empirical results in Section 5 show the predicted overconfidence survives in two deployed operators with finite archives and approximate samplers. I do not find this assumption internally inconsistent or unaddressed: it is a stated scope condition, not a hidden flaw. The machine-checked proofs, controlled capacity-matched experiments, and explicit near-null qualifications give independent support. The only residual risk is quantitative: how quickly the finite-archive approximation converges to the population law, which the paper does not characterize theoretically.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the practice of training a generative prior for an inverse problem on an archive of legacy reconstructions rather than on true images (\"prior laundering\"). In the population limit of an exact posterior-sample archive and a nonparametric refit, the curated prior is shown to equal the old regularizer advanced by one EM step (Theorem 4.1); on the operator's blind subspace, the curated prior's conditional is frozen at the regularizer's (Theorem 4.2). In the linear-Gaussian case the paper derives a closed-form coverage shortfall (Theorem 4.3), proves that no measurement-side statistic can detect or correct it when the likelihood factors through the forward operator (Theorem 4.4), and shows that single-best archives collapse the blind credible interval to zero width (Theorem 4.6). A new measurement channel de-freezes exactly the directions it resolves (Corollary 4.7). The statements are carried to near-null and nonlinear operators, and demonstrated on seismic Born imaging and groundwater flow with diffusion and normalizing-flow priors, comparing a truth-trained oracle with a legacy-trained curated prior.","tokens_in":25473,"tokens_out":10387,"duration_ms":101851,"significance":"If the results hold, the paper provides a structural, provable mechanism for a widely suspected data crime: handcrafted regularizer assumptions re-emerge with the epistemic authority of data and cannot be detected from the measurements themselves. Strengths include a core derivation that is simple enough to be checked by hand and is additionally claimed machine-checked in Lean 4, released code for all experiments, explicit near-null bounds, and a numerical protocol that gates on model capacity, so the deployed undercoverage is attributed to the training target rather than underfitting. The practical recommendation—a blind-subspace report card computed from the operator alone—is cheap and operational. The main limitation is the acknowledged population idealization behind Theorem 4.1, but it is stated as a scope condition, and the empirical effect survives in finite, approximate deployments.","major_comments":[],"minor_comments":[{"comment":"The phrase \"a frozen mean error δ≠0 only worsens it\" following (4.4) is not true for r = sρ/s⋆ > 1: C is maximized at δ = 0 for any fixed r, so a large mean error can push coverage below nominal even when the prior is wider than the truth. This is a miss-centering failure rather than the width-type overconfidence of main interest, and (4.3) already covers it; I suggest rewording.","section":"§4.2, Eq. (4.3)"},{"comment":"The paper does not quantify the finite-archive/architecture deviation from the population identity (4.1). It is honest about the idealization, but adding an explicit open-problem sentence in §4.1 would sharpen the scope, especially because the experiments are protocol-matched comparisons rather than direct estimates of Eq. (4.1).","section":"§4.1 and §5.3"},{"comment":"The zero-width collapse is proven for deterministic variational maps. Real MAP pipelines using stochastic or randomized optimizers are not exactly deterministic; conditioning on the algorithmic randomness preserves the absence of within-y posterior spread, but the statement could say this explicitly to avoid an over-literal reading.","section":"Theorem 4.6"},{"comment":"The abstract's undetectability claim is correct for a genuine blind subspace, but in the near-null regime Theorem A.3 makes the shortfall detectable in principle above n⋆ surveys. The body is careful about this distinction; an abstract-level qualifier such as \"on exactly blind directions\" would prevent overquoting.","section":"Abstract and §4.2"}],"recommendation":"accept","confidential_remarks":"I see no load-bearing error. The weakest point is the population idealization in Theorem 4.1, but the authors are explicit about it and the experiments support transfer to finite, approximate archives. I could not independently verify the Lean development from the paper text alone, but the code is released and the formal claim is scoped. The paper is likely to draw attention and is appropriate for the journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I've read the paper, and the short version is: it does what it claims, and the claims are more solid than the abstract suggests. The core identity — that a posterior-sample archive from a legacy regularizer yields, in the population limit, exactly the regularizer advanced by one EM step — is a neat observation, and the blind-fiber freeze (Theorem 4.2) is the real workhorse: the reweight is fiber-constant, so the conditional on the blind subspace is untouched. The coverage shortfall (Theorem 4.3) and the zero-width collapse for single-best archives (Theorem 4.6) follow cleanly. Having the formal core machine-checked in Lean 4 with no sorry makes this unusually trustworthy for a theory paper with heavy notation.\n\nThe experiments are careful. The matched architectures and capacity gating address the obvious confounder that the prior is just too weak to fit the archive. The groundwater example with a correctly specified regularizer is a nice control: it isolates curation from misspecification, and the collapse still shows up. The bootstrap intervals and release of code are good practice.\n\nThe soft spots are in proportion. The biggest is that the entire population law — the exact one-EM-step identity — depends on the archive containing exact posterior samples from the legacy posterior under the deployment likelihood, plus a nonparametric refit. Real archives are approximate. The authors say this in Section 4.1 and the experiments are designed to show the effect survives approximation, but there is no finite-sample theory. That's a real gap, though not a fatal one. A second, related caveat: the 'no measurement-side check' claim is exactly true only for an exact kernel. The paper is upfront about this and gives the near-null crossover in Theorem A.3, but that crossover depends on the noise level and singular values, so in practice the 'undetectable' part is a matter of degree for real operators. Again, they qualify this clearly.\n\nThe paper is honest about its limitations. It doesn't oversell. The recommendation — report which directions the operator resolves — is achievable and cheap.\n\nWho should read it: anyone working with learned priors in imaging or inverse problems, especially those using legacy reconstruction archives; also people working on calibration and simulation-based inference. It deserves a serious referee. I would accept it, with the caveat that the finite-sample convergence question should be flagged for future work rather than block publication.","headline":"A solid, honest paper that identifies a real mechanism for overconfidence in archive-trained priors; the exact identity is population-level, and the paper is worth a serious referee.","tokens_in":25987,"tokens_out":4243,"would_cite":true,"duration_ms":39004,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","65J22"],"pacs":[],"model":"deepseek-v4-flash","headline":"Training a learned prior on an archive of past reconstructions—prior laundering—freezes the old method's assumptions on the directions a survey cannot resolve, and the resulting overconfidence cannot be detected from the measurements.","keywords":["prior laundering","learned priors","inverse problems","Bayesian uncertainty quantification","blind subspace","coverage of credible intervals","EM algorithm","generative models"],"falsifier":"Build a linear–Gaussian inverse problem with a nontrivial null space, construct an archive by exact posterior sampling from a legacy Gaussian regularizer under the true measurement law, refit the curated prior nonparametrically to convergence, and score its blind-fiber credible intervals against a truth that is tighter along the null space. The paper predicts the blind conditional is exactly the regularizer's and coverage equals C = Φ(δ + z s_ρ/s_⋆) − Φ(δ − z s_ρ/s_⋆); observing a moved blind conditional, or nominal coverage while s_ρ < s_⋆, refutes the mechanism. Separately, any statistic of","tokens_in":25107,"feed_emoji":"🎯","tokens_out":13678,"duration_ms":105073,"temperature":0.7,"pith_summary":"Learned generative priors that regularize ill-posed inverse problems are often trained not on true images—scarce in seismic and medical imaging—but on archives of past reconstructions. This paper proves that such 'prior laundering' makes the reported uncertainty inherit the old method's assumptions on exactly the directions the measurements cannot resolve: the operator's blind subspace. Averaged over measurements, a posterior-sample archive is the old regularizer advanced one expectation–maximization step; that step moves belief toward the truth on resolved directions and freezes the blind conditional at the regularizer's. The blind-fiber coverage shortfall has a closed form, no goodness-of-fit or self-consistency check can reveal it, and single-best archives collapse the blind credible interval to zero width; the pattern reproduces on deployed seismic and groundwater imaging against truth-trained controls. The stakes: in data-scarce imaging, the uncertainty an archive-trained prior reports can be silent, structural overconfidence that the measurements themselves cannot certify.","feed_headline":"Training on old reconstructions hides overconfidence no data can see","feed_subtitle":"That one EM step moves belief only where the operator sees; blind directions stay frozen at the old assumption.","key_machinery":"The load-bearing object is the data-averaged posterior map T[ρ](x) = E_{y∼p⋆}[π(x|y)] = ρ(x) E_{y∼p⋆}[p(y|x)/ρ(y)] — the classical EM/NPMLE multiplicative update (Vardi–Lucy–Richardson) — applied to the legacy regularizer ρ. Curating on a posterior-sample archive computes exactly this map (Theorem 4.1). The update's reweight depends on x only through the forward image F(x), which makes it constant on each blind fiber; that fiber-constancy is the mechanism behind the freeze (Theorem 4.2), the closed-form coverage law (Theorem 4.3), the non-identifiability of the shortfall (Theorem 4.4), and the zero-width collapse under single-best archives, where flatness of the data term on each fiber leave","core_discovery":"An archive of legacy reconstructions is, in the population limit, exactly one expectation–maximization step of the old regularizer toward the truth (Theorem 4.1). Because the likelihood reaches the unknown only through the forward operator, the step's reweight is constant along every blind fiber, so the curated prior's blind conditional equals the regularizer's—frozen no matter how much curation data (Theorem 4.2). Blind credible intervals then follow C = Φ(δ + z s_ρ/s_⋆) − Φ(δ − z s_ρ/s_⋆), under-covering whenever the inherited spread s_ρ is tighter than the truth's s_⋆ (Theorem 4.3); two truths differing only there induce identical data laws, so no measurement-side statistic can detect or","pith_inferences":["The one-EM-step identification suggests a concrete diagnostic the authors do not state: compare an archive-trained prior's blind-fiber spread with the legacy regularizer's—a close match is direct evidence the freeze is operating, while a mismatch reflects model capacity or sampling error rather than data support.","The coverage law is monotone in the ratio of inherited to true spread, which implies a testable ordering: sparse or smoothness-regularized archives (which suppress variation everywhere) should produce sharper under-coverage than conservative or over-regularized ones; the paper's two experiments follow this ordering but do not claim it as a universal law.","The undetectability result transfers to any amortized Bayesian pipeline whose training targets are outputs of a surrogate—not just imaging—so the audit (report which directions the measurements resolve) carries over to other non-identified inverse problems where 'ground truth' is itself a reconstruction."],"forward_implications":["A credible interval an archive-trained prior reports on any direction the operator leaves unresolved is the old method's belief in new clothing: its coverage is fixed by whether that belief is tighter than the truth, not by anything the data say.","No goodness-of-fit test, held-out check against recorded data, or simulation-based calibration can certify an archive-trained prior's blind-subspace uncertainty—a clean calibration report does not mean the uncertainty was earned from data.","Archives that store a single best reconstruction per survey—the common practice in seismic and medical imaging—produce zero-width credible intervals on blind directions; the estimator has nothing to say about them, and reporting that honestly would require discarding the uncertainty readout.","The repair the theory licenses is structural rather than statistical: compute the resolved/blind split from the operator alone, flag blind-subspace intervals as prior-supplied, and add a genuinely new measurement channel (e.g., well logs) to de-freeze exactly the directions it resolves."],"fun_headline_variants":["Prior laundering makes uncertainty overconfident, and no data can tell","Old reconstructions secretly freeze prior uncertainty into blind overconfidence","Learned priors from legacy data inherit undetectable overconfidence","Blind directions in priors hide overconfidence from all data checks","One EM step lands prior belief, but only where the operator sees"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central theorems assume the archive behaves like an exact i.i.d. sample from the legacy regularizer's own posterior under the deployment measurements, so that the empirical archive converges to the one-EM-step population law; real archives hold approximate, heuristic, or single-best reconstructions, and the paper itself flags this as a population idealization (Section 4.1).","fun_headline_variants_meta":{"raw":{"variants":["Prior laundering makes uncertainty overconfident, and no data can tell","Old reconstructions secretly freeze prior uncertainty into blind overconfidence","Learned priors from legacy data inherit undetectable overconfidence","Blind directions in priors hide overconfidence from all data checks","One EM step lands prior belief, but only where the operator sees"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1252,"prompt_tokens":814,"completion_tokens":438,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":558,"tokens_out":438,"duration_ms":5346,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:54:17.566131+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a linear–Gaussian inverse problem with a nontrivial null space, construct an archive by exact posterior sampling from a legacy Gaussian regularizer under the true measurement law, refit the curated prior nonparametrically to convergence, and score its blind-fiber credible intervals against a truth that is tighter along the null space. The paper predicts the blind conditional is exactly the regularizer's and coverage equals C = Φ(δ + z s_ρ/s_⋆) − Φ(δ − z s_ρ/s_⋆); observing a moved blind conditional, or nominal coverage while s_ρ < s_⋆, refutes the mechanism. Separately, any statistic of","supporting_citations":[],"review_version":1}