{"id":"29261b53-3f83-47cd-89c6-12a08d1c287a","arxiv_id":"2608.06893","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"The optimal bridge reference is v = x*(T) P, proportional to the destroyed-information spectrum, but real image statistics break this prediction and favor white noise.","lead":"This paper derives the optimal reference noise spectrum for Schrödinger bridge models in a linear-Gaussian setting, showing it should be proportional to the information destroyed by the sensor. The theory works exactly in Gaussian problems, but real image statistics invert the predicted ordering and favor white noise.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's 'every optimal spectrum is P_k-proportional' requires a singleton arg-min of Phi; uniform-grid uniqueness is open for general T and non-uniform grids already refute it, so the abstract overclaims.","rationale":"The reader's strongest claim already inserts 'whenever arg min Phi is a singleton,' and the reader's weakest assumption is the Gaussian per-mode model. I agree that the Gaussianity assumption is the empirical boundary and that the FFHQ studies honestly locate the inversion. But the most load-bearing technical premise inside the paper's own model is exactly that singleton condition. Without it, even with Gaussian per-mode conditionals and exact drift, the headline statement is not a theorem. The paper proves uniqueness only for a finite range of uniform T and shows the condition fails on admissible non-uniform grids, so the unqualified abstract claim is not supported. The concrete Sturm-sequence check would settle whether the uniform-grid claim is true for all T; if it fails, the central theorem needs a major caveat, and if it passes, the verdict remains CONDITIONAL because the paper must state the qualifications and supply artifacts. Thus I recommend no change from the reader's conditional verdict: the concern is real but addressable.","tokens_in":37535,"tokens_out":16183,"duration_ms":154415,"concrete_test":"Use exact rational arithmetic to compute the derivative of H(x)=D0/P from Eq. (5) on the uniform grid and count positive roots of its numerator polynomial via Sturm sequences for each T=121,...,1000, using the same method applied for T<=120 in App 6.4. If any T has two or more global minima of Phi, instantiate two modes with P=(1,4) and place v1/P1 and v2/P2 at two distinct minimizers; if that allocation is a global optimizer, the 'every optimal spectrum is proportional' claim is false. Separately, on the Prop. 2 two-cluster grid rho=(1/(1+C),1/2,1), solve Eq. (74) for a C>=100 at which the two local minima have exactly equal H; an exact rational equality would immediately produce a non-proportional global optimizer.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is Theorem 3(iii): the conclusion v*_k = x* P_k for every mode follows only if the mode-independent profile Phi has a singleton global minimizer. Proposition 2 shows how fragile this is: uniqueness is proved on uniform grids only for 2<=T<=120, numerically certified for a few larger T, and extending the proof to general T remains an open problem. On any grid where Phi has two global minima, the sum objective can put different modes in different minima, giving non-proportional global optimizers; the paper itself gives a two-cluster T=3 construction with at least two local minima and a two-mode instance with v2/v1 of order 10^6 while P2/P1=4. Because the abstract states 'every optimal noise spectrum is proportional to Pk' without the singleton-arg-min and uniform-grid qualifications, the advertised 'calculation from the known degradation model' is stronger than the theorem. The result also depends on a shared, non-adaptive grid; under the z-optimal schedule of Theorem 4(a) all colors are equivalent, so the P_k law is not an invariant prediction of the model but a statement about a fixed schedule convention.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops PRISM, a theory for designing the reference process in Schrödinger bridge image-restoration models. In a linear-Gaussian per-mode degradation model, the authors characterize which time-varying Gaussian references remain exactly tractable (those with commuting instantaneous covariances), prove an invisibility principle (with exact drift and unlimited solver steps the terminal law converges to the true posterior independently of the reference), and derive the exact finite-step terminal KL in closed form. The central structural result is that the per-mode KL depends only on the dimensionless ratio x = v/P, so the objective separates; whenever the mode-independent profile Phi has a unique global minimizer, every global optimizer satisfies v_k = x*(T) P_k, with x*(T) = (2 ln T)^{-1/2}(1+o(1)) on uniform grids. A z-identity shows that color and temporal schedule are exchangeable under per-mode schedule adaptation, and extensions cover drift errors, ridge regularization, and budget constraints. Experiments on exact recursions and learned Gaussian models confirm the predicted orderings and closed-form floors. On FFHQ, the distortion-perception trade-off and spectral localization transfer, but white noise outperforms the matched P_k-proportional reference; a pre-registered training-regime study and a 2x2 mechanism study trace the inversion to the non-Gaussian per-mode statistics of real images.","tokens_in":37692,"tokens_out":11811,"duration_ms":101847,"significance":"If the results are correct, this is a significant contribution to bridge-reference design: it replaces heuristic sweeps with a closed-form calculation in the Gaussian regime, provides a rigorous finite-step objective with explicit asymptotic laws, and carefully delineates where the Gaussian theory breaks down on real data. The paper is unusually rigorous for an ML submission: the main theorems have appendix proofs with exact recursions and explicit error bounds, numerical certificates are used to corroborate asymptotic claims, and the experiments include pre-registered predictions, common random numbers, and controlled mechanism studies. The learned Gaussian experiments quantitatively confirm the theoretical floors and orderings. The main weakness is that the headline claim is stated more strongly than the theorems support: the proportionality v_k = x* P_k requires a unique minimizer of Phi, which is proved only for uniform grids with 2 <= T <= 120 and numerically certified for a few larger T, and it fails on non-uniform grids. The paper also depends on a fixed shared schedule; under the z-optimal schedule of Theorem 4(a) all colors are equivalent.","major_comments":[{"comment":"The abstract states that PRISM proves 'every optimal noise spectrum is proportional to P_k' and that the optimal constant is x*(T) = (2 ln T)^-1/2(1+o(1)). This overstates what Theorem 3 establishes. Theorem 3(iii) gives v*_k = x*(grid,T) P_k only when arg min Phi is a singleton. Proposition 2 shows that on uniform grids uniqueness is proved only for 2 <= T <= 120, numerically certified for T <= 200 and a few larger values, and is false on some non-uniform grids, on which different modes can occupy different valleys of Phi and yield non-proportional optimizers. Thus the unqualified 'every optimal' claim in the abstract and conclusion is stronger than the proven result. Please qualify the claim with the conditions of a fixed shared grid and a unique minimizer of the mode-independent profile, and state the current proof range for uniqueness.","section":"Abstract; §3.3, Theorem 3(iii); §6.4, Proposition 2"},{"comment":"The paper's central proportionality claim rests on the singleton-arg-min property, but the appendix does not provide the promised certificates. The text says that for each integer 2 <= T <= 120 an explicit polynomial P_T in Z[x] is derived whose positive roots are exactly the critical points of H, with one sign change, and that numerical certification covers T up to 200 and T in {500, 1000, 5000}; however, no polynomial, certificate, or generating code is included. Without these data the reader cannot verify the uniqueness condition on which the 'every optimal' statement depends. Please provide the certificates or a script that generates them, and describe the exact algorithm and precision used for the numerical certification beyond T = 120.","section":"§6.4, 'Uniqueness (Proposition 2)'"}],"minor_comments":[{"comment":"The abstract should clarify that the proportionality law v_k = x*(T) P_k applies to a fixed shared schedule. Theorem 4(a) shows that under per-mode schedule adaptation the schedule-optimal deficit is color-independent, so the P_k law is a statement about a schedule convention rather than an invariant model prediction. The paper does state this in Section 3.3, but the abstract and introduction should carry the same qualification to avoid misleading readers.","section":"Abstract; §3.3, Theorem 4(a)"},{"comment":"The matched reference on FFHQ uses x*(T=50)=0.404 for all evaluated NFE values (5, 10, 20, 50, 100), but Proposition 1 gives a T-dependent constant. At NFE values other than 50, the matched reference is therefore not at the theory's predicted optimal scale for that budget. Please state this explicitly, or use per-NFE constants when the goal is to test the predicted ordering at each budget.","section":"§4.3 and §7.2"},{"comment":"The paper's statement that the pre-registered study 'refutes ridge whitening as the explanation' is carefully qualified in Section 7.6 by the observation that low-frequency fingerprint error persists at 300k steps and 'the ridge-whitening law remains untested at true convergence.' This is honest, but the abstract's phrasing may be read as a stronger refutation; consider adding the convergence caveat to the abstract or conclusion.","section":"§7.6, 'Convergence audit'"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a machine learning journal and the main mathematical derivations appear sound. The key issue is that the abstract and conclusion overclaim the proportionality result relative to Theorem 3 and Proposition 2; this is fixable with explicit qualifications. I would also ask the authors to make the uniqueness certificates available, since the 'every optimal' claim depends on them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Take this paper seriously, but referee the abstract with a pen. The core contribution is real: in the linear-Gaussian per-mode model, PRISM converts bridge reference selection from a heuristic choice into a calculation. Theorem 1's commuting-covariance characterization is clean; Theorem 2's invisibility result is a useful negative; the closed-form finite-step KL and the P_k-proportional optimum with x*(T) = (2 ln T)^-1/2 are derived, not fitted. I checked the appendix enough to see the proofs are complete and the numerics are exact recursion checks rather than Monte Carlo. The learned Gaussian experiments confirm the ordering; the FFHQ section is unusually honest, with a pre-registered prediction that fails and a 2x2 study tracking the inversion to non-Gaussian per-mode statistics. That is good scientific behavior.\n\nSoft spots, in proportion. The abstract's 'every optimal noise spectrum is proportional to P_k' is stronger than Theorem 3(iii), which requires arg min Phi to be a singleton. That condition is real: non-uniform grids give counterexamples, and even on uniform grids uniqueness is proved only for 2<=T<=120, with numerical checks for larger T. So the 'calculation from the known degradation model' should be phrased with those qualifiers. It is a fixable overstatement, not a hole in the argument. Similarly, the theory deliberately idealizes independence and Gaussianity per mode; the authors are upfront that FFHQ inverts the ordering and that this is the limit. I would ask for code/data/checkpoints before accepting the empirical claims as reproducible, but the experiments are already described in enough detail to be credible.\n\nCitation pattern: prior work on colored noise is empirical, and the paper locates itself properly. I don't see a circularity problem: the optimal reference is derived from the degradation model, then verified independently.\n\nWho gets value: anyone working on Schrödinger bridge restoration or noise schedule design. It deserves a serious referee. My recommendation: send it to review, and require the abstract to match the theorem conditions and a reproducibility statement.","headline":"PRISM is a genuine derivation of bridge reference design in a linear-Gaussian model, with careful experiments and honest boundary statements; the main caveat is that the headline result needs the singleton-argmin/uniform-grid qualifiers the abstract omits.","tokens_in":38259,"tokens_out":1733,"would_cite":true,"duration_ms":15907,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Bridge noise optimum is the destroyed-information spectrum","keywords":["Schrödinger bridge","reference process design","noise color spectrum","inverse problem restoration","finite-step sampling","Wiener filter","Gaussian model","distortion-perception tradeoff"],"falsifier":"Train the same bridge model on synthetic images whose per-mode conditional distributions are exactly Gaussian with the same spectrum and degradation as the real-image experiments; PRISM predicts the matched reference must beat white noise at low NFE. If white noise still wins, the Gaussian per-mode assumption is not the operative failure; if matched wins, the real-image inversion is caused by non-Gaussianity, confirming the paper's mechanism study.","tokens_in":37275,"feed_emoji":"🧮","tokens_out":7322,"duration_ms":62070,"temperature":0.7,"pith_summary":"PRISM asks a question that Schrödinger bridge models previously left to intuition: when a bridge restores a degraded signal, what noise should the reference process inject at each spatial frequency? In a linear-Gaussian model of each frequency, where the observation is $x_1 = h x_0 + n$ and the posterior residual variance is $P_k$, the paper proves that the reference is invisible in the infinite-step, exact-drift limit, so reference choice matters only under a finite solver budget. For a fixed number $T$ of solver steps, the finite-step objective separates by mode, and every global optimizer is proportional to the destroyed-information spectrum, $v_k^* = x^*(T) P_k$, with the same mode-independent constant $x^*(T) = (2\\ln T)^{-1/2}(1+o(1))$. It also proves a $z$-identity making noise color and temporal schedule exchangeable, and that regularization shifts the optimum toward white noise; experiments confirm these closed-form predictions in Gaussian settings, while a controlled study on real face images shows the matched reference stops winning and traces the breakdown to non-Gaussian per-mode statistics.","feed_headline":"Bridge noise optimum is the destroyed-information spectrum","feed_subtitle":"In Gaussian settings the paper derives the optimal reference spectrum from the known degradation model, replacing hand-tuning.","key_machinery":"The machinery is the per-mode linear-Gaussian model plus four exact structures: the Wiener residual spectrum $P_k = S_k N_k/(h_k^2 S_k + N_k)$, which measures how much information the sensor destroyed; Theorem 1's characterization that exactly tractable time-varying references are precisely those with pairwise-commuting instantaneous covariances, equivalently a fixed modal basis that decouples the bridge into independent scalar modes; the closed-form variance-deficit $D_0 = vP^3 \\sum_{i=1}^T (\\Delta\\rho_i)^2 / (\\rho_i \\varphi(\\rho_i) \\varphi(\\rho_{i-1})^2)$; and the $z$-identity $z(\\rho) = v\\rho/\\varphi(\\rho)$, which converts the deficit into $P \\sum_i (\\Delta z_i)^2/z_i$ and proves the color-schedule exchangeability. The scale symmetry $\\mathrm{KL}(v,P) = \\Phi(v/P)$ is the algebraic heart of the argument: it forces every global optimizer to share one mode-independent constant $x^*(T)$.","core_discovery":"The central claim is a pair of theorems about reference design in Gaussian Schrödinger bridges. First, with the exact drift and an unlimited number of solver steps, the reference is invisible: every admissible reference variance $v>0$ converges to the true posterior, so any reason to prefer one reference must come from finite steps, model error, or finite data. Second, for a fixed $T$-step budget the per-mode KL separates and is scale-invariant, depending on the reference only through $x = v/P$; whenever $\\arg\\min \\Phi$ is a singleton, every global optimizer is $v_k^* = x^*(T) P_k$, proportional to the Wiener residual spectrum $P_k$, with $x^*(T) = (2\\ln T)^{-1/2}(1+o(1))$ on uniform grids. The paper also proves that color and temporal schedule are interchangeable through a $z$-identity, that a fixed total budget bends the exponent from $P_k$ to $P_k^2$, and that shared ridge regularization shifts the optimizer toward white noise with $v_k^* = x^*(nP_k)P_k$.","pith_inferences":["Because the proportionality law uses only the spectra $S_k$, $h_k$, and $N_k$, it should extend to any linear inverse problem with a known degradation operator; testing it on inpainting or compressed sensing would be a direct check.","The $z$-identity suggests that temporal schedule and noise color can be optimized independently even outside the Gaussian regime, since schedule adaptation already removed part of the white-versus-matched gap in the experiments; whether $z$-optimal grids close the real-image gap at larger step budgets is a testable extension.","The experimental finding that mild $P_k$-coloring helps at low NFE while white noise wins at high NFE implies a practical recipe the theory itself does not claim: use partial coloring in the few-step regime and white otherwise; that is an empirical extrapolation.","A deeper open question the paper raises is whether, for non-Gaussian per-mode conditionals, the optimal reference is still some functional of those conditional distributions; measuring the reference that minimizes empirical KL per mode on real images would begin to answer it."],"forward_implications":["Reference design becomes a calculation in the Gaussian regime: from the known signal, blur, and noise spectra, the optimal reference is $v_k^* = x^*(T)P_k$ with $x^*(T) = (2\\ln T)^{-1/2}(1+o(1))$, replacing a hyperparameter sweep.","In the exact-drift, unlimited-step limit no reference is better than any other, so reported color advantages must be re-examined for step-budget or schedule confounds.","Noise color and temporal scheduling are exchangeable: with per-mode-adapted schedules discretization error is color-blind, and the unique optimal schedule is given by the explicit recurrence $2a_{i+1} = 1 + a_i^2$, $a_1 = 0$, with a per-mode KL floor of $4/T^2$.","Under a fixed total noise budget the optimal allocation bends from $P_k$ to $P_k^2$, so equal-budget comparisons, the usual empirical protocol, silently change the optimization problem and should be reported separately.","Shared ridge regularization whitens the optimal reference: with effective sample size $nP_k$, the optimum is $v_k^* = x^*(nP_k)P_k$, so weakly observed frequencies need disproportionately more reference noise."],"supporting_citations":[{"why":"Supplies the image-to-image Schrödinger bridge sampler whose reference design is the paper's object of study.","marker":"[31]"},{"why":"Establishes the diffusion Schrödinger bridge framework the paper builds on.","marker":"[9]"},{"why":"Formulates the bridge-matching training procedure whose reference spectrum PRISM optimizes.","marker":"[38]"},{"why":"Provides the Gaussian conditioning formula used to derive all conditional bridge laws and terminal KL expressions.","marker":"[1]"},{"why":"Supplies the simultaneous diagonalization theorem behind the commuting-reference characterization.","marker":"[18]"},{"why":"Gives the Gaussian KL divergence formula that defines the finite-step objective.","marker":"[7]"},{"why":"Provides Descartes' rule of signs used in the exact uniqueness certificates for the optimal scale.","marker":"[2]"},{"why":"Supplies total-variation bounds used in the asymptotic expansion of the proportionality constant.","marker":"[36]"}],"fun_headline_variants":["Optimal bridge noise equals destroyed information","Reference noise optimal is the Wiener residual spectrum","In Gaussian bridges, optimal noise is proportional to P_k","PRISM proves optimal reference: spectrum of destroyed information","Finite-step bridge reveals optimal noise spectrum"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every spatial frequency is an independent scalar Gaussian conditional on the observation, with known spectra $S_k$, $h_k$, $N_k$ and an exact or scale-equivariant drift; the paper's own real-image experiments show the matched-reference prediction fails once per-mode conditionals become non-Gaussian.","fun_headline_variants_meta":{"raw":{"variants":["Optimal bridge noise equals destroyed information","Reference noise optimal is the Wiener residual spectrum","In Gaussian bridges, optimal noise is proportional to P_k","PRISM proves optimal reference: spectrum of destroyed information","Finite-step bridge reveals optimal noise spectrum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1620,"prompt_tokens":1060,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":490}},"tokens_in":676,"tokens_out":560,"duration_ms":4853,"temperature":1.0,"reasoning_tokens":490,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:28:59.139768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same bridge model on synthetic images whose per-mode conditional distributions are exactly Gaussian with the same spectrum and degradation as the real-image experiments; PRISM predicts the matched reference must beat white noise at low NFE. If white noise still wins, the Gaussian per-mode assumption is not the operative failure; if matched wins, the real-image inversion is caused by non-Gaussianity, confirming the paper's mechanism study.","supporting_citations":[{"cited_title":"Diffusion Schrödinger bridge with applications to score-based generative modeling.Advances in Neural Information Processing Systems, 34:17695–17709, 2021","cited_arxiv_id":null,"evidence_quote":"Establishes the diffusion Schrödinger bridge framework the paper builds on."},{"cited_title":"Diffusion Schrödinger bridge matching.Advances in Neural Information Processing Systems, 36:62183–62223, 2023","cited_arxiv_id":null,"evidence_quote":"Formulates the bridge-matching training procedure whose reference spectrum PRISM optimizes."},{"cited_title":"Anderson.An Introduction to Multivariate Statistical Analysis","cited_arxiv_id":null,"evidence_quote":"Provides the Gaussian conditioning formula used to derive all conditional bridge laws and terminal KL expressions."},{"cited_title":"Horn and C.R","cited_arxiv_id":null,"evidence_quote":"Supplies the simultaneous diagonalization theorem behind the commuting-reference characterization."},{"cited_title":"Cover and J.A","cited_arxiv_id":null,"evidence_quote":"Gives the Gaussian KL divergence formula that defines the finite-step objective."},{"cited_title":"Springer, 2006","cited_arxiv_id":null,"evidence_quote":"Provides Descartes' rule of signs used in the exact uniqueness certificates for the optimal scale."}],"review_version":2}