{"id":"3ee87d68-1d0c-47f2-a922-69d7ca7e1c8b","arxiv_id":"2607.21987","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An alternating FWI-plus-diffusion scheme (SGDS) improves synthetic seismic velocity reconstruction over L2 and TV-regularized FWI on GeoFWI, Marmousi, Overthrust, and Sigsbee2A benchmarks.","lead":"This paper compares three ways to couple a pretrained diffusion model (a learned geology prior) with physics-based full waveform inversion, and finds that an alternating split scheme—seismic-data updates followed by diffusion denoising—reconstructs synthetic velocity models better than L2 or TV-regularized inversion. A generalist might read it to see how learned priors can regularize a strongly nonlinear, PDE-constrained inverse problem while keeping the wave-equation solver","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark gains may be artifacts of deployment-mode selection; the learned prior's contribution is not isolated from the deployment strategy.","rationale":"The reader's weakest assumption exactly identifies the deployment-mode transfer issue as the most fragile link. My stress-test converges on the same concern. The paper itself provides partial evidence of this fragility: the Overthrust ablation (Table 7) shows that the improvement depends heavily on the SDEdit initialization and the FWI+TV seed, and §6.3 concedes that deployment choice matters. Since the central claim's benchmark-scale portion is not robustly supported without deployment ablations, the original CONDITIONAL verdict remains appropriate. A concrete deployment-mode ablation with a trivial-prior control would settle whether the learned prior, rather than the patching scheme, drives the reported gains.","tokens_in":85,"tokens_out":1098,"duration_ms":24823,"concrete_test":"Run the Marmousi and Overthrust experiments with the deployment mode fixed to each of the four variants (direct patch, downsample-single, downsample-multi, columnwise), using the same SGDS schedule and baselines, and report mean±std PSNR/SSIM over at least 3 random seeds per variant. If SGDS does not beat L2+TV in most variants—or if the best variant's margin collapses when compared to the same variant applied to a trivial prior (e.g., a fixed smoothing operator)—then the headline benchmark gain is deployment-selection artifact, not learned-prior benefit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—SGDS improves on L2 and TV baselines in synthetic benchmarks—rests most heavily on the Marmousi and Overthrust results, where a 100×100 GeoFWI patch prior is transplanted to 321×901/401×901 models via patchwise/downsample/columnwise deployment. The paper admits (§6.3) that 'deployment choice matters' and that 'simple patchwise use can introduce stitching or scale artifacts,' yet the final headline numbers (Marmousi PSNR 21.01, Overthrust 25.36) are single runs using deployment modes selected after observing results. No error bars across deployment variants are given, and no ablation separates the effect of the learned diffusion prior from the effect of the specific patching/denoising schedule. If the chosen deployment mode itself manufactures the improvement—e.g., by smoothing or scale-matching that coincidentally benefits the metric—then the improvement is not attributable to the diffusion prior. This is a load-bearing gap because the abstract's claim of 'SGDS improves... relative to conventional L2 and TV regularized FWI' is not restricted to in-distribution GeoFWI cases; it explicitly includes benchmark-scale transfer. The theoretical local-stability bound (Eq. 12) is also asserted assuming an approximate projection and Lipschitz/coercivity without verification, but that does not sink the empirical claim; the deployment dependency does.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using a pretrained diffusion model as a learned regularizer inside PDE-constrained full waveform inversion (FWI). Three guidance strategies are compared: MPGD (FWI steps injected into diffusion inference), SDEdit (denoising an FWI seed from an intermediate noise level), and SGDS (alternating FWI likelihood updates with diffusion denoising). The core empirical claim is that SGDS improves reconstruction quality over standard L2 FWI and L2+TV FWI in clean and moderately noisy synthetic settings. Evidence includes an 18-case GeoFWI study (Table 6), benchmark-scale Marmousi and Overthrust tests, a noise-robustness study, and a Sigsbee2A salt stress test. The paper also presents a local 'projected-gradient' interpretation of SGDS via Eq. (12). Code and data links are provided.","tokens_in":18421,"tokens_out":5064,"duration_ms":55006,"significance":"If the empirical claims are robust, the paper would provide a practical recipe for using unconditional diffusion priors in strongly nonlinear wave-equation inversion, with useful cost characterization (Table 4) and an honest discussion of deployment and out-of-distribution issues. The in-distribution GeoFWI results, with mean±std over 18 cases, support a modest claim that SGDS outperforms the two classical baselines and the two other diffusion couplings, although the margins are not tested for significance. The larger benchmark-scale claims are less secure, because the deployment strategy is selected after seeing results and is not ablated or repeated. The theoretical result (Eq. 12) is asserted rather than proved. The paper is transparent about several limitations and provides reproducible code and data, which is a strength.","major_comments":[{"comment":"The benchmark-scale Marmousi/Overthrust gains are not attributable to the learned prior because the deployment mode (direct patch, downsample-single, downsample-multi, columnwise) is selected per benchmark after observing results, and the reported numbers are single runs. The paper itself states that 'the deployment choice matters' and that 'simple patchwise use can introduce stitching or scale artifacts' (§6.3). Without an ablation that varies the deployment mode while holding the diffusion prior fixed, or error bars over deployment variants, the 21.01 dB Marmousi and 25.36 dB Overthrust results could be artifacts of the chosen patching/smoothing strategy rather than of the diffusion prior. This is load-bearing for the abstract's benchmark-transfer claim.","section":"§5.4, §5.5, §6.3"},{"comment":"The local stability bound in Eq. (12) is asserted as a 'standard projected-gradient expansion' but no proof, explicit constant definitions, or validation of the assumptions are given. The assumptions that the denoiser P_θ is an approximate projection with small defect δ and that ∇J is locally Lipschitz/coercive along tangent directions are not verified for the trained denoiser or the FWI objective. As stated, Eq. (12) is a heuristic. Either provide a rigorous derivation under explicit and checkable conditions, or explicitly label the bound as an interpretive analogy rather than a theoretical guarantee.","section":"§2.1.2, Eq. (12)"},{"comment":"The diffusion start time σ1=350 is selected from Table 5 using the same family of GeoFWI test cases that later appear in the 18-case statistics of Table 6. No separate validation split is described for hyperparameter selection. Since SGDS's advantage depends on this tuned noise level, the reported mean improvements may be inflated by test-set selection. Please describe the selection procedure, or provide a clear split between cases used for tuning and cases used for final evaluation.","section":"§5.3, Tables 5–6"},{"comment":"The conditional likelihood gradient is written as ∇_xt ||y − Ru(x̂0(xt))||², but the adjoint-state expression in Eq. (21) computes the gradient with respect to the model x̂0, not with respect to xt. The RHS of Eq. (22) therefore appears to be ∇_{x̂0} log p_t(y|x̂0), missing the Jacobian ∂x̂0/∂xt. This is technically incorrect as a derivation of the conditional score. Please clarify whether an approximation is intended (e.g., ignoring the Tweedie Jacobian) and state the resulting bias, or correct the chain rule.","section":"§3.3, Eqs. (20)–(22)"}],"minor_comments":[{"comment":"The existence/convergence theorem for TV regularization is stated without proof or citation. It is standard, but a reference would help readers verify the exact conditions (e.g., the weak* sequential closedness of F).","section":"§2.1.1, Theorem 1"},{"comment":"The text says 'Tikhonov (1963) [3]', but reference [3] is Engl and Ramlau's encyclopedia entry, not Tikhonov's 1963 paper. The citation should be corrected or supplemented.","section":"§1, References"},{"comment":"The phrase 'In additional dissertation experiments...' is vague and not reproducible. Please cite the dissertation or remove the sentence, since it concerns a claimed effect on salt recovery.","section":"§5.1"},{"comment":"MPGD wall time is reported, but its PDE-call count is listed as 'n/a'. This makes the cost comparison incomplete, since the number of PDE solves is the main driver of FWI cost. Please instrument and report it.","section":"Table 4"},{"comment":"There are minor formatting issues: equation numbers (23) is inline within Algorithm 3, some references to 'L 2' have irregular spacing, and Table 7 reports results from the 'Ap3/downsample-multi' setting without a definition of 'Ap3'.","section":"Global"}],"recommendation":"major_revision","confidential_remarks":"The in-distribution GeoFWI experiment is the paper's most defensible contribution; the benchmark transfer claim needs a deployment ablation and preferably repeated runs. If the authors cannot isolate the prior's contribution from the deployment mode, they should restrict the abstract's claim to the GeoFWI setting. The theoretical section should be revised to avoid asserting Eq. (12) as a proven bound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this paper if you care about learned regularizers in nonlinear PDE-constrained inversion. What it actually contributes is a systematic, head-to-head comparison of three guidance mechanisms (MPGD, SDEdit, SGDS) inside the same FWI loop, with a common GeoFWI-trained prior and careful controls. That comparison was previously missing. The 18-case GeoFWI statistics (Table 6) are the strongest part: SGDS beats L2 and TV on mean PSNR, SSIM, and relative L2 error across held-out salt/fault/layer models, with standard deviations reported. The noise-degradation table is honest, including SGDS failing at 5 dB. The Overthrust ablation (Table 7) is a real attempt to separate the SDEdit seed contribution from the alternating loop, and the paper explicitly labels Sigsbee2A a 'workflow-level stress test' rather than a clean method comparison. Code and data are available, which adds credibility.\n\nNow the soft spots, in proportion. First, Eq. (12) is asserted as a 'standard projected-gradient expansion' but it is not proved or numerically verified; the assumptions on the denoiser and the objective are just stated. The paper itself calls this a 'local interpretation' in Table 1, so it is not a fatal flaw, but the presentation overclaims. Rewrite it as heuristic or provide evidence. Second, the Marmousi and Overthrust headline numbers are single runs, and the deployment mode (downsample-single vs columnwise, etc.) was chosen after seeing results. The paper admits deployment matters and can cause stitching artifacts. That admission is honest, but it leaves open the possibility that part of the benchmark gain comes from the patching schedule rather than the learned prior. The Overthrust ablation only covers one deployment setting; there are no error bars over deployment variants. This is the main weakness. The central in-distribution claim does not depend on deployment, so it survives. I would not call the transfer claim an artifact, but it needs an ablation where the same deployment mode is applied to a trivial prior to isolate the learned component.\n\nThird, the benchmark-scale improvement over L2 is huge on Overthrust (PSNR 12.73 to 25.36), which makes me want to see the actual convergence and whether the baselines were tuned fairly. Table 8 is sparse on baseline detail. Minor point: trained checkpoints are not included; the authors say they can be regenerated, but for a reproducibility-focused geophysics audience that is a real barrier.\n\nBottom line: this is a solid, useful empirical study. It deserves a serious referee and likely acceptance after revisions that (i) release checkpoints, (ii) add deployment-mode ablations and error bars on benchmarks, and (iii) either prove the local contraction bound or downgrade it to an interpretation. I would bring it to reading group, and I would cite it as a reference for diffusion-prior FWI baselines.","headline":"A genuinely useful empirical comparison of diffusion-prior couplings for nonlinear FWI, with a believable in-distribution result; the benchmark-scale transfer claims are real but need deployment-mode ablations before you trust them.","tokens_in":18922,"tokens_out":2393,"would_cite":true,"duration_ms":29488,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pretrained diffusion model of geological velocity patches can act as a learned regularizer for full waveform inversion, and an alternating split-Gibbs coupling beats standard L2 and total-variation baselines on synthetic benchmarks.","keywords":["full waveform inversion","diffusion generative models","learned regularization","seismic velocity inversion","score-based denoising","alternating optimization","geological prior","inverse problems"],"falsifier":"Run SGDS on a single benchmark, such as Marmousi, with all four deployment modes—direct patch, downsample-single, downsample-multi, and columnwise denoising—keeping everything else fixed, and report per-mode PSNR with repeated runs. If the advantage over TV disappears or flips sign under any deployment choice, or if the winning mode was selected after seeing results, the benchmark-transfer claim is called into question.","tokens_in":17905,"feed_emoji":"🌍","tokens_out":3927,"duration_ms":43071,"temperature":0.7,"pith_summary":"This paper argues that a pretrained diffusion model of geological velocity patches can serve as a practical learned regularizer inside the full waveform inversion (FWI) loop, without being retrained on seismic data. It compares three ways of coupling diffusion denoising with PDE-based FWI updates and finds that the split Gibbs scheme—alternating a data-misfit FWI step with a diffusion-prior denoising step—is the most consistent. On held-out synthetic models, benchmark-scale Marmousi and Overthrust tests, and noisy data down to 10 dB, this scheme improves reconstruction quality relative to standard L2 and total-variation-regularized FWI. The core claim is that learned diffusion priors can stabilize precisely the parts of FWI that classical regularizers handle poorly, at modest extra computational cost, at least on synthetic benchmarks.","feed_headline":"Diffusion prior beats classical regularizers in seismic inversion","feed_subtitle":"Alternating FWI and denoising steps lifts Marmousi and Overthrust reconstruction quality on synthetic benchmarks.","key_machinery":"The load-bearing mechanism is the denoising map of a diffusion model pretrained on 100×100 geological velocity patches, used as a learned prior through Tweedie's estimate of the clean model. In Split Gibbs Diffusion Sampling, each outer cycle solves a penalized FWI subproblem for data consistency, adds controlled noise, then applies the denoiser; the coupling strength is set by the noise level schedule. This alternation is what lets the diffusion prior repair weakly constrained structure without overwhelming the PDE-constrained likelihood update.","core_discovery":"The central claim is that separating the likelihood and prior updates—rather than threading FWI gradients into every diffusion step—makes diffusion-guided inversion stable enough to outperform classical regularizers in a nonlinear wave-equation setting. In 18 held-out synthetic inversions, the split Gibbs scheme reaches mean PSNR 21.56 ± 3.51 versus 19.96 ± 3.91 for plain L2 FWI; on Marmousi PSNR rises from 17.44 to 21.01, on Overthrust from 12.73 to 25.36, and it remains best at 20 and 10 dB noise before degrading sharply at 5 dB. The paper interprets the alternating update as a projected-gradient-like step in which the denoiser acts as an approximate projection onto a learned manifold of g","pith_inferences":["The benchmark-scale gains may depend substantially on the chosen patch deployment; a reader should treat the Marmousi and Overthrust numbers as workflow-level results, not pure evidence about the learned prior, until deployment modes are ablated with error bars.","A natural extension is adaptive noise scheduling pegged to misfit reduction, which the paper leaves for future work; it could remove the manual tuning of the diffusion start time.","Because the prior was trained only on one synthetic geological family, the method's value on genuinely out-of-distribution geology is untested; a cost-matched comparison against a diffusion model trained on the target benchmark would separate prior quality from coupling mechanism.","The local contraction argument assumes the denoiser is close to an exact projection; measuring that projection defect on real denoisers would give a quantitative stability margin."],"forward_implications":["If correct, learned diffusion priors can be added as practical regularizers alongside classical ones for FWI, improving salt, fault, and deep-layer recovery.","SGDS retains its advantage over classical baselines at moderate noise levels (20 and 10 dB) on synthetic Marmousi data, motivating noise-robust likelihood models for extreme noise.","Benchmark-scale transfer shows that deployment mode—direct patch, downsample, or columnwise denoising—is a tunable part of the algorithm, not a trivial detail.","The alternating-scheme interpretation offers a local stability rationale for why splitting works, suggesting principled ways to schedule noise levels.","Diffusion-guided FWI adds computational overhead (about 4.1× wall time in the reported cases) but remains cheaper than tightly coupled guidance strategies."],"fun_headline_variants":["Split Gibbs diffusion beats classical regularizers in seismic inversion","Alternating diffusion steps lift FWI quality on synthetic benchmarks","Diffusion-guided FWI: separating prior and likelihood boosts PSNR","Learned diffusion prior stabilizes wave-equation inversion"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmark-scale improvements rest on transferring a 100×100 patch-trained diffusion prior to larger, structurally distinct models through hand-picked deployment modes; if that transfer, rather than the learned prior itself, produces the reported gains, the central claim weakens.","fun_headline_variants_meta":{"raw":{"variants":["Split Gibbs diffusion beats classical regularizers in seismic inversion","Alternating diffusion steps lift FWI quality on synthetic benchmarks","Diffusion-guided FWI: separating prior and likelihood boosts PSNR","Learned diffusion prior stabilizes wave-equation inversion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1245,"prompt_tokens":729,"completion_tokens":516,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":448}},"tokens_in":473,"tokens_out":516,"duration_ms":5582,"temperature":1.0,"reasoning_tokens":448,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:07:31.517770+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run SGDS on a single benchmark, such as Marmousi, with all four deployment modes—direct patch, downsample-single, downsample-multi, and columnwise denoising—keeping everything else fixed, and report per-mode PSNR with repeated runs. If the advantage over TV disappears or flips sign under any deployment choice, or if the winning mode was selected after seeing results, the benchmark-transfer claim is called into question.","supporting_citations":[],"review_version":1}