{"id":"ec3700a6-9136-44a0-82a0-813daa26e863","arxiv_id":"2608.05763","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A frozen pretrained seismic waveform encoder, SeisLM, defines an FWI misfit that broadens the basin of attraction and improves early-stage inversion on Marmousi and Overthrust benchmarks.","lead":"This paper uses a pretrained seismic AI model, SeisLM, as a fixed feature extractor to compare modeled and measured seismic traces during full waveform inversion. In synthetic tests on the Marmousi and Overthrust models, this feature-based objective avoids cycle skipping better than the standard waveform-difference objective and provides better starting models for conventional refinement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pretraining attribution is not tested in the actual FWI benchmarks: the random-encoder control appears only in the 1D and layerwise diagnostics, so the benchmark gains could in principle come from the architecture and preprocessing rather than SeisLM's pretrained weights.","rationale":"The reader's weakest assumption is domain transfer from earthquake waveforms to exploration-scale acoustic data. That is a real concern, and the paper itself flags it in Section 2.1. However, the benchmarks actually constitute a transfer test, and they succeed on the tested synthetic cases. The more direct load-bearing gap is attribution: none of the FWI inversions include a randomly initialized encoder, even though the abstract and the 1D diagnostics explicitly use that comparison to argue that pretraining, not architecture, produces the broader basin. The 1D and layerwise results are good evidence, but they are data-space sweeps; the Marmousi and Overthrust experiments are the actual claims of practical utility, and there the random-encoder control is absent. The adjoint-source derivation appears correct, the hybrid normalization is clearly described, and the benchmark results are internally consistent. This concern does not invalidate the paper; it identifies a missing control that should be run before the pretraining attribution is finalized. I would keep the conditional verdict, now conditioned on that control.","tokens_in":16926,"tokens_out":12285,"duration_ms":138024,"concrete_test":"Run the Marmousi stage-1 workflow with the same Devito modeling, time-space mask, L-BFGS settings, and stage-2 L2 refinement, using three objectives: (i) the frozen pretrained SeisLM encoder, (ii) the same SeisLM architecture with randomly initialized weights, and (iii) the conventional L2 objective, with at least five random seeds for (ii). If the stage-1 relative velocity error and final stage-2 model from (ii) are statistically indistinguishable from (i), the pretraining attribution fails; if (ii) tracks the L2 baseline, the attribution is confirmed.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the pretrained SeisLM weights, not merely the encoder architecture, are responsible for the cycle-skipping robustness. The paper's only direct pretrained-versus-random comparison appears in the 1D time-shift diagnostic (Section 3.1, Fig. 3) and the layerwise sensitivity profile (Fig. 4). The Marmousi and 2D/3D Overthrust inversions compare the SeisLM feature loss only against L2 and a scheduled hybrid; no randomly initialized encoder is run as an FWI objective. If a random feature map with the same standardization and replication preprocessing also produced stage-1 models from which L2 refinement succeeds, then the abstract's contrast between the pretrained encoder and a randomly initialized encoder would not be established in the actual setting where the method is claimed to work. The 1D evidence makes this unlikely, but it is not a substitute for the benchmark control, because the model-space optimization path can interact with the feature landscape differently than a data-space shift sweep.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a feature-space misfit for full waveform inversion (FWI) in which modeled and observed seismic traces are compared through the frozen encoder of SeisLM, a pretrained seismic foundation model. The SeisLM feature loss is differentiated with respect to the modeled data to obtain an adjoint source compatible with the standard adjoint-state framework, including the derivatives of per-trace standardization and the channel-replication operator. The authors also test a scheduled hybrid of the feature loss and the conventional L2 objective. Time-shift diagnostics in a 1D toy example and a layerwise sensitivity analysis on Marmousi data indicate that the pretrained encoder produces a broader, smoother misfit basin than both the raw L2 objective and an encoder with the same architecture but randomly initialized weights. Two-stage FWI experiments on Marmousi and 2D/3D SEG/EAGE Overthrust models compare L2, SeisLM-feature, and hybrid workflows, reporting that the SeisLM objective improves early-stage background-model recovery and provides better starting models for subsequent L2 refinement, while the hybrid works well in Marmousi but less well in 2D Overthrust.","tokens_in":17101,"tokens_out":6238,"duration_ms":59440,"significance":"If the central claim holds, the paper makes a useful contribution: it shows that a frozen, self-supervised seismic foundation model pretrained on earthquake waveforms can define an FWI misfit that mitigates cycle skipping without introducing trainable parameters during inversion. The adjoint-source derivation in Section 2.4 is presented carefully and is technically sound, and the paper includes reproducible public code and data availability, a clear computational-cost table, and honest discussion of limitations. The strongest evidence for the pretraining attribution is the 1D time-shift landscape and the layerwise sensitivity profile, both of which include a randomly initialized encoder control. The main weakness is that this control is not carried into the actual FWI benchmarks, so the benchmark improvements are not directly shown to require the pretrained weights rather than the architecture and preprocessing. The 3D experiment also alters the objective by temporal averaging without a direct comparison to the full sequence-level loss.","major_comments":[{"comment":"The central attribution of the FWI improvements to the pretrained encoder is not tested in the actual inversion benchmarks. The only pretrained-versus-random comparisons are the 1D time-shift landscape (Figure 3) and the layerwise sensitivity profile (Figure 4) in Section 3.1; the Marmousi and 2D/3D Overthrust inversions in Sections 3.2–3.3 compare L2, SeisLM, and hybrid only. Because the optimization path in model space can interact with a feature landscape differently from a data-space shift sweep, the benchmark gains could in principle be produced by the architecture and preprocessing alone rather than by SeisLM's pretrained weights. Please run the same stage-1 FWI with a randomly initialized encoder of identical architecture, at least for Marmousi and 2D Overthrust and ideally for the 3D test, or explicitly limit the pretraining-attribution claim to the Section 3.1 diagnostics.","section":"§3.1–3.3"},{"comment":"The 3D Overthrust experiment uses a modified objective: the last-layer feature matrices are averaged over time before comparison, whereas Equation (3) and the 2D experiments use the full feature sequence. No comparison between the full and averaged representations is provided, so the 3D result establishes a property of the temporally averaged feature loss, not of the sequence-level objective used elsewhere. The Discussion concedes this point ('the 2D and 3D results do not provide a direct comparison'). To support the claim that the benefit extends to 3D with the original objective, add at least a 2D experiment comparing full versus averaged feature losses, or a 3D comparison on a subset of shots.","section":"§3.3.2"},{"comment":"The random-encoder control is based on a single random initialization, and the text repeatedly qualifies it as 'the randomly initialized encoder tested here.' Since this control is load-bearing for the claim that pretraining rather than architecture creates the smooth basin, one initialization is insufficient to establish the general contrast with randomly initialized parameters. Report the range or distribution over several seeds, for example the mean and min/max envelope, for both the time-shift and layerwise diagnostics.","section":"§3.1, Figs. 3–4"}],"minor_comments":[{"comment":"Equation (3) defines the feature loss for the full feature sequence, but the 3D experiment uses a temporally averaged variant that is never written out; please add an explicit definition of the averaged loss and state where it replaces Equation (3).","section":"§2.2 and §3.3.2"},{"comment":"The citation 'A. Brougois et al. [1990]' for the Marmousi model appears to contain a misspelled author name; please verify the canonical author list.","section":"References"},{"comment":"The 1D diagnostic is a useful controlled test, but it would be clearer to specify whether the raw L2 misfit is computed on standardized traces or on the raw traces, since the feature loss always includes standardization; if the L2 baseline is unstandardized, state this explicitly to avoid ambiguity in the comparison.","section":"§3.1"},{"comment":"The loss curves are normalized per-stage and include line-search trial evaluations, while velocity-error curves include only accepted iterates; this is explained in the text, but one sentence in each caption noting this distinction would help readers avoid misreading curve lengths.","section":"Figures 7, 10, 14"},{"comment":"The claim that the feature loss 'changes only the data-comparison term' is accurate, but the 3D temporal averaging is an additional change to the objective; please make clear in this section that the 3D comparison is specifically for the averaged variant.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a geophysical journal and the code/data availability statement is a strength. The main risk is that the abstract's pretraining-attribution claim is broader than what the benchmark experiments directly test; the missing random-encoder FWI control and the unaveraged-versus-averaged 3D comparison are both fixable within the scope of a revision. I do not see signs of circularity or inappropriate use of benchmark data: the encoder is frozen and pretrained on unrelated earthquake waveforms, and no parameters are fitted to the FWI benchmarks. The recommendation is major revision rather than rejection because the central derivation is sound and the missing controls are experimental additions rather than fundamental modeling errors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious look. The core idea is genuinely new: instead of training or updating a network on an FWI objective, the authors freeze SeisLM, a self-supervised seismic foundation model, and use its feature space as the misfit. The adjoint source comes from backpropagating through the encoder, and the derivation in Section 2.4 is standard and correct, including the per-trace standardization and the replication operator. That part is clean.\n\nThe empirical work is also better than most. The 1D time-shift diagnostic and the layerwise sensitivity profile directly address the question that matters: is it the pretrained weights or just the architecture? They show the pretrained transformer blocks widen the basin and that depth alone does not. The benchmarks are consistent: feature loss beats L2 in Marmousi and Overthrust, the hybrid helps in Marmousi but hurts in Overthrust, and L2 stalls in 3D while the feature loss builds a useful background. The paper is honest about the limits—domain shift, no noise or field data, temporal averaging in 3D—and does not oversell.\n\nThe stress-test concern has some teeth but I think it is a request, not a flaw. The random-encoder control appears only in the 1D and layerwise diagnostics, not in the actual FWI benchmarks. So in principle the Marmousi and Overthrust gains could come partly from the standardization, replication, or just a feature-map geometry that is unrelated to pretraining. The 1D evidence makes that unlikely, but it is not a substitute for running a random-encoder inversion on one of the 2D benchmarks. A referee should ask for that.\n\nThe other gaps are real but not fatal. There is no comparison with non-learned robust misfits like optimal transport, dynamic time warping, or envelope-based objectives. Those are the natural alternatives, and the paper would be stronger with one. The 3D experiment also changes the loss (temporal averaging of features) relative to 2D, so the reader cannot cleanly attribute the 3D success to the same mechanism. Minor.\n\nBottom line: the paper deserves a serious referee. The idea is interesting, the math is solid, and the evidence is fairly presented. I would ask for the random-encoder benchmark control and one robust-misfit comparison before accepting, but I would not desk-reject this.","headline":"A genuinely new way to build an FWI misfit from a frozen pretrained seismic encoder, with solid experiments and one missing control that a referee should request.","tokens_in":17681,"tokens_out":1356,"would_cite":true,"duration_ms":14206,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that comparing modeled and observed seismic traces in the frozen feature space of a pretrained seismic foundation model yields a full-waveform-inversion objective with a broader basin of attraction, mitigating cycle…","keywords":["full waveform inversion","cycle skipping","seismic foundation model","feature-space misfit","adjoint-state method","pretrained encoder","SeisLM","velocity model building"],"falsifier":"Run the 2D Overthrust stage-1 workflow to 100 iterations with three random seeds, replacing the frozen pretrained SeisLM encoder with a randomly initialized encoder of identical architecture; if the relative velocity error of the random-encoder workflow is equal to or better than the pretrained one, the claim that pretraining, rather than architecture alone, creates the broadened basin is falsified.","tokens_in":16697,"feed_emoji":"🌍","tokens_out":13765,"duration_ms":112617,"temperature":0.7,"pith_summary":"Full waveform inversion (FWI) recovers subsurface velocity models by matching modeled to observed seismic data, but the conventional $L_2$ waveform misfit is trapped by cycle skipping when the starting model is inaccurate. This paper proposes replacing that comparison, during early-stage inversion, with a discrepancy in the feature space of SeisLM, a seismic foundation model pretrained on unlabeled earthquake waveforms; the frozen encoder maps standardized modeled and observed traces to contextualized feature vectors, and the squared difference of those features serves as the misfit. Time-shift diagnostics show the pretrained feature loss has a broader and smoother basin around the correct alignment than either the $L_2$ objective or the feature loss of a randomly initialized encoder, and the layerwise diagnostics locate the widened basin in the pretrained transformer blocks. In Marmousi and 2D/3D Overthrust tests, this feature-space misfit produces starting velocity models from which conventional $L_2$ refinement succeeds, whereas $L_2$ alone stalls or converges to inaccurate models. The paper's intended result is that pretrained seismic representations are useful as early-stage FWI objectives, not as complete replacements for waveform-domain misfits.","feed_headline":"Frozen seismic AI model eases full-waveform inversion cycle skipping","feed_subtitle":"Earthquake-trained encoder makes the FWI misfit smoother; L2 refinement then recovers the velocity model.","key_machinery":"The central object is the SeisLM encoder, a masked-prediction transformer pretrained on unlabeled three-component earthquake waveforms and kept frozen during inversion. Its input is a per-trace standardized single-component trace replicated across three channels, and its output is a sequence of 240-dimensional contextualized feature vectors, roughly one quarter as many as the input time samples. The misfit is $J_{\\mathrm{SeisLM}}(m)=\\tfrac12\\sum_s\\sum_r\\|F_{\\hat\\theta}(\\tilde d_{s,r}(m))-F_{\\hat\\theta}(\\tilde d^{\\mathrm{obs}}_{s,r})\\|_F^2$, and the adjoint source is the chain-rule product of the per-trace standardization Jacobian, the channel-replication transpose, and the encoder Jacobian applied to the feature residual. This construction is what lets a frozen, FWI-agnostic network define a differentiable misfit without training or fine-tuning; the empirical mechanism is that the pretrained transformer features increase monotonically over a much wider range of time shifts than the data-space $L_2$ distance, while a randomly initialized encoder of the same architecture does not.","core_discovery":"The central claim is that an FWI objective defined in the frozen feature space of a pretrained seismic foundation model has a wider basin of attraction than the conventional waveform $L_2$ objective, and that this is enough to make cycle-skipping-prone $L_2$ refinement succeed afterwards. The SeisLM encoder, pretrained by masked prediction on three-component earthquake waveforms, is kept frozen; each single-component acoustic trace is standardized over time and replicated across three channels, and the misfit is the squared Frobenius norm of the difference between modeled and observed feature matrices. Because the encoder is differentiable, the feature residual backpropagates through the standardization, replication, and encoder into an adjoint source that plugs directly into the standard adjoint-state FWI workflow. The paper reports that the pretrained feature loss produces a broader and smoother basin than both $L_2$ and the feature loss of a randomly initialized network, that in Marmousi the feature-loss and hybrid workflows give comparable improvements over $L_2$, that in 2D Overthrust the pure feature loss outperforms both $L_2$ and the hybrid, and that in 3D Overthrust $L_2$ stalls near a linear-gradient start while the feature loss builds a background from which $L_2$ refinement recovers the main structures. The intended role of the feature loss is therefore to steer early-stage inversion into the basin of attraction of the waveform misfit, not to replace the waveform misfit entirely.","pith_inferences":["A testable extension the paper leaves implicit is an adaptive hand-off rule: monitor the agreement between feature-space and waveform gradients, or the slope of the feature loss, and switch to $L_2$ only once the two objectives point in similar directions.","Because per-trace standardization removes absolute amplitudes, the feature loss probably sacrifices legitimate amplitude information; it should therefore be most reliable for kinematics-dominated early stages and less suitable for recovering amplitude-bearing parameters such as density or attenuation.","The 3D experiment's temporal averaging suggests the robust signal is distributed across the feature sequence rather than locked to individual arrivals; directly comparing full-sequence and averaged feature losses on the affordable 2D Overthrust geometry would test this.","If the cross-domain transfer from earthquake to exploration waveforms is as robust as these benchmarks suggest, the same frozen encoder could serve as a generic misfit for other seismic time-series inverse problems, though the paper does not test these settings."],"forward_implications":["Because the feature loss backpropagates into an adjoint source, a frozen pretrained encoder can be dropped into an existing adjoint-state FWI workflow without modifying the wave-equation solver or training a network.","The 2D Overthrust hybrid result implies that adding $L_2$ too early can undo the feature loss's protection, so any successful schedule must be conditioned on the current model's accuracy rather than fixed at 100 iterations.","The layerwise shift diagnostics imply that the widened basin lives in the deeper pretrained transformer blocks; using early convolutional features alone would likely retain the oscillatory behavior of $L_2$.","The 3D experiment indicates that temporally averaged feature vectors retain enough information to guide inversion, although the paper does not directly compare full-sequence versus averaged features.","The measured cost table implies that encoder forward and backward passes dominate the added overhead, and bfloat16 roughly halves full-gradient time in 3D without changing the model-error trajectory."],"supporting_citations":[{"why":"Supplies SeisLM, the frozen pretrained encoder whose feature space defines the proposed misfit.","marker":"Liu et al. [2024]"},{"why":"Provides the masked-prediction pretraining framework that SeisLM's architecture follows.","marker":"Baevski et al. [2020]"},{"why":"Establishes the adjoint-state full waveform inversion formulation into which the feature-loss adjoint source is inserted.","marker":"Tarantola [1984]"},{"why":"Defines conventional L2 FWI and the cycle-skipping failure mode the paper targets.","marker":"Virieux and Operto [2009]"},{"why":"Comparative study supporting the view that kinematic misfits are most useful in early FWI stages.","marker":"Pladys et al. [2021]"},{"why":"Introduces perceptual losses computed in pretrained feature spaces, the conceptual template for the SeisLM feature loss.","marker":"Johnson et al. [2016]"}],"fun_headline_variants":["Frozen seismic AI encoder widens full-waveform inversion basin","Pretrained seismic network smooths FWI misfit to beat cycle skipping","Frozen foundation model guides early FWI, L2 refinement finishes","Hybrid feature loss reduces cycle skipping in waveform inversion","Seismic AI feature loss widens attraction basin for FWI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that SeisLM's encoder, pretrained on 30-second three-component earthquake waveforms at 100 Hz, still represents shorter, single-component, acoustic, exploration-frequency traces in a way that avoids the cycle-skipping trap after per-trace standardization and channel replication; if that transfer fails, the feature loss has no advantage over a randomly initialized encoder.","fun_headline_variants_meta":{"raw":{"variants":["Frozen seismic AI encoder widens full-waveform inversion basin","Pretrained seismic network smooths FWI misfit to beat cycle skipping","Frozen foundation model guides early FWI, L2 refinement finishes","Hybrid feature loss reduces cycle skipping in waveform inversion","Seismic AI feature loss widens attraction basin for FWI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000363,"raw_usage":{"total_tokens":2069,"prompt_tokens":1166,"completion_tokens":903,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":782,"completion_tokens_details":{"reasoning_tokens":816}},"tokens_in":782,"tokens_out":903,"duration_ms":7047,"temperature":1.0,"reasoning_tokens":816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T23:56:39.275160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the 2D Overthrust stage-1 workflow to 100 iterations with three random seeds, replacing the frozen pretrained SeisLM encoder with a randomly initialized encoder of identical architecture; if the relative velocity error of the random-encoder workflow is equal to or better than the pretrained one, the claim that pretraining, rather than architecture alone, creates the broadened basin is falsified.","supporting_citations":[{"cited_title":"IEEE Transactions on Geoscience and Remote Sensing , volume=","cited_arxiv_id":null,"evidence_quote":"Supplies SeisLM, the frozen pretrained encoder whose feature space defines the proposed misfit."},{"cited_title":"2020 , publisher=","cited_arxiv_id":null,"evidence_quote":"Provides the masked-prediction pretraining framework that SeisLM's architecture follows."},{"cited_title":"Geophysics , volume =","cited_arxiv_id":null,"evidence_quote":"Establishes the adjoint-state full waveform inversion formulation into which the feature-loss adjoint source is inserted."},{"cited_title":"On cycle-skipping and misfit function modification for full-wave inversion: Comparison of five recent approaches , journal =","cited_arxiv_id":null,"evidence_quote":"Comparative study supporting the view that kinematic misfits are most useful in early FWI stages."}],"review_version":1}