{"id":"65213893-ac56-4008-bbb9-b35509f644c9","arxiv_id":"2411.14182","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A symmetric variational autoencoder separates coherent crustal signal from per-earthquake noise in receiver functions, producing cleaner virtual receiver functions than standard averaging.","lead":"This paper uses a symmetric variational autoencoder to strip noise-like 'nuisance effects' from receiver functions, the seismic signals used to image crustal structure. The method generates cleaner 'virtual' receiver functions from all earthquakes, including noisy ones, and could sharpen images of subduction zones and fault systems.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplementary Section 1 reports that virtual RFs lose deep converted phases (beyond ~10 s) for thicker crust because in-bin moveout breaks coherency; this directly undercuts the claim that crustal effects are preserved.","rationale":"The reader's CONDITIONAL verdict is confirmed and sharpened by evidence the reader did not emphasize: the supplementary material to this paper explicitly reports that for thicker-crust stations in the synthetic experiment, \"surface multiples beyond 10 s are not effectively obtained... due to the increase in depth of the crustal interface, which results in greater variability of the converted phases within a bin, thereby causing a loss of coherency.\" This is a direct admission that the core coherency assumption—the basis for the accumulated crustal code q̄_j in Eq. 7 and the virtual RF in Eq. 10—breaks down as structure depth grows. Since the central claim is that virtual RFs preserve crustal effects while minimizing nuisance, this is precisely a regime where the claim fails: the single latent code must blur or average the incoherent phases. The real-data MNCC (Eq. D.2) cannot reveal this failure because it measures agreement with the raw RFs, not with true structure, so the field-data demonstrations do not rescue the deep-structure claim. I agree with the reader that coherency within bins is the weakest assumption; the supplementary provides an internal, concrete failure case. The other concerns raised by the reader—nuisance-code transfer across bins (Eq. 11) and the self-referential metric—are real but secondary: they affect generality and evaluation, while the coherency failure affects the core mechanism. The proposed test—windowed NCC on a synthetic model with a deeper Moho and a mantle discontinuity—would settle whether deep crustal effects are preserved. If the test shows deep-phase loss, the paper should restrict its claims to shallow crustal imaging or adapt binning to target depth; the low-SNR coverage improvement may still stand. The verdict remains CONDITIONAL: the method is promising and the synthetic/real demonstrations are valuable, but the central claim needs restriction and a depth-dependent validation before full acceptance.","tokens_in":28006,"tokens_out":5767,"duration_ms":52212,"concrete_test":"Re-run the Section 4 synthetic experiment with a thicker crust (e.g., Moho at 60 km) and at least one deeper discontinuity (e.g., a mantle interface at 80–100 km), keeping the same bin sizes (10° backazimuth, 10° epicentral distance) and training setup. Compute windowed NCC (Eq. D.1) separately for the 0–10 s and >10 s windows for virtual, linearly averaged, and phase-weighted averaged RFs. If the virtual RFs do not outperform linear averaging in the >10 s window, or if the deeper Pds/PpS phases are visibly attenuated relative to the true RFs, then the central claim that crustal effects are preserved fails exactly in the regime flagged by the supplementary as losing coherency. Report per-station, per-window NCC values so that the depth-dependence of the method's benefit is quantified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central mechanism of SymVAE is that RFs within a backazimuth–epicentral-distance bin share coherent crustal effects, so the symmetric encoder can accumulate a single crustal code q̄_j (Eq. 7) and generate a virtual RF preserving those effects (Eq. 10). The paper's own supplementary results contradict this assumption exactly where it matters: for Stations 4–6, where the crust is thicker, \"surface multiples beyond 10 s are not effectively obtained... due to the increase in depth of the crustal interface, which results in greater variability of the converted phases within a bin, thereby causing a loss of coherency\" (Supplementary Section 1). This is not a peripheral detail: deeper crust and mantle discontinuities have larger differential moveout across a fixed bin, and the Discussion concedes that \"as travel time differences increase for deeper mantle structures, the earthquakes must be nearer,\" while the bin sizes used (10°×10° in Sec. 4; 8°×5° in Sec. 5.1) are fixed. When coherency fails, the single accumulated code q̄_j cannot represent the crustal effects; the virtual RF will blur or average them, losing the very signal the method claims to preserve. The real-data MNCC metric (Eq. D.2) cannot detect this loss because it measures agreement with the raw RFs, not fidelity to true deep structure; a smooth virtual RF can score highly simply by resembling the average of noisy inputs. Thus the strongest claim—minimal nuisance while preserving crustal effects—is unproven and, by the authors' own supplement, false for the deeper-structure regime. The method may still improve shallow crustal imaging and enable use of low-SNR events, but the claimed preservation of deep crustal effects is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a symmetric variational autoencoder (SymVAE) approach to denoise teleseismic receiver functions (RFs). Receiver functions are grouped into backazimuth–epicentral-distance bins per station, and the network is trained to disentangle a coherent crustal code shared within each bin from earthquake-specific nuisance codes. After training, a virtual RF is generated per bin by decoding the accumulated crustal code with an optimally chosen nuisance code, selected by minimizing the KL divergence between the accumulated crustal posterior and the posterior of the candidate virtual RF. The method is tested on synthetic RFs with realistic source signatures and ambient noise, where virtual RFs achieve higher normalized correlation with true RFs than linear or phase-weighted averaging. It is then applied to real data in Cascadia and southern California, with qualitative comparisons to previous studies and a quantitative mean normalized correlation coefficient (MNCC) relative to the raw RFs in each bin. The authors claim that virtual RFs contain minimal nuisance effects while preserving crustal effects, and that the method can use all available earthquakes regardless of signal quality.","tokens_in":28348,"tokens_out":4199,"duration_ms":42976,"significance":"If the central claim holds, the method would be a useful unsupervised tool for RF imaging, particularly for temporary stations and uneven earthquake distributions, and it would improve backazimuth–slowness coverage without discarding low-SNR data. The paper has clear strengths: the synthetic benchmark is externally grounded against true RFs; the comparison includes linear and phase-weighted averaging; the two real-data applications cover geologically distinct settings; and the authors state that code and data are open-access. The main limitations are that the real-data evaluation metric is self-referential (it measures agreement with the same raw RFs used for training and generation), and the supplementary material reports a loss of deep converted phases for thicker crust, which directly undercuts the claim that crustal effects are preserved. The significance is therefore conditional: the method may be useful, but the current evidence does not yet establish that virtual RFs preserve deep crustal signals better than averaging in real data.","major_comments":[{"comment":"The paper's own synthetic results show that virtual RFs fail to recover surface multiples beyond about 10 s at Stations 4–6 because greater crustal depth increases the variability of converted phases within a bin, causing a loss of coherency. This directly contradicts the core assumption in Section 2.2 that crustal effects are coherent within fixed backazimuth–epicentral-distance bins, and it also conflicts with the central claim in Sections 3.2 and 7 that virtual RFs preserve crustal effects. Since deeper crust and mantle structures are precisely where moveout variation across a fixed bin is largest, the claim must be scoped, or the binning and accumulation scheme must be adapted (for example, by slowness-dependent moveout correction or narrower bins for later arrivals).","section":"Supplementary Section 1"},{"comment":"The real-data quality metric MNCC is computed between each virtual RF and the raw RFs in the same datapoint that was used to train the network and to accumulate the crustal posterior. This metric therefore measures self-consistency with the training distribution rather than fidelity to true crustal structure; a smooth virtual RF close to the average of noisy inputs can score highly without preserving genuine converted phases. The consistently higher MNCC of virtual RFs relative to linearly averaged RFs is thus not, on its own, evidence of enhanced crustal information in real data. An independent evaluation is needed, such as comparison with known interface depths, holdout stations or events, synthetic ground truth corrupted with real noise, or consistency checks against independent geophysical constraints.","section":"Appendix D.2, Eq. (D.2)"},{"comment":"The optimization that selects the optimal nuisance code uses the same encoder that defines Q(q|rj), so minimizing the KL divergence in Eq. (9) ensures that the virtual RF's crustal posterior matches the accumulated posterior, but it does not by itself guarantee that nuisance effects are minimized in any absolute sense. In addition, Eq. (11) reuses one optimal nuisance code for all datapoints of a station, which assumes that nuisance effects are transferable across backazimuth and epicentral distance; this assumption is not justified or tested. I suggest diagnostics such as repeating the virtual-RF generation with different nuisance-code initializations and reporting the spread of the resulting RFs, and comparing station-level results when the nuisance code is taken from a bin other than the highest-count bin.","section":"Section 3.2, Eqs. (9)–(11)"},{"comment":"The synthetic setup states that the epicentral distance spans from 50° to 60° with 10°-interval bins, which yields one epicentral-distance bin, yet the text reports '36 × 6 × 6 synthetic datapoints'. The count should be 36 × 6 × 1, or else the number of epicentral-distance bins must be stated explicitly. This inconsistency affects the reproducibility of the training dataset description.","section":"Section 4"}],"minor_comments":[{"comment":"The notation Q(q | rj, μq) appears after Eq. (7) and is inconsistent with the earlier notation Q(q | rj); the extra parameter μq is not defined and should be removed or clarified.","section":"Section 3.1, Eq. (7)"},{"comment":"The caption text appears to mislabel the subplots: it says '(b) shows a reconstructed radial RF ... related to a station with a 27 km thick crust' while the figure panel labels and the surrounding text refer to different panel letters; the correspondence between panels (a)–(d) and the described crustal thicknesses should be corrected.","section":"Figure 2"},{"comment":"The union expression contains a corrupted symbol ('station »') instead of the station index κ; this should be fixed.","section":"Section 3.2, Eq. (12)"},{"comment":"The metric in Eq. (D.2) is written for yopt_j while the virtual RF in the main text is denoted yvirt_j in Eq. (10); the notation should be unified.","section":"Appendix D.2"},{"comment":"The statement that training was stopped at 190 epochs 'where the reconstruction loss of both the training and testing dataset converged' is vague; providing learning curves or a convergence criterion would improve reproducibility.","section":"Appendix E.1"},{"comment":"The phase-weighted averaging results are reported only for power ν = 0.8; a brief sensitivity statement for ν would clarify how the comparison depends on this choice.","section":"Section 4 / Figure 4"},{"comment":"There are minor typographical spacing issues, such as 'V ancouver Island' and 'normalised' in Figure 13; these should be corrected in the final version.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of a geophysical imaging journal and the proposed method is potentially useful. My main concern is that the real-data evaluation is circular and the supplementary results explicitly report loss of deep converted phases for thicker crust; these are load-bearing issues for the central claim that virtual RFs preserve crustal effects. The paper is not fatally flawed, but it needs substantial revision, including an independent real-data validation strategy and a scoping of the claims, before I can recommend acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a sensible new application of the authors' own SymVAE architecture to receiver-function nuisance reduction, with a clever KL-based scheme for generating 'virtual' RFs. The synthetic benchmark against known truth is the right way to evaluate, and it does show genuine improvement over linear and phase-weighted averaging for shallow converted phases. But the paper's strongest claim—that crustal effects are preserved even for deeper structures—is undercut by their own supplementary material, and the real-data quantitative metric is self-referential.\n\nWhat's actually new: the grouping of RFs into backazimuth–epicentral-distance bins, the accumulation of a single crustal code per bin via a conjunction of posteriors, and the KL-optimization of a shared nuisance code to produce virtual RFs. Using all earthquakes regardless of SNR is a real practical advantage, as the Cascadia and southern California examples show improved coverage. The sanity checks (neighboring-station similarity, split-bin consistency) are well chosen.\n\nWhere it goes soft. Supplementary Section 1 explicitly states that for the thicker-crust stations (4–6), surface multiples beyond 10 s are not obtained because greater depth increases in-bin moveout variability and destroys coherency. That is exactly the regime where the method's core assumption fails, and the Discussion concedes that deeper structures demand closer earthquakes while using fixed bin widths. So the claim of 'preserving crustal effects' is only established for shallow interfaces. The real-data MNCC (Eq. D.2) correlates the virtual RF against the same raw RFs used to train and generate it; that measures self-consistency, not fidelity to true structure, and a smooth virtual RF can score well simply by resembling the average of noisy inputs. The synthetic comparison is the honest metric, and it already shows the deep-structure failure. Finally, the data-availability statement says all codes are open-access, but there is no repository link; for a deep-learning method paper, missing code is a significant reproducibility gap.\n\nProportion: the central idea is plausible and the synthetic benchmark is executed cleanly. The flaw is in the over-broad claim, not in the method's basic machinery. A focused revision that narrows the scope to shallow crustal imaging, or adds adaptive binning and a demonstration on deeper synthetic models, would strengthen it considerably. An independent real-data validation (e.g., comparing to H-k estimates or known Moho depths from independent studies) would also be needed to make the quantitative claims credible.\n\nRecommendation: worth a serious referee—the method is new and potentially useful, and the synthetic evaluation is honest. But it is not ready to pass as is. I'd send it to review with the expectation of major revision, and I'd insist on code release and a re-framed scope.","headline":"A genuine and potentially useful application of the authors' own SymVAE to receiver-function denoising, but the central claim of preserving deep crustal effects is undercut by their own supplementary results, and the real-data metric is self-referential.","tokens_in":28902,"tokens_out":2756,"would_cite":false,"duration_ms":25189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A symmetric variational autoencoder that separates shared crustal effects from earthquake-specific nuisance in receiver-function gathers can generate virtual receiver functions that outperform linear and phase-weighted averaging.","keywords":["Receiver functions","Symmetric variational autoencoder","Coherent crustal effects","Nuisance reduction","Deconvolution","Crustal structure","Subduction zone","Backazimuth coverage"],"falsifier":"Generate synthetic receiver functions from a model with a dipping Moho whose dip angle is varied, and compare each bin's virtual receiver function with the true synthetic receiver function. If the normalized correlation coefficient of the virtual receiver function falls below that of phase-weighted averaging in the bins where converted-phase arrival times vary most within the bin, the coherency assumption is falsified.","tokens_in":27813,"feed_emoji":"🌍","tokens_out":16141,"duration_ms":133755,"temperature":0.7,"pith_summary":"Receiver functions — seismic traces built by deconvolving the radial or transverse component with the vertical component to isolate waves converted at crustal interfaces — are contaminated by earthquake-specific nuisance effects that survive bandpass filtering and are not Gaussian. This paper claims that a symmetric variational autoencoder can learn to disentangle these nuisance effects from the coherent crustal response in a bin of similar earthquakes, and that decoding the accumulated crustal code with an optimized nuisance code produces a 'virtual' receiver function with minimal nuisance but preserved crustal signal. On synthetic data with real source signatures and ambient noise, and on dense networks at the Cascadia subduction zone and in southern California, the virtual receiver functions show clearer converted phases and higher correlation with ground truth or reference stacks than linear or phase-weighted averaging. The practical payoff claimed is that all earthquakes — including low signal-to-noise events that are normally discarded — can be used, improving coverage of earthquake arrival directions and distances for sparse or temporary stations.","feed_headline":"Autoencoder cleans seismic traces without discarding noisy quakes","feed_subtitle":"Virtual traces expose slab and Moho phases more clearly than averaging, so sparse networks can use all events.","key_machinery":"The key machinery is the symmetric variational autoencoder with a partitioned latent space. A symmetric encoder $h_q$ maps each receiver function in a bin to a Gaussian posterior over crustal effects; these posteriors are combined by the conjunction rule $Q(q\\mid r_j) \\propto \\prod_i Q(q\\mid r^i_j)$ (Eq. 7), so the shared crustal code becomes sharper as more earthquakes are included. A separate nuisance encoder $h_p$ maps each RF to an earthquake-specific posterior, and a decoder $f$ reconstructs RFs from sampled pairs $(\\hat{q}_j, \\hat{p}^i_j)$. To form a virtual RF, the authors decode $\\bar{q}_j$ together with a nuisance code $\\hat{m}$ that minimizes the KL divergence between the bin's accumulated crustal posterior and the crustal information in the candidate RF — a direct measure of how much crustal information would be lost by replacing the bin with that candidate. This structure is what lets nuisance be reduced without assuming Gaussian statistics.","core_discovery":"The paper's central claim is that receiver functions can be generated, not just averaged, once a symmetric variational autoencoder (SymVAE) has partitioned its latent space into a crustal component and a nuisance component. For each bin of earthquakes with similar arrival direction and distance, the symmetric encoder accumulates per-earthquake crustal posteriors into a single crustal code $\\bar{q}_j$; after training, a nuisance code $\\hat{m}$ is chosen by minimizing $D_{\\mathrm{KL}}(Q(q\\mid r_j)\\,\\|\\,Q(q\\mid y^{\\hat m}_j))$, and the virtual RF is $y^{\\mathrm{virt}}_j = f(\\bar{q}_j, \\hat{m}^{\\mathrm{virt}})$ (Eq. 10). The authors claim these virtual RFs contain less nuisance while retaining crustal effects, and that they outperform linear averaging and phase-weighted averaging on both synthetic and real data. They also claim that because the network learns the nuisance distribution instead of assuming it is Gaussian, every available earthquake can be used regardless of signal quality, yielding denser coverage of arrival directions and epicentral distances.","pith_inferences":["A natural extension the paper leaves implicit is to feed the virtual receiver functions into standard $H$–$\\kappa$ stacking or common-conversion-point migration; cleaner waveforms should sharpen Moho depth and $V_p/V_s$ estimates wherever bins are sparse.","The paper reuses a single optimal nuisance code for all bins at a station; a station-by-station cross-validation on synthetic models with strong anisotropy would reveal how much structure this transferability assumption could distort.","The same disentanglement recipe could be applied to S, SKS, or PKP receiver functions, where per-event nuisance is even more severe because the source signature is less well constrained, a direction the paper notes but does not test.","A direct test of the coherency assumption would be to halve the backazimuth bin size at one station and recompute virtual receiver functions; if converted-phase arrival times shift by more than the pulse width, the original bins were smearing distinct crustal signals."],"forward_implications":["All teleseismic earthquakes, regardless of signal quality, can be included in receiver-function analysis, improving coverage of arrival directions and epicentral distances for temporary and sparse stations.","Virtual transverse receiver functions display polarity reversals more clearly than averaged transverse receiver functions, making anisotropy and dipping-interface analysis feasible in data that would normally be too noisy.","The method, trained jointly across stations with no labels, carries over to a new geological region without retuning, as shown by running the same hyperparameters on Cascadia and southern California.","Virtual receiver functions from Cascadia resolve the slab's top negative contrast and the two deeper positive contrasts consistently across stations, while virtual receiver functions from southern California show sharp converted-phase delay changes across the San Andreas and Jacinto fault zones."],"supporting_citations":[{"why":"Introduces the symmetric variational autoencoder and the KL-divergence criterion for generating virtual receiver functions.","marker":"[Bharadwaj, 2024, SymV AE]"},{"why":"Provides the variational autoencoder and evidence-lower-bound training objective that SymVAE extends.","marker":"[Kingma and Welling, 2014]"},{"why":"Supplies the Cascadia slab phase model, reference receiver-function picks, and the comparison dataset for the subduction-zone application.","marker":"[Bloch et al., 2023]"},{"why":"Defines phase-weighted stacking, one of the two conventional baselines the virtual receiver functions are compared against.","marker":"[Schimmel and Paulssen, 1997]"},{"why":"Water-level deconvolution is used to compute synthetic and real receiver functions, generating the nuisance effects the method aims to remove.","marker":"[Clayton and Wiggins, 1976]"},{"why":"PyRaysum software is used to generate the anisotropic crustal impulse responses for the synthetic validation.","marker":"[Bloch and Audet, 2023]"},{"why":"Supplies the conjunction-of-information-states principle used by the symmetric encoder to accumulate crustal information across receiver functions in a bin.","marker":"[Tarantola, 2005]"},{"why":"Provides the southern California Moho geometry reference against which the virtual receiver-function profiles are interpreted.","marker":"[Ozakin and Ben-Zion, 2015]"}],"fun_headline_variants":["Symmetric autoencoder cleans receiver functions, uses all quakes","Virtual seismograms expose crustal phases with less noise","Deep learning disentangles crustal signal from seismic noise","Symmetric VAE improves crustal images using all noisy quakes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that earthquakes arriving from similar directions and distances produce receiver functions whose crustal signals are essentially the same, so a single accumulated code can represent all of them; the paper validates this coherency only on simple synthetic models, and its own supplementary results show the virtual receiver functions degrade for deeper interfaces where converted-wave arrival times vary more within a bin.","fun_headline_variants_meta":{"raw":{"variants":["Symmetric autoencoder cleans receiver functions, uses all quakes","Virtual seismograms expose crustal phases with less noise","Deep learning disentangles crustal signal from seismic noise","Symmetric VAE improves crustal images using all noisy quakes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":3118,"prompt_tokens":1075,"completion_tokens":2043,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":691,"completion_tokens_details":{"reasoning_tokens":1975}},"tokens_in":691,"tokens_out":2043,"duration_ms":13143,"temperature":1.0,"reasoning_tokens":1975,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:26:40.645251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate synthetic receiver functions from a model with a dipping Moho whose dip angle is varied, and compare each bin's virtual receiver function with the true synthetic receiver function. If the normalized correlation coefficient of the virtual receiver function falls below that of phase-weighted averaging in the bins where converted-phase arrival times vary most within the bin, the coherency assumption is falsified.","supporting_citations":[{"cited_title":"On extracting coherent seismic wavefield using variational symmetric autoencoders","cited_arxiv_id":"2411.15613","evidence_quote":"Introduces the symmetric variational autoencoder and the KL-divergence criterion for generating virtual receiver functions."},{"cited_title":"Pyraysum: Software for modeling ray-theoretical plane body-wave propagation in dipping anisotropic media","cited_arxiv_id":null,"evidence_quote":"Supplies the Cascadia slab phase model, reference receiver-function picks, and the comparison dataset for the subduction-zone application."},{"cited_title":"Source shape estimation and deconvolution of teleseismic body waves","cited_arxiv_id":null,"evidence_quote":"Water-level deconvolution is used to compute synthetic and real receiver functions, generating the nuisance effects the method aims to remove."},{"cited_title":"Pyraysum: Software for modeling ray-theoretical plane body-wave propagation in dipping anisotropic media","cited_arxiv_id":null,"evidence_quote":"PyRaysum software is used to generate the anisotropic crustal impulse responses for the synthetic validation."}],"review_version":1}