{"id":"30f9232f-a7b2-4c3f-90dc-171d417efbd2","arxiv_id":"2411.15613","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A symmetric variational autoencoder trained on groups of teleseismic P waves can extract and synthesize coherent earthquake source signals without deconvolution.","lead":"This paper presents a variational autoencoder, SymAE, that separates seismic waveforms into a shared source signal and per-waveform path noise, then generates cleaned virtual seismograms. It demonstrates the method on deep-focus earthquake P waves, extracting source information without traditional deconvolution.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 9's conditional independence is unvalidated for real path correlations; if shared path energy leaks into s, the same-encoder objectives of eqs. 19 and 21 will not reveal the failure.","rationale":"The reader's weakest assumption correctly identifies eq. 9 as the load-bearing premise, and I agree that the self-referential evaluation through the same encoder can conceal its failure. The concern is not an internal inconsistency; it is an identifiability gap between the generative assumptions and realistic teleseismic path correlations. It is also directly addressable with a synthetic experiment that includes a common path component, which the current synthetic validation omits. The paper otherwise presents a coherent variational derivation, a plausible disentanglement mechanism, and qualitative demonstrations on synthetic and real data, so the appropriate outcome remains a conditional acceptance pending this external validation rather than rejection. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":15941,"tokens_out":5549,"duration_ms":57312,"concrete_test":"Generate a second synthetic suite mirroring Figs. 1e-1h in which each event j has a common path factor q_j shared by all receivers (e.g., x_k = s_j * q_j * r_k + noise, with q_j drawn independently of s_j and r_k receiver-specific). Train the same SymAE with time-shift transformer and recover s_j by the eq. 21 optimization. Then measure (i) normalized correlation and RMSE between the recovered y_j and the true band-limited s_j, and (ii) correlation between recovered y_j and q_j, comparing runs with and without q_j contamination. If recovered y_j correlates significantly with q_j, or source recovery degrades (e.g., Pearson r below 0.9) only when q_j is present, eq. 9 is the load-bearing assumption and the central claim fails in this regime. Repeat over several random seeds and report mean/std.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the eq. 21 optimized virtual seismogram is enriched in coherent source information and less influenced by path scattering. This claim inherits the generative factorization of eq. 9, P(x_k,x_l|s,p_k,p_l)=P(x_k|s,p_k)P(x_l|s,p_l). For teleseismic P waves from one event, receiver waveforms share path-related structure (near-source velocity heterogeneity, common mantle attenuation, source-side scattering) that has no dedicated latent variable. When eq. 9 is violated, the precision-weighted accumulation in eqs. 17-18 has nowhere to put that shared path energy except the shared code s, so the 'source' code is actually source plus common path. The synthetic test uses independent receiver-specific path effects and therefore cannot detect this. Moreover, both the informativeness metric H in eq. 19 and the eq. 21 optimization target are computed with the very encoder Q(s|.) that was trained under eq. 9; a contaminated s makes H and KL look small even when the virtual seismogram is not source-dominated. Fig. 11 is qualitative, with no external source time function or quantitative recovery error, so it does not break the circularity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a variational symmetric autoencoder (SymAE) that disentangles a set of seismograms into a shared source code and per-seismogram path codes. It introduces a KL-divergence-based informativeness metric (Eq. 19) and an optimization procedure (Eq. 21) to generate virtual seismograms enriched in coherent source information. The method is applied to synthetic data with known source time functions and to teleseismic P-wave records from deep-focus earthquakes, with comparisons to traditional deconvolution. The central claim is that SymAE extracts coherent source information without deconvolution, enabling analysis of complex earthquakes with multiple rupture episodes.","tokens_in":16163,"tokens_out":5466,"duration_ms":45810,"significance":"If the claims are substantiated, the paper offers a novel nonlinear stacking/representation framework for seismic source extraction, building on prior SymAE work and extending it with a practical informativeness criterion and a latent-space optimization for virtual seismograms. The variational formulation in Sections 2-3 is standard and internally consistent, and the synthetic experiment with known sources is a well-posed falsifiable test. The paper also states that all codes and data are open access, which is a strength for reproducibility, though no repository is provided. However, the current experimental validation is insufficient to establish the central claims, particularly because the synthetic results are only visually compared and the real-data comparison is qualitative.","major_comments":[{"comment":"The conditional independence assumption P(x_k, x_l | s, p_k, p_l) = P(x_k | s, p_k) P(x_l | s, p_l) is load-bearing for the precision-weighted accumulation in Eqs. (17)-(18). For teleseismic P waves, receiver waveforms share path-related structure beyond the source (e.g., near-source heterogeneity, common mantle attenuation, source-side scattering) that has no dedicated latent variable in the model. When Eq. (9) is violated, shared path energy has nowhere to go except the shared code s, so the learned 'source' code is actually source plus common path. The synthetic experiment uses independent receiver-specific path effects and therefore cannot detect this failure. I recommend adding a synthetic experiment with correlated or partially shared path components across receivers and measuring whether the recovered s matches the true source under those conditions.","section":"§3, Eq. (9)"},{"comment":"Both the informativeness metric H and the optimal virtual seismogram objective in Eq. (21) are evaluated using the same trained encoder Q(s|x) that was fit under Eq. (9). If s contains shared path energy, these self-evaluations will not reveal the failure; H will be small even when the virtual seismogram is not source-dominated. The synthetic validation in Fig. 9 is only visual: no quantitative recovery metric (e.g., correlation or L2 error between the optimized virtual seismogram and the true band-limited source), no error bars, and no random-seed variability. Please add quantitative metrics with multiple seeds and, if possible, an out-of-sample decoder or a separate validation encoder to break the circularity.","section":"§3.3-3.5, Eqs. (19)-(21)"},{"comment":"The real-data comparison against SCARDEC deconvolution is purely qualitative. The claim that SymAE 'captures complex features due to multiple events' is supported only by visual inspection. There is no external source time function, no quantitative similarity metric, and no uncertainty analysis. The statement that 'the half-durations of the source functions extracted using SymAE closely align with those determined from the raw seismograms' is not quantified. Please provide a quantitative comparison (e.g., cross-correlation, misfit, or at least a table of half-durations with uncertainties) and, if available, compare with independent source models for these events.","section":"§5, Fig. 11"},{"comment":"The text contradicts itself on the sign of H. Section 3.3 states that seismograms with lower H values are more informative (Eq. 19), while Section 3.4 states that 'seismograms that exhibit higher H values (refer to eq. 19) are considered informative' and Fig. 5's caption repeats this. This reversal is not cosmetic: it governs the interpretation of Figs. 4-6 and the redatuming results. Please reconcile the definition and apply it consistently throughout.","section":"§3.3-3.4 and Fig. 5"}],"minor_comments":[{"comment":"The display of the constraints in Eq. (21) is garbled; please write the zero-mean and unit-variance constraints clearly (e.g., sum_n (y_j[m])_n = 0 and (1/D) sum_n ((y_j[m])_n - mean)^2 = 1).","section":"§3.5, Eq. (21)"},{"comment":"The heatmap description is ambiguous: the text says 'higher values after normalization indicate greater similarity' while the caption says lighter colors indicate lower information loss; these are opposite unless the normalization inverts the KL values. Please clarify which quantity is plotted.","section":"§3.4, Fig. 6"},{"comment":"The paper states that SymAE is trained on 'roughly 5000 displacement seismograms (all components)' but does not specify the number of receivers, the time window, preprocessing details, network architecture (layer counts, latent dimensions), or the values of hyperparameters α and γ. Please add these details for reproducibility.","section":"§5"},{"comment":"The data availability section says all data and codes are open access, but no repository or DOI is provided. Please include a link or accession code.","section":"§8"},{"comment":"There are several reference formatting errors, e.g., the Weaver 2001 entry is garbled ('RL Weaver. Lobkis 0 i. Ultrasonic without a sources Thermal fluctuation correlations at MHz frequencies...'). Please correct the reference list.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a strong conceptual core, but the validation is currently insufficient to support the abstract's strong claims. The contradictory statements about H and the lack of quantitative metrics should be addressed before publication. The conditional-independence concern is a correctness risk that can be partly mitigated by targeted synthetic experiments. I recommend major revision rather than rejection because the issues are addressable within the manuscript's scope and the method may be sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick take: this paper has a real idea. Bharadwaj extends the SymAE/ML-VAE machinery to earthquake source imaging, adding a KL-based informativeness metric (eq. 19) and an optimal virtual-seismogram objective (eq. 21) that together let you pull a coherent source time function from a collection of noisy teleseismic P waveforms without deconvolution. That's worth something: deconvolution-based source imaging struggles with multi-rupture events, and the real-data comparisons in Fig. 11 show the method producing sharper source functions for events like Bonin and Spain than a standard technique. The variational derivation is standard and internally consistent, and the time-shift transformer (Section 4) is a sensible fix for misalignment.\n\nThe soft spots are real, but they are mostly validation gaps, not fatal flaws. The synthetic experiments compare extracted sources to ground truth only by eye (Fig. 9). No correlation coefficient, no error bars, no seed variability. It would take an afternoon to add those. There's also no comparison against a simple non-linear stacking baseline, which undercuts the 'generalizes stacking' claim. And despite the Data Availability section saying all code and data are open access, there's no repository link or instruction—that's a practical problem for reproducibility.\n\nThe stress-test concern about eq. 9's conditional independence is legitimate but maybe slightly overplayed. Yes, if receiver waveforms share path structure beyond the source—near-source heterogeneity, common mantle attenuation—the model has nowhere to put it except the shared code. And yes, the informativeness metric uses the same encoder, so a contaminated source code could fool the self-evaluation. But that's a limitation of the framework, not a hidden contradiction. The synthetic test with independent path effects does ground the basic machinery; it just doesn't test the correlated-path scenario. The author should acknowledge this explicitly and, better, run a synthetic experiment with correlated path noise or a receiver-function-style grouping.\n\nMy recommendation: with the quantitative metrics, a baseline comparison, and the code released, this belongs in a good seismology or machine-learning-for-geophysics journal after revision. As is, it's a conditional accept rather than a desk reject. Send it to a referee who knows both the source-imaging literature and VAE practice—they'll be able to judge whether the absence of external validation is a dealbreaker for the specific claim of multi-rupture resolution. I'd bring it to our reading group, though I'd pair it with a classic ML-VAE paper to give context.","headline":"A genuinely useful new formulation for extracting source time functions without deconvolution, but the paper undersells its own contribution by leaning on qualitative validation and skipping baselines.","tokens_in":16728,"tokens_out":5008,"would_cite":true,"duration_ms":40498,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A variational symmetric autoencoder disentangles shared earthquake source information from station-specific scattering, and reconstructs clean virtual seismograms without deconvolution.","keywords":["symmetric autoencoder","variational autoencoder","disentangled representation","coherent wavefield extraction","earthquake source time function","nonlinear stacking","virtual seismogram","time-shift transformer"],"falsifier":"Take the synthetic setup of Fig. 1 and add a common scattered coda, identical across all receivers in a group, on top of the per-receiver path effects; if the recovered source time function changes or contains that common coda, or if the eq. (19) ranking starts preferring contaminated traces, the conditional-independence assumption is broken and the extraction is not purely source information.","tokens_in":15699,"feed_emoji":"🌊","tokens_out":8703,"duration_ms":75137,"temperature":0.7,"pith_summary":"The paper sets out to show that a symmetric autoencoder (SymAE) can extract the coherent part of a seismic wavefield, namely the part truly shared across recordings, and separate it from per-receiver nuisance such as path scattering and noise. The model treats each group of seismograms as generated from a common source code $s$ and individual path codes $p_k$, with waveforms assumed independent given $s$. After training, a Kullback-Leibler divergence $H[x_j^k]$ between the accumulated source posterior and each individual posterior ranks how informative a seismogram is, and a latent-space optimization produces virtual seismograms with enhanced signal-to-noise ratio. Applied to teleseismic P-wave records of deep-focus earthquakes, the method recovers source time functions with multiple rupture episodes that standard deconvolution smooths out. A reader would care because the approach promises data-driven extraction of earthquake source information without assuming known path Green's functions.","feed_headline":"Neural net extracts earthquake source without deconvolution","feed_subtitle":"SymAE separates shared source from station scattering; rebuilds clean virtual seismograms","key_machinery":"The central object is the Symmetric Autoencoder (SymAE), a variational autoencoder whose latent space is split into a coherent code $s$ shared by a whole group of waveforms and nuisance codes $p_k$ for each waveform. The load-bearing mechanism is the conditional independence assumption $P(x_j^k, x_j^l | s, p_k, p_l) = P(x_j^k | s, p_k) P(x_j^l | s, p_l)$ (eq. 9), which lets the encoder factorize the approximate posterior and accumulate source information across receivers as a product of Gaussians; the resulting precision-weighted mean and variance (eqs. 17-18) are a probabilistic generalization of nonlinear stacking. A time-shift transformer module, built on spatial-transformer-style localization networks, aligns waveforms before encoding so the model can handle timing variations without explicit cross-correlation.","core_discovery":"The central claim is that SymAE learns a disentangled latent representation in which one component, $s$, accumulates all information about the earthquake source shared by the seismograms in a group, while the remaining components $p_k$ carry waveform-specific nuisance. The accumulation is done by multiplying the per-seismogram approximate posteriors, which for Gaussians reduces to precision-weighted averages (eqs. 17-18). The paper defines an informativeness measure based on the KL divergence between $Q(s|X_j)$ and $Q(s|x_j^k)$ (eq. 19), shows that low-divergence seismograms carry reliable source content, and solves an optimization (eq. 21) for the nuisance code that minimizes this divergence while enforcing sparsity and normalization, thereby generating a virtual seismogram for each earthquake. On synthetic data the extracted source signatures match the ground-truth sources, and on real deep earthquakes the SymAE-derived source time functions are clearer and capture complex multi-event ruptures that traditional deconvolution obscures, all without an explicit deconvolution step.","pith_inferences":["Editorial inference: If eq. (9) is violated, for example when receivers share near-source structure or common mantle heterogeneity, that correlated path energy will leak into $s$, and because the same encoder scores informativeness, the leakage may not be visible in the KL metric; an external comparison against back-projection or moment-tensor inversions would expose it.","Editorial inference: The same accumulation-by-conjunction recipe applies directly to receiver-based groups: treating a station's waveforms across many earthquakes as the group would let SymAE extract a coherent receiver function or site response without stacking individual teleseisms.","Editorial inference: A natural stress test is to add a synthetic common scattered coda to all receivers in a group; if the recovered source time function changes or the informativeness ranking shifts toward contaminated traces, the method's core independence assumption is falsified.","Editorial inference: The sparsity weight in eq. (21) controls a trade-off between temporal resolution and noise suppression; tuning it per target frequency band may let seismologists tailor virtual source functions to specific rupture processes."],"forward_implications":["Analysis of complex earthquakes with multiple rupture episodes becomes possible without deconvolution, because the source code $s$ is learned from the waveform population rather than inverted from a known path response.","The KL informativeness score (eq. 19) gives a principled way to rank, weight, or discard seismograms according to source content, potentially improving receiver-function and ambient-noise stacking workflows.","Virtual seismograms formed by swapping path codes between similar earthquakes (eq. 20) provide a data-driven redatuming tool, so a clear recording from one event can help visualize a smaller or more obscured event with a comparable source.","Source-similarity structure recovered from the information-loss heatmap (Fig. 6) aligns with known rupture complexity, suggesting the model can serve as an unsupervised grouping tool for earthquakes."],"supporting_citations":[{"why":"Supplies the grouped-observation variational framework and the Gaussian conjunction equations used to accumulate source posteriors.","marker":"[Bouchacourt et al., 2018]"},{"why":"Introduces the original Symmetric Autoencoder for redatuming, which this work extends with an informativeness metric and virtual-seismogram optimization.","marker":"[Bharadwaj et al., 2022]"},{"why":"Foundational variational autoencoder formulation that provides the ELBO, encoder-decoder structure, and reparameterization used throughout.","marker":"[Kingma and Welling, 2013]"},{"why":"Supplies the concept of conjunction of independent information states, which motivates the product-of-Gaussians accumulation in eq. (16).","marker":"[Tarantola, 2005]"},{"why":"Provides the SCARDEC deconvolution baseline against which SymAE-derived source time functions are compared on real earthquakes.","marker":"[Vallée et al., 2011]"},{"why":"Documents the doublet rupture complexity of the 2018 Fiji deep earthquakes, used to validate the information-loss heatmap similarities.","marker":"[Jia et al., 2020]"},{"why":"Spatial transformer networks form the basis for the time-shift transformer module that aligns waveforms before encoding.","marker":"[Jaderberg et al., 2015]"},{"why":"Attend-infer-repeat networks inspire the localization and repeated processing structure used in the time-shift transformer.","marker":"[Eslami et al., 2016]"}],"fun_headline_variants":["AI source extraction skips deconvolution","Disentangled autoencoder cleans earthquake signals","Variational SymAE pulls source without deconvolution","Seismic source separation via symmetric autoencoder","No deconvolution needed for source imaging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that once the shared earthquake source is known, the recordings at different stations carry no additional shared information, so any leftover common path, noise, or site effect would be folded into the learned source and silently treated as source signal.","fun_headline_variants_meta":{"raw":{"variants":["AI source extraction skips deconvolution","Disentangled autoencoder cleans earthquake signals","Variational SymAE pulls source without deconvolution","Seismic source separation via symmetric autoencoder","No deconvolution needed for source imaging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2629,"prompt_tokens":1012,"completion_tokens":1617,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":1549}},"tokens_in":628,"tokens_out":1617,"duration_ms":10797,"temperature":1.0,"reasoning_tokens":1549,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:06:05.235884+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the synthetic setup of Fig. 1 and add a common scattered coda, identical across all receivers in a group, on top of the per-receiver path effects; if the recovered source time function changes or contains that common coda, or if the eq. (19) ranking starts preferring contaminated traces, the conditional-independence assumption is broken and the extraction is not purely source information.","supporting_citations":[],"review_version":1}