{"id":"d94a1e7e-2ee5-4ea1-9c66-24db9c70d39e","arxiv_id":"2504.14830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A single graph mixture density network trained on random 2D velocity models and station layouts can produce posterior velocity maps for synthetic and real seismic datasets of different sizes and scales.","lead":"A team trained a graph-based neural network to turn seismic travel time readings into maps of underground velocity, then showed the same network can be applied to new regions with different sizes and station layouts. The paper demonstrates fast probabilistic imaging, but only for simple 2D problems and without quantitative benchmarks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comparable solutions' claim is unsupported quantitatively: comparisons are visual only, and the authors' own Section 3.3 concedes the MDN and McMC results differ, with no metric or calibration check reported.","rationale":"The reader's verdict of CONDITIONAL is appropriate. I agree with the reader that the scaling/representativeness assumption in Section 3.1 is load-bearing and untested against data generated outside the 2D path-average model. However, the more immediate and decisive gap is that the paper's core quantitative claim, comparable solutions, is never quantified; the comparisons are visual only, and the authors explicitly concede differences likely due to insufficient training. This is not an internal contradiction that destroys the method, but it is a missing measurement at the center of the claim. My proposed test would settle it directly. I therefore keep the CONDITIONAL verdict rather than upgrading or rejecting: the method may work, but the submission currently lacks the evidence needed to accept it. I flagged the manuscript's own self-reflective limitation statements (Sections 3.3 and 3.4) rather than dismissing them, as instructed. No ad hominem intended; the concern is about the evidence, not the authors. The lack of code/data and fixed 11x11 grid further limit the all-problems claim, but those are secondary to the missing quantitative validation.","tokens_in":11055,"tokens_out":7523,"duration_ms":74803,"concrete_test":"Run a quantitative comparison on the South England, Southwest China, and synthetic 20x20 km tests: per-period RMSE and mean absolute error between the graph-MDN and McMC/SVGD posterior mean maps, mean absolute difference of posterior standard deviation maps, and a calibration measure such as the fraction of McMC/SVGD posterior samples inside the MDN 68% and 95% credible intervals (or the energy distance between the two predictive distributions). Report Monte Carlo convergence of the reference methods. If MDN-reference discrepancies are comparable to run-to-run McMC variability, the comparable claim is supported; if they are larger, the paper must quantify and train further before the claim can stand.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract, Section 5) is that graph MDNs produce solutions 'comparable' to McMC/SVGD in seconds. In Sections 3.3-3.4 the evidence is side-by-side visual inspection of mean and standard deviation maps. No numerical discrepancy, calibration, or coverage metric is given. The authors themselves state in Section 3.3 that the McMC maps show higher magnitude of velocity variations and lower uncertainty and attribute this to insufficient training, and Section 3.4 repeats a similar caveat. Because the central quantity of the paper is comparable posterior estimates, an admitted, unquantified difference in exactly that quantity leaves the main claim unestablished. A second, independent load-bearing assumption is the Section 3.1 claim that path-averaged velocity is scale-invariant and that real surface-wave data obey the same 2D fast-marching model used for training; the real-data comparisons are against Bayesian inversions using the same kind of 2D model, so they cannot test this domain assumption. Either untested layer would, if wrong, invalidate the claim that the network provides the true Bayesian posterior without retraining.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a graph mixture density network (graph MDN) for 2D seismic travel-time tomography, trained on synthetic velocity models with random station distributions, and applies it to problems with different scales and acquisition geometries. The central claim is that a single trained network can provide posterior estimates 'comparable' to traditional Bayesian methods (McMC, SVGD) in seconds for unseen sites, as demonstrated on synthetic tests and two real datasets (South England Love-wave group velocities and Southwest China Rayleigh-wave phase velocities). The authors report computational costs and discuss limitations, including a fixed regular parameterization grid and possible insufficient training.","tokens_in":11289,"tokens_out":2583,"duration_ms":26657,"significance":"If the central claim holds, the work would provide a fast, amortized approximation to Bayesian seismic tomography that generalizes across acquisition geometries and scales, which is a practically valuable contribution. The method's strengths include a flexible graph representation of travel-time data, random training geometries that go beyond fixed-array networks, and two real-data demonstrations with reported runtime comparisons. However, the claim of 'comparable' posterior solutions is currently supported only by qualitative visual inspection, and the underlying domain assumptions about 2D ray theory and scale invariance are not tested against data that violate them. The paper would be strengthened substantially by quantitative accuracy/calibration metrics and a more carefully scoped statement of applicability.","major_comments":[{"comment":"The central claim that graph MDNs provide 'comparable' solutions to McMC and SVGD is supported only by side-by-side visual comparison of mean and standard deviation maps. No quantitative discrepancy measures (e.g., mean absolute difference, correlation, structural similarity) or posterior calibration metrics (e.g., credible-interval coverage, probability integral transform histograms) are reported. The authors themselves state in §3.3 that McMC results show 'higher magnitude of velocity variations and lower uncertainty' and attribute the difference to 'insufficient training of the network'; this admitted, unquantified discrepancy concerns exactly the posterior mean and uncertainty that the paper claims to reproduce. I ask for quantitative comparisons on the synthetic examples (where the true model is known) and, if possible, calibration diagnostics on the real-data posteriors.","section":"§3.3 and §3.4, Figs. 5–8"},{"comment":"The load-bearing assumption that a network trained on 2D fast-marching travel times on an 11×11 grid can be applied to real surface-wave data is not tested. The argument that 'path-averaged velocity remains the same for a given velocity structure when the scale of the study area and receiver distribution change proportionally' holds only under strict 2D ray theory and proportional scaling of the whole structure. Real surface-wave dispersion is sensitive to depth-dependent 3D structure, and the real-data comparisons against McMC and SVGD use the same 2D path-average forward model, so they cannot reveal a mismatch between the training distribution and real physics. The paper needs an explicit statement of the domain of validity, or a test on synthetic data generated from 3D/strongly 2.5D structures, to support the claim that the network output approximates the true Bayesian posterior for real data.","section":"§3.1 and §3.4"},{"comment":"The title and abstract claim that graph MDNs can solve 'all seismic tomographic problems,' but the method is restricted to 2D travel-time surface-wave tomography with a fixed regular 11×11 parameterization, a limitation acknowledged in §4. This overstatement is load-bearing because the 'all problems' claim is a central selling point. I recommend revising the title, abstract, and conclusion to state clearly the scope (2D path-averaged surface-wave travel-time tomography with a given parameterization), and to reframe the broader generalization to 3D or full-waveform problems as future work rather than as part of the demonstrated claim.","section":"Title and Abstract; also §4 Discussion"}],"minor_comments":[{"comment":"The caption for Figure 6 lists periods of 5 s, 10 s, 15 s, 20 s, 25 s, and 30 s, but the text and Figure 5 describe the South England example at periods of 4 s, 6 s, 8 s, 10 s, 12 s, and 14 s. Please correct the mismatch so that the comparison panels correspond to the same periods.","section":"Figure 6 caption"},{"comment":"There is a typo in §3.3: 'sendimentary thickness' should be 'sedimentary thickness'. Also, the introduction's claim that Bayesian methods 'generally require expensive computational cost' is accurate but would benefit from a citation to a representative comparison; the later runtime table partly addresses this.","section":"§1 and §3.3 text"},{"comment":"The definition of the graph transformer layer would be easier to follow if the softmax argument were explicitly subscripted per layer and if the concatenation symbol were defined in the text; currently the notation 'C c=1' is formatted ambiguously.","section":"§2.3, Eq. (4)"},{"comment":"The paper does not include a data availability or code availability statement. Given that the training data are synthetic and the architecture is described in detail, releasing the network implementation would materially strengthen reproducibility, which several journals now expect.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The core idea is a natural extension of the authors' prior graph MDN work [32], and the new element is the random-geometry/scale generalization rather than a fundamentally new architecture. The paper's main weakness is evidential: the headline claim of 'comparable' posteriors is not quantified, and the admitted differences in §3.3 are exactly the quantities that matter. I also see a potential fit issue: the title and abstract promise universality that the method does not deliver, which is likely to draw criticism from readers. I would encourage the editor to require quantitative validation before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is training one graph MDN on random station layouts and proportional scales, using normalized coordinates and path-averaged velocity as edge features, then running it on real data at different sizes. That is a real extension of the fixed-geometry graph MDN paper [32], and the speed numbers are striking: 0.6 s per period against hours for McMC/SVGD. The two real-data examples, South England and Southwest China, are appropriate tests because both differ from the training distribution in station count and scale. Credit is also due for honest writing: the authors explicitly flag that MDNs are hard to train in high dimension and that their results show small discrepancies.\n\nBut the soft spots are real. The central claim that graph MDNs produce 'comparable' posterior estimates is supported only by side-by-side figures. No RMSE, no coverage or calibration statistic, no quantitative comparison of mean or standard deviation fields between the network and McMC/SVGD. More importantly, Section 3.3 states that McMC shows higher magnitude of velocity variations and lower uncertainty, and the authors attribute that to insufficient training. That is an admission that the main claim is not yet demonstrated. Section 3.4 repeats the same caveat. It may be fixable with longer training, more mixture components, or better loss weighting, but that is not shown here.\n\nTwo further issues. First, scope: the abstract says 'all seismic tomographic problems', but the trained network is a fixed 11x11 grid for 2D path-averaged surface waves. The discussion acknowledges the grid limitation, yet the title and abstract overstate. Second, the scale-invariance assumption: path-averaged velocity is indeed invariant under proportional scaling of coordinates and travel times, but that only holds if the same 2D ray model remains appropriate at all scales. The real-data comparisons use the same 2D model for the Bayesian inversions, so they cannot validate the domain assumption against 3D or anisotropic effects. That should be stated as a limitation. Also, Figure 5 lists MDN periods 4-14 s while Figure 6's caption lists McMC at 5-30 s; likely a caption mistake, but it needs fixing.\n\nNo code or data are provided, so independent verification is limited. The self-citation to [32] is appropriate since this is a direct extension. Bottom line: worthwhile extension with real potential, but the evidence for the headline claim is qualitative and partly self-undermined. A serious referee should engage, and revisions should add quantitative posterior diagnostics, preferably calibration tests on synthetic out-of-distribution cases, and rein in the 'all problems' language. I'd bring it to a reading group as a good example of a promising method whose validation lags its ambition.","headline":"A genuinely useful extension of graph MDNs to variable geometry and scale, with two real-data demos, but the 'comparable to McMC/SVGD' claim rests on visual comparison and the authors' own admitted discrepancies.","tokens_in":11788,"tokens_out":2719,"would_cite":true,"duration_ms":25997,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a graph mixture-density network, trained once on random 2D synthetic velocity models and their travel times, can estimate Bayesian posterior distributions for seismic tomography problems at other scales, station…","keywords":["seismic tomography","graph neural networks","mixture density networks","Bayesian inversion","surface wave tomography","travel time tomography","uncertainty quantification","deep learning"],"falsifier":"Generate a synthetic test set with a full 3D or anisotropic wave-equation forward model (or use real data known to contain strong off-path or depth-dependent sensitivity), then compare the trained network's posterior against a Monte Carlo inversion that uses the same correct forward model; if the network's mean or uncertainty falls outside the Monte Carlo posterior, the claim that one network solves all tomographic problems fails for that class.","tokens_in":10826,"feed_emoji":"🌍","tokens_out":10324,"duration_ms":86309,"temperature":0.7,"pith_summary":"The paper sets out to show that a single neural network can solve seismic tomography problems it was never trained on, across different region sizes, station layouts, and data counts. The authors train a graph mixture density network on random 2D velocity models in a small 10 km by 10 km region with 30 randomly placed stations, using fast-marching travel times as data. Once trained, the same network is applied without retraining to larger synthetic problems, to Love-wave group-velocity data from 15 stations in South England, and to Rayleigh-wave phase-velocity data from 57 stations in Southwest China. In each case the network returns posterior mean and standard-deviation maps in about 0.6 seconds per period, and the paper reports these are broadly consistent with posteriors from Markov chain Monte Carlo and Stein variational gradient descent, which take hours. If correct, this would make fast, uncertainty-aware seismic imaging available for arbitrary new regions without bespoke Bayesian computation.","feed_headline":"Returns uncertainty-aware velocity maps for new sites in seconds","feed_subtitle":"A graph mixture-density network trained once on synthetic models generalizes to new stations, scales, and sites.","key_machinery":"The central object is the graph mixture density network (MDN): a graph neural network that outputs a mixture-of-Gaussians probability density instead of a point estimate. Seismic stations are graph nodes with normalized coordinates as node features; each station pair with a travel-time measurement defines an edge whose feature is the path-averaged velocity (travel time divided by inter-station distance). Graph transformer layers update node features using attention over neighbors, a global mean pooling layer aggregates nodes, and a linear layer outputs the mixture coefficients, means, and standard deviations of 10 Gaussian kernels, giving a diagonal-covariance approximation to the posterior over the 11 by 11 velocity grid. Training minimizes the negative log-likelihood of the mixture, with random dropping of stations and travel times so the network learns to handle variable data sizes. The scale-invariance property does the transfer work: proportional scaling of region and receivers preserves path-averaged velocity, so a network trained at one scale can be applied at another.","core_discovery":"On the paper's own terms, the central discovery is that acquisition geometry and scale do not have to be fixed at training time for a neural-network tomographic solver. Because travel-time data are encoded as a graph whose node features are normalized station coordinates and whose edge features are path-averaged velocities, the representation is scale-invariant: for a given velocity structure, scaling the region and station layout proportionally leaves the path-averaged velocity unchanged. A graph MDN trained on random small-region synthetic models therefore transfers to larger regions, different station numbers, and real data at other sites. The paper applies the unchanged network to 20 x 20 km synthetic tests, South England group-velocity tomography at 4–14 s periods, and Southwest China phase-velocity tomography at 5–30 s periods, and reports that the resulting posterior means and standard deviations are broadly similar to those produced by McMC and SVGD, with differences attributed to insufficient training of the MDN or incomplete convergence of the samplers.","pith_inferences":["A natural next test is calibration: check whether the MDN's posterior standard deviations match the actual frequency of containing the true model when tested on many out-of-distribution synthetic models; the paper shows visual similarity to Bayesian baselines but does not report a quantitative calibration metric.","The transfer claim rests on period-by-period path-averaged velocity being scale-invariant; one could extend the same network to continuous dispersion curves by adding period as a node or edge feature, a step the paper does not take.","For real applications, prior choice matters: the posterior is conditioned on the training prior of 1.8–4.0 km/s on an 11x11 grid, so regions with velocities outside that range or needing finer spatial resolution may require retraining or a different architecture.","If the approach does generalize in practice, the largest payoff would be in rapid aftershock or volcano monitoring, where a pre-trained network could update probabilistic velocity images as data stream in, something the current hours-long Bayesian samplers cannot do."],"forward_implications":["A network trained once on random synthetic models can produce posterior mean and uncertainty maps for other 2D surface-wave tomography problems with different numbers of stations and different region sizes, without retraining.","Posterior inference becomes near-instant: about 0.6 seconds per period compared with roughly 16 hours for McMC and 3.9 hours for SVGD in the paper's examples.","Uncertainty maps can flag well-constrained anomalies and poorly covered regions in new datasets, just as full Bayesian inversions do.","The same approach could in principle be extended to 3D body-wave tomography or full waveform inversion by treating sources and receivers (or waveforms) as graph nodes and node features, though the authors note larger networks would be required."],"supporting_citations":[{"why":"Introduces the graph MDN architecture and training scheme for seismic tomography that this paper extends to variable scales and geometries.","marker":"[32]"},{"why":"Provides the fast marching method used to simulate inter-station travel times for training and synthetic test data.","marker":"[22]"},{"why":"Supplies the South England Love-wave group-velocity data used as the first real-data application.","marker":"[10]"},{"why":"Supplies the Southwest China Rayleigh-wave phase-velocity data used as the second real-data application.","marker":"[16]"},{"why":"Defines the adaptive Metropolis-Hastings algorithm used as the McMC baseline for the South England comparison.","marker":"[12]"},{"why":"Defines Stein variational gradient descent, the baseline method used for the Southwest China comparison.","marker":"[15]"},{"why":"Provides the adaptive Metropolis sampling implementation used to run the McMC chains.","marker":"[24]"},{"why":"Defines graph mixture density networks, the architectural foundation the paper builds on.","marker":"[8]"}],"fun_headline_variants":["Graph MDN generalizes across scales and sites for seismic tomography","Seismic tomography in seconds for any geometry via graph MDN","One trained graph MDN solves varied seismic tomography problems","Uncertainty-aware velocity maps from graph MDN in seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Real surface-wave travel times must behave like the synthetic 2D path-averaged ray-model data used in training, so that scaling the region and receivers proportionally leaves the underlying velocity structure unchanged and no 3D, depth-dependent, or anisotropic sensitivity enters.","fun_headline_variants_meta":{"raw":{"variants":["Graph MDN generalizes across scales and sites for seismic tomography","Seismic tomography in seconds for any geometry via graph MDN","One trained graph MDN solves varied seismic tomography problems","Uncertainty-aware velocity maps from graph MDN in seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1339,"prompt_tokens":924,"completion_tokens":415,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":346}},"tokens_in":540,"tokens_out":415,"duration_ms":4020,"temperature":1.0,"reasoning_tokens":346,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:40:47.737339+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a synthetic test set with a full 3D or anisotropic wave-equation forward model (or use real data known to contain strong off-path or depth-dependent sensitivity), then compare the trained network's posterior against a Monte Carlo inversion that uses the same correct forward model; if the network's mean or uncertainty falls outside the Monte Carlo posterior, the claim that one network solves all tomographic problems fails for that class.","supporting_citations":[{"cited_title":"Rapid Bayesian Seismic Tomography using Graph Mixture Density Networks","cited_arxiv_id":"2412.08080","evidence_quote":"Introduces the graph MDN architecture and training scheme for seismic tomography that this paper extends to variable scales and geometries."},{"cited_title":"Multiple reflection and transmission phases in complex layered media using a multistage fast marching method","cited_arxiv_id":null,"evidence_quote":"Provides the fast marching method used to simulate inter-station travel times for training and synthetic test data."},{"cited_title":"Transdimensional love-wave tomography of the British Isles and shear-velocity structure of the east Irish Sea Basin from ambient-noise inter- ferometry","cited_arxiv_id":null,"evidence_quote":"Supplies the South England Love-wave group-velocity data used as the first real-data application."},{"cited_title":"The community velocity model v","cited_arxiv_id":null,"evidence_quote":"Supplies the Southwest China Rayleigh-wave phase-velocity data used as the second real-data application."},{"cited_title":"An adaptive metropolis algorithm","cited_arxiv_id":null,"evidence_quote":"Defines the adaptive Metropolis-Hastings algorithm used as the McMC baseline for the South England comparison."},{"cited_title":"Probabilistic programming in python using pymc3","cited_arxiv_id":null,"evidence_quote":"Provides the adaptive Metropolis sampling implementation used to run the McMC chains."},{"cited_title":"Graph mixture density networks","cited_arxiv_id":null,"evidence_quote":"Defines graph mixture density networks, the architectural foundation the paper builds on."}],"review_version":1}