{"id":"1e06b97d-7415-4029-824b-a181acea5227","arxiv_id":"2412.08080","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Graph mixture density networks can rapidly estimate Bayesian posterior uncertainty in seismic tomography even when the number of travel time measurements varies between problems.","lead":"A team applied graph mixture density networks to seismic travel time tomography, allowing the number of measurements to vary between different inversions. The trained network produces posterior uncertainty maps in seconds, comparable to what Markov chain Monte Carlo takes days to compute.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim requires that a 10-kernel diagonal-Gaussian graph MDN approximates high-dimensional posteriors, but the paper's evidence is only qualitative and it explicitly reports uncertainty mismatch with McMC; no quantitative accuracy test supports 'accurate and rapid.'","rationale":"The paper's goal is clear: a graph MDN should deliver posterior pdfs for seismic tomography with variable data sizes, at quality comparable to McMC. The most load-bearing condition is therefore that the MDN's output distribution, a mixture of 10 diagonal-covariance Gaussians in 81 or 90 model dimensions, can represent the true posterior sufficiently well after training. The paper's own discussion and figures indicate that this condition is not fully met: high-dimensional MDN limitations are cited, and the graph MDN uncertainties are consistently larger than McMC uncertainties in the real-data study. Moreover, the comparison to McMC is qualitative throughout, relying on visual similarity of mean and standard-deviation maps rather than any quantitative metric. This is not a claim about internal inconsistency or fraud; the method and speed improvements are plausible, and the graph representation for variable data sizes is a reasonable contribution. But the conclusion 'accurate and rapid solutions' is stronger than the evidence. A held-out quantitative comparison of posterior means, marginal standard deviations, and credible-interval coverage against McMC would settle whether the discrepancy is a minor calibration issue or a fundamental representational failure. The reader's verdict already conditions acceptance on quantitative posterior comparisons and public artifacts, and this stress-test supports that position without moving it to accept or reject.","tokens_in":20614,"tokens_out":3539,"duration_ms":40901,"concrete_test":"On the synthetic test set (the held-out 10%), for each test model compute the posterior mean and per-cell marginal standard deviation from graph MDN and from the same McMC chain setup used in Sec. 3. Report (a) the mean absolute difference between graph MDN and McMC marginal standard deviations, normalized by the McMC standard deviation, and (b) the empirical coverage of 90% and 95% credible intervals derived from graph MDN marginals over the true test models. If the normalized standard-deviation bias exceeds roughly 10% in a majority of cells, or if coverage deviates substantially from nominal levels, the claim of accurate posterior approximation fails. Repeating the same comparison on the real-data posterior maps, with McMC as reference, would also test calibration in the actual application.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Sec. 6) is that trained graph MDNs provide 'accurate and rapid solutions to seismic tomographic problems with variable sizes of data.' For this to hold, the network's output distribution must faithfully approximate the posterior in 81-90 dimensions. Eq. (3) restricts each Gaussian kernel to a diagonal covariance, and only 10 such kernels are used. The paper itself acknowledges the resulting mismatch: Sec. 5 says 'results obtained using graph MDNs show higher uncertainty estimates than those obtained using McMC because of difficulties of training MDNs in high dimensionality'; Sec. 3 notes that high-standard-deviation features visible in McMC are not clearly observed in graph MDN, and that 'the two solutions can have different detailed structures.' In the real-data application (Figs. 10-11), the McMC standard-deviation maps are generally smaller than the graph MDN maps. Thus the load-bearing condition, that the finite diagonal mixture is expressive and trainable enough to capture the posterior's shape, is exactly where the paper admits weakness. The validation is entirely qualitative: no quantitative metric, such as per-cell standard-deviation error, empirical coverage, or Kullback-Leibler divergence, is computed between graph MDN and McMC posteriors. Without such a metric, 'comparable posterior pdfs' is an unsupported assertion, and a posterior whose marginal variances are systematically overstated is not an accurate posterior pdf even if its means are close.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes graph mixture density networks (graph MDNs) for Bayesian seismic travel-time tomography with variable numbers of travel-time data. Travel-time observations between station pairs are represented as a graph, and a graph transformer with edge features feeds a mixture-density-network head that outputs a Gaussian mixture approximation to the posterior over a fixed grid of velocity parameters. The method is demonstrated on a synthetic 2D travel-time tomography problem with 16 receivers and a 9×9 velocity grid, where the number of station pairs is randomly reduced, and on a real Love-wave group-velocity tomography study in South England at six periods with data of varying size. In both cases the posterior means and standard deviations are compared visually with McMC reference solutions, and computational timings are reported. The authors conclude that graph MDNs provide accurate and rapid posterior approximations for variable data sizes.","tokens_in":20953,"tokens_out":3956,"duration_ms":38611,"significance":"If the claims were fully substantiated, the method would be a useful contribution: it removes the fixed-input-size restriction of standard MDNs by using graph neural networks, and it reports orders-of-magnitude speedups after training. The explicit graph representation of station pairs, the edge-dropping training strategy, and the demonstration on both synthetic and real data are strengths. The paper also provides clear network configurations and runtime figures. However, the central claim of 'accurate' posterior approximation is not supported by quantitative validation: the comparison with McMC is visual only, and the paper itself reports that graph-MDN uncertainties are systematically larger than McMC's. The significance therefore depends on whether the authors can supply calibration or coverage metrics and address the expressiveness limitation of a 10-kernel diagonal Gaussian mixture in 81–90 dimensions.","major_comments":[{"comment":"The conclusion that graph MDNs provide comparable posterior pdfs to McMC is based entirely on qualitative visual comparison of posterior mean and standard deviation maps. No quantitative metric (e.g., per-cell mean or standard-deviation error, empirical coverage of the McMC marginals by the MDN predictive intervals, or a divergence measure such as KL) is reported. This matters because the paper notes in §3 that high-standard-deviation features present in the McMC maps (e.g., the four lobes around the central anomaly in Figs 4f and 5f) are not clearly visible in the graph-MDN maps, and in §4 the McMC standard deviations are 'generally smaller than that obtained using graph MDN' (p. 22). A systematic overestimation of variance is not a benign discrepancy: it changes the interpretation of the posterior. I would ask the authors to add quantitative agreement metrics between the graph-MDN and McMC posteriors, and also to report the test-set negative log-likelihood or calibration of the predictive distribution on the held-out 10% of training pairs.","section":"§3 and §4, Figs 4–7 and 10–11"},{"comment":"The load-bearing approximation is that a mixture of 10 Gaussian kernels with diagonal covariances, Eq. (3), is sufficiently expressive and trainable to represent the posterior in an 81- or 90-dimensional model space. The paper cites universal approximation properties of mixture models, but that is an asymptotic statement. The authors explicitly concede in §5 that 'results obtained using graph MDNs show higher uncertainty estimates than those obtained using McMC because of difficulties of training MDNs in high dimensionality.' That admission directly undercuts the accuracy claim, because an inflated covariance is not an accurate posterior. I would like to see a discussion or demonstration that the chosen number of components and the training budget are adequate: for example, a sensitivity test with more than 10 kernels, a comparison on a lower-dimensional problem where exact posteriors are known, or diagnostics (posterior predictive checks, coverage probabilities) that show the approximation error is acceptable. Without this, the claim in §6 that graph MDNs 'provide accurate approximations of posterior pdfs' is not justified.","section":"§2.2, Eq. (3); §5, first paragraph"}],"minor_comments":[{"comment":"The phrase 'We use ReLU activate functions for each graph transformer layer or linear layer' should read 'We use ReLU activation functions for each graph transformer layer or linear layer'.","section":"§2.3"},{"comment":"The reference to Kingma & Welling (2013) contains a typo: 'Auto-encoding variational Byes' should be 'Auto-encoding variational Bayes'.","section":"References"},{"comment":"The reported inference times of 0.6 s per data set should clarify whether this figure includes graph construction and data pre-processing or only the network evaluation; this would make the cost comparison with McMC more transparent.","section":"§3 and §4, timing comparisons"},{"comment":"The paper only varies the number of station pairs in the experiments, not the number of stations; although §5 discusses training with variable station distributions as a possibility, no experiment demonstrates that scenario, so the broad 'variable sizes of data' phrasing in §6 is stronger than what is shown.","section":"§5 vs. §6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a geophysics journal and addresses a practical limitation of MDN-based inversion. My main concern is that the evidence for the central accuracy claim is qualitative and the authors' own statements in §4 and §5 indicate systematic uncertainty overestimation. The fixed receiver geometry in both experiments and the lack of a station-variability test also make the 'variable sizes of data' claim narrower than implied. I would encourage adding quantitative posterior-validation metrics and a sensitivity analysis of the mixture-size/high-dimensional limitation before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the Zhang et al. graph MDN tomography paper. The genuinely new piece: they encode travel-time data as a graph with stations as nodes and inter-station travel times as edge features, so a trained MDN accepts variable numbers of station pairs. Prior MDN and INN tomography used fixed-size input vectors; this is a real extension, demonstrated on both synthetic edge-drop scenarios and real Love-wave data with different path counts across periods. That part is solid and useful.\n\nThe paper does a few things well beyond the gimmick: the synthetic tests are reasonably designed (200k training samples, test structures outside the training distribution, a second random-structure case), the McMC benchmark is heavy (six chains, a million samples), and the authors are transparent about the known weakness of MDNs in high dimensions. They don't hide that their uncertainties are consistently larger than McMC's.\n\nThe soft spot is the central claim. The conclusion says the network provides 'accurate approximations of posterior pdfs,' but there is no quantitative comparison anywhere: no per-cell standard-deviation error, no empirical coverage, no KL divergence, nothing. The evidence is side-by-side images and 'broadly consistent' statements. The one quantitative-ish observation they do report goes against them: in Section 5 and the real-data section they state graph MDN uncertainties are higher than McMC's. An uncertainty estimate that is systematically inflated is not an accurate posterior, even if the means look fine. It might still be useful for relative or conservative uncertainty, but 'accurate' is not supported.\n\nAlso note the 'variable data sizes' is mostly variable edge counts on a fixed station array with fixed grid parameterization; they acknowledge this in the discussion. That limits the practical generality, though it is still the right first step.\n\nNo code or data is provided, which makes the qualitative validation harder to check. For a methods paper in this area, shipping the graph MDN code and trained weights would materially raise confidence.\n\nBottom line: this is a legitimate incremental contribution that deserves referee time. I would send it out, but I'd ask for quantitative posterior diagnostics (coverage on synthetic data, calibration plots, per-cell error vs McMC) and a fair handling of the 'accurate' wording. For citation: I'd cite it as an example of variable-size neural Bayesian tomography, but not as evidence that MDN posteriors are faithful in high dimensions. Bring it to reading group only if you're specifically interested in GNN surrogates; otherwise a skim suffices.\n\nRecommendation: yes to peer review, with the expectation of major revision or at minimum a much more guarded conclusion.","headline":"A useful graph-MDN extension for variable-size tomography, but the 'accurate' claim outruns the evidence: validation is visual only and the paper's own uncertainty mismatch is unquantified.","tokens_in":21432,"tokens_out":3886,"would_cite":true,"duration_ms":36148,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a single trained graph mixture density network can deliver posterior mean and uncertainty maps for seismic travel-time tomography in about 0.6 seconds, matching Markov chain Monte Carlo quality even when the number…","keywords":["seismic tomography","Bayesian inference","mixture density networks","graph neural networks","uncertainty quantification","travel time tomography","surface wave tomography","Markov chain Monte Carlo"],"falsifier":"Run the same graph MDN on a tomographic problem with a larger model dimension, such as a 3D grid or a 2D grid with more than 100 cells, and compare posterior statistics against a converged McMC or Hamiltonian Monte Carlo reference; if the standard-deviation maps diverge systematically or the network cannot represent correlations between neighboring cells, the claim of accurate rapid Bayesian tomography fails. A cheaper discriminating test is to train on data with strongly multimodal posteriors or with correlated model parameters and check whether the diagonal-Gaussian mixture reproduces the reference covariance.","tokens_in":20416,"feed_emoji":"🌍","tokens_out":9092,"duration_ms":80015,"temperature":0.7,"pith_summary":"Seismic tomography produces images of subsurface velocity, but the Bayesian methods that also quantify image uncertainty are slow: Markov chain Monte Carlo (McMC) typically takes many hours to weeks per inversion. Mixture density networks (MDNs) make Bayesian inversion fast by training a network to output a probability distribution over models, but earlier versions require a fixed number of data points, so they fail when station-pair or travel-time counts vary. This paper proposes a graph MDN, which encodes travel-time data as a graph—stations as nodes, travel times as edge features—and combines graph neural networks with MDNs so that one trained network accepts data of any size. On a synthetic 2D travel-time tomography problem and on real Love-wave group-velocity tomography of South England, the trained graph MDN produces posterior means and standard deviations broadly consistent with McMC, at a cost of roughly 0.6 seconds per new data set after training. The paper concludes that graph MDNs can provide accurate and rapid Bayesian solutions to seismic tomographic problems with variable data sizes.","feed_headline":"One trained network answers seismic tomography in 0.6 seconds","feed_subtitle":"Graph mixture density networks give posterior velocity maps for variable station-pair data at a fraction of McMC cost.","key_machinery":"The central object is the graph mixture density network, a neural network that takes a graph as input and outputs the parameters of a Gaussian mixture model approximating the posterior probability density p(m|d). Travel-time data are encoded as a graph: stations are nodes with coordinate features, and each station pair with a measured travel time becomes an undirected edge whose feature is the path-averaged velocity derived from the travel time and inter-station distance. The network stacks graph transformer layers, which aggregate node and edge features through attention coefficients, with graph linear layers; the updated node features are pooled by a global mean, and a final linear layer emits the mixture coefficients, means, and diagonal standard deviations of ten Gaussian kernels. During training, station-pair edges are randomly dropped at each iteration (rate 0.1 in the synthetic case, 0.6 for the real data), teaching the network to handle variable data sizes. The load-bearing identity is the graph representation itself: because predictions depend only on graph structure, the same trained network applies to any subset of the available station pairs.","core_discovery":"The central claim is that the graph MDN removes the fixed-input limitation of neural-network-based Bayesian inversion: because the network operates on a graph rather than a fixed-length vector, it can ingest any subset of station pairs and still output an approximate posterior pdf over the velocity model. The validation uses 200,000 simulated velocity models whose travel times were computed with the fast marching method, then compares the network's posterior mean and standard deviation maps against adaptive Metropolis-Hastings McMC runs of six chains of one million samples each. On the synthetic problem (81 model parameters, 16 receivers, full and randomly thinned data), both methods recover the same central low-velocity anomaly, show similar edge effects, and place higher uncertainty where ray coverage is poor. On the real data (six periods, 90 model parameters), the graph MDN reproduces the known geological features of South England and follows the McMC uncertainty pattern, although the magnitudes of the uncertainties differ because MDNs become difficult to train accurately in high dimensions. The paper's conclusion is that for problems requiring many similar tomographic inversions with different numbers of data, a trained graph MDN provides accurate posterior estimates in seconds.","pith_inferences":["Editorial inference: the architecture suggests a route toward a single array-wide tomographic network trained once on random station subsets and then used for any data availability pattern, including station additions; the paper notes this possibility but does not demonstrate it.","Editorial inference: the same graph encoding could be applied to other network-based inverse problems with pairwise measurements, such as electrical resistivity or hydrological tomography, provided a forward simulator exists to generate training pairs.","Editorial inference: the reported agreement is limited to posterior means and standard deviations; the paper does not test whether the full posterior shape, including multimodality or correlations between distant cells, is captured, so the phrase 'comparable posterior pdfs' should be read as applying to these statistics.","Editorial inference: a natural stress test would increase the number of Gaussian kernels or replace the mixture with a graph-based normalizing flow and measure how the agreement with McMC changes as model dimension grows; the paper itself names invertible neural networks as a possible improvement."],"forward_implications":["After training, a graph MDN answers a new tomographic problem in about 0.6 seconds, making near-real-time subsurface monitoring and time-lapse imaging feasible for problems that previously required 16 to 21 hours of McMC computation per data set.","The same trained network can be applied to data sets with different numbers of station pairs, so missing or noisy measurements do not require retraining or a new inversion.","The graph abstraction carries over to other geophysical inverse problems in which data are organized as pairs of stations or sources and receivers, such as body-wave travel-time tomography.","When many inversions are needed (for example, one per period or per time window), the one-time training cost is amortized and the method becomes more efficient than running McMC for every data set."],"supporting_citations":[{"why":"Introduces graph mixture density networks, the method this paper adapts to seismic tomography.","marker":"Errica et al. 2021"},{"why":"Probabilistic neural network 2D travel-time tomography that the synthetic experiment mirrors and whose fixed-data-size limitation graph MDNs remove.","marker":"Earp & Curtis 2020"},{"why":"Invertible neural network Bayesian inversion used as the methodological and numerical comparison baseline for the synthetic and real experiments.","marker":"Zhang & Curtis 2021"},{"why":"Supplies the mixture density network formulation and the negative-log-likelihood training objective used by the graph MDN.","marker":"Bishop 2006"},{"why":"Demonstrates MDN-based geophysical inversion and the training approach that this work extends to graph inputs.","marker":"Meier et al. 2007a"},{"why":"Provides the South England Love-wave data, prior bounds, and geological reference structures used in the real-data validation.","marker":"Galetti et al. 2017"},{"why":"Fast marching method used to compute travel times for the 200,000 training models.","marker":"Rawlinson & Sambridge 2004"},{"why":"Adaptive Metropolis algorithm used to generate the McMC reference posteriors.","marker":"Haario et al. 2001"},{"why":"Provides the McMC implementation used for the comparison runs.","marker":"Salvatier et al. 2016"}],"fun_headline_variants":["Graph MDNs give fast Bayesian seismic tomography for variable data","Seismic uncertainty in seconds: graph mixture density networks","Variable-size seismic data? Graph MDNs handle it fast","Graph MDN tomography: Bayesian posteriors at McMC quality, in seconds","One network, any data size: graph MDNs for seismic tomography"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that a mixture of ten Gaussian kernels with no inter-parameter covariance can faithfully represent a posterior in an 81- or 90-dimensional model space; if that approximation is too crude, the uncertainty maps are biased and the agreement with McMC will not survive harder or higher-dimensional problems.","fun_headline_variants_meta":{"raw":{"variants":["Graph MDNs give fast Bayesian seismic tomography for variable data","Seismic uncertainty in seconds: graph mixture density networks","Variable-size seismic data? Graph MDNs handle it fast","Graph MDN tomography: Bayesian posteriors at McMC quality, in seconds","One network, any data size: graph MDNs for seismic tomography"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000778,"raw_usage":{"total_tokens":3479,"prompt_tokens":1026,"completion_tokens":2453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":2381}},"tokens_in":642,"tokens_out":2453,"duration_ms":16292,"temperature":1.0,"reasoning_tokens":2381,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:13:42.473702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same graph MDN on a tomographic problem with a larger model dimension, such as a 3D grid or a 2D grid with more than 100 cells, and compare posterior statistics against a converged McMC or Hamiltonian Monte Carlo reference; if the standard-deviation maps diverge systematically or the network cannot represent correlations between neighboring cells, the claim of accurate rapid Bayesian tomography fails. A cheaper discriminating test is to train on data with strongly multimodal posteriors or with correlated model parameters and check whether the diagonal-Gaussian mixture reproduces the reference covariance.","supporting_citations":[],"review_version":1}