{"id":"110917f9-c725-4888-bb25-add0d65363e2","arxiv_id":"2501.14301","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A perspective review arguing that seismic tomography models differ because of subjective choices, and that this diversity should be embraced as a Community Monte Carlo sampling of uncertainty.","lead":"This paper explains how seismic tomography works and argues that the many subjective choices in building Earth models make their uncertainties larger than formal error estimates suggest. It proposes that the community treat the diversity of published models as a Monte Carlo ensemble to quantify those uncertainties.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 6.3's Community Monte Carlo lacks an operational criterion for 'sufficiently diverse'; existing models share data and conventions, so their spread cannot validate the claim that uncertainties are larger than estimated.","rationale":"The reader's verdict already identifies the representativeness of the model ensemble as the weakest assumption. My stress-test concurs and sharpens it: the problem is not merely that the quiz slices are hand-picked, but that the paper provides no formal characterization of the space being sampled. The strong claim (Section 1) that true uncertainties exceed formal error estimates is presented as a consequence of the model differences, but that inference is valid only if the models are independent samples over subjective choices. The paper's own Section 6.3 concedes the condition without supplying a method to achieve or test it. This is a circularity risk: to know whether an ensemble is sufficiently diverse, one already needs a measure of the uncertainty the ensemble is meant to estimate. The concrete test proposed—quantifying shared data and conventions among the five models—would at least determine whether the illustrative evidence is consistent with the paper's own diversity requirement. If the test shows heavy overlap, the central claim remains an unsupported assertion and the Community Monte Carlo proposal is underspecified; the paper should either present a sampling design (e.g., coordinated factorial variations of parameterization, regularization, misfit function, crustal correction, anisotropy scaling) or explicitly reframe the claim as a hypothesis. This does not undermine the educational value of Sections 3-5, and it does not require rejecting the paper; it supports the reader's CONDITIONAL verdict, now conditioned on providing an operational definition of sample space and diversity. No concerns about internal mathematical consistency were found; the tutorial derivations (Sections 4.1-4.2) are standard and correct.","tokens_in":22407,"tokens_out":4875,"duration_ms":46397,"concrete_test":"Use the model descriptions in Section 5 and Table 1 to compute, for each of the ten pairs among S40RTS, SEMUCB-WM1, SPiRaL, GLAD-M35, and REVEAL, the overlap in input data (e.g., fraction of shared earthquake-station recordings and period bands) and the provenance of starting models and parameterization families. Define independence as less than, say, 30% shared data and different parameterization families. If any pair exceeds this threshold, the five-model ensemble fails the paper's own diversity condition, meaning the observed spread is a lower bound on subjective-choice uncertainty and cannot validate the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central argument requires that the dispersion among tomographic models samples the space of reasonable subjective choices. The paper itself states the condition in Section 6.3: 'Community Monte Carlo can map out the true uncertainties only when the samples are sufficiently diverse.' But no definition of this sample space or of 'sufficiently diverse' is given, and the five illustrative models are not independent draws from it. Section 5 and Table 1 show that S40RTS, SEMUCB-WM1, SPiRaL, GLAD-M35, and REVEAL share the same global network data, largely the same 1-D reference models (e.g., PREM), and overlapping parameterization families (spherical harmonics, splines, spectral-element meshes). Their differences therefore reflect a convenience sample shaped by community norms, not a randomized or stratified sample of the subjective-choice space. Consequently, the observation that models differ (Section 2, Figs. 1-2) does not by itself establish the Section 1 claim that uncertainties are 'significantly larger than estimated by individual practitioners.' That claim would require knowing how much of the subjective-choice space remains unsampled; the paper provides no such bound. The proposal to 'produce more different tomographic models' is an appeal for more of the same unstructured diversity until an operational protocol for sampling subjective choices is specified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a pedagogical review and perspective on global seismic tomography aimed at non-specialists. It introduces the main data types (surface waves, body waves), the mathematical framework of travel-time tomography (Eqs. 1–13), finite-frequency and full-waveform approaches, and then argues that a large number of reasonable but subjective modeling choices (damping, parameterization, crustal corrections, misfit functions, etc.) accumulate into model differences that are larger than formal uncertainty analyses suggest. The authors support this claim by presenting five recent tomographic models (S40RTS, SEMUCB-WM1, SPiRaL, GLAD-M35, REVEAL) and a quiz based on visual comparison of slices. They conclude that producing similar models should not be the goal; instead, they propose a 'Community Monte Carlo' effort in which diverse groups deliberately generate dissimilar but data-fitting models, and they discuss how such ensembles could propagate meaningful uncertainties into geodynamic modeling. The paper ends with a glossary that satirically deflates jargon such as 'high-resolution'.","tokens_in":22648,"tokens_out":4800,"duration_ms":44588,"significance":"If the central argument is accepted, it offers a concrete community-level response to a recognized problem: model differences in tomography are often treated as noise rather than as signal about uncertainty. The pedagogical parts (ray theory, damping, resolution matrix, FWI scaling) are accurate and well suited for the target audience. The paper also deserves credit for explicitly stating the limitation of its own proposal ('Community Monte Carlo can map out the true uncertainties only when the samples are sufficiently diverse') and for citing established ensemble practices in climate and ice-sheet modeling (CMIP, ISMIP6). However, the paper is a perspective rather than a quantitative study: the claim that uncertainties are 'significantly larger' than formal estimates is supported by visual inspection of five hand-picked models and by citation, not by a direct comparison to formal uncertainty bounds, and the proposal lacks an operational definition of the sample space of subjective choices. These gaps are load-bearing because they concern exactly what the paper asks the community to adopt.","major_comments":[{"comment":"The Community Monte Carlo proposal lacks an operational definition of 'sufficiently diverse.' The paper asserts that groups 'sample the actual range of plausible models' and that current model differences already map out uncertainties, but Table 1 shows that the five models share largely the same global-network data and overlapping parameterization families (spherical harmonics, splines, spectral-element meshes) and common reference models such as PREM. Without a defined sample space of subjective choices and a coverage criterion, the proposal is unfalsifiable, and the current five-model ensemble cannot be shown to be an unbiased sample. I suggest specifying the axes of subjective choice (the paper already lists several in Section 4.3), a minimal experimental design (e.g., independent groups inverting common data with deliberately varied but justifiable choices), and a diagnostic for coverage, for example comparing the ensemble spread to the formal posterior uncertainty reported for GLAD-M35 (Cui et al., 2024).","section":"Section 6.3"},{"comment":"The central assertion that 'the uncertainties in seismic tomography are significantly larger than estimated by individual practitioners based on statistical analyses' (Section 1) is not quantitatively supported. The evidence presented is visual: five models are shown to differ in Figs. 1-2, and the paper states that this 'indicates that uncertainties are large.' Yet the models differ in data selection, method, and parameterization simultaneously, so the displayed spread cannot be attributed to subjective choices alone, and no comparison is made to the formal uncertainty bounds published for at least one of the models (GLAD-M35 is described in the reference list as including uncertainty quantification). A direct comparison between the inter-model spread and the formal posterior standard deviations would provide a concrete test of the 'significantly larger' claim; without it, the claim should be explicitly framed as a testable hypothesis rather than an established conclusion.","section":"Sections 1 and 2"},{"comment":"The paper states that 'it is impossible to exactly quantify how much lower [the true quality] is' relative to misfit-based measures. This admission is honest but weakens the paper's own central claim. If the magnitude cannot be quantified, the paper should say what evidence would suffice to establish the direction and rough magnitude of the effect and whether the 'Community Monte Carlo' proposal is intended to produce that evidence. Currently, the paper alternates between asserting that uncertainties are 'significantly larger' and denying that the difference can be quantified, leaving the reader without a clear sense of what would confirm or refute the thesis.","section":"Section 6.2"}],"minor_comments":[{"comment":"The section heading appears as 'SEISMIC DA TA' with an inserted space; this is presumably a typesetting artifact and should be corrected to 'SEISMIC DATA'.","section":"Section 3 heading"},{"comment":"The phrase 'comparison of different tomographic model' is missing the plural 's'; it should read 'tomographic models.'","section":"Section 1"},{"comment":"The title uses 'High-resolution' while the Glossary explicitly advises against using the term 'unless a precise definition can be provided.' If this is intentional irony, a footnote or a sentence in the Introduction would make the intent clearer; if not, the title is in tension with the paper's own advice.","section":"Title and Section 7"},{"comment":"The captions do not state whether all slices are plotted with the same color scale and amplitude range. Differences in color scaling across the five models could visually amplify or diminish apparent differences; stating that the color scale is common (or not) is important for interpreting the quiz.","section":"Figures 1 and 2"},{"comment":"The Fresnel-zone width formula w ≈ 0.5*sqrt(v*l/f) is presented without derivation or citation; since the paper is aimed at non-experts, adding a brief reference or a one-sentence explanation would be helpful.","section":"Section 4.2.2"},{"comment":"The reference for SPiRaL (Simmons et al., 2021) lacks page numbers (the entry ends at 'Geophys. J. Int. 227'); several other entries also omit DOIs or final page ranges, which is untypical for a journal submission.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a perspective/review paper rather than a primary research article, so I evaluated it on its own terms. The central claim and the main proposal are coherent and potentially influential, but both need to be sharpened: the 'significantly larger' claim needs a quantitative anchor (e.g., comparison with GLAD-M35's UQ), and the Community Monte Carlo proposal needs an operational definition of diversity. The authors are honest about the limitations, which is a strength, but that honesty also exposes the gap between the strength of the language in the abstract and the evidence provided. I do not see any issue of citation ethics: the self-citations to REVEAL and related work are appropriate in context, and the argument does not depend on those works being correct. The paper fits the journal's scope if the journal welcomes perspective pieces; if not, the editors may want to consider commissioning or repositioning it as a discussion paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a perspective/review, not a new method. The authors argue that tomographic uncertainties are larger than formal estimates because of subjective choices, and they propose a Community Monte Carlo to map true uncertainty by collecting diverse models. The argument is coherent, and the paper is well written. But the central claim is not quantitatively pinned down, and the proposal lacks an operational criterion for what counts as a sufficiently diverse sample.\n\nThe paper does several things well. The tutorial sections on waves, travel-time tomography, regularisation, resolution, and full-waveform inversion are accurate and clearly explained for non-specialists. The quiz with unlabeled slices is an effective device for making the bias point concrete. The glossary is refreshingly honest, especially the entries on 'high-resolution' and 'checkerboard tests'. And the authors are upfront that their suggested practice already exists in climate and ice-sheet modeling, citing CMIP and ISMIP6. So the novelty is real but modest: they are re-framing ensemble modeling for seismology, not inventing it.\n\nThe soft spots are exactly where the stress-test note lands. Section 6.3 says Community Monte Carlo works only if samples are sufficiently diverse, but no definition or protocol is given. The five models used for illustration share the same global network data, largely the same 1-D reference model, and overlapping parameterizations, so their spread is a convenience sample shaped by community norms. That spread does not by itself prove that uncertainties are larger than individual estimates; it only shows that these particular teams made different choices. Without a bound on unsampled subjective choices, the headline claim remains a plausible assertion, not a demonstrated fact.\n\nThat said, the paper is honest about this condition, and the proposal is a reasonable call to action. The lack of a protocol is a gap a referee can ask them to address, not a fatal flaw. For a perspective piece, this deserves serious engagement.\n\nI'd send it to review. A good referee could push for a more explicit definition of the sample space, maybe a worked example, and a discussion of how to avoid shared-convention bias. The paper is for anyone who uses tomographic models and wants to understand their reliability; they'll get a clear, fair introduction. As a researcher, I'd cite it as a reference for the Community Monte Carlo idea, but I'd be careful about citing the 'uncertainties are larger' claim without a footnote.","headline":"A readable, honest perspective on tomographic uncertainty; the Community Monte Carlo proposal is a good idea that needs a sharper definition of sample diversity.","tokens_in":23200,"tokens_out":2961,"would_cite":true,"duration_ms":26559,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Seismic tomography's published models of Earth's mantle diverge more than their formal error analyses indicate, and the paper argues the field should embrace this divergence as a Community Monte Carlo ensemble rather than pursue a single…","keywords":["seismic tomography","uncertainty quantification","full-waveform inversion","mantle tomography","Community Monte Carlo","subjective modeling choices","geodynamic inference"],"falsifier":"A controlled community experiment would settle the claim: have several independent groups invert the same waveform data with freely chosen regularization, parameterization, and crustal corrections, then compare the spread of their resulting models with each group's formal resolution-based uncertainty. If the cross-group spread is no larger than the internal error estimates, the paper's central premise that subjective choices dominate tomographic uncertainty would be falsified.","tokens_in":22217,"feed_emoji":"🌍","tokens_out":9189,"duration_ms":79072,"temperature":0.7,"pith_summary":"This paper, written for researchers who use tomographic images without building them, argues that the uncertainty attached to seismic tomography is larger than the formal error bars of any single model suggest. The reason is that many individually small, reasonable choices—how to regularize the inversion, how to parameterize and discretize the Earth, how to treat the crust, what misfit function to use, how to scale poorly constrained parameters—compound into large model differences. The authors demonstrate this with a quiz in which five modern global models, built with different data and methods including full-waveform inversion, disagree on basic mantle features such as subducting slabs and mantle plumes. They conclude that converging on a single 'best' model is the wrong goal, and instead propose a Community Monte Carlo approach in which many deliberately diverse models, all explaining the seismic data, are assembled and propagated into geodynamic inferences with meaningful uncertainties.","feed_headline":"Seismic tomography's true uncertainty is the artist's choice","feed_subtitle":"Model disagreements are real uncertainty, so the field should chase diverse ensembles, not one best image.","key_machinery":"The load-bearing concept is the 'little choices' catalog of Section 4.3: regularization, inclusion or exclusion of model parameters, discretization and parameterization, crustal corrections, misfit functions, and empirical scaling of unmodeled parameters such as density and anisotropy. These choices, each defensible on its own, act as compounding degrees of freedom that make the set of plausible models much wider than any single inversion's resolution analysis admits. The proposed remedy, Community Monte Carlo, treats independent research groups as Monte Carlo samples over this subjective-choice space: each group's model is one draw, and the ensemble of dissimilar, data-fitting models maps the true uncertainty. The paper stresses the condition that the sampling only works when the models are sufficiently diverse.","core_discovery":"On the paper's own terms, the central assertion is that the spread among independently published tomographic models is itself the most honest estimate of tomographic uncertainty, and that this spread is substantially larger than the misfit reductions and resolution matrices of individual inversions indicate. The paper traces that excess uncertainty to the 'artistic component' of tomography: the unavoidable, reasonable, and partly subjective decisions made at every stage of model construction. Because these decisions differ across research groups, the models they produce agree on widely reproduced features such as the plate-tectonic fabric of the upper mantle and the two large low-shear-velocity provinces at the base of the mantle, but disagree on the presence, sharpness, and position of features like slabs and plumes. The paper therefore advocates that the community stop aiming for similar models and instead deliberately cultivate a diverse ensemble of data-fitting models—Community Monte Carlo—which can serve as input to geodynamic simulations that carry genuine seismic uncertainties.","pith_inferences":["Editorial inference: one could operationalize diversity by measuring the statistical independence of the ensemble, for example through pair-wise differences between models, and use that to compute an effective sample size for the Community Monte Carlo.","Editorial inference: the same logic transfers to other data-assimilation fields: any inverse problem with many defensible algorithmic choices should report ensemble spread across independently constructed pipelines, not just pipeline-internal covariance.","Editorial inference: a targeted experiment varying one subjective choice at a time—say, regularization norm or crustal treatment—while holding data and method fixed would quantify how much each 'little choice' contributes, testing the paper's claim that the effects compound.","Editorial inference: if the proposal is adopted, a natural deliverable would be a probability-style statement for a given structure—such as '95 percent of data-fitting models contain a slab here'—which is more actionable for non-specialists than a single image with resolution lengths."],"forward_implications":["Users of tomographic models should treat the divergence among independently produced models as a first-order uncertainty estimate rather than picking a single 'best' image.","Funding agencies and community efforts should reward methodological diversity and the production of intentionally different models, because convergence would hide uncertainty rather than reduce it.","Geodynamic inversions that assimilate tomographic structure should be run repeatedly, once per ensemble member, so that uncertainty in the seismic images propagates into an ensemble of plausible mantle-flow histories.","Formal resolution tests and misfit statistics remain necessary but are not sufficient: they describe data coverage and data fit for a fixed set of choices, not the true range of defensible models.","As data volumes grow and computational costs fall, the practical barrier to regular Monte Carlo sampling shrinks, making the complementary combination of regular and Community Monte Carlo increasingly feasible."],"supporting_citations":[{"why":"S40RTS is one of the five compared models and the baseline travel-time and ray-theory result whose disagreement with waveform-inversion models motivates the uncertainty argument.","marker":"(Ritsema et al., 2011)"},{"why":"SEMUCB-WM1 is the whole-mantle waveform-inversion model used in the quiz and comparison, representing a different methodological lineage.","marker":"(French and Romanowicz, 2014)"},{"why":"SPiRaL provides the multiresolution travel-time model used in the comparison, adding another independent set of data and parameterization choices.","marker":"(Simmons et al., 2021)"},{"why":"GLAD-M35 is the joint P and S full-waveform-inversion model in the comparison and an example of a modern model with its own uncertainty quantification.","marker":"(Cui et al., 2024)"},{"why":"REVEAL is the data-adaptive full-waveform-inversion model in the comparison, showing how numerical mesh choices change the resulting image.","marker":"(Thrastarson et al., 2024)"},{"why":"It supplies systematic evidence that approximations and arbitrary choices in geophysical inversions materially change recovered images, underpinning the artistic-component claim.","marker":"(Valentine and Trampert, 2016)"},{"why":"It is the foundational Monte Carlo method for sampling solutions to inverse problems that Community Monte Carlo extends to the space of subjective modeling choices.","marker":"(Mosegaard and Tarantola, 1995)"},{"why":"It reviews Monte Carlo sampling in geophysical inverse problems and supports the paper's claim that regular Monte Carlo is challenging but feasible.","marker":"(Sambridge and Mosegaard, 2002)"}],"fun_headline_variants":["Seismic model spread is the real uncertainty, not misfit","Tomographic models differ more than error bars suggest","Embrace diverse seismic models, not a single best image","Subjective picks drive seismic model disagreements","Seismic tomography's artistic choices are real uncertainty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the ensemble of community models being genuinely independent: if groups share starting models, data selections, and parameterization conventions, the observed model spread will underestimate the true uncertainty, and Community Monte Carlo will not map it out.","fun_headline_variants_meta":{"raw":{"variants":["Seismic model spread is the real uncertainty, not misfit","Tomographic models differ more than error bars suggest","Embrace diverse seismic models, not a single best image","Subjective picks drive seismic model disagreements","Seismic tomography's artistic choices are real uncertainty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000762,"raw_usage":{"total_tokens":3355,"prompt_tokens":892,"completion_tokens":2463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":2389}},"tokens_in":508,"tokens_out":2463,"duration_ms":18259,"temperature":1.0,"reasoning_tokens":2389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:14:17.117652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled community experiment would settle the claim: have several independent groups invert the same waveform data with freely chosen regularization, parameterization, and crustal corrections, then compare the spread of their resulting models with each group's formal resolution-based uncertainty. If the cross-group spread is no larger than the internal error estimates, the paper's central premise that subjective choices dominate tomographic uncertainty would be falsified.","supporting_citations":[{"cited_title":"Deuss, H","cited_arxiv_id":null,"evidence_quote":"S40RTS is one of the five compared models and the baseline travel-time and ray-theory result whose disagreement with waveform-inversion models motivates the uncertainty argument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SEMUCB-WM1 is the whole-mantle waveform-inversion model used in the quiz and comparison, representing a different methodological lineage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SPiRaL provides the multiresolution travel-time model used in the comparison, adding another independent set of data and parameterization choices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GLAD-M35 is the joint P and S full-waveform-inversion model in the comparison and an example of a modern model with its own uncertainty quantification."},{"cited_title":"van Herwaarden , S","cited_arxiv_id":null,"evidence_quote":"REVEAL is the data-adaptive full-waveform-inversion model in the comparison, showing how numerical mesh choices change the resulting image."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies systematic evidence that approximations and arbitrary choices in geophysical inversions materially change recovered images, underpinning the artistic-component claim."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is the foundational Monte Carlo method for sampling solutions to inverse problems that Community Monte Carlo extends to the space of subjective modeling choices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It reviews Monte Carlo sampling in geophysical inverse problems and supports the paper's claim that regular Monte Carlo is challenging but feasible."}],"review_version":1}