{"id":"8aeb7f0c-f325-425d-860c-8bfabf6d5d90","arxiv_id":"2605.30541","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces SubsurfaceGen, a procedural generator for field-scale 3D velocity models and seismic data, releases a dataset of 4276 2D slices from 42 models across six geological settings, and evaluates neural operators and encoder-decoders on wavefield prediction and velocity inversion with ","lead":"SubsurfaceGen is a GPU-accelerated procedural generator that creates large 3D subsurface velocity models and corresponding seismic data for training machine learning models in full waveform inversion. A smart generalist might read it to understand how better synthetic data could improve AI tools used in carbon storage, energy exploration, and earthquake hazard assessment.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption correctly flags the realism/diversity question, but because the paper is framed as a tool+dataset release rather than a claim that the generated models are proven to be indistinguishable from real geology, that assumption is not load-bearing for the stated contribution. No internal inconsistency or missing derivation is evident from the abstract and described experimental design.","tokens_in":1809,"tokens_out":266,"duration_ms":13371,"concrete_test":"Re-run the encoder-decoder inversion experiments on the released dataset after replacing one SubsurfaceGen-generated geological setting with an independent public 3D model (e.g., SEAM Phase I) of comparable size; if the reported failure modes disappear or change character, the claim that the procedural models surface representative field-scale issues would require qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a procedural generator plus released dataset for training ML FWI models at field scale. The argument does not rest on a single unverified mathematical step or hidden assumption that would invalidate the released artifacts or the reported experiments; the generator and paired data are the deliverable, and the OOD hold-out experiments are presented as illustrative rather than as a formal proof of generalization.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SubsurfaceGen, a GPU-accelerated procedural generator for creating 3D velocity models and seismic data at field scale. It releases a paired dataset consisting of 4,276 2D velocity slices, 5 s wavefields, and 8 s shot gathers drawn from 42 realistic 3D models (each 10 km x 10 km x 6.19 km at 10 m resolution) spanning six geological settings relevant to carbon sequestration and hydrocarbon exploration (four generated with SubsurfaceGen, two from prior sources). The authors evaluate neural operators on wavefield prediction and encoder-decoders on end-to-end velocity inversion, with one geological setting held out for out-of-distribution testing, claiming that the experiments surface failure modes at field scale.","tokens_in":1902,"tokens_out":542,"duration_ms":22507,"significance":"If the generated models prove sufficiently realistic and diverse, the work supplies a much-needed large-scale resource for ML-based FWI that addresses documented shortcomings of Marmousi, SEAM, and OpenFWI in spatial extent, temporal extent, and geological variety. The explicit release of the generator, code, and paired dataset constitutes a concrete, reusable contribution that can support reproducible research and systematic identification of ML limitations in realistic settings.","major_comments":[{"comment":"Dataset construction paragraph: the central claim that the procedurally generated models are 'realistic' and capable of surfacing meaningful field-scale failure modes rests on an unverified assumption of physical and geological fidelity; the manuscript provides no quantitative validation such as velocity histograms, spatial autocorrelation statistics, or direct comparison against real field surveys.","section":"Dataset construction"},{"comment":"Evaluation setup: the OOD experiments hold out one geological setting, but the text does not specify which setting is held out nor demonstrate that the held-out geology differs in a manner that genuinely tests generalization rather than merely presenting a new realization of similar geology.","section":"Evaluation setup"}],"minor_comments":[{"comment":"Abstract: the phrases '5 s wavefields' and '8 s shot gathers' are ambiguous; clarify whether these refer to simulation duration, recording length, or another quantity.","section":"Abstract"},{"comment":"Introduction: the discussion of limitations of Marmousi, SEAM, and OpenFWI would benefit from explicit numerical comparisons (e.g., lateral extent, depth, number of distinct geological regimes) rather than qualitative statements.","section":"Introduction"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment and constructive feedback. We address each major comment below.","responses":[{"response":"We acknowledge that the manuscript does not include quantitative statistical comparisons such as velocity histograms, spatial autocorrelation, or direct comparisons to real field surveys. Realism claims rest on the use of established procedural methods for the six geological settings. In revision we will add velocity histograms and basic statistical summaries of the generated models to support the claims.","revision_made":"yes","referee_comment":"[Dataset construction] Dataset construction paragraph: the central claim that the procedurally generated models are 'realistic' and capable of surfacing meaningful field-scale failure modes rests on an unverified assumption of physical and geological fidelity; the manuscript provides no quantitative validation such as velocity histograms, spatial autocorrelation statistics, or direct comparison against real field surveys."},{"response":"We agree the held-out setting should be named and the geological distinctions motivating the OOD test should be made explicit. In the revised manuscript we will identify the held-out setting and add a brief description of its distinguishing features relative to the training settings.","revision_made":"yes","referee_comment":"[Evaluation setup] Evaluation setup: the OOD experiments hold out one geological setting, but the text does not specify which setting is held out nor demonstrate that the held-out geology differs in a manner that genuinely tests generalization rather than merely presenting a new realization of similar geology."}],"tokens_in":1466,"tokens_out":321,"duration_ms":18498,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that SubsurfaceGen and its released dataset address the shortage of realistic training data for machine learning on full waveform inversion at actual field scales.\n\nThe generator produces 3D velocity models at 10 km by 10 km by 6.19 km with 10 m resolution, and the paper ships 4276 paired 2D slices that include wavefields and shot gathers. It covers six geological settings relevant to carbon sequestration and hydrocarbon work, mixing four procedurally generated ones with two drawn from earlier sources. The evaluations test neural operators on wavefield prediction and encoder-decoders on inversion, with one setting held out.\n\nThis is genuinely new at the stated scale and diversity. Previous resources like Marmousi or OpenFWI are smaller in extent or narrower in geology, so the public release of both the generator and the paired data is the concrete step forward.\n\nThe soft spot is the reliance on procedural rules for realism. If the generated models miss important field features such as complex faulting or fine layering, the failure modes shown in the ML tests may not carry over to real surveys. The abstract gives no quantitative error analysis or direct comparison metrics, so the strength of the OOD results is hard to judge without the full numbers.\n\nThis paper is for people building ML models for FWI who need larger training sets. Readers working on neural operators or inversion will get immediate use from the data. It deserves a serious referee because the scale and the public artifacts are substantive even if the generator's fidelity requires scrutiny in review.\n\nI would send it to peer review.","headline":"This is a data and tool release that gives ML researchers a bigger, more diverse set of field-scale seismic models than prior benchmarks, with OOD splits included.","tokens_in":2431,"tokens_out":402,"would_cite":true,"duration_ms":19769,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"SubsurfaceGen generates field-scale 3D velocity models and paired seismic data to train and test ML methods for full waveform inversion.","keywords":["procedural generation","full waveform inversion","velocity models","seismic data","machine learning","subsurface imaging","neural operators","dataset"],"falsifier":"If models trained on the SubsurfaceGen dataset show the same performance and failure modes on real seismic surveys as models trained only on smaller existing datasets like Marmousi or OpenFWI, the claim that the new data impacts ML-based FWI would not hold.","tokens_in":2713,"feed_emoji":"🗺️","tokens_out":724,"duration_ms":31610,"temperature":0.7,"pith_summary":"The paper claims that existing datasets for machine learning in full waveform inversion are too small, lack geological variety, and do not reach field scales, limiting their usefulness. It introduces SubsurfaceGen as a generator that creates 3D velocity models at 10 km by 10 km by 6.19 km with 10 m resolution along with matching wavefields and shot gathers. The work releases a dataset drawn from 42 such models across six geological settings and uses it to test neural operators and encoder-decoder models with one setting held out. Experiments show that these models encounter specific problems at realistic scales that smaller datasets do not reveal.","feed_headline":"Generator creates 42 field-scale 3D models for seismic ML","feed_subtitle":"Dataset of 4,276 velocity slices and wavefields from 10 km models reveals failure modes in neural FWI methods at realistic scale.","key_machinery":"SubsurfaceGen, a GPU-accelerated procedural generator that builds physically realistic 3D velocity models and corresponding seismic data at field scale.","core_discovery":"SubsurfaceGen produces 3D velocity models spanning 10 km x 10 km laterally and 6.19 km deep at 10 m resolution, together with 5 s wavefields and 8 s shot gathers. The released dataset contains 4,276 2D velocity slices from 42 models in six geological settings, supporting tests of wavefield prediction and end-to-end velocity inversion with out-of-distribution evaluation on held-out settings.","pith_inferences":["If the generated models capture essential geological features, the same generator could supply training data for related geophysical tasks such as travel-time tomography.","Larger volumes of paired data from this approach might allow systematic study of how model size and geological complexity interact in learned inversion methods.","The procedural nature of the generator suggests it could be extended to produce time-lapse or anisotropic models for testing more advanced inversion scenarios."],"forward_implications":["Neural operators for wavefield prediction and encoder-decoders for velocity inversion can be evaluated at spatial and temporal scales matching actual field applications.","Holding out entire geological settings enables direct measurement of out-of-distribution performance in subsurface imaging tasks.","Paired velocity and seismic data allow quantitative assessment of end-to-end inversion accuracy at 10 m resolution over 10 km extents.","The dataset construction from six distinct geological settings supports targeted testing for applications such as carbon sequestration and hydrocarbon exploration."],"fun_headline_variants":["SubsurfaceGen creates 42 field-scale 3D velocity models and wavefields","Field-scale seismic dataset with 4276 slices from 42 3D models","42 realistic 3D models support ML tests for field-scale velocity inversion","GPU generator produces 5s wavefields and 8s shot gathers at 10m res"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The procedurally generated models are realistic and diverse enough that training and testing on them reveals the actual limitations ML models will face with real field data.","fun_headline_variants_meta":{"raw":{"variants":["SubsurfaceGen creates 42 field-scale 3D velocity models and wavefields","Field-scale seismic dataset with 4276 slices from 42 3D models","42 realistic 3D models support ML tests for field-scale velocity inversion","GPU generator produces 5s wavefields and 8s shot gathers at 10m res"]},"model":"grok-4.3","cost_usd":0.012041,"raw_usage":{"total_tokens":5279,"prompt_tokens":710,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":120412000,"prompt_tokens_details":{"text_tokens":710,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4483,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":710,"tokens_out":86,"duration_ms":34242,"temperature":1.0,"reasoning_tokens":4483,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T08:51:31.553281+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If models trained on the SubsurfaceGen dataset show the same performance and failure modes on real seismic surveys as models trained only on smaller existing datasets like Marmousi or OpenFWI, the claim that the new data impacts ML-based FWI would not hold.","supporting_citations":[],"review_version":1}