{"id":"a7e7edff-99a7-4857-9231-5117477e0900","arxiv_id":"2508.16640","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A conditional neural field plus latent diffusion model jointly generates geological models and flow responses, enabling zero-shot Bayesian inversion for CO2 storage.","lead":"This paper adapts a generative AI framework, CoNFiLD-geo, to infer underground rock properties and CO2 plume movements in carbon storage reservoirs from sparse monitoring data. It reports that a single pretrained model can do both forward prediction and inverse uncertainty quantification without retraining.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"All validation is in-distribution: every test case uses fields sampled from the same Gaussian random model family as training, so the 'real-world' and zero-shot generalization claims are untested.","rationale":"The stress-test identifies the same weakest assumption as the reader: the absence of any out-of-distribution or real-data validation. This is load-bearing because the entire inverse-modeling capability rests on the learned prior; if the prior does not cover the test geology, the 'posterior' is meaningless. No internal inconsistency in the derivations was found; the DPS-style likelihood guidance and CNF-LDM architecture are coherent. The reader's CONDITIONAL verdict is appropriate: the method is a promising contribution but its central claims are not yet supported. No adjustment is needed.","tokens_in":47769,"tokens_out":5091,"duration_ms":57531,"concrete_test":"Train the Case-1 model exactly as described (Gaussian permeability, correlation length 80 m, 2000 realizations). Then condition on 16×16 low-resolution CO2 saturation from test fields drawn from (i) a Gaussian model with correlation length 160 m, and (ii) a channelized model (e.g., from a training image such as Stanford VI). Compute SSIM/RMSE of inferred permeability and saturation against the reference. If performance degrades by more than ~30% relative to the in-distribution test, the zero-shot generalization claim fails and the real-world claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that CoNFiLD-geo can perform zero-shot conditional generation for 'real-world' GCS scenarios. But every validation case is in-distribution: Case 1 permeability fields are drawn from a Gaussian covariance model with fixed mean/correlation (Supp. 2.2); Case 2 uses the Sleipner geometry but samples permeability from another Gaussian model (Supp. 2.3), and the 'observations' are downsampled/probed versions of the same simulated reference; Case 3 depth and thickness are also Gaussian random fields (Supp. 2.4). No test uses permeability, geometry, or operating conditions outside the training distribution, and no actual field data are used. The posterior sampler is only as good as the learned prior p(Φ); if a new site's geology lies off the prior manifold, the 'posterior' samples are dominated by the prior and are not trustworthy. The Discussion implicitly concedes this, stating the model 'implicitly learns physical priors through data-intensive training' and proposing physics-based constraints to reduce that dependence. Thus the key evidence for the headline claim—generalization to unseen real-world geology—is absent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CoNFiLD-geo, a two-stage generative framework for geological carbon storage. A conditional neural field (SIREN with full-projection conditioning) encodes spatiotemporal fields Φ=(M,U) into a compact latent sequence z0, and a latent diffusion model learns the joint distribution p(z0). At inference, observations Ψ are incorporated through a Tweedie/DPS-style guided score, yielding conditional samples from p(Φ|Ψ) without task-specific retraining. The method is tested on three cases: 2D heterogeneous CO2 drainage, a single-layer Sleipner-like model with synthetic Gaussian permeability, and a stratigraphically complex model with Gaussian-random depth and thickness. Conditional generation is evaluated with visual comparisons, SSIM, and RMSE for low-resolution seismic data, sparse wells, multi-source monitoring, and missing-data restoration; an external comparison against the deterministic forward model U-FNO appears only in the Supplementary Notes.","tokens_in":48031,"tokens_out":4930,"duration_ms":65185,"significance":"If the generalization and calibration claims were substantiated, the framework would be a useful unified inverse solver and forward surrogate for GCS, with the mesh-agnostic CNF and the zero-shot flexibility over observation operators being genuine strengths. The methodological building blocks (CNF+LDM, Tweedie/DPS guidance) are standard, and the application to joint geomodel-response inversion is a natural and timely extension. However, all quantitative evidence is in-distribution synthetic data: one held-out trajectory per case, an ensemble of 10 generated samples, no inverse-modeling baseline, and no test with a distribution shift in geology, grid, or operating conditions. The 'real-world' and 'zero-shot generalization' claims are therefore currently promises rather than demonstrated results. The paper is clearly written, the derivations are largely standard, and the commitment to release data/code is a positive feature.","major_comments":[{"comment":"The central generalization claim is not supported by the validation design. The abstract and §1 advertise zero-shot conditional generation for 'real-world' GCS scenarios, and §2.3 calls the Sleipner case 'field-scale'/'realistic'. But the test permeability fields in Case 2 are sampled from the same Gaussian covariance model used to generate training data (Supp. Note 2.3), the observations are downsampled or masked versions of the same simulated reference, and Case 3 depth/thickness are likewise Gaussian random fields drawn from the same family as the training set (Supp. Note 2.4). No case evaluates a shift in geology, grid resolution, or operational parameters. Since the posterior sampler inherits the learned prior p(Φ), the reported success on these in-distribution tasks cannot be read as evidence for generalization to unseen real-world geology. The Discussion (end of §3) implicitly con","section":"§2.2–2.4, Supplementary Notes 2.2–2.4"},{"comment":"The claimed advantages over existing inverse modeling methods are not benchmarked. The only external baseline is U-FNO, a deterministic forward surrogate compared in the Supplementary Notes for forward prediction tasks. There is no comparison against established inverse methods (e.g., ES-MDA, MCMC with the numerical simulator, or a conditional GAN/VAE/diffusion baseline) on the same test cases. The manuscript's claims of 'superior efficiency, generalization, scalability, and robustness' for data assimilation require at least one inverse baseline measuring posterior mean/median accuracy, uncertainty calibration, and wall-clock time. Without this, the contribution relative to existing GCS inversion workflows is not quantified.","section":"§2.2–2.4, Supplementary Note 4"},{"comment":"The statistical evidence for uncertainty quantification is thin. The ensemble size is 10 throughout the manuscript, the main text shows a single reference trajectory per case, and the reported SSIM/RMSE statistics are means and standard deviations over these 10 samples. No calibration metrics (coverage, reliability diagram, interval score) are reported, and no multiple test trajectories are aggregated. The claim that CoNFiLD-geo provides reliable uncertainty quantification in 'real-time' is therefore not supported by the present experiments. Reporting results over more independent test fields and adding a calibration check would materially strengthen the posterior-sampling claim.","section":"Figs. 2–5, §2.2–2.4"},{"comment":"The likelihood variance σ_c^2 is a critical tuning parameter but no procedure for setting it is given. The posterior width and the strength of the data-guidance gradient scale inversely with σ_c^2. In Case 3 the seismic data are said to contain 5% noise, and Fig. 6 perturbs observations by 0%, 10%, and 30% noise, but there is no statement of how σ_c^2 was chosen for these experiments or any sensitivity analysis with respect to it. Since the paper's UQ is Bayesian, the observation-noise model should be specified and its influence on the reported posterior intervals assessed.","section":"§4.4, Eqs. (24), (29)–(30)"}],"minor_comments":[{"comment":"The terminology 'real-world GCS scenarios' is misleading for Case 2 because only the geometry is realistic while the permeability is synthetic and drawn from the training distribution. Consider using 'field-inspired geometry' or 'synthetic field-scale model' throughout.","section":"§1, §2.3"},{"comment":"The factor of 2 in the gradient is written explicitly in Eq. (30) but not in Eq. (29); the notation is understandable but could be clarified by writing the gradient of the squared norm explicitly in Eq. (29).","section":"Eq. (29) and Eq. (30)"},{"comment":"Small typos: 'relatioships' (Eq. 1), 'trival' (after Eq. 9), 'uncontional' (Eq. 22), 'orginal' (§4.2), 'CoNFoLD-geo' (Supp. Note 6.3).","section":"Throughout"},{"comment":"The shaded regions labeled 'standard deviation' are computed over only 10 samples. State this explicitly in the figure captions or text, and consider showing individual samples or box plots for the key metrics.","section":"Figs. 2–5"},{"comment":"The statement that CoNFiLD-geo is 'largely comparable' to U-FNO in the fully observed case is fine, but the Figure S8 caption should indicate which rows are CoNFiLD-geo samples and which are U-FNO predictions for clarity.","section":"Supplementary Note 4.4"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern in the reader's report lands: the validation is entirely in-distribution and the headline 'zero-shot / real-world' generalization is not evidenced. I recommend major revision rather than reject because the methodological core appears sound and the missing evidence could in principle be supplied: out-of-distribution tests, an inverse baseline, and calibration metrics. If the authors choose to keep the real-world framing, a substantial rewrite of the claims is needed; otherwise the paper would be better framed as a demonstration of a flexible generative framework on synthetic GCS scenarios."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the joint encoding of geomodel and reservoir response in one CNF latent space is a real step beyond the authors' prior CoNFiLD work, and the zero-shot posterior sampling pipeline is cleanly derived and clearly written. But the paper validates entirely in-distribution, with synthetic geology drawn from the same Gaussian random fields used for training, so the \"real-world\" claims are not backed by the evidence shown.\n\nWhat's new: previous diffusion-based GCS papers either modeled states only or required retraining for each observation scenario. CoNFiLD-geo learns a shared latent space over M(x) and U(x,t), letting one pretrained model handle both forward prediction and Bayesian inversion from sparse, noisy observations. That is a legitimate contribution. The bias/uncertainty decomposition in Figure 6 (reconstruction, sampling, and alignment bias) is also more candid than most papers in this space.\n\nSoft spots, in proportion: the stress-test note is correct. Every test case samples permeability (and in Case 3 depth/thickness) from the same Gaussian covariance family as the training set, and the \"observations\" are downsampled or masked versions of the same simulated reference. The Sleipner case uses the real geometry but not real geology or real monitoring data. The Discussion implicitly concedes this when it says the model relies on data-intensive priors and needs physics-based constraints. That is an honest acknowledgment, but it should have shaped the headline claims. Second, there is no inverse baseline — no ES-MDA, no MCMC, no data-space inversion. The U-FNO comparison is forward-only. Third, quantitative results come from one held-out trajectory per case with 10 samples; no aggregate statistics over many test realizations. And code/data are only \"available upon acceptance,\" which is not a release. None of these are fatal, but together they mean the evidence is suggestive, not demonstrative.\n\nWho this is for: anyone working on generative inverse modeling for subsurface flow, and GCS modelers who want a fast prior-based inference tool. The framework is transferable beyond CO2 storage.\n\nRecommendation: send it to peer review. A good referee should require out-of-distribution tests (different correlation lengths, channelized geology, different well controls), a real inverse baseline, and actual release of code and data. With those, this could be a solid contribution; without them, the usefulness for real sites remains unproven.","headline":"Solid extension of the authors' CoNFiLD framework to GCS with a genuinely new joint latent encoding, but the validation stays entirely in-distribution, so the real-world claims outrun the evidence.","tokens_in":48513,"tokens_out":1783,"would_cite":true,"duration_ms":25610,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single pretrained model learns geology and CO2 flow jointly, then inverts sparse, noisy, or incomplete monitoring data into consistent geology–response pairs by Bayesian posterior sampling, with no per-task retraining.","keywords":["generative modeling","geological carbon sequestration","inverse modeling","uncertainty quantification","latent diffusion model","conditional neural field","Bayesian posterior sampling","zero-shot conditional generation"],"falsifier":"Condition the pretrained Case-1 model on CO2 saturation snapshots from a permeability field with double the training correlation length (160 m instead of 80 m), or from a channelized non-Gaussian facies model, and compare the inferred permeability and plume evolution against a fresh two-phase-flow simulation. The central claim fails if the posterior samples drift from the reference while the model's uncertainty bands stay narrow — i.e., if the reported confidence is systematically wrong for conditioning data drawn outside the training distribution. (This test awaits the data and code, which th","tokens_in":47626,"feed_emoji":"🛢️","tokens_out":16833,"duration_ms":174278,"temperature":0.7,"pith_summary":"The paper claims that inverse modeling for geological carbon sequestration — inferring hidden permeability or reservoir geometry from what monitoring can see — can be compressed into a single sampling pass of one pretrained generative model. CoNFiLD-geo learns the joint distribution of the geomodel $M(x)$ and its CO$_2$ flow response $U(x,t)$ inside a shared latent space: a conditional neural field encodes both into the same latent trajectory, a latent diffusion model learns that joint distribution, and at inference a likelihood gradient steers the denoiser toward whatever data are available, from a single well to a coarse seismic plume image. Because the steering happens at sampling time, the same network serves as forward surrogate and inverse solver, and each new monitoring configuration is handled without retraining. The payoff, demonstrated on a 2D synthetic reservoir, the Sleipner field site, and an unstructured-grid reservoir with brine production, is a posterior ensemble with quantified uncertainty in under a minute to a few minutes, compared with 5–30 minutes for one high-fidelity simulation.","feed_headline":"Sparse well data recovers full CO2 reservoir fields without retraining","feed_subtitle":"A single pretrained diffusion model turns sparse, noisy readings into permeability and CO2 plume fields, with uncertainty.","key_machinery":"The load-bearing object is the joint latent trajectory $z_0=[L_1,\\ldots,L_{N_t}]$: each column $L_t$ is the conditional-neural-field code for one time step of the combined field $\\Phi(x,t)=[M(x),U(x,t)]$. The CNF is a coordinate-based network — a sinusoidal representation network (SIREN) — modulated by full-projection conditioning in an auto-decoding formulation, making it mesh-agnostic and queryable at arbitrary locations. The latent diffusion model learns $p(z_0)$; the zero-shot step is the guided score $s_{\\theta^\\ast}+\\nabla_{z_\\tau}\\log p(\\Psi|z_\\tau)$, where the likelihood gradient is computed by automatic differentiation through the decoder using Tweedie's formula. That gradient — com","core_discovery":"Geology and flow share one latent representation. CoNFiLD-geo concatenates the geomodel $M(x)$ with the responses $U(x,t)$ into a joint field $\\Phi=[M,U]$, encodes it with a conditional neural field into a latent trajectory $z_0$, and trains a diffusion model on that trajectory, so the prior $p(z_0)$ carries the physics linking parameters to solutions. At inference, observations enter as a likelihood through the differentiable decoder — Tweedie's formula plus a Jensen approximation — steering sampling toward the posterior over the whole joint field; conditioning on any observed slice constrains the rest, and a batch of samples is an uncertainty-quantified answer. Validation spans three setti","pith_inferences":["The zero-shot claim is tested strictly inside the training distribution — every validation field is a new draw from the same Gaussian random-field statistics used for training, and every observation is derived from the simulated reference field — so the genuine stress test is deployment on a site whose geology, grid, or operational parameters differ from the training set.","The joint-latent trick is not CO2-specific: the same architecture should transfer to geothermal, hydrogen storage, or groundwater inverse problems, where the physics enters only through the training data and the differentiable likelihood.","The paper's own proposal to add PDE-residual constraints at the sampling stage is the most natural extension: it would attack the 'alignment bias' it identifies between parameter and solution spaces, and could reduce the volume of training simulations needed.","The 20-well (0.5% data) reconstruction suggests an empirical limit worth mapping: how few probes, and of which variables, suffice for a trustworthy posterior depends on how compressible the prior geology is — a question this framework makes directly measurable."],"forward_implications":["One pretrained CoNFiLD-geo is both a forward surrogate and an inverse solver: unconditional generation acts as a fast numerical emulator (~20 s per field) and conditional posterior sampling (48–192 s) performs data assimilation, against 5–30 minutes for a single high-fidelity simulation.","New monitoring configurations — different well counts and positions, coarser seismic images, noisier or partially damaged records — are handled by the same network at inference time, with no task-specific retraining.","Uncertainty quantification becomes a batch operation: an ensemble of posterior (geology, response) pairs is produced in one generation run, and the spread shrinks as observations become more informative, as Bayesian reasoning expects.","Because the CNF is mesh-agnostic, the same framework works on unstructured grids and complex stratigraphy, where CNN-based latent methods are limited or fail outright.","Under sparse or low-resolution permeability inputs, CoNFiLD-geo matches or beats the deterministic U-FNO surrogate on forward prediction while also returning uncertainty bands; the trade-off is slightly lower accuracy than U-FNO when the permeability is fully observed."],"supporting_citations":[{"why":"The CoNFiLD framework being extended; supplies the conditional-neural-field plus latent-diffusion design and the Bayesian conditional sampling machinery.","marker":"[57]"},{"why":"CoNFiLD-inlet; supplies the full-projection conditioning scheme that modulates the SIREN layers from the latent code.","marker":"[58]"},{"why":"Denoising diffusion probabilistic models; the diffusion training objective and reverse-process parameterization used in the latent space.","marker":"[51]"},{"why":"SIREN; the sinusoidal representation network that parameterizes the conditional neural field.","marker":"[77]"},{"why":"DeepSDF; the auto-decoding formulation used to optimize the latent codes during CNF pretraining.","marker":"[65]"},{"why":"Diffusion posterior sampling; the likelihood-gradient guidance that makes zero-shot conditional generation work in the latent space.","marker":"[80]"},{"why":"Tweedie's formula; the identity used to derive the clean-latent estimate that links the likelihood to the noisy latent during sampling.","marker":"[81]"},{"why":"The Sleipner 2019 benchmark model; supplies the realistic Utsira L9 stratigraphy for the field-scale case.","marker":"[70]"},{"why":"TOUGH2; the finite-volume simulator whose governing equations and outputs generate the training and reference data.","marker":"[76]"},{"why":"U-FNO; the deterministic neural-operator surrogate used as the baseline in the forward-modeling comparison.","marker":"[29]"}],"fun_headline_variants":["Diffusion model turns sparse well data into full CO2 fields","Zero-shot diffusion for CO2 reservoir uncertainty quantification","Joint latent diffusion links geology and flow for CO2 modeling","Pretrained diffusion model assimilates CO2 data without retraining","Latent diffusion captures geology and flow for carbon storage"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The entire demonstration is in-distribution: the validation permeability fields (and, in Case 3, depth and thickness) are drawn from Gaussian random fields with the same covariance parameters as the training set, and the 'observations' are downsampled, masked, or perturbed versions of those same simulated reference fields (Supplementary Notes 2.2–2.4); if a real site's geology, grid, or operational parameters shift away from that synthetic family, the posterior samples the mo","fun_headline_variants_meta":{"raw":{"variants":["Diffusion model turns sparse well data into full CO2 fields","Zero-shot diffusion for CO2 reservoir uncertainty quantification","Joint latent diffusion links geology and flow for CO2 modeling","Pretrained diffusion model assimilates CO2 data without retraining","Latent diffusion captures geology and flow for carbon storage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000881,"raw_usage":{"total_tokens":3658,"prompt_tokens":770,"completion_tokens":2888,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2808}},"tokens_in":514,"tokens_out":2888,"duration_ms":23921,"temperature":1.0,"reasoning_tokens":2808,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:26:02.033135+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Condition the pretrained Case-1 model on CO2 saturation snapshots from a permeability field with double the training correlation length (160 m instead of 80 m), or from a channelized non-Gaussian facies model, and compare the inferred permeability and plume evolution against a fresh two-phase-flow simulation. The central claim fails if the posterior samples drift from the reference while the model's uncertainty bands stay narrow — i.e., if the reported confidence is systematically wrong for conditioning data drawn outside the training distribution. (This test awaits the data and code, which th","supporting_citations":[],"review_version":1}