{"id":"31ac57cf-1c57-4c8e-8200-414b71807b4d","arxiv_id":"2607.18557","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Agents4GEOS is an MCP-based agentic framework that translates natural language into validated GEOS simulation decks and generated a 200-case dataset for training a GNN plume surrogate.","lead":"A research group built an AI-agent system that turns plain-English requests into ready-to-run simulations of CO2 storage using the open-source GEOS simulator, and used it to create training data for a graph-neural-network surrogate. The one demonstrated success is a reproduction of a published PUNQ-S3 benchmark, but the paper provides no code, no data, and no numerical error bars, so the scale of the claimed capability is hard to verify.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation claim lacks quantitative comparison: single qualitative match, untested modeling assumptions (dead-oil, depth shift), and unresolved injection-rate discrepancy between sources.","rationale":"The paper presents an interesting agentic framework, but the central claim of validated simulation fidelity is supported only by a single session with qualitative agreement. The reader's weakest assumption focused on the modeling equivalence; this stress-test concurs but broadens the concern: even if the equivalence held, the paper provides no quantitative error metrics or full-field comparison against the ECLIPSE/MRST references. The injection-rate ambiguity between [11] and [12] further undermines the choice of reference data. These issues are addressable with artifact release and numeric validation, so conditional acceptance remains appropriate. No reason to reject outright, but the claim should not be taken at face value.","tokens_in":10556,"tokens_out":5140,"duration_ms":58774,"concrete_test":"Obtain the published ECLIPSE and MRST reference saturation fields for PUNQ-S3 (e.g., from [11]/[12] or their supplementary data) and compute quantitative error metrics (RMSE, L2 error over all active cells and time snapshots) between the GEOS output produced by Agents4GEOS and each reference. Additionally, run two controlled sensitivity experiments using the same agent-built deck: (a) replace dead-oil with a compositional CO2-brine model, and (b) restore the original 2340 m depth, keeping all other parameters fixed; then compare the observation-point curves and global saturation fields. If the quantitative error is small and the match degrades substantially under either perturbation, the modeling-equivalence assumption is load-bearing; if the error is large even for the nominal case, the validation claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 4) that 'validation against published ECLIPSE and MRST results confirms the fidelity of the GEOS simulations produced by Agents4GEOS' is not substantiated by the evidence presented. The only validation shown is Figure 7, a comparison of gas saturation versus time at three observation points, described as 'matching perfectly the results of [12]'—with no numerical error metrics, no comparison of full 3D saturation fields, and no explicit identification of which reference results come from ECLIPSE versus MRST. The reproduction itself relies on two modeling alterations recommended by the agent and accepted by the user: modeling the system as immiscible dead-oil instead of compositional CO2–brine, and relocating the formation from ≈2340 m to 840 m. Section 3.1 explicitly states the relocation 'is assumed not to interfere with the exploration of the main aspects of the flow dynamics,' an assumption that is never tested. Moreover, the two source papers disagree on injection rate by over an order of magnitude; the agent chose the literal target value, but the sensitivity of the match to this choice is not explored. Thus, the claim that Agents4GEOS produces physically consistent datasets suitable for GNN training rests on a single anecdotal, qualitative agreement, not a quantitative benchmark. If the dead-oil/depth approximations alter plume migration away from the three sampled observation cells—or if the injection-rate discrepancy resolves differently—the validation collapses. This is load-bearing because the entire value proposition (large validated training sets) depends on the fidelity of GEOS simulations being real and reproducible beyond one session.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Agents4GEOS, a Model Context Protocol (MCP) based multi-agent system with 52 deterministic tools, eleven slash-command agents and four fresh-context subagents, plus knowledge modules distilled from more than 200 GEOS input files. The stated goal is to lower the barrier to producing schema-valid, physically consistent GEOS simulation decks from natural-language prompts. The central demonstration is a single conversational session that reproduces a PUNQ-S3 no-hysteresis CO2 storage benchmark from a one-sentence request, followed by the generation of 200 permeability-perturbed simulations, training of the Plumecast GNN surrogate, and a visual comparison of predicted and simulated CO2 saturation at 100 years. The paper claims in Section 4 that 'validation against published ECLIPSE and MRST results confirms the fidelity of the GEOS simulations produced by Agents4GEOS.'","tokens_in":10916,"tokens_out":6233,"duration_ms":73410,"significance":"If the validation claim can be made quantitative and the modeling assumptions tested, the contribution is significant: it would reduce a real bottleneck in generating large training datasets for GNN surrogates on unstructured meshes. The architecture has genuine strengths: a strict separation between agents (decisions), tools (computation), and knowledge modules (domain patterns); structured JSON contracts for subagent handoffs; an independent fresh-context reviewer; and grounding of every returned quantity in actual computation. These are valuable design choices worth publishing. However, the evidence that the produced GEOS simulations reproduce the published benchmarks is currently mainly qualitative, and the manuscript itself flags the two most load-bearing limitations: the dead-oil/depth modeling assumption and the endpoint-fitted reversible relative-permeability curve. Those limitations need to be resolved before the central claim can be accepted.","major_comments":[{"comment":"The sentence 'Validation against published ECLIPSE and MRST results confirms the fidelity ...' is the central evidence, but the only benchmark comparison is a three-point saturation-versus-time figure described as 'matching perfectly.' No error metrics are reported, no curve is identified as coming from ECLIPSE ([12]) versus MRST ([11]), and no spatial comparison of full 3D saturation fields is shown. Please add quantitative metrics (e.g., RMSE or max error per observation point), overlay the reference curves on Figure 7, and state explicitly which source each curve is from. As written, the reader cannot distinguish a genuine quantitative match from a favorable visual impression.","section":"Section 4 and Figure 7"},{"comment":"Two modeling changes are introduced before the benchmark comparison: the compositional CO2-brine system is replaced by immiscible dead-oil, and the formation is relocated from approximately 2340 m to 840 m. The text states this 'is assumed not to interfere with the exploration of the main aspects of the flow dynamics,' but the assumption is never tested. The injection-rate discrepancy between the two sources (18 m3/day per well versus 0.15 pore volumes in 10 years) is also unresolved; the agent kept the literal value, but the sensitivity of the results to the alternative value is not explored. Because these choices directly condition the validation claim, the paper should quantify their impact, for example by running the original depth/compositional configuration or a rate-sensitivity study.","section":"Section 3.1"},{"comment":"The 'no-hysteresis signature' in Figure 7 is largely imposed by construction. The agent fits a reversible Brooks-Corey curve using only the published endpoints (Swc = 0.31, Sg,max = 0.69), and a reversible curve cannot exhibit hysteresis. The paper acknowledges that the curve is 'an endpoint fit rather than the raw published curve,' but then Figure 7 tests only that fit, not the physical fidelity of the simulation. Please compare the fitted curve against the full published drainage relative-permeability data over the saturation range, and examine the sensitivity of the saturation histories to the fitted exponent and endpoints.","section":"Section 3.1"},{"comment":"The Plumecast surrogate comparison is qualitative only: no error metric is reported, and the dataset-generation protocol is underspecified. In particular, the permeability sampling range and distribution, the construction of the 100/100 train/test split, and whether the displayed case belongs to the test set are not stated. Since the paper presents the 200-simulation dataset as a contribution for GNN training, quantitative evaluation over the test set (e.g., RMSE of CO2 saturation fields, plume-area error as a function of time) is needed. Note also that comparing a GEOS-trained surrogate with GEOS validates surrogate accuracy, not the physical fidelity of GEOS itself; the latter must rest on the benchmark comparison.","section":"Section 3.2 and Figure 11"}],"minor_comments":[{"comment":"Typo: '18 rm3/day' should be '18 m3/day'. Also, Figures 9 and 10 use 'PUNQ-3D' while the text uses 'PUNQ-S3'; please make the nomenclature consistent.","section":"Section 3.1"},{"comment":"The placeholder '(ADD REFS)' appears after 'PyG and PGT'; unresolved references must be completed before publication.","section":"Section 3.2"},{"comment":"The caption says 'cf. Figure 6 of [12]', but Section 4 claims validation against both ECLIPSE and MRST results. Specify in the caption which reference curve (ECLIPSE or MRST) is being compared, and whether Figure 7 includes data from both sources.","section":"Figure 7 caption"},{"comment":"Figures 9 and 10 are not discussed in the main text. Either reference them explicitly in Section 3.2 or remove them to avoid dangling figures.","section":"Figures 9 and 10"},{"comment":"Given the paper's emphasis on open-source software and reproducibility, please add a code/data availability statement with a repository URL and version/commit identifier for Agents4GEOS, the GEOS version used, and the 200-run dataset generation seeds.","section":"General"},{"comment":"Minor typo: in the decision-gate text, 'COz-brine' should be 'CO2-brine'.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The reader's conditional assessment matches my reading: the platform is plausible and the architectural choices are interesting, but the validation claim is not quantitatively supported and the manuscript itself acknowledges the key unresolved points. The gaps are fixable with additional experiments and metrics, so I recommend major revision rather than rejection. A supplementary transcript of the agent session or a link to a public repository would substantially strengthen the reproducibility of the demonstration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a genuinely useful engineering description with a validation claim that outruns the evidence. The architecture — agents decide, tools compute, knowledge modules encode expertise — is a sensible separation, and the 52 MCP tools plus knowledge modules distilled from 200+ GEOS files are a real, reusable contribution. The session walkthrough is the best part: it shows failures (well-depth misalignment, dead-oil component naming) and the agent diagnosing them, and it explicitly reports the injection-rate discrepancy between the two source papers instead of hiding it. That honesty is rare and should be credited.\n\nThe soft spot is the central validation. Section 4 says 'validation against published ECLIPSE and MRST results confirms the fidelity' but the only evidence is a single figure of gas saturation versus time at three observation points, described as matching 'perfectly' with no error metrics, no field-level comparison, and no breakdown of which reference results come from which simulator. The reproduction itself depends on two untested modeling shifts — dead-oil instead of compositional CO2–brine, and moving the formation from roughly 2340 m to 840 m — and the paper acknowledges the second is assumed not to interfere. The agent kept the literal injection rate even though the two sources disagree by an order of magnitude; the match could be sensitive to that choice. These are not fatal by themselves, but they mean the 'fidelity' claim is not yet substantiated. The surrogate comparison has the same interpolation problem: Plumecast is trained on GEOS data and then 'validated' against GEOS, so it mostly demonstrates that the surrogate fits its training distribution.\n\nThere is also a placeholder in the text ('ADD REFS'), and the paper is light on reproducibility artifacts — no code/data release is mentioned, so the single session cannot be independently checked.\n\nOn balance, this is not a takedown: the framework is plausible, the description is clear, and the dataset-generation pipeline is a useful demonstration. What it needs is a quantitative evaluation across multiple sessions, release of the artifacts, and a direct numerical comparison against the published references at the chosen modeling assumptions. That is exactly what a serious referee could ask for. I would send it to review, but with the expectation of major revision. The audience is researchers building LLM-based tools for subsurface simulation and anyone relying on synthetic GEOS data for surrogate training; they should treat the validation as provisional until the artifacts appear.","headline":"Well-engineered, honest engineering description whose central validation claim outruns the evidence.","tokens_in":11432,"tokens_out":2989,"would_cite":false,"duration_ms":32070,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Agents4GEOS turns plain-English requests into validated GEOS CO2-sequestration simulations and trains a GNN surrogate that tracks plume migration.","keywords":["agentic AI","GEOS simulator","CO2 sequestration","graph neural networks","surrogate models","reservoir simulation","Model Context Protocol","physics-informed ML"],"falsifier":"Run both the agent-generated dead-oil GEOS deck and a full compositional CO2-brine GEOS deck on the same mesh and depth, and compare the saturation fields at the three PUNQ-S3 observation cells over the 500-year horizon; a clear difference in plume shape, crest saturation, or flank draining times would falsify the claim that the approximation preserves the main flow dynamics.","tokens_in":10485,"feed_emoji":"🤖","tokens_out":7155,"duration_ms":70912,"temperature":0.7,"pith_summary":"This paper claims that a multi-agent AI system called Agents4GEOS can lower the barrier to building physically consistent, validated multiphysics simulations of subsurface CO2 storage. The system takes a natural-language request, decomposes it into specialized subagents, and produces a schema-valid GEOS input file through real computations of fluid properties, meshes, and XML assembly — not free-form text. As its main demonstration, the agents reproduce the PUNQ-S3 no-hysteresis CO2-sequestration benchmark from a one-sentence prompt, and the resulting GEOS runs are compared against published ECLIPSE and MRST results. The same pipeline generates 200 permeability-varying cases on which the Plumecast graph neural network is trained, and the surrogate's 100-year CO2 saturation fields match the high-fidelity GEOS ground truth well. If correct, this enables large, validated simulation datasets and surrogate models for uncertainty quantification and many-query reservoir workflows, at a fraction of the usual manual effort.","feed_headline":"One prompt yields validated CO2 simulations for GNN training","feed_subtitle":"Agents4GEOS reproduces a published benchmark and trains a plume-predicting surrogate from 200 agent-built cases.","key_machinery":"The load-bearing mechanism is the strict separation between the agent layer (eleven slash-command agents plus four fresh-context subagents that return typed JSON contracts), the tool layer (52 stateless MCP tools grouped into six scientific domains and backed by real computation libraries such as pyResToolbox and PyVista), and the knowledge modules (seven Python modules encoding GEOS field names, fluid models, cross-references, sanity rules, unit conventions, formatting, and preprocessing, distilled from an audit of 200+ official GEOS input files). Orchestration patterns — pipeline, fan-out, feedback loop, and quality contract — coordinate these pieces, with a fresh-context independent revie","core_discovery":"Agents4GEOS is an agentic framework built on the Model Context Protocol in which agents plan, tools compute, and knowledge modules encode domain expertise. The paper's central demonstration is that a user asking, in plain English, to reproduce a published PUNQ-S3 CO2-sequestration study can receive, in a single session, a schema-valid, physics-checked GEOS XML deck that reproduces the published saturation behavior at the three reference observation points — with the system openly flagging an injection-rate discrepancy it could not reconcile between the two source papers. The paper further claims that the 200-case dataset produced this way is validated against published ECLIPSE and MRST resul","pith_inferences":["If the dead-oil/depth-relocation equivalence generalizes, agent-driven approximations could become a standard, reported practice in benchmark reproduction — but the unvalidated equivalence is a risk that should be tested on other reservoirs before broad claims about fidelity are made.","The same separation-of-concerns architecture could transfer to other XML-driven multiphysics simulators beyond GEOS, such as thermal-hydraulic or geomechanics codes, whenever schema-valid input generation and dataset curation are bottlenecks.","The system's habit of surfacing the unreconciled injection-rate discrepancy rather than silently picking a number is a model for agentic scientific tools that prioritize transparency over apparent smoothness.","A direct stress test: generate the same PUNQ-S3 case with hysteresis and compare the agent's relative-permeability endpoint fit against the raw published curves — if the fit fails to reproduce hysteresis behavior, the fidelity claim narrows to the no-hysteresis setting."],"forward_implications":["Natural-language-to-GEOS workflows could let domain experts author and audit complex multiphysics simulations without hand-writing hundreds of lines of XML.","Agent-built datasets, validated against published benchmarks, provide a controlled source of training data for GNN surrogates of CO2 plume migration.","The Plumecast results suggest physics-informed features — transmissibility-aware edges, tabulated relative permeabilities, Kozeny-Carman porosity — improve long-horizon surrogate accuracy and reduce false plume spreading.","The reusable learning loop that persists runtime-error lessons into knowledge modules reduces the chance of repeating the same mistakes across future simulation decks.","The system's capability-tier model routing indicates that cost-aware agent architectures can keep token budgets manageable while maintaining reliability on complex tasks."],"fun_headline_variants":["AI agents auto-build validated CO2 simulation datasets","Plain-English prompts drive CO2 simulations for GNN training","Agentic framework reproduces benchmark CO2 studies","MCP-based agents simplify CO2 simulation setup"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That representing the PUNQ-S3 reservoir as an immiscible dead-oil fluid (ignoring gas dissolution) and shifting it from roughly 2340 m to 840 m depth does not materially change the CO2 plume and water dynamics that the benchmark is meant to capture.","fun_headline_variants_meta":{"raw":{"variants":["AI agents auto-build validated CO2 simulation datasets","Plain-English prompts drive CO2 simulations for GNN training","Agentic framework reproduces benchmark CO2 studies","MCP-based agents simplify CO2 simulation setup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00078,"raw_usage":{"total_tokens":3264,"prompt_tokens":706,"completion_tokens":2558,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":2497}},"tokens_in":450,"tokens_out":2558,"duration_ms":17960,"temperature":1.0,"reasoning_tokens":2497,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:01:38.707161+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run both the agent-generated dead-oil GEOS deck and a full compositional CO2-brine GEOS deck on the same mesh and depth, and compare the saturation fields at the three PUNQ-S3 observation cells over the 500-year horizon; a clear difference in plume shape, crest saturation, or flank draining times would falsify the claim that the approximation preserves the main flow dynamics.","supporting_citations":[],"review_version":1}