{"id":"0ea1ee69-e5ee-4dbb-ad16-78cafba06c12","arxiv_id":"2506.14597","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An MLP surrogate of CFD embedded in a particle filter estimates methane source location and emission rate in transient atmospheric flows, with accuracy close to full CFD but at far lower per-evaluation cost.","lead":"This paper combines a neural network surrogate of computational fluid dynamics with a particle filter to infer gas source location and emission rate from sensor measurements. It reports accuracy comparable to full CFD simulations at much lower per-evaluation cost, validated on a controlled methane release dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-data validation rests on an unvalidated 2D uniform-wind CFD; selected ideal-window tests do not establish comparable accuracy.","rationale":"The reader's weakest assumption correctly identifies that the CFD simulator is not independently validated against measured concentration fields. My stress-test agrees and sharpens it: the validation in Table 1 uses the same wind time series used for training, and covers only two selected 'ideal wind' periods, so it does not establish that the CFD (and hence the surrogate) generalizes to the conditions where the method would be deployed. This is load-bearing because the entire accuracy argument flows through the CFD surrogate: if the 2D, uniform-wind CFD is biased, then the surrogate inherits that bias, and the reported improvement over the Gaussian plume baseline in Table 2 may be an artifact of a misspecified simulator rather than a genuine modeling gain. The paper also does not directly compare a full CFD-based SIR inversion to the surrogate-based inversion, so the 'comparable accuracy to full CFD' claim for source localization is not directly supported. These concerns do not refute the methodological contribution, but they make the central empirical claim conditional on additional validation. The reader's verdict is already CONDITIONAL, and my analysis does not move it; it reinforces the need for a held-out time-period validation and a direct full-CFD comparison.","tokens_in":9825,"tokens_out":7831,"duration_ms":86404,"concrete_test":"Select one full Chilbolton release period or time window not used in Section 4.1 (or split the available data by time, not by source location), train the surrogate on CFD runs from the other period(s), and evaluate the surrogate/CFD predictions against the actual 20-s averaged measurements at all seven sensors over the held-out period. For a stricter test, also run the Gaussian plume baseline on the same held-out period. If the surrogate's MAPE exceeds roughly 20-30% or the Gaussian plume matches/beats it, the claim of comparable accuracy to real dispersion fails. Additionally, report whether the observation model uses point concentrations or path-integrated concentrations; if point values are compared to path-averaged Chilbolton data, recompute the MAPE with line-integrated model outputs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim ('comparable accuracy to full CFD and Gaussian plume models') rests on the assumption that the CFD solver used to generate surrogate training data is a faithful model of Chilbolton dispersion. Section 3.1 states the CFD is run in 2D with a spatially uniform, temporally varying wind, and Section 4.1 validates it only on two chosen 'ideal wind' test cases (10 and 15 minutes), using the same wind time series that generated the training data. Because the surrogate is trained to match this CFD, any systematic CFD error is inherited by the inversion. Table 1 compares model predictions to real sensor data, but only at true source locations and only on selected windows; it does not validate the CFD on held-out time periods or on non-ideal conditions. Moreover, the paper never runs a full CFD-based SIR inversion, so 'comparable accuracy to full CFD' for source localization (the Table 2 result) is inferred only from surrogate-vs-CFD prediction error, not from a direct inversion comparison. If the CFD is biased, the 5.82 m localization result may reflect simulator error rather than a real advantage over the Gaussian plume baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian source-inversion framework for gas emissions in which a multilayer perceptron (MLP) is trained to emulate a CFD-based concentration model, and the resulting surrogate is embedded in a sequential importance resampling (SIR) particle filter. The method is applied to the Chilbolton controlled methane release dataset, where the authors report surrogate predictions with MAPE close to the full numerical solver and lower than a Gaussian plume model, and a source-localization error of 5.82 m versus 11.09 m for the plume baseline. The paper also includes a synthetic obstructed-flow study with time-varying emission rates. The central claims are that the surrogate provides accuracy comparable to full CFD at a fraction of the computational cost and enables real-time, near-real-time inversion.","tokens_in":10073,"tokens_out":5509,"duration_ms":58905,"significance":"The core idea—replacing expensive CFD evaluations inside a Monte Carlo inversion with a trained neural emulator—is timely and sensible, and the paper demonstrates a working instance on a real controlled-release dataset with clearly reported numbers. The state-space formulation is standard, and the comparison against an independent Gaussian plume baseline is a useful design choice. However, the headline claims outrun the evidence: the reported end-to-end runtime of 83.3 minutes for 10 minutes of data contradicts the 'real-time' framing, the surrogate is trained on CFD simulations driven by the same wind conditions used for evaluation, only two 'ideal wind' windows are tested, and no full CFD-based inversion is run. The paper's value will be substantially higher if these issues are addressed by reframing the claims and adding temporal held-out validation.","major_comments":[{"comment":"The reported end-to-end runtime of 83.3 minutes to process 10 minutes of Chilbolton data, including roughly 8 minutes of CFD data generation and training per MLP window, is incompatible with the abstract's 'real-time' and 'near-real-time' claims and with Section 4.2's statement that the framework enables 'real-time spatio-temporal inference'. The comparison to the Gaussian plume baseline is only a factor-of-two saving (83.3 vs 173.6 minutes), not the 'orders-of-magnitude faster runtimes' asserted in the abstract and conclusion.","section":"Section 4.2, Table 2"},{"comment":"The MLP is trained on CFD simulations that use the same wind boundary time series as the two 'ideal wind' evaluation windows, and the CFD itself is a 2D approximation with a spatially uniform, temporally varying wind. As a result, the close agreement between MLP and CFD in Table 1 is expected by construction, and the paper does not validate the CFD against measured concentration fields on held-out time periods or non-ideal conditions. The claim of 'comparable accuracy to full CFD' therefore rests on an unvalidated simulator and on interpolation over source locations only, not over flow conditions.","section":"Sections 3.1 and 4.1"},{"comment":"No full CFD-based SIR inversion is ever run, so the abstract's 'comparable accuracy to full CFD solvers' for source localization is inferred only from surrogate-versus-CFD prediction errors, not from a direct inversion comparison. If the CFD is biased, the 5.82 m localization result may reflect simulator error rather than a genuine advantage over the Gaussian plume baseline; a full CFD inversion on at least one window, or a sensitivity analysis with perturbed CFD inputs, is needed to support the claim.","section":"Section 4.2"},{"comment":"The synthetic obstructed-flow results are presented only as posterior density plots and qualitative statements, with no quantitative metrics, no error bars, and no comparison to a baseline such as the Gaussian plume model or a full CFD inversion. Since the abstract claims 'robustness in complex environments', the main text should report numerical localization errors, emission-rate tracking errors, and runtimes for each of the three scenarios.","section":"Section 5"}],"minor_comments":[{"comment":"The notation s_k:t and beta_k:t is introduced only verbally as 'history'; please define the window length k and the exact dependence of the observation model on it, since the particle filter state described in the text contains only the current emission rate and location.","section":"Section 2, Eq. (1)"},{"comment":"The phrase 'ground truth' is used for the CFD output; this overstates the status of a numerical simulator. Suggest using 'simulator reference' or 'high-fidelity simulation output' throughout.","section":"Section 3.1"},{"comment":"MAPE is not defined in the table or text; please state the formula and explain how zero or near-zero concentration readings are handled, since these can dominate the percentage error.","section":"Table 1"},{"comment":"The 'mean distance from all particles at the SIR last iteration' is not a standard posterior summary; please also report the posterior mean location, a credible region, and the standard deviation across independent filter runs.","section":"Section 4.2"},{"comment":"The training-set sizes are reported as 484 simulations in Section 4.1 and 499 CFD-based training simulations in Section 5; please explain the difference or correct the inconsistency.","section":"Sections 4.1 and 5"},{"comment":"The figures would be easier to interpret with color bars, axis labels with units, and quantitative annotations; the claim in Figure 1 that one posterior is 'closer to the true source location' should be supported by the numerical metric in Table 2.","section":"Figures 1 and 3"},{"comment":"Several load-bearing details, including CFD solver settings, supplementary results on long CFD runs, and the vertical Gaussian plume correction, are deferred to a supplement that was not available for review; please ensure the supplement is included or move essential details into the main text.","section":"Supplementary Materials"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a relevant problem and the surrogate-plus-particle-filter pipeline is reasonable, but the real-time claims are contradicted by the paper's own runtimes and the validation is weakened by using the same wind conditions for training and evaluation. I would ask for a revised version that reframes the performance claims, adds at least one held-out temporal validation or a full CFD inversion comparison, and reports quantitative results for the synthetic scenarios. The supplement is essential and should be made available to reviewers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version. The paper builds a real pipeline—MLP surrogate of time-dependent CFD, embedded in an SIR particle filter, retrained on sliding windows—and tests it on controlled methane release data. That combination is genuinely new relative to the literature it cites, and the headline result is a 5.82 m vs 11.09 m mean localization error against a Gaussian plume SIR baseline. That is worth paying attention to.\n\nWhat it does well: the surrogate training is straightforward (MLP mimics CFD, trained on sampled source locations), the sensor-prediction comparison in Table 1 is honest—MLP is close to the solver (13.38% vs 13.28% MAPE on Source 1) while being far cheaper—and the simulated obstructed-flow experiments add evidence that the method tracks time-varying emissions, with a candid note about lag after abrupt changes.\n\nThe soft spots are real but not disqualifying. First, the 'real-time' framing is contradicted by the paper's own Table 2: 83.3 minutes to process 10 minutes of data. The 'orders of magnitude faster' refers to per-evaluation surrogate cost vs CFD, not the end-to-end inversion. That claim needs rewording or a streaming-mode experiment. Second, the validation window is narrow: two 'ideal wind' periods, no held-out time intervals or non-ideal conditions. The CFD itself is 2D with spatially uniform, temporally varying wind, and is never checked against independent measured concentration fields; the surrogate inherits any CFD bias. There is also no full-CFD SIR inversion run, so the abstract's 'comparable accuracy to full CFD' for source localization is inferred from Table 1's prediction errors, not directly measured. Third, the 5.82 m result comes from a single particle filter run with no error bars or repeated seeds. Fourth, no code or data is released, which limits reproducibility given the large training budget (90 CPUs, 250 GB memory).\n\nNone of this is a load-bearing flaw. The method is sensible and the empirical work is real. But the claims in the abstract outrun the evidence as written.\n\nWho is this for: anyone working on Bayesian inversion for atmospheric emissions, especially practical source localization on industrial sites. It deserves a serious referee: the novelty is clear, the data are real, and the issues are fixable with a revised framing, more validation periods, repeated runs for error bars, and a release of code/data. I would engage with it if I were reviewing for a journal.\n\nRecommendation: send it to peer review.","headline":"A sensible surrogate-in-the-loop Bayesian inversion with a real data win over a plume baseline, but the real-time claim outruns the reported runtimes and the validation is narrower than the abstract implies.","tokens_in":10575,"tokens_out":2928,"would_cite":false,"duration_ms":29734,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural-network surrogate for computational fluid dynamics, embedded in a sequential Monte Carlo filter, locates methane sources from sparse sensor measurements in unsteady wind fields at a fraction of the solver's cost.","keywords":["gas emission inversion","source localization","Bayesian state-space model","particle filter","CFD surrogate","multilayer perceptron","methane monitoring","spatio-temporal inference"],"falsifier":"Run the trained surrogate and the CFD solver against the full Chilbolton sensor record for a held-out release period: if the surrogate matches the CFD but the CFD misses the measured concentration time series by large margins when local wind deviates from the uniform assumption, the surrogate-based posterior will be biased even though it emulates its teacher perfectly. A direct check is to compute the particle-filter localization error on a known ground-truth release under non-uniform wind and see whether the 5.82 m error degrades beyond the Gaussian plume baseline.","tokens_in":9631,"feed_emoji":"💨","tokens_out":7439,"duration_ms":71057,"temperature":0.7,"pith_summary":"The paper proposes a way to find and quantify gas leaks in near real time from sparse, high-frequency concentration measurements. It replaces the expensive computational-fluid-dynamics solver inside a Bayesian particle filter with a fast neural-network surrogate trained on CFD output, so thousands of candidate source locations can be scored in milliseconds. On the Chilbolton controlled methane-release dataset, the surrogate-based filter locates the source to within about 5.8 metres, compared to about 11.1 metres for a Gaussian plume filter, while cutting total inversion time roughly in half. The method also tracks time-varying emission rates and works in simulated obstructed flows, which suggests it could support continuous industrial emissions monitoring.","feed_headline":"Methane leaks found to 5.8 m in half the compute time","feed_subtitle":"A neural CFD stand-in in a Bayesian filter locates leaks faster and more accurately than plume models.","key_machinery":"The load-bearing object is the trained MLP surrogate $C_{\\mathrm{MLP}}(\\tilde{x}, \\tilde{y})$, which maps candidate source coordinates to the 20-second-averaged concentration expected at each sensor and is trained by regression on CFD runs of the time-dependent Navier-Stokes and advection-diffusion equations. It sits inside the SIR particle filter's likelihood evaluation, where the measurement model is evaluated once per particle per time step. Sliding time windows keep the surrogate tied to the current wind regime, and the random-walk state equation maintains particle diversity. The surrogate is what turns a computationally prohibitive forward model into a millisecond evaluation.","core_discovery":"The central claim is that a multilayer perceptron trained on outputs of a two-dimensional CFD solver can stand in for the physics within a sequential importance resampling particle filter, making Bayesian inference of methane source location and emission rate feasible in unsteady wind fields. The observation model $\\hat{d}_t = C(\\dot{x},\\dot{y},\\dot{z}|\\tilde{x},\\tilde{y},\\tilde{z})\\, s_{\\kappa:t} + \\beta_{\\kappa:t} + \\epsilon_t$ is retained, but $C$ is evaluated by the MLP instead of the numerical solver; separate MLPs are trained on sliding time windows so each sees roughly homogeneous wind. Validation on the Chilbolton releases shows the surrogate's concentration predictions closely match the CFD solver (MAPE 13.38% versus 13.28% on Source 1) and its inverted source location is more accurate than the Gaussian plume baseline (5.82 m versus 11.09 m), at less than half the total compute. The paper also demonstrates on synthetic obstructed scenarios that the filter recovers hidden sources and follows fluctuating emission rates.","pith_inferences":["The same pattern of training a static surrogate on an expensive simulator and embedding it in a sequential Monte Carlo filter should transfer to other sparse-sensor inverse problems, such as groundwater contaminant source identification or indoor pollutant tracing.","Because the surrogate is retrained for each wind window, the framework implicitly treats the wind field as quasi-static; conditioning a single surrogate on wind speed and direction as additional inputs could remove the retraining overhead and allow true online drift.","The reported accuracy gain over Gaussian plume models suggests that transient physics captured by CFD matters even at a flat, open site; if that holds, plume-model-based inversions elsewhere may be underestimating localization uncertainty.","The delayed posterior adjustment after Source 2's sharp emission-rate drop means the raw filter is better suited to gradual leakage than sudden events; a regime-switching layer would be needed for detecting abrupt leak onset or shutdown."],"forward_implications":["Continuous monitoring becomes practical: because each likelihood evaluation costs milliseconds, the filter can ingest high-frequency sensor streams and update source estimates minute by minute.","The CFD cost is paid once offline to create training data, so replacing the plume model as the inner loop more than halves total inversion time in the reported experiments.","The particle posterior provides a full distribution over source location and emission rate, not just a point estimate, so uncertainty can be propagated into downstream decisions.","Synthetic obstructed-flow experiments show the filter localizes sources occluded by obstacles and tracks increasing, decreasing, and fluctuating emission rates, extending the approach beyond flat open terrain."],"supporting_citations":[{"why":"Supplies the SIR particle filter algorithm that performs the sequential Bayesian update.","marker":"[9]"},{"why":"Provides the differentiable CFD solver used to generate surrogate training data.","marker":"[13]"},{"why":"Supplies the Chilbolton controlled-release methane measurements used for validation.","marker":"[11]"},{"why":"Defines the Gaussian plume inversion baseline the surrogate-based filter is compared with.","marker":"[22]"},{"why":"Supplies the Gaussian plume dispersion model used as the analytical baseline.","marker":"[30]"}],"fun_headline_variants":["Neural CFD stand-in pinpoints methane leaks in half the time","MLP surrogate speeds up methane source finding 2x","Deep learning replaces CFD for real-time gas leak location","Bayesian filter with neural surrogate finds methane leaks faster","Methane leaks located to 5.8 m in half the compute time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole chain inherits its physics from a two-dimensional CFD model driven by a spatially uniform, time-varying wind, and the paper does not compare that model's concentration predictions against independently measured plume fields, so the surrogate can only be as faithful as the simulator it learns from.","fun_headline_variants_meta":{"raw":{"variants":["Neural CFD stand-in pinpoints methane leaks in half the time","MLP surrogate speeds up methane source finding 2x","Deep learning replaces CFD for real-time gas leak location","Bayesian filter with neural surrogate finds methane leaks faster","Methane leaks located to 5.8 m in half the compute time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001339,"raw_usage":{"total_tokens":5429,"prompt_tokens":919,"completion_tokens":4510,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":4425}},"tokens_in":535,"tokens_out":4510,"duration_ms":29130,"temperature":1.0,"reasoning_tokens":4425,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:50:22.449279+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained surrogate and the CFD solver against the full Chilbolton sensor record for a held-out release period: if the surrogate matches the CFD but the CFD misses the measured concentration time series by large margins when local wind deviates from the uniform assumption, the surrogate-based posterior will be biased even though it emulates its teacher perfectly. A direct check is to compute the particle-filter localization error on a known ground-truth release under non-uniform wind and see whether the 5.82 m error degrades beyond the Gaussian plume baseline.","supporting_citations":[{"cited_title":"Novel approach to nonlinear/non- Gaussian Bayesian state estimation","cited_arxiv_id":null,"evidence_quote":"Supplies the SIR particle filter algorithm that performs the sequential Bayesian update."},{"cited_title":"ΦFlow (PhiFlow): differentiable simulations for PyTorch, TensorFlow and Jax","cited_arxiv_id":null,"evidence_quote":"Provides the differentiable CFD solver used to generate surrogate training data."},{"cited_title":"Methane emissions: remote mapping and source quantification using an open-path laser dispersion spectrometer.Geophysical Research Letters, 47(10), 2020","cited_arxiv_id":null,"evidence_quote":"Supplies the Chilbolton controlled-release methane measurements used for validation."},{"cited_title":"Probabilistic Inversion Modeling of Gas Emissions: A Gradient-Based MCMC Estimation of Gaussian Plume Parameters","cited_arxiv_id":"2408.01298","evidence_quote":"Defines the Gaussian plume inversion baseline the surrogate-based filter is compared with."},{"cited_title":"The mathematics of atmospheric dispersion modeling.Siam Review, 53(2):349– 372, 2011","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian plume dispersion model used as the analytical baseline."}],"review_version":1}