{"id":"aacf008f-ea87-4763-a79c-82ba2cafe4ee","arxiv_id":"2501.03406","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A neural network reconstructs gust-encounter flow fields and lift from sparse pressure sensors, separating aleatoric and epistemic uncertainties in a low-dimensional latent space.","lead":"This paper uses machine learning to reconstruct the air flow around an airfoil and the lift it produces from a small number of noisy pressure sensors, while also estimating how uncertain those predictions are. The method separates uncertainty from noisy measurements and uncertainty from limited training data, which could support more reliable sensor-based gust-response estimation for aircraft control.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported accuracy may reflect interpolation within gust cases already present in training: Sections 2.2 and 2.5 describe an 80/20 snapshot split without specifying a case-level partition, so test snapshots from the same gust encounter may be nearly identical to training snapshots.","rationale":"The reader identifies the train/test split as the weakest assumption, and I agree: it is the single most load-bearing issue because it determines whether the empirical results support the paper's central generalization claim. The paper's strongest claim, that from 11 noisy pressure measurements the model reconstructs vorticity and lift under extreme gust conditions with well-calibrated uncertainties, is only meaningful if test states come from gust cases absent from training. The text in Sections 2.2 and 2.5 specifies only an 80/20 data split, and because data are enumerated as 78,225 snapshots from 105 cases, the default reading is a snapshot-level split. Under that split, each gust case contributes 745 highly correlated snapshots, so a random partition places near-duplicate states in both training and test sets. This is data leakage by temporal correlation and shared gust parameters, and it inflates all reported log-likelihood values and apparent calibration in Figures 8-13. The concern is not that the methodology is internally inconsistent; the modeling framework is coherent and based on reasonable existing techniques. The problem is that the evidence as presented does not yet establish generalization to new gust encounters. The fix is straightforward: repeat the evaluation with a case-level split, and additionally report empirical coverage rather than only log-likelihood. If the authors confirm that their split was already case-level, the main concern disappears and remaining issues such as calibration metrics and baselines become secondary. I therefore recommend keeping the reader's CONDITIONAL verdict: the paper should be accepted only after the split protocol is clarified and the held-out-case evaluation is provided.","tokens_in":21285,"tokens_out":2500,"duration_ms":27422,"concrete_test":"Retrain and evaluate the full pipeline with a case-level split: randomly partition the 100 gust cases by their sampled parameters (G, y_o/c, 2R/c) into 80 training and 20 held-out cases, keep all 5 base cases in training, and ensure no snapshot from a held-out gust case is used in training either the autoencoder (Section 2.2) or the estimator F_p (Section 2.5). Recompute the results reported in Figures 10 and 13 for the held-out cases, including the pixel-wise vorticity log-likelihood, lift RMSE, and empirical 95% coverage rates for both aleatoric and epistemic intervals. If coverage remains near 95% and log-likelihoods stay comparable to the reported values, the concern is resolved; if performance degrades substantially, the current results overstate generalization to unseen gust encounters.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that sparse pressure measurements generalize to gust-encounter aerodynamics, with uncertainty intervals that contain the true lift and vorticity. This requires evaluation on gust conditions not seen during training. The dataset consists of 105 independent trajectories (5 base cases and 100 gust cases) with 745 temporally correlated snapshots each. Sections 2.2 and 2.5 state only that 80% of the data is used for training and 20% for validation/testing; they do not state that the split is performed at the case level. If the split is random across the 78,225 snapshots, then for virtually every test snapshot there is a training snapshot from the same gust case with the same gust strength, position, radius, and angle of attack, at a nearby time. Because the inputs are continuous pressure time series over a deterministic trajectory, the test input is nearly identical to a training input, and the reported log-likelihoods and 95% intervals in Figures 8-13 measure interpolation within known encounters rather than generalization to new gust parameters. The abstract's claim of accurate prediction 'under extreme conditions' and the proposed online-estimation use case depend on the latter. The correct evaluation is a leave-gust-cases-out split, applied to both the estimator and the lift-augmented autoencoder, since the autoencoder is also trained on the same snapshot pool. Without this clarification, the strength of the empirical evidence for the central claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a two-stage deep-learning pipeline for reconstructing two-dimensional vorticity fields and lift coefficients from 11 noisy surface pressure measurements (a 33-dimensional input) in unsteady gust-encounter flow around a NACA 0012 airfoil at Re = 100. A lift-augmented autoencoder compresses the high-dimensional flow fields into a three-dimensional latent space, and a separate MLP maps stacked pressure readings and sensor coordinates to the mean and covariance of a multivariate Gaussian in that latent space, trained with a heteroscedastic negative log-likelihood loss and Monte Carlo dropout. The paper also analyzes sensor informativeness through eigenmodes of the measurement-space Gramian and presents aleatoric and epistemic uncertainty estimates for latent variables, lift, and vorticity fields, supported by selected log-likelihood values and qualitative comparisons of predicted and reference trajectories.","tokens_in":21534,"tokens_out":5052,"duration_ms":49437,"significance":"If the central claims withstand proper evaluation, the framework is a useful contribution: it combines a nonlinear lift-augmented autoencoder with heteroscedastic latent-space regression and MC dropout, which is a computationally attractive route to sensor-based flow estimation with uncertainty. The sensor-sensitivity analysis via the measurement-space Gramian is physically interpretable and adds value beyond a purely black-box regression. The paper is also transparent about the network architecture and gives explicit equations for the uncertainty model. However, the current evidence is weakened by the apparent snapshot-level train/test split, the lack of quantitative uncertainty calibration, and the self-referential choice of inference-time noise directions. The significance is therefore conditional on fixing these evaluation issues.","major_comments":[{"comment":"The data split is not defined at the level of gust cases. The manuscript states only that \"eighty percent of the data is allocated for training, while the remaining twenty percent is used for validation and testing\" (Section 2.2, repeated in Section 2.5). The dataset consists of 105 independent trajectories (5 base cases and 100 gust cases) with 745 temporally correlated snapshots each, so a random snapshot-level split places test snapshots from the same gust encounter, with the same gust strength, position, radius, and angle of attack, at nearby times in the training set. Because the pressure time series are deterministic functions of the trajectory, the test inputs are then nearly identical to training inputs, and the reported log-likelihoods and 95% intervals in Figures 8 through 13 measure interpolation within known encounters rather than generalization to unseen gust conditions. The autoencoder in Section 2.2 is trained on the same snapshot pool, so its decoder is likewise contaminated. Please re-run the entire pipeline with a leave-gust-cases-out split, state the number of cases in training, validation, and testing, and re-report all quantitative results.","section":"Sections 2.2 and 2.5"},{"comment":"The central claim that the true lift and vorticity fall within the 95% uncertainty intervals is not supported by a quantitative coverage statistic. The only quantitative metrics are average per-pixel log-likelihoods for a small number of selected cases (for example, \"ll\" values in Figures 8, 10, 12, and 13), and the \"within the uncertainty bounds\" statements are made qualitatively. Please report, on the case-level held-out set, the empirical fraction of time instants at which the true CL lies inside the 95% interval, the analogous pixel-level coverage for the vorticity field, and the same statistics for the aleatoric and epistemic distributions separately. A calibration curve or reliability diagram would also help. Without these numbers, the uncertainty-quantification claim in the abstract is not quantitatively established.","section":"Sections 3.2.2 and 3.2.3, Figures 8, 10, 12, 13"},{"comment":"The inference-time noise used to evaluate aleatoric uncertainty is generated as η ~ ζ U_r, where U_r contains the dominant eigenvectors of the measurement-space Gramian C_x computed from the Jacobian of the same network F_p that produces the predictive covariance. This choices makes the reported aleatoric intervals a measure of the network's sensitivity in its own most sensitive directions rather than of a physical sensor-noise process. Moreover, the training-time data augmentation is described as generic Gaussian injection, so the train and test noise models are not matched. The coverage of the uncertainty intervals under independent per-sensor noise or under a physically motivated correlated-noise model is not reported. Please evaluate the same estimators under i.i.d. sensor noise and at least one correlated-noise model, and report the resulting coverage and log-likelihoods; this is important for the stated real-world online-estimation use case.","section":"Section 2.5 and Section 3.2.2"},{"comment":"The latent dimension is fixed at l = 3 with the statement that its appropriateness \"will be discussed below,\" but no reconstruction-error analysis, explained-variance metric, or comparison across latent dimensions appears in Section 3.1. Since the estimator and the uncertainty quantification operate entirely in this latent space, a demonstration that the 3D latent representation is adequate (for example, reconstruction error on held-out gust cases and sensitivity of results to l = 2, 3, 4) is needed to support the \"low-order\" and \"key physics\" claims.","section":"Section 2.2"}],"minor_comments":[{"comment":"The predictive distributions are described as Gaussian, but the distribution of the mean over dropout passes is not necessarily Gaussian; please clarify whether the reported intervals are computed from the Gaussian forms in Equations (12) and (13) or from the empirical samples, and justify the Gaussian approximation.","section":"Equations (12) and (13)"},{"comment":"The text says that stacking sensor x and y positions \"in a global reference frame\" helps the network distinguish angles of attack, but the airfoil geometry and sensor coordinates are fixed for each case; please clarify exactly how the coordinates vary with angle of attack and why this is not just a one-hot encoding of the five cases.","section":"Section 2.3"},{"comment":"The latent-space axes in Figure 3 are unlabeled; please state which latent coordinate is shown on each axis and what the units are, since the paper later refers to ξ1, ξ2, and ξ3.","section":"Figure 3"},{"comment":"The captions report \"ll\" as an average pixel-wise log-likelihood, but it is not clear whether the average is over pixels only or also over time snapshots; please define the exact averaging procedure in the captions or in the text.","section":"Figures 8, 10, 12, 13"},{"comment":"The loss-balance coefficient β = 0.05 and the regularization and dropout hyperparameters are stated without a sensitivity analysis; please report the range of values explored or state explicitly that the results are insensitive to these choices.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The core methodological idea is sound and the paper fits the journal's scope, but the empirical evidence for the central claim depends critically on the train/test split. If the authors have actually performed a case-level split, they should state it clearly and report the case counts; if not, the reported accuracy and uncertainty intervals are likely inflated by temporal leakage. The uncertainty quantification also needs quantitative coverage statistics rather than selected examples. These issues are fixable within the scope of the manuscript, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version: this is a serious paper on a useful problem, and the engineering is mostly careful, but the evidence for the headline claim—that sparse pressure readings reconstruct gust flows they haven't seen—is not established as written. The train/test split is described as a random 80/20 over 78,225 snapshots, not over the 105 gust cases. If test snapshots come from the same gust trajectories that appear in training, the reported log-likelihoods are interpolation within a known encounter, not generalization. The paper never says the split is stratified by case, so a reader can't verify the core claim. That's the main thing I'd want fixed before believing the numbers.\n\nWhat's genuinely new and good: combining the lift-augmented autoencoder with heteroscedastic latent-space regression and MC dropout for gust-encounter flows is a sensible combination, and the time-resolved Gramian-based sensor importance analysis is a nice contribution—it actually gives physical insight into which sensors matter at different phases of a gust interaction. The latent-space visualizations are informative, and the distinction between aleatoric and epistemic uncertainty is handled correctly in the mathematics. The underlying DNS data looks solid. The citation pattern is fine; they build on the right prior work, and the self-citation for the flow solver is legitimate.\n\nWhere I'd push back: the aleatoric noise model is self-referential. They inject noise aligned with the dominant eigenvectors of the Gramian computed from the network's own Jacobian, then evaluate the network's uncertainty on that same noise. That tells you how sensitive the network is to its own most sensitive directions, but it's not a realistic sensor-noise test. The paper also doesn't report any calibration metric like coverage across the full test set; we get selected examples and per-image log-likelihoods, which can look fine while missing systematic under-coverage. And there are no baselines—no deterministic MLP, no linear regression, no Kalman filter—so we can't tell how much the UQ machinery actually buys you.\n\nNone of these are fatal to the idea. They're fixable: split by case, retrain, report coverage, and add a baseline. If those come back positive, this would be a useful paper for people working on sensor-based estimation and gust load alleviation. As it stands, I'd send it to review, because the framework is well-motivated and the sensor analysis is worth refereeing, but I'd expect major revision.","headline":"Useful framework, but the central generalization claim is undermined by an unspecified train/test split and a self-referential noise model.","tokens_in":22120,"tokens_out":1779,"would_cite":false,"duration_ms":16970,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"From 11 noisy pressure sensors, a neural network reconstructs a gust-disturbed airfoil's vorticity field and lift, with separate aleatoric and epistemic uncertainty bounds that contain the true values in the tested cases.","keywords":["gust encounter aerodynamics","sparse pressure measurements","flow field reconstruction","lift coefficient estimation","aleatoric uncertainty","epistemic uncertainty","Monte Carlo dropout","lift-augmented autoencoder"],"falsifier":"Hold out complete gust cases—entire triples of gust strength $G$, position $y_o/c$, and radius $2R/c$—from training, then evaluate whether the 95% predictive intervals for lift and per-pixel vorticity contain the simulation truth on those unseen cases; if coverage falls well below 95% or any test snapshot shares a gust case with the training data, the central claim is not established.","tokens_in":21017,"feed_emoji":"🌪️","tokens_out":11041,"duration_ms":91171,"temperature":0.7,"pith_summary":"This paper claims that a neural network can reconstruct the two-dimensional vorticity field and the lift coefficient of an airfoil in a gust encounter from just 11 noisy surface pressure sensors, and can say how confident it is. The key idea is to compress each high-dimensional flow field into three latent variables with a lift-augmented autoencoder, then train a small MLP to map pressure readings plus sensor coordinates to a mean and a covariance for those latent variables. Sampling from the resulting latent distribution and decoding gives vorticity and lift with separate aleatoric (measurement-noise) and epistemic (model) uncertainties; in the reported cases the true values sit inside the 95% intervals and most per-pixel log-likelihoods are positive. If the claim holds, online sensor-based flow estimation for gust encounters becomes cheap enough for real-time use, because the expensive decoder is trained once and the online estimator is a small network.","feed_headline":"11 sensors reconstruct gust flows with truth in error bars","feed_subtitle":"A three-variable latent model maps noisy pressure data to full flow fields with separate model and noise uncertainties.","key_machinery":"The load-bearing object is the three-dimensional latent representation $\\boldsymbol{\\xi}$ of the flow, learned by a lift-augmented autoencoder whose decoder also predicts lift. Around that latent space the paper builds a measurement-space Gramian $C_x = \\mathbb{E}[\\nabla f(x)^T \\nabla f(x)]$ to rank sensor directions, and a predictive distribution $\\pi(y|x) = \\mathcal{N}(\\mu, \\Sigma)$ over latent states with $\\Sigma = LL^T$ parameterized by a lower-triangular Cholesky factor. Monte Carlo dropout turns the estimator into a stochastic map, so $T$ forward passes yield the aleatoric distribution (mean of the predicted covariances) and the epistemic distribution (covariance of the predicted means). This is the machinery that lets 11 sensors produce field-level reconstructions with calibrated confidence intervals.","core_discovery":"The central discovery is that the physics of gust-airfoil interaction survives compression to a three-dimensional latent space, and that the sensor-to-flow relationship can be learned as a distribution over that space rather than over the full field. Using 11 pressure sensors stacked with their coordinates as a 33-dimensional input, the estimator learns the mean and a Cholesky-decomposed covariance of a multivariate normal in latent space under a heteroscedastic negative log-likelihood loss. Active dropout during inference supplies the epistemic component as the covariance of the predicted means, while the averaged predicted covariances supply the aleatoric component, and decoding samples from either distribution gives vorticity and lift with uncertainty bands that contain the reference values in the reported tests. The measurement-space Gramian of the estimator Jacobian has two dominant eigenmodes carrying over 99% of the energy, and the most informative sensors migrate from suction-side mid-chord to trailing-edge sensors as vortices shed and as positive or negative gusts pass.","pith_inferences":["A decisive check the paper leaves implicit is case-level holdout: train on some gust parameter triples and test on others, because the described 80/20 split is not stated to be stratified by gust case.","The nearly tangential aleatoric uncertainty ellipses suggest that adding a temporal filter or state-space prior over successive latent snapshots could shrink the noise-driven bounds beyond the paper's per-snapshot inference.","The reweighting of dominant sensors as the gust crosses the airfoil hints at adaptive sensing: a sensor suite whose weights track the leading eigenmode could keep information content high through the whole encounter.","The paper's own hypothesis about large uncertainty in high-gradient regions implies a testable extension: at higher Reynolds numbers, where shear layers fragment, the 95% coverage should be re-measured to see whether the latent-space assumption still holds."],"forward_implications":["Flow reconstruction and load estimation from sparse noisy pressure become feasible online, because the estimator is a small MLP in a three-dimensional latent space rather than a full flow solver.","Separating aleatoric from epistemic uncertainty tells an operator whether to improve sensor accuracy or collect more training data, since the two uncertainties have different dominant directions in latent space.","The measurement-space Gramian gives a quantitative sensor-placement criterion: the dominant eigenmodes identify which sensors carry the most information at each phase of vortex shedding and gust passage.","If the split generalizes to unseen gust parameters, the 95% intervals demonstrated here would support safety-relevant decisions such as gust load alleviation with calibrated confidence.","The same pretrained-decoder-plus-latent-uncertainty pipeline can be reused for other distributed measurements, since only the small sensor-to-latent map must be retrained."],"supporting_citations":[{"why":"Supplies the lift-augmented autoencoder that compresses the high-dimensional vorticity field to a three-dimensional latent manifold.","marker":"Fukami and Taira (2023)"},{"why":"Proves the dropout-as-Bayesian-approximation result that justifies treating Monte Carlo dropout as variational inference for the network weights.","marker":"Gal and Ghahramani (2016a)"},{"why":"Introduces Monte Carlo dropout for predictive model uncertainty, the procedure used here for stochastic forward passes at inference.","marker":"Gal and Ghahramani (2016b)"},{"why":"Establishes the combined aleatoric-plus-epistemic heteroscedastic regression loss that the paper adapts to the latent space.","marker":"Kendall and Gal (2017)"},{"why":"Provides the immersed-layers Cartesian-grid Navier-Stokes solver that generated the 78,225 gust-encounter snapshots used for training and testing.","marker":"Eldredge (2022)"},{"why":"Defines the Taylor vortex model whose strength, position, and radius parameterize the gust disturbances in the simulations.","marker":"Taylor (1918)"},{"why":"Supplies the Jacobian-based Gramian construction used to define the measurement-space sensitivity analysis.","marker":"Quinton and Rey (2024)"},{"why":"Earlier sparse-pressure machine-learning load estimation for gust encounters, providing the nearest baseline this work extends to latent-space uncertainty quantification.","marker":"Chen et al. (2024)"}],"fun_headline_variants":["11 sensors reconstruct gust flows with truth in error bars","Sparse pressure taps rebuild gust flow via 3D latent space","Dual uncertainty UQ for gust-airfoil reconstruction from 11 sensors","Low-order latent space turns sparse pressure into flow and lift","Calibrated uncertainty in gust flow reconstruction from 11 taps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the reported test accuracy reflects generalization to gust encounters the network has not seen, which requires that the 80/20 split be made at the level of entire gust cases rather than random snapshots.","fun_headline_variants_meta":{"raw":{"variants":["11 sensors reconstruct gust flows with truth in error bars","Sparse pressure taps rebuild gust flow via 3D latent space","Dual uncertainty UQ for gust-airfoil reconstruction from 11 sensors","Low-order latent space turns sparse pressure into flow and lift","Calibrated uncertainty in gust flow reconstruction from 11 taps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000266,"raw_usage":{"total_tokens":1593,"prompt_tokens":912,"completion_tokens":681,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":594}},"tokens_in":528,"tokens_out":681,"duration_ms":6445,"temperature":1.0,"reasoning_tokens":594,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:53:58.068597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out complete gust cases—entire triples of gust strength $G$, position $y_o/c$, and radius $2R/c$—from training, then evaluate whether the 95% predictive intervals for lift and per-pixel vorticity contain the simulation truth on those unseen cases; if coverage falls well below 95% or any test snapshot shares a gust case with the training data, the central claim is not established.","supporting_citations":[{"cited_title":"and Taira, K","cited_arxiv_id":null,"evidence_quote":"Supplies the lift-augmented autoencoder that compresses the high-dimensional vorticity field to a three-dimensional latent manifold."},{"cited_title":"and Gal, Y","cited_arxiv_id":null,"evidence_quote":"Establishes the combined aleatoric-plus-epistemic heteroscedastic regression loss that the paper adapts to the latent space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the immersed-layers Cartesian-grid Navier-Stokes solver that generated the 78,225 gust-encounter snapshots used for training and testing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Taylor vortex model whose strength, position, and radius parameterize the gust disturbances in the simulations."},{"cited_title":"E., Fukami, K., and Taira, K","cited_arxiv_id":null,"evidence_quote":"Earlier sparse-pressure machine-learning load estimation for gust encounters, providing the nearest baseline this work extends to latent-space uncertainty quantification."}],"review_version":1}