{"id":"31a6552c-bd6f-44d5-9499-64b2366419ca","arxiv_id":"2509.10195","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"First DRL-based active flow control on a 3D separated wing reports 21% drag reduction at Re=1000, but the lift-oscillation improvement of 124% is internally inconsistent.","lead":"A deep reinforcement learning agent controls jets on a three-dimensional NACA0012 wing at high angle of attack and low Reynolds number, reporting a 21% drag reduction and a claimed reduction in lift oscillations. The work is an early test of DRL on 3D separated flows, but the oscillation metric is reported as 124%, which is impossible for an RMS reduction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported ΔCl,rms = −124% is physically impossible and is repeated as a central result; the quantitative claims cannot be accepted without correction and raw data.","rationale":"The reader's weakest_assumption—single unseeded PPO run—is a valid methodological concern, but I identify a more decisive and more easily falsifiable flaw: the reported ΔCl,rms = −124% is internally impossible. The reader's rationale also notes this metric is impossible, so there is partial agreement, but the formal weakest_assumption field points to seed variance rather than the impossible number. The qualitative observation of actuation frequency-locking to the shedding frequency (St = 0.64) is plausible and could survive, but the central quantitative claims as stated—21% mean drag reduction and elimination of lift oscillations—are not credible without correction of the RMS metric and access to raw data or multiple seeds. Therefore REJECT remains the appropriate verdict; nothing in this stress test changes the reader's rejection.","tokens_in":7224,"tokens_out":4802,"duration_ms":53446,"concrete_test":"Recompute from the raw Cl(t) and Cd(t) signals behind Fig. 5: select the statistically stationary controlled and baseline windows, compute Cl,rms for each, and form ΔCl,rms = (Cl,rms,ctrl − Cl,rms,base)/Cl,rms,base × 100. Because RMS is nonnegative, any value below −100% is impossible. If the corrected value is −55% or +124%, the published −124% is a sign/denominator error and the text must be corrected; if no raw signals can be produced, the central result is unverifiable.","verdict_should_be":"REJECT","load_bearing_attack":"The central quantitative result in §3 is that the learned policy reduces lift-oscillation RMS by ΔCl,rms = −124% and mean drag by ΔCd = −21%, restated in §4. The RMS reduction is not merely surprising; as a relative change it is mathematically impossible. Since rms_base ≥ 0 and rms_ctrl ≥ 0, (rms_ctrl − rms_base)/rms_base cannot be less than −100%. A reported −124% implies a sign or denominator error, a normalization inconsistency (e.g. comparing with the wrong baseline window), or a typo propagated into the abstract and conclusions. This is an internal inconsistency, not a disagreement with prior work. The paper provides no raw time series or code to disambiguate, so the numbers as published cannot be verified. Even the more plausible −21% drag reduction rests on the same uncontrolled analysis and a single unseeded PPO run, so the headline 'success' is not established independently.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a deep reinforcement learning (DRL) framework for active flow control (AFC) around a three-dimensional NACA0012 wing at Re=1,000 and AoA=20°. The environment is a GPU-accelerated spectral-element CFD solver (SOD2D) coupled to the TF-Agents DRL library through Redis/SmartSim, with a multi-environment, multi-agent setup. The agent actuates two spanwise-distributed jet pairs (three pseudo-environments) with a reward combining drag reduction and lift-oscillation penalty. The authors report that the learned policy reduces mean drag by 21% and lift-coefficient RMS by 124%, shifts the vortex-shedding Strouhal number from 0.56 to 0.64, and produces an actuation signal locked to the shedding frequency. They interpret these results as evidence that DRL can discover effective AFC strategies on three-dimensional separated wings.","tokens_in":7494,"tokens_out":2311,"duration_ms":27904,"significance":"If the reported results are correct, this would be a useful demonstration of DRL-based active flow control on a three-dimensional wing with separation, extending previous two-dimensional studies. The computational framework—GPU-accelerated solver, multi-environment MARL training, and reward decomposition—is a constructive engineering contribution, and the validation of the baseline against published airfoil data is a positive feature. However, the central quantitative claims are not presently credible. The reported ΔCl,rms = −124% is mathematically impossible for a nonnegative RMS quantity, and the drag reduction rests on a single unseeded PPO run with no statistical characterization. The scientific conclusion therefore cannot be accepted as stated, despite the soundness of the general methodology.","major_comments":[{"comment":"The reported lift-oscillation reduction ΔCl,rms = −124% is physically impossible. Since Cl,rms is nonnegative by definition, the relative reduction (Cl,rms,ctrl − Cl,rms,base)/Cl,rms,base cannot be less than −100%. This value is repeated in the Conclusions and must indicate a sign error, a denominator error, or a comparison against an inconsistent baseline window. The manuscript gives no raw time series or computation formula to resolve the discrepancy. This is a load-bearing error: the first central quantitative claim is invalid as stated.","section":"§3, Fig. 5 and §4"},{"comment":"The reported 21% mean-drag reduction is based on a single trained policy obtained from one unseeded PPO run. DRL training is stochastic, and the paper does not report multiple seeds, error bars, or run-to-run variability. Without this information, the 21% reduction cannot be distinguished from the variance of a single training trajectory, which is a known failure mode in RL flow control. The absence of a sensitivity study for the hyperparameters α, β, and the actuator velocity bounds further weakens the claim that the observed performance is robust rather than a lucky initialization.","section":"§3, Figs. 4–5; §2.1"},{"comment":"The paper states that after the transient period 'the root-mean-square of the signal has been noticeably reduced' but the reported −124% implies a reduction beyond complete elimination of oscillations. Even setting aside the arithmetic, the manuscript does not specify how Cl,rms is computed over the controlled and baseline windows (sampling period, number of vortex-shedding cycles, detrending, or window alignment). A proper window-matched RMS calculation with confidence intervals is needed to support any quantitative comparison of oscillation amplitudes.","section":"§3, Fig. 5 and text near 'root-mean-square'"}],"minor_comments":[{"comment":"Typo: 'limitted' should be 'limited' in the Introduction. Also, the sentence 'in the present work we propose to advance AFC-DRL applicability' reads awkwardly; consider revising.","section":"Abstract/Introduction"},{"comment":"Reference [17] is attributed to 'F. Schulman' but the correct first author is John Schulman for 'Proximal Policy Optimization Algorithms'.","section":"References"},{"comment":"Figures 4 and 5 show training curves and temporal signals, but no error bars or confidence intervals are provided. Adding shading or error bands would help assess convergence and variability. Also, the caption of Fig. 3 lists panels (a)–(d), but panels (a)–(c) are cited in the text; please ensure the panel definitions match.","section":"Figures"},{"comment":"The term 'flow-throughs' is used to describe the transient period, but a precise definition (e.g., convective time units c/U∞) would improve clarity. The paper uses 'time units' without defining the nondimensionalization.","section":"§2.2"}],"recommendation":"reject","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is the first DRL active flow control on a 3D separated wing section, extending the multi-environment DRL framework to spanwise pseudo-environments on a NACA0012 at Re=1000, AoA=20. The paper is honest about its own scope: low Reynolds, laminar separated flow, far from aviation. The engineering integration (SOD2D + TF-Agents via Redis/SmartSim) is real and the training curves show the reward improving over 67 episodes. Baseline validation against Gupta et al. and Kouser et al. for Cl, Cd, St looks reasonable. The qualitative finding that the learned policy locks to the shedding frequency (St 0.56→0.64) and produces a more organized wake is plausible and interesting.\n\nThe soft spots are concentrated in the quantitative claims. ΔCl,rms = −124% is impossible for a positive RMS quantity; it appears in §3 and is repeated in §4. That is an internal inconsistency, not a matter of interpretation, and with no raw time series or code, the reader cannot tell whether it's a typo, a wrong baseline window, or a normalization mistake. Even the more plausible −21% drag reduction comes from a single unseeded PPO run, with no error bars or sensitivity analysis of α and β. Given that the reward explicitly maximizes drag reduction and penalizes lift oscillations, the reported improvements are the optimization objective; the real question is whether the learned policy achieves them robustly, and that question is not answered.\n\nI would not publish the numbers as they stand. But the underlying demonstration is worth a serious referee: the application to a 3D wing is a genuine step for the DRL-flow-control community, and the infrastructure is reusable. The paper needs a corrected metric, multiple seeds, and ideally a release of the deterministic-run time series. If those are added, it could be a solid contribution to the subfield.","headline":"First DRL AFC on a 3D wing, but an impossible ΔCl,rms = −124% and a single unseeded run make the central numbers unreliable.","tokens_in":7962,"tokens_out":2084,"would_cite":false,"duration_ms":22633,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep reinforcement learning discovers a jet control for a three-dimensional separated wing that cuts mean drag by 21% and suppresses lift oscillations.","keywords":["deep reinforcement learning","active flow control","NACA0012 wing","flow separation","lift oscillations","drag reduction","proximal policy optimization","multi-agent reinforcement learning"],"falsifier":"Re-run the same DRL training with five different random seeds and measure the distribution of drag reduction. If the 21% reduction is not consistently reproduced (e.g., some seeds give no reduction or a drag increase), the central claim collapses. Alternatively, compute ΔCl,rms directly from the raw lift signals as (Cl,rms,ctrl − Cl,rms,base)/Cl,rms,base: if the controlled RMS is not smaller than the baseline RMS, the '−124%' claim fails on its own terms.","tokens_in":7188,"feed_emoji":"✈️","tokens_out":3834,"duration_ms":41640,"temperature":0.7,"pith_summary":"The paper aims to show that deep reinforcement learning (DRL) can discover effective active flow control for a three-dimensional wing in a separated-flow regime. A PPO-based agent controls a pair of jets on a NACA0012 wing at Reynolds number 1,000 and 20 degrees angle of attack, where vortex shedding and three-dimensional wake structures are fully developed. By the end of training, the learned strategy reduces mean drag by 21% and, as reported, cuts the root-mean-square lift oscillations by 124% relative to the uncontrolled baseline. The agent learns a periodic forcing that locks to the vortex-shedding frequency and delays shear-layer breakdown, leaving a more organized, less three-dimensional wake. This matters because it extends DRL flow control from 2D and canonical cases to a 3D separated wing, a step toward realistic aerodynamic applications.","feed_headline":"AI jet control cuts 3D wing drag 21%","feed_subtitle":"Reinforcement learning learns a periodic blowing/suction strategy that stabilizes the separated wake at Re=1,000.","key_machinery":"The key machinery is the closed-loop DRL control loop: a Proximal Policy Optimization (PPO) agent receives 270 witness-point pressure values (three spanwise slices per pseudo-environment) and outputs a jet velocity, applied to a front/rear jet pair with opposite mass flow to ensure instantaneous mass conservation. The reward function rewards drag reduction and penalizes lift oscillations. Training is accelerated by ten parallel CFD simulations, each split into three pseudo-environments (30 trajectories per step), and by a GPU-accelerated spectral-element solver (SOD2D) communicating with the TF-Agents DRL library through a Redis/SmartSim in-memory layer. The learned policy is then run determ","core_discovery":"The paper's central claim is that a DRL agent, using only local surface pressure information from witness points, learns a jet actuation strategy that nearly eliminates lift fluctuations and reduces drag by 21% on a three-dimensional NACA0012 wing at Re=1,000 and AoA=20°, where vortex shedding and three-dimensional wake structures are fully developed. The learned control maintains a negative (suction) mean on the leading-edge jet and a positive (blowing) mean on the rear jet, with an oscillation frequency matching the vortex-shedding Strouhal number St=0.64. This periodic in-phase forcing extends the leading-edge shear layer, delays vortex shedding, and reduces the spanwise disorder of the w","pith_inferences":["As stated, the −124% reduction in Cl,rms exceeds the mathematically possible range for a reduction relative to baseline; the authors likely use a different normalization, and a direct comparison of raw RMS values would resolve the actual oscillation suppression.","The reported 21% drag reduction comes from a single unseeded training run; repeating training with multiple random seeds would establish whether the learned policy is representative or merely a favorable initialization.","Because the actions lock to the vortex-shedding frequency, a natural test is to compare the learned nonlinear policy against a simple sinusoidal forcing at St=0.64; this would quantify the added value of the policy's extra harmonic components.","The spanwise periodic pattern of actuation may depend on the periodic boundary conditions and spanwise domain length; testing in a wider domain would probe the robustness of the control in a more realistic configuration."],"forward_implications":["If the result is correct, DRL can discover effective active-flow-control strategies for three-dimensional separated wings, not just 2D airfoils or cylinders.","The learned control, which locks to the shedding frequency, suggests that DRL can identify the dominant flow instability and exploit it with a periodic forcing that includes additional harmonic components.","The framework's scalability—ten parallel simulations and thirty simultaneous trajectories—demonstrates a practical path toward higher Reynolds numbers and more complex geometries.","The suppression of lift oscillations and drag reduction could translate into improved aerodynamic efficiency and reduced structural vibration in real applications.","The method's success at Re=1,000 invites testing at higher Reynolds numbers, where separation control is more challenging and industrially relevant."],"fun_headline_variants":["DRL learns jet control to cut 3D wing drag 21%","Deep RL jet control stabilizes 3D wing flow, drag down 21%","Using surface pressure, RL cuts 3D wing drag 21%","Reinforcement learning cuts 3D wing drag 21% with jet control","AI learns periodic jets to reduce 3D wing drag 21%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claim rests on the assumption that a single unseeded PPO training run is representative, so the observed 21% drag reduction and lift-oscillation suppression are not just run-to-run variance.","fun_headline_variants_meta":{"raw":{"variants":["DRL learns jet control to cut 3D wing drag 21%","Deep RL jet control stabilizes 3D wing flow, drag down 21%","Using surface pressure, RL cuts 3D wing drag 21%","Reinforcement learning cuts 3D wing drag 21% with jet control","AI learns periodic jets to reduce 3D wing drag 21%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001194,"raw_usage":{"total_tokens":4731,"prompt_tokens":684,"completion_tokens":4047,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":3945}},"tokens_in":428,"tokens_out":4047,"duration_ms":32199,"temperature":1.0,"reasoning_tokens":3945,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:04:56.300121+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same DRL training with five different random seeds and measure the distribution of drag reduction. If the 21% reduction is not consistently reproduced (e.g., some seeds give no reduction or a drag increase), the central claim collapses. Alternatively, compute ΔCl,rms directly from the raw lift signals as (Cl,rms,ctrl − Cl,rms,base)/Cl,rms,base: if the controlled RMS is not smaller than the baseline RMS, the '−124%' claim fails on its own terms.","supporting_citations":[],"review_version":1}