{"id":"de193154-8305-4535-a60f-cf2bbd0931eb","arxiv_id":"2505.00222","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An AI-driven pipeline co-optimizes underwater glider hull shape and control to maximize lift-to-drag ratio, producing fabricated gliders that outperform a torpedo baseline in pool tests.","lead":"This paper introduces an automated design pipeline that co-optimizes underwater glider hull shapes and control using a neural-network fluid surrogate. The resulting 3D-printed gliders were tested in water and achieved higher lift-to-drag efficiency than a traditional torpedo-shaped glider.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pool-measured η=2.5 is compared to a CFD-only literature value for a different torpedo glider; no matched baseline test is reported, so the headline outperformance claim is not yet supported.","rationale":"The reader's weakest assumption focuses on surrogate generalization, but the more load-bearing issue is the uncontrolled baseline comparison in Section III-D. The wind tunnel provides some support for the surrogate at 9° AoA, and the convex hull limitation is a standard design-space restriction; however, the headline claim asserts a real-world efficiency advantage, and that assertion rests on comparing a measured value to a CFD-only literature value from a different vehicle. The paper even had an appropriate control available (the 'traditional' shell in Fig. 5a) but did not report its measured η. This omission makes the central claim unproven as stated, while remaining addressable through a matched experiment. The reader's CONDITIONAL verdict is therefore appropriate; my concern does not change that classification, so the verdict stays UNCHANGED.","tokens_in":9010,"tokens_out":7692,"duration_ms":77353,"concrete_test":"Perform a matched baseline experiment: mount the standard torpedo-style shell from Ref. [31] (or a geometry matching Ref. [19]) on the same internal hardware, run the identical pool gliding protocol with repeated dives, and compute η from measured horizontal and vertical speeds. If the torpedo shell achieves η ≥ 2.5 under identical conditions, the central claim fails; if it achieves η close to the reported 0.3, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Section III-D is not controlled: the two-wing glider's η=2.5 is measured in a pool, with real surface details, holes, and manufacturing alterations, while the baseline η=0.3 from Ref. [19] is a CFD simulation of a different torpedo-shaped glider. The paper's own Section IV acknowledges a threefold simulation-to-reality gap (η=7 simulated vs 2.5 measured), showing that the simulated objective and the measured quantity diverge substantially. The modular hardware (Fig. 5) includes a 'traditional' shell [31], yet no pool-test data for that shell are reported. Consequently, the 8× improvement is not established under matched conditions; discrepancies in size, speed, Reynolds number, or test protocol could account for part or all of the gap. Without a same-protocol baseline, the headline claim of 'significantly outperforms' lacks the necessary control.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an end-to-end automated framework for designing underwater glider hull shapes. The method combines a deformation-cage shape parameterization, a neural-network surrogate for lift and drag coefficients trained on OpenFOAM simulation data, and CMA-ES optimization to maximize the lift-to-drag ratio η across different angles of attack. The authors validate the surrogate in a wind tunnel for one optimized shape at four angles of attack, report a dynamics simulation, fabricate two optimized hull shells (a two-wing design for 9° and a four-wing design for 30°) around a modular internal hardware assembly, and conduct swimming-pool gliding tests. The pool tests yield inferred lift-to-drag ratios of η=2.5 and η=2.4 for the two designs, which the authors compare to a literature CFD value of η=0.3 for a torpedo-shaped glider, claiming significant outperformance. The paper also includes an appendix derivation showing that glider energy efficiency per distance is inversely proportional to cd/cl, hence proportional to η.","tokens_in":9215,"tokens_out":8048,"duration_ms":76894,"significance":"If the framework's claims are supported, the work would be a valuable contribution to computational robot design: it demonstrates a differentiable, surrogate-based pipeline from a reduced-order shape representation to physical fabrication, and the wind-tunnel validation of the surrogate is a concrete strength. The modular hardware system with exchangeable shells is also a practical contribution. However, the headline claim of superior efficiency over a prior design is not yet rigorously established, the reported surrogate accuracy is presented in a misleading way, and the optimality claims are restricted to a curated design space that is not clearly described in the abstract. The strongest contribution is the demonstration of a fully automated design loop, not the specific efficiency comparison.","major_comments":[{"comment":"The comparison of pool-measured η=2.5 (two-wing) and η=2.4 (four-wing) to the literature CFD value η=0.3 from Ref. [19] is not a controlled comparison. The baseline is a different torpedo-shaped glider evaluated with a different method (CFD) and no matched experimental protocol involving the same hardware, pool, and test conditions. The paper's own Section IV reports a simulation-to-reality gap of approximately threefold for the two-wing glider (simulated η=7 vs measured η=2.5), so the measured value is not the optimized objective value. Without a same-protocol baseline, for example the 'traditional' shell [31] shown in Fig. 5a tested in the same pool with the same internal hardware, the claim that the designs 'significantly outperform' the previous standard design is not supported.","section":"Section III-D and Section IV"},{"comment":"The reported 'average of 4.5% error' in the wind-tunnel validation is computed as the mean signed error (+0.34) divided by the maximum simulated η (7.53), rather than as a conventional relative or mean-absolute-percentage error. This normalization makes the error appear much smaller than it is in the low-angle-of-attack regime. At 0° AoA, the error is +0.87 against a measured wind-tunnel value of 1.10, corresponding to a relative error of roughly 79%. Thus the validation statement overstates the surrogate's accuracy, and this misreporting is load-bearing because the wind-tunnel experiment is the primary evidence that the surrogate can replace CFD in the optimization loop.","section":"Section III-A and Table I"},{"comment":"The surrogate is validated in the wind tunnel for only one fabricated shape, the 9° optimal hull, while the framework claims to discover optimal shapes across a range of angles of attack, including the 30° four-wing design. No independent hydrodynamic validation is provided for the four-wing design or for any other optimized shape, and the simulated performance of the four-wing design is not reported. The generalization of the surrogate to the broader optimized design space is therefore not established, which weakens the claim that the framework 'discovers a wide range of optimal, non-trivial glider designs'.","section":"Sections II-D, III-A, and III-D"},{"comment":"The optimization is explicitly constrained to the convex hull of the 20 manually curated base shapes ('we constrain this search to fall inside the convex hull defined by the base shapes'), so the 'optimal' designs are optimal only within that curated subspace. The abstract and introduction do not state this restriction, and the introduction even says prior designs 'have not yet approached what could be considered the globally optimal configuration', implying a global search. The assumption that this convex hull contains the relevant high-efficiency designs is neither justified nor tested. The optimality claims should be framed as relative to the chosen design space.","section":"Section II-E and Abstract/Introduction"}],"minor_comments":[{"comment":"The text states 'the work required to pump water is directly proportional to η', but the appendix derivation shows the work per distance is proportional to cd/cl, i.e., inversely proportional to η. The conclusion to maximize η is correct, but the wording is backwards.","section":"Section II-A"},{"comment":"The pool-test results report horizontal and vertical speeds and inferred η values without uncertainty quantification, number of runs, or details of the test protocol (e.g., pool depth, glide distance, starting conditions). Adding this information would strengthen the quantitative claims.","section":"Section III-D"},{"comment":"The caption and the 'Overall' row do not define how the '+4.50%' error is computed. Clarify whether the reported error is the mean signed error, mean absolute error, or a normalized quantity.","section":"Table I"},{"comment":"The OpenFOAM setup is described only as echoing 'characteristic values found in sea waters'; the Reynolds number, mesh resolution, turbulence model, and boundary conditions are not specified, which limits reproducibility of the surrogate training data.","section":"Section II-D"},{"comment":"The dynamics simulation is presented as validation of the surrogate in a dynamic setting, but no quantitative comparison between simulated and measured trajectories or glide velocities is reported. Consider rephrasing this as a qualitative demonstration or adding quantitative comparison.","section":"Section III-B"}],"recommendation":"major_revision","confidential_remarks":"The manuscript describes an interesting and timely framework, and the hardware realization is a notable step. However, the experimental comparison that supports the central efficiency claim is not controlled, and the surrogate validation error is reported in a manner that flatters the results. The authors should either add a matched baseline experiment (e.g., testing the traditional shell with the same hardware in the same pool) or significantly temper the outperformance claim. The 4.5% error reporting should be corrected. I do not see a fundamental flaw in the method itself, so major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chen et al. build an automated pipeline: cage-based shape parameterization over 20 curated base shapes, a shared neural surrogate for lift-to-drag across angles of attack, then CMA-ES co-optimization of shape and control. The wind tunnel check of the 9-degree design (average 4.5% error in lift-to-drag) and the modular hardware with two fabricated shells are real evidence. The appendix derivation linking glider work-per-distance to cd/cl is clean. This is a coherent, well-engineered paper, and the integrated result is genuinely new relative to the cited literature, which mostly tunes low-dimensional hulls or optimizes shape and control separately.\n\nThe soft spots are all around the central efficiency claim. Section III-D compares η=2.5 measured in the pool for the two-wing glider against η=0.3 from Ref. [19], which is a CFD study of a different torpedo-shaped glider. That is not a matched baseline. No pool data are given for a 'traditional' shell on the same modular hardware, even though Figure 5 shows one. The pool measurements have no run counts or uncertainty. And Section IV honestly reports the sim-to-real gap: simulation said η=7, the pool said 2.5. The authors attribute it to surface details and manufacturing holes; that is plausible, but it means the surrogate's predictions for optimized shapes are accurate in the wind tunnel at model scale and much less accurate for the actual flooded, detailed vehicle. So the claim that the designed gliders 'significantly outperform' the previous standard design is not yet supported to the level the text states. It is a plausible claim; the paper simply lacks the controlled experiment needed to establish it.\n\nTwo smaller notes. The design space is the convex hull of 20 manually curated shapes, so 'optimal' is relative to that hull, and the paper does not discuss how restrictive that is despite the cage's expressiveness. No code or data are released, so reproducibility is limited, though the wind tunnel table helps.\n\nWho is this for? Anyone working on computational robot co-design or underwater glider hull optimization. It deserves a serious referee. The main fix is not new theory; it is a matched baseline: run the same modular vehicle with the traditional shell and one of the optimized shells in the same pool protocol, report repeats and spread, and frame the comparison around that. With that addition the central claim becomes credible. I would treat the paper as a conditional accept if I were the editor.","headline":"Clever co-design pipeline with real hardware, but the headline efficiency claim compares a pool measurement against a different vehicle's CFD number, so the outperformance is not yet controlled.","tokens_in":9733,"tokens_out":1631,"would_cite":true,"duration_ms":16863,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated pipeline that co-optimizes hull shape and control with a neural fluid surrogate is claimed to yield underwater gliders whose measured lift-to-drag ratio beats conventional torpedo designs.","keywords":["underwater glider","lift-to-drag ratio","neural fluid surrogate","deformation cage","co-optimization of shape and control","computational design","hydrodynamic shape optimization"],"falsifier":"Measure the lift-to-drag ratio of the fabricated four-wing design at its 30-degree operating angle in the same wind tunnel used for the 9-degree design; if the measured $\\eta$ does not exceed the torpedo baseline's $\\eta = 0.3$ by the margin the surrogate predicts, the surrogate's generalization to unseen shapes is not established.","tokens_in":8813,"feed_emoji":"🌊","tokens_out":11094,"duration_ms":101401,"temperature":0.7,"pith_summary":"The paper sets out to replace manual trial-and-error hull design for underwater gliders with an automated, end-to-end computational pipeline. It argues that the reason existing gliders look alike is not manufacturing but the lack of design tools that can explore non-trivial shapes and evaluate their fluid performance cheaply. The proposed workflow represents a hull by a low-dimensional deformation cage, predicts its lift and drag with a neural network trained on computational fluid dynamics data, and co-optimizes the shape and the angle of attack to maximize the lift-to-drag ratio $\\eta = c_l/c_d$. The authors report that two fabricated designs from this pipeline reach measured effective efficiencies of $\\eta = 2.5$ and $\\eta = 2.4$, compared with $\\eta = 0.3$ for a standard torpedo-shaped glider, so the practical upshot would be a path to more energy-efficient ocean sampling vehicles.","feed_headline":"AI-surrogate pipeline out-glides torpedo-shaped designs","feed_subtitle":"Measured glider efficiency of 2.5 beats the torpedo shape's 0.3 in wind-tunnel and pool tests.","key_machinery":"The load-bearing mechanism is the pairing of a three-dimensional deformation cage with a four-layer MLP (multilayer perceptron) fluid surrogate. The deformation cage takes offsets of cage handles as input and produces a deformed mesh of an initial ellipsoid; the paper curates 20 base shapes and interpolated morphs so that every geometry in the dataset shares the same low-dimensional parameterization. The neural surrogate maps cage parameters plus angle of attack to drag and lift coefficients, replacing a CFD solve with a fast differentiable evaluation. A covariance matrix adaptation evolution strategy (CMA-ES) then maximizes $\\eta = c_l/c_d$ within the convex hull of the training shapes, and the selected optimum is exported directly to CAD for three-dimensional printing. The argument's force comes from moving the expensive fluid solve into training-data generation, leaving the design loop itself cheap enough to iterate over many shapes and angles.","core_discovery":"On its own terms, the paper's central discovery is that a differentiable neural fluid surrogate, coupled with a compact geometry parameterization, is enough to navigate a design space of hull shapes that would be impractical to search with full computational fluid dynamics at every iteration. The authors demonstrate this by optimizing the lift-to-drag ratio $\\eta$, which their appendix shows is the sole shape-dependent factor in buoyancy-engine work per distance, across angles of attack from $-30^\\circ$ to $+30^\\circ$. The optimizer returns a family of non-torpedo hulls, two of which were fabricated as exchangeable shells on a common internal hardware assembly. Wind-tunnel measurements on the 9-degree design agree with the surrogate to an average 4.5% error, and pool tests give effective lift-to-drag ratios of $\\eta = 2.5$ for the two-wing and $\\eta = 2.4$ for the four-wing glider, both said to outperform the torpedo baseline at $\\eta = 0.3$.","pith_inferences":["Inference: the same cage-plus-surrogate recipe should transfer to other lift-driven vehicles, such as aerial gliders or underwater helicopters, whenever the objective can be written as a ratio of aerodynamic coefficients; the paper's contribution is the workflow, not a claim limited to gliders.","Inference: the reported sim-to-real gap for the two-wing design ($\\eta = 7$ in simulation versus $2.5$ measured) suggests that adding a shear-stress or surface-roughness penalty to the surrogate could recover most of the lost efficiency, and that the shape optimizer may already be near the practical optimum for smooth hulls.","Inference: because the search is restricted to the convex hull of 20 curated shapes and the paper notes the cage handles thin shapes poorly, the claimed optimum is relative to a curated design space; a representation allowing thin or toroidal geometries could plausibly shift the optimum further.","Inference: a testable extension is to optimize a single hull for a distribution of angles of attack or for robustness to currents, rather than emitting a separate optimal shape per angle, which would make the designs more useful in unsteady ocean environments."],"forward_implications":["A glider's travel distance per unit of buoyancy-engine work is set by $\\eta$, so if the reported $\\eta = 2.5$ transfers to real missions, each ballast cycle carries the glider several times farther than the $\\eta = 0.3$ torpedo shape for the same energy.","Because the pipeline exports fabrication-ready CAD, the design-to-prototype loop can be closed much faster than manual trial and error, and the same internal hardware can be re-shelled for different missions.","Wind-tunnel agreement within 4.5% suggests the surrogate can stand in for CFD throughout the optimization loop, which is what makes the co-design of shape and angle of attack tractable.","The co-optimization perspective implies that a hull optimized for one operating angle may be suboptimal at another, so choosing the angle of attack is part of the design problem rather than a post-hoc controller setting."],"supporting_citations":[{"why":"Supplies the torpedo-shaped baseline whose lift-to-drag ratio $\\eta = 0.3$ is the comparison for the optimized gliders.","marker":"[19]"},{"why":"Introduces the deformation-cage reduced-order geometry representation that defines the glider design space.","marker":"[20]"},{"why":"CFD solver used to generate ground-truth lift and drag coefficients for surrogate training.","marker":"[24]"},{"why":"Covariance matrix adaptation evolution strategy used to maximize lift-to-drag ratio over the cage parameters.","marker":"[27]"},{"why":"Supplies the wind-tunnel facility specifications used for validation measurements of the surrogate predictions.","marker":"[28]"},{"why":"Open-source glider hardware used as the traditional configuration against which optimized shells are mounted and compared.","marker":"[31]"}],"fun_headline_variants":["AI-designed gliders beat torpedo shape in efficiency","Neural surrogate guides glider design to 2.5 efficiency","AI co-optimizes shape and control for efficient gliders","Differentiable fluid model yields non-torpedo hulls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the neural fluid model trained on simulated data for 20 curated hull shapes and their interpolated variants also predicts the performance of the newly optimized shapes; if that generalization fails, the claimed efficiency advantage over conventional torpedo designs does not follow.","fun_headline_variants_meta":{"raw":{"variants":["AI-designed gliders beat torpedo shape in efficiency","Neural surrogate guides glider design to 2.5 efficiency","AI co-optimizes shape and control for efficient gliders","Differentiable fluid model yields non-torpedo hulls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1270,"prompt_tokens":953,"completion_tokens":317,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":247}},"tokens_in":569,"tokens_out":317,"duration_ms":3571,"temperature":1.0,"reasoning_tokens":247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:47:26.813512+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the lift-to-drag ratio of the fabricated four-wing design at its 30-degree operating angle in the same wind tunnel used for the 9-degree design; if the measured $\\eta$ does not exceed the torpedo baseline's $\\eta = 0.3$ by the margin the surrogate predicts, the surrogate's generalization to unseen shapes is not established.","supporting_citations":[{"cited_title":"Computational fluid dynamics study of the hydrodynamic characteristics of a torpedo-shaped underwater glider,","cited_arxiv_id":null,"evidence_quote":"Supplies the torpedo-shaped baseline whose lift-to-drag ratio $\\eta = 0.3$ is the comparison for the optimized gliders."},{"cited_title":"Openfoam: Open source cfd in research and industry,","cited_arxiv_id":null,"evidence_quote":"CFD solver used to generate ground-truth lift and drag coefficients for surrogate training."},{"cited_title":"The cma evolution strategy: a comparing review,","cited_arxiv_id":null,"evidence_quote":"Covariance matrix adaptation evolution strategy used to maximize lift-to-drag ratio over the cage parameters."},{"cited_title":"AeroAstro","cited_arxiv_id":null,"evidence_quote":"Supplies the wind-tunnel facility specifications used for validation measurements of the surrogate predictions."},{"cited_title":"Williams","cited_arxiv_id":null,"evidence_quote":"Open-source glider hardware used as the traditional configuration against which optimized shells are mounted and compared."}],"review_version":1}