{"id":"efabffc0-0d6a-4e9b-a21c-8f97019dcf21","arxiv_id":"2411.15130","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A PPO-trained policy tracks aerobatic 3D trajectories for a simulated bird-inspired flapping-wing robot, with stability argued from a fitted linear model rather than a formal proof.","lead":"A UC Berkeley team trained a reinforcement learning policy in simulation to control a bird-sized flapping-wing robot with five actuated joints, tracking loops, turns, and other aerobatic paths. The paper reports agile multimodal flight in simulation, but the robot hardware is not built and the stability claim rests on a fitted linear model, not a proof.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stability conclusion overreaches: BIBO stability of fitted LTI transfer functions (Sec. IV.A) does not imply the asymptotic stability of the nonlinear closed-loop system claimed in Sec. V.","rationale":"The paper's principal theoretical claim is that the RL-controlled flapping-wing robot is asymptotically stable (Sec. V). The only support is the BIBO stability of three identified LTI transfer functions (Eqs. 9-11) and the observation that all poles lie in the LHP (Fig. 6). This support is logically insufficient: an approximate LTI fit to input-output data does not certify properties of the underlying nonlinear system; BIBO stability does not imply state convergence; and the system operates on a periodic orbit, so the relevant notion is orbital stability, not classical asymptotic stability. I therefore identify the stability overreach as the single most load-bearing concern. The reader's weakest assumption, aerodynamic fidelity, is a valid external-validity caveat that the paper itself acknowledges in Sec. III.E; it threatens transfer to hardware but does not undermine the internal simulation demonstration as directly as the stability overclaim does. The proposed perturbation test would settle whether the asymptotic stability claim holds even within the simulation where the claim is made. Because the tracking demonstration itself appears plausible and the stability claim can be revised or withdrawn without invalidating the central learning result, the reader's CONDITIONAL verdict remains appropriate rather than moving to REJECT.","tokens_in":12223,"tokens_out":6698,"duration_ms":64652,"concrete_test":"Run a closed-loop disturbance-rejection experiment in the same MuJoCo environment: from steady 3.8 m/s cruising, inject a one-time offset of +0.75 m in position and +1 m/s in each body velocity, then record position and joint states for 10 s. Repeat across at least 50 random seeds and across the randomized dynamics in Table III. If trajectories do not converge back to the pre-perturbation periodic orbit in all cases, the Sec. V asymptotic-stability claim is false. As an analytical complement, compute the eigenvalues of the Poincare return map of the nominal flapping orbit; orbital asymptotic stability requires the non-trivial eigenvalues to lie strictly inside the unit circle.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central formal claim is that the RL-controlled flapping-wing robot is asymptotically stable, stated in the Conclusion (Sec. V). The only support is the BIBO stability of three identified LTI transfer functions (Eqs. 9-11), inferred from poles in the left-half plane (Fig. 6). This evidence is logically insufficient in three ways. First, an approximate LTI fit to input-output data does not certify properties of the underlying nonlinear system; the reported MSE of 5.629e-5 is a fitting error on the identification data, not a certificate. Second, BIBO stability concerns bounded-input/bounded-output behavior, not convergence of internal states; it does not imply asymptotic stability. Third, the system operates on a periodic flapping orbit, so the relevant notion is orbital stability of the limit cycle, which cannot be read off from the pole locations of a position-to-position transfer function. The paper's own phase portraits in Sec. IV.B confirm periodic motion, making 'asymptotic stability' in the classical sense an ill-posed claim. The stability statement therefore overreaches from a local, data-driven input-output property to a nonlinear asymptotic-stability conclusion, and this is the load-bearing weak point of the paper's formal narrative.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a model-free reinforcement learning framework for trajectory tracking of a five-joint, bird-inspired flapping-wing robot simulated in MuJoCo. The policy receives joint, attitude, and velocity observations, a history window, and a look-ahead trajectory buffer, and outputs target joint positions passed through a low-pass filter and a joint PD controller. Training uses a three-stage curriculum plus dynamics and aerodynamic randomization. The authors evaluate tracking on longitudinal, lateral, and aerobatic trajectories, test robustness to winds and aerodynamic coefficient randomization, perform linear system identification on the closed-loop input-output behavior, and present phase portraits of wing joints. The central claims are that the RL policy achieves multimodal agile tracking and that the closed-loop system is asymptotically stable.","tokens_in":12501,"tokens_out":4036,"duration_ms":40330,"significance":"The empirical simulation results are a useful contribution: they show that a single model-free policy can track diverse trajectories, including loops and turns, on a high-degree-of-freedom flapping-wing model, and the ablation-style randomization study gives a first-order sensitivity ranking of aerodynamic coefficients. The paper is transparent about the simulation-only status (the hardware platform is still in design) and about the simplifying stateless aerodynamics. Those strengths make the tracking claim credible within the simulator. The theoretical stability claim, however, is not supported by the presented evidence, and the manuscript should be revised to align the formal statements with what the analysis actually shows.","major_comments":[{"comment":"The conclusion that the closed-loop system is 'asymptotically stable' is not supported. The only evidence is the left-half-plane pole locations of three identified LTI transfer functions (Eqs. 9–11), which establish bounded-input/bounded-output stability for those fitted linear models. BIBO stability of an approximate input-output model does not imply asymptotic stability of the underlying nonlinear closed-loop system, nor does it address orbital stability of the periodic flapping orbit that the phase portraits in Sec. IV.B exhibit. The text in Sec. IV.A.2 correctly limits the claim to 'locally input-output stable'; the Conclusion should state the same limitation, or the authors should provide a nonlinear stability analysis (e.g., Lyapunov or contraction arguments) or a Poincaré-section analysis.","section":"IV.A.2, V"},{"comment":"The system-identification section does not report validation on held-out data, excitation conditions, or uncertainty bounds. The MSE of 5.629e-5 appears to be a fit error on the identification input-output pairs; as such it does not certify that the low-dimensional LTI model captures the closed-loop dynamics outside the particular fitted trajectory. Please report train/test MSE, the identification signal (e.g., chirp or random step sequence), and the operating region over which the fit is valid, and restrict stability conclusions to that region.","section":"IV.A.1"},{"comment":"All flight and robustness results are obtained with MuJoCo's stateless ellipsoid fluid model, whose coefficients are hand-tuned to match only the designed glide lift-to-drag ratio. The paper itself acknowledges this as a simplifying approximation and states that the physical robot is still in design. The tracking and robustness claims therefore should be framed as properties of the simulation model, not of a physical flapping-wing aircraft; an unqualified statement such as 'achieve stable flight' in the abstract may be read as a real-world claim. A concrete step would be to evaluate the trained policy in a higher-fidelity unsteady aerodynamic solver (e.g., UVLM or CFD) or on the physical platform once available, and to report differences in tracking error and success rate.","section":"II.C, III.E"}],"minor_comments":[{"comment":"The caption lists 'wing pitch angle' twice; the first occurrence should probably be 'wing flap angle'.","section":"Fig. 4 caption"},{"comment":"'MoJoCo' should be 'MuJoCo' in the paragraph on aerodynamic randomization.","section":"III.E"},{"comment":"The phrase 'Unsteady V ortex Latex Method' appears to be a typo for 'Unsteady Vortex Lattice Method'.","section":"I.B"},{"comment":"The sentence 'the input of the closed system is determined by the input of the policy network' is unclear; please clarify how the desired trajectory is converted into the LTI input u and what exactly is measured as the output.","section":"IV.A.1"},{"comment":"No code or simulation environment is provided; a public release would improve reproducibility and allow independent assessment of the system-identification and stability analysis.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's main contribution is empirical, and its formal stability claim is not supported by the presented LTI analysis. The paper should be returned for revision that tempers the stability language, adds validation details for the system identification, and clearly scopes the results to the simulation model. I do not see concerns about prior disclosure or inappropriate citation patterns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is what you should know. This paper trains a PPO policy for a five-actuator bird-scale ornithopter model in MuJoCo and shows it can track cruise, climb, glide, dive, turn, loop, and Immelmann maneuvers. That is a legitimate extension of the RL flapping-wing literature, and the qualitative evidence (trajectory plots, phase portraits, maneuver snapshots) is consistent with the claim that the policy tracks in this simulator. The authors are also honest about their modeling: they state the MuJoCo fluid model is stateless, that coefficients are hand-tuned to match a gliding lift-to-drag ratio, and that the hardware does not exist yet. Credit where due.\n\nThe soft spot is the stability narrative. The conclusion says the closed-loop system is asymptotically stable. The support is a system-ID step that fits three LTI transfer functions (Eqs. 9-11) and checks pole locations in the LHP. That gives at most local BIBO stability of a fitted linear model, not asymptotic stability of the nonlinear closed-loop system. Worse, the phase portraits in Sec. IV.B show periodic orbits, so the relevant notion is orbital stability of a limit cycle, not asymptotic stability toward an equilibrium. The paper's own fit has MSE 5.629e-5, but that is a fitting error on the identification data, not a certificate. So the stability claim overreaches from a data-driven input-output property to a nonlinear theoretical property. That is the main load-bearing flaw.\n\nThe other soft spots are minor in comparison: no code or data release, and the aerodynamic model fidelity is untested. Those are standard for a simulation-only paper at this stage, but they make the practical impact speculative. The training recipe is established PPO with domain randomization, so novelty is modest though the platform and curriculum are new.\n\nOverall, the core simulation demonstration is defensible and the overclaim is addressable. I would send this to peer review with a request to fix the stability language and to add at least a reproducibility artifact or a sensitivity analysis beyond the success-rate plot, and I would not cite it in my own work yet.","headline":"Solid RL flapping-wing tracking demo in MuJoCo with an overreaching asymptotic-stability claim that should be softened.","tokens_in":13011,"tokens_out":2369,"would_cite":false,"duration_ms":22575,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A model-free RL policy tracks aerobatic 3D trajectories for a flapping-wing robot in simulation.","keywords":["flapping-wing robot","ornithopter","reinforcement learning","trajectory tracking","PPO","MuJoCo","domain randomization","system identification"],"falsifier":"Measure the actual lift and drag of the physical flapping-wing platform across the flapping frequencies used by the policy (4–6 Hz) and compare them against MuJoCo predictions under identical kinematics; a substantial discrepancy would indicate that the policy is exploiting a simulation artifact. A more direct test is to deploy the trained policy on the physical robot and check whether it can sustain stable straight-line flight and track a simple trajectory at all.","tokens_in":12000,"feed_emoji":"🐦","tokens_out":3149,"duration_ms":30204,"temperature":0.7,"pith_summary":"The paper claims that a model-free reinforcement-learning controller can make a high-degree-of-freedom, bird-inspired flapping-wing robot track complex 3D trajectories, including loops, sharp turns, and roll-off-the-bottom maneuvers, while also maintaining stable, periodic flapping and spontaneously switching between flight modes. This is shown entirely in a MuJoCo simulation where the aerodynamic model is a stateless ellipsoid/inertia approximation with hand-tuned coefficients. The authors add that the closed-loop system is locally input-output stable, based on linear transfer functions identified from simulated input-output data, and that the controller remains robust under wind disturbances and randomized aerodynamic conditions. If correct, this would mean that model-free RL can substitute for hand-derived aerodynamic models in ornithopter control, a task that classical model-based methods have only partially solved.","feed_headline":"One RL policy flies a flapping-wing robot through loops and turns","feed_subtitle":"Simulated bird-sized ornithopter tracks 3D paths and switches flight modes on its own, with stability shown via identified linear models.","key_machinery":"The load-bearing machinery is the MuJoCo simulation environment with its stateless aerodynamic model, in which lifting bodies are modeled as ellipsoids with five aerodynamic force contributions (added mass, viscous drag, Magnus lift, Kutta lift, and viscous resistance) and the main body is modeled with a simplified inertia model. The fluid coefficients are manually tuned to match a designed lift-to-drag ratio at gliding, and this model is embedded in a curriculum-based PPO training pipeline with three progressively harder stages (constant forward flight, climbing/diving, then turning and arbitrary maneuvers) followed by domain randomization of masses, inertias, aerodynamic coefficients, added mass and inertia, and wind velocity. The identified third-order LTI transfer functions in x, y, and z are the device used to argue closed-loop stability.","core_discovery":"The central claim is that a single model-free RL policy, trained with PPO through a curriculum and heavy domain randomization, can act as a trajectory-tracking controller for a simulated 11-DoF flapping-wing robot (5 actuated wing and tail joints plus a 6-DoF floating base). The policy outputs target joint positions at 50 Hz that are low-pass filtered and passed to a low-level PD controller running at 250 Hz, and this closed loop tracks straight, climbing, diving, gliding, turning, loop, and roll-off-the-bottom maneuvers. The stability argument is built from system identification: a third-order linear time-invariant transfer function is fit in each spatial axis to the closed-loop input-output behavior, and the identified poles all lie in the left-half plane, indicating bounded-input bounded-output (BIBO) stability and minimum-phase behavior. Phase portraits of the wing flap and pitch joints show closed periodic orbits during forward flight, climbing, and turning, which the paper interprets as stable and periodic joint action patterns.","pith_inferences":["An untested but implicit implication is that the same approach would transfer to physical hardware, but the stateless aerodynamics model omits unsteady effects such as leading-edge vortices and wing flexibility, so the demonstrated tracking and stability may not survive on a real robot.","A testable extension would be to compare the learned flapping frequency (4–6 Hz) and wing kinematics against measured data from biological birds or existing ornithopters of similar size; a mismatch would indicate that the policy exploits simulation artifacts rather than physical aerodynamics.","The BIBO stability of the identified local linear model does not by itself establish asymptotic stability of the full nonlinear stochastic closed-loop system; a stronger certificate, such as a Lyapunov function on the original dynamics, would be needed to make the stability conclusion robust.","The paper's sensitivity analysis shows the policy is most affected by the Kutta lift coefficient, which suggests that improving the fidelity of that single aerodynamic term could be the highest-impact step toward real-world transfer."],"forward_implications":["If the RL control claim holds, model-free reinforcement learning could become a practical alternative to model-based control for bird-sized flapping-wing platforms, removing the need for analytic aerodynamic models.","The curriculum-plus-domain-randomization training scheme may transfer to other high-degree-of-freedom, underactuated flying robots, including morphing-wing and bat-like platforms.","The stability-analysis method, fitting low-dimensional linear models to the closed-loop input-output behavior of a learned policy, could be reused to certify other learned flight controllers beyond this platform.","The single policy's spontaneous switching between flight modes suggests that a single learned controller can replace hand-tuned supervisory logic that selects among separate controllers for cruise, climb, dive, and turn.","The policy's robustness to randomized aerodynamic coefficients and wind in simulation points toward a potentially deployable controller once the sim-to-real gap for flapping-wing aerodynamics is closed."],"supporting_citations":[{"why":"MuJoCo is the physics engine that computes the multi-body dynamics and provides the stateless fluid force models used for training and evaluation.","marker":"[38]"},{"why":"Supplies the aerodynamic force and moment formulas for the ellipsoid and inertia fluid models, including added mass, drag, Magnus, Kutta, and viscous terms.","marker":"[36]"},{"why":"Proximal Policy Optimization is the RL algorithm used to train the trajectory-tracking policy.","marker":"[41]"},{"why":"The system identification approach for extracting low-dimensional linear models from closed-loop RL systems is the method used to analyze closed-loop stability.","marker":"[42]"},{"why":"The energy-minimization reward design that encourages gliding over flapping and discourages unrealistically high flapping frequencies is based on this work.","marker":"[40]"},{"why":"The domain randomization ranges for body mass, inertia, and center of mass follow the practice established in this prior hierarchical RL work.","marker":"[39]"}],"fun_headline_variants":["RL policy tracks loops and turns on a flapping-wing robot","Single RL controller enables agile flapping-wing flight","Learning-based control stabilizes a bird-inspired ornithopter","Reinforcement learning unlocks flapping-wing agility","One RL policy flies a flapping-wing bird through tricks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire demonstration rests on the assumption that MuJoCo's stateless aerodynamic model, with fluid coefficients manually tuned to match a gliding lift-to-drag ratio, faithfully represents real flapping-wing flight; if that model is not representative of unsteady flapping-wing aerodynamics, the learned tracking and stability conclusions do not transfer to a physical robot.","fun_headline_variants_meta":{"raw":{"variants":["RL policy tracks loops and turns on a flapping-wing robot","Single RL controller enables agile flapping-wing flight","Learning-based control stabilizes a bird-inspired ornithopter","Reinforcement learning unlocks flapping-wing agility","One RL policy flies a flapping-wing bird through tricks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1392,"prompt_tokens":893,"completion_tokens":499,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":421}},"tokens_in":509,"tokens_out":499,"duration_ms":5693,"temperature":1.0,"reasoning_tokens":421,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:27:24.622462+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the actual lift and drag of the physical flapping-wing platform across the flapping frequencies used by the policy (4–6 Hz) and compare them against MuJoCo predictions under identical kinematics; a substantial discrepancy would indicate that the policy is exploiting a simulation artifact. A more direct test is to deploy the trained policy on the physical robot and check whether it can sustain stable straight-line flight and track a simple trajectory at all.","supporting_citations":[{"cited_title":"Mujoco: A physics engine for model-based control,","cited_arxiv_id":null,"evidence_quote":"MuJoCo is the physics engine that computes the multi-body dynamics and provides the stateless fluid force models used for training and evaluation."},{"cited_title":"Whole-body simulation of realistic fruit fly locomotion with deep reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the aerodynamic force and moment formulas for the ellipsoid and inertia fluid models, including added mass, drag, Magnus, Kutta, and viscous terms."},{"cited_title":"Bridging model- based safety and model-free reinforcement learning through system identification of low dimensional linear models,","cited_arxiv_id":null,"evidence_quote":"The system identification approach for extracting low-dimensional linear models from closed-loop RL systems is the method used to analyze closed-loop stability."},{"cited_title":"Minimizing energy consumption leads to the emergence of gaits in legged robots,","cited_arxiv_id":null,"evidence_quote":"The energy-minimization reward design that encourages gliding over flapping and discourages unrealistically high flapping frequencies is based on this work."},{"cited_title":"Hierarchical reinforcement learning for precise soccer shooting skills using a quadrupedal robot,","cited_arxiv_id":null,"evidence_quote":"The domain randomization ranges for body mass, inertia, and center of mass follow the practice established in this prior hierarchical RL work."}],"review_version":1}