{"id":"c0f0af24-806b-41db-a9a2-f9cdcaaa829a","arxiv_id":"2504.21586","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"A domain-randomized neural network policy trained in simulation races both a 3-inch and a 5-inch quadcopter in the real world, and randomization level trades speed for sim-to-real robustness.","lead":"A single neural network trained with randomized physics in simulation can fly two very different racing quadcopters in the real world, a first for drone racing. The work shows that adding randomization during training trades a little speed for much better transfer across drone platforms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generalization claim is an interpolation result: both test drones lie inside the randomization ranges built from their identified parameters; no holdout platform tests 'various types'.","rationale":"The paper contains a genuine positive result: one network, trained with domain randomization, flies two physically different real quadcopters, and 0%-randomization policies fail to transfer, providing a clean negative control. That supports an existence proof of a robust interpolating policy. My concern is not primarily about model fidelity or statistical power; even a simplified model is acceptable if the real flights demonstrate transfer. The load-bearing issue is scope: the randomization distributions were built from the test platforms' identified parameters (Sec. II-C.1), so the real tests are in-sample relative to the training distribution. For a claim of generalization across 'various types' or 'any platform', a held-out platform is required. The reader identified the overreach in the rationale but placed the weakest-assumption weight on model accuracy; I make the missing holdout the central concern. This does not change the verdict: conditional acceptance remains appropriate, with the holdout experiment as the decisive next check.","tokens_in":9408,"tokens_out":18908,"duration_ms":214958,"concrete_test":"Deploy the already-trained general network on a third quadcopter (e.g., a 4-inch or 7-inch craft) whose identified parameters fall outside the Table II randomization ranges, with no retraining. Report gate-passing counts, episode reward, and crash outcomes over the same 12-s figure-eight task. If the unchanged network completes the track, the 'various types' claim is supported; if it fails, the conclusion should be narrowed to interpolation over the designed randomization ranges.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing step is the inference from 'works on the 3-inch and 5-inch drones' to 'various types of quadcopters' (abstract) and 'any platform' (conclusion). The randomization ranges in Table II were explicitly designed to encompass the parameters of these two drones (Sec. II-C.1), so both real tests lie inside the support of the training distribution. Success on them demonstrates robust interpolation within a hand-chosen parameter box, but not extrapolation to a platform whose parameters fall outside that box. Since the claimed contribution is cross-platform generality, this missing holdout is the weakest point of the central argument. The paper itself flags related limitations (gyroscopic effects ignored in Sec. II.A; only three 12-s flights per condition in Sec. IV-C; online adaptation 'did not achieve the desired improvements' in Sec. V), but these do not undermine the two-drone existence proof; the unverified extrapolation does.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a single PPO-trained neural network policy for quadcopter racing that maps the vehicle state and gate information directly to motor commands. The policy is trained in a parametric simulator with domain randomization over dynamic parameters, and the authors compare it with fine-tuned policies trained on identified 3-inch and 5-inch drone models at 0-30% randomization. The main empirical claims are that the general network flies both real drones through a figure-eight gate track at up to about 10 m/s, that 0% randomization fails to transfer while increasing randomization improves robustness but lowers speed, and that the general policy is slightly slower than the best fine-tuned policies but transfers across platforms. The paper reports 1000-rollout simulation evaluations, real flights with three trials per condition, and an optimal-control lap-time comparison.","tokens_in":9553,"tokens_out":4758,"duration_ms":48338,"significance":"If the results hold, the paper provides a useful demonstration that a single end-to-end racing policy can transfer across two substantially different quadcopter sizes, with an important negative control (0% randomization) and an openly available implementation. The novelty relative to earlier domain-randomization quadrotor work (e.g., Molchanov et al. [16]) is the extension to high-speed racing and the systematic ablation of randomization level. However, the central generalization claim is demonstrated only as interpolation: both test drones lie inside the randomization ranges that were designed from their identified parameters. The evidence base is small, and the model omits effects that matter at high speed, so the broad 'any platform' claim requires a holdout platform or a more cautious wording.","major_comments":[{"comment":"The generalization claim goes beyond what the experiments establish. Table II and the text state that the uniform randomization ranges were 'designed to encompass parameters from both sizes'; both test quadcopters therefore lie inside the support of the training distribution. The real-world successes on the 3-inch and 5-inch drones demonstrate robust interpolation inside a hand-picked parameter box, not extrapolation to 'various types' of quadcopters or 'any platform' as claimed in the abstract and conclusion. To support the headline contribution, the authors should either add a holdout platform whose identified parameters fall outside the training ranges, or substantially reword the generalization claims.","section":"Sec. II-C.1 / Abstract / Conclusion"},{"comment":"The sim-to-real transfer rests on a parametric model that the paper itself acknowledges is simplified: gyroscopic effects are ignored, the moment of inertia is estimated indirectly, and motor saturation, blade flapping, and other high-speed aerodynamic effects are absent. With only three 12-second real flights per condition and no reported statistical analysis, the quantitative claims -- for example, that reward increases as randomization decreases from 30% to 10%, or that 0% randomization fails -- are vulnerable to run-to-run noise. The paper should report confidence intervals or additional trials for the real flights, and should temper claims about the model's adequacy for platforms other than the two tested.","section":"Sec. II-A and Sec. IV-C"},{"comment":"The statement that 'at 0% randomization, the drone no longer passes through the gates' is internally inconsistent with the data in the same tables: the 0% policy passes 5, 5, and 7 gates on the 3-inch drone and 8, 13, and 8 gates on the 5-inch drone in real flight. The correct observation is that the episode reward collapses relative to simulation and the flights do not complete the intended track. This matters because the 0% condition is the central negative control for the domain-randomization argument; the paper should state precisely what failure criterion defines 'fails to transfer'.","section":"Sec. IV-C, Tables IV and V"}],"minor_comments":[{"comment":"The footnote contains typos: 'inderectly' should be 'indirectly' and 'throught' should be 'through'.","section":"Sec. II-A, footnote 1"},{"comment":"The notation pgi, vgi, λgi, and the superscripts on pgi+1 and ψgi+1 should be defined in the text; the reference-frame notation is not explained before use.","section":"Sec. II-D, Eq. (3)"},{"comment":"The abbreviation 'ep rew' is not defined; define it as 'episode reward' in the captions to improve readability.","section":"Tables IV and V"},{"comment":"The cap on kl during fine-tuning is described only in prose; the fine-tuning description and Table II should state the cap explicitly so that the procedure is fully reproducible.","section":"Sec. II-C.2"},{"comment":"The optimal-control comparison omits drag terms, as the authors note, and the claim that the RL path is 'relatively far away from the edges of the gates' is not quantified; a trajectory-overlay figure or a distance metric would support this point.","section":"Sec. IV-D"}],"recommendation":"major_revision","confidential_remarks":"The work is a reasonable incremental contribution for a robotics venue: it combines established domain-randomization ideas with a new racing application and a clean ablation of randomization levels. The main risk is that the published claims (abstract and conclusion) overstate the generality of a two-drone interpolation result; the revision should either add a holdout platform or scale back those claims. I see no circularity or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper actually delivers its core empirical result: one PPO policy, trained with domain randomization on a parametric sim, flies both a 3-inch and a 5-inch race quad through a figure-eight gate track in the real world at up to ~10 m/s. That is a clean existence proof for cross-platform high-speed racing, and as far as I know it is new. The 0% randomization negative control is the right kind of experiment: fine-tuned policies crash or miss gates while the DR policy completes laps. The DR sweep from 0% to 30% gives a coherent story, supported by 1000-rollout simulations and three real flights per condition. Architecture selection is more thorough than most papers bother with, and they openly report that adding input history or parameter conditioning did not help. Code is public. That all earns real credit.\n\nThe soft spots are mostly about scope. The randomization ranges in Table II were explicitly chosen to encompass the identified parameters of the two drones (Sec. II-C.1: \"Uniform distributions were designed to encompass parameters from both sizes\"). So the two real tests are interpolation inside a hand-picked parameter box, not extrapolation. The abstract's \"various types of quadcopters\" and the conclusion's \"any platform\" overstate the evidence. Two platforms, both inside the training distribution, do not justify that. The real-flight stats are also thin: three runs per condition, no error bars, battery draining across the three runs, and no significance testing. The model is simplified (gyroscopic effects ignored, no motor saturation or blade flapping), and the time-optimality analysis shows the reward function does not truly minimize lap time, which they acknowledge. One citation looks wrong: [24] is about a vibration energy harvester, not Moongel damping performance. Minor but sloppy.\n\nNone of these break the central claim. The two-drone existence proof holds up. What does not hold up is the scope of the language. A good revision would temper \"various types\" to \"two tested platforms,\" add a proper limitation paragraph on interpolation, and either add a third platform outside the randomization box or justify why that is not needed.\n\nWho is this for? Anyone working on drone racing, sim-to-real transfer, or robust control policy generalization. I would cite it as the first high-speed cross-platform DR demonstration, with the interpolation caveat noted. It deserves a serious referee: the empirical core is sound, the negative controls are informative, and the weaknesses are fixable in revision. Send it to review.","headline":"A solid two-drone existence proof for cross-platform DR racing, but the 'any platform' rhetoric outsells the data: both tested drones sit inside the randomization box.","tokens_in":10125,"tokens_out":1902,"would_cite":true,"duration_ms":21960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single neural network trained with domain randomization can race both a 3-inch and a 5-inch quadcopter through a real figure-eight gate course, while per-drone policies trained with no randomization fail to transfer from simulation to…","keywords":["drone racing","reinforcement learning","domain randomization","sim-to-real transfer","reality gap","quadcopter control","neural network controller","cross-platform generalization"],"falsifier":"Take the trained general network and fly it on a third quadcopter whose identified parameters fall inside the randomization ranges, such as a 4-inch racer, under the same track and measurement setup; if it passes substantially fewer gates or crashes within the same 12-second episodes, the claimed cross-platform generalization is refuted. In simulation, the equivalent test is to evaluate the general policy on parameter sets drawn from the same distributions but held out from training; if the crash rate rises far above the reported 1.2–5.4%, the policy has memorized the randomization ranges rather than generalizing.","tokens_in":9181,"feed_emoji":"🚁","tokens_out":6807,"duration_ms":71143,"temperature":0.7,"pith_summary":"This paper claims that one neural-network controller can race two physically different quadcopters—a 3-inch and a 5-inch racer—by training in a simulator whose drone parameters are randomized and then deploying the same network unchanged on both real aircraft. The controller reads only the current state and information about the current and next gate, and it outputs four motor commands directly, with no per-platform tuning. Real flights reached roughly 10 m/s and completed the gate course on both platforms, while separately fine-tuned policies trained with zero randomization failed to transfer to the real drones. The finding matters because, if it holds, platform-specific controller identification and training can be replaced by one randomized policy, at the cost of a modest speed penalty.","feed_headline":"One net races two very different drones through the same course","feed_subtitle":"Domain-randomized training lets one controller race 3-inch and 5-inch quadcopters; per-drone policies with no randomization fail to…","key_machinery":"The load-bearing object is the parametric quadcopter model of Section II-A: a first-order motor model with force and moment coefficients estimated from manual flight data and normalized by maximum motor speed, so the same equations describe both airframes. The paper adds a domain-randomization layer: training samples each coefficient from uniform distributions whose ranges cover both the 3-inch and 5-inch parameter sets, so the policy cannot memorize one platform. The 20-dimensional observation—state expressed in the current gate frame plus the next gate's position and orientation—is what lets the same network perform guidance and control end-to-end.","core_discovery":"The central claim is that domain randomization suffices to bridge the sim-to-real gap for high-speed racing across distinct quadcopters. The general policy was trained with proximal policy optimization on a randomized parametric model in which motor limits, thrust and torque coefficients, and time constants are sampled from uniform ranges spanning both aircraft; the resulting three-layer, 64-unit ReLU network takes a 20-dimensional observation and directly outputs normalized motor commands. In three 12-second real flights per aircraft, this one network navigated the seven-gate figure-eight on both the 3-inch drone (31 gates passed, mean speed 6.31–6.38 m/s, max about 10.4–10.6 m/s) and the 5-inch drone (46 gates, mean speed about 7.8 m/s, max about 9.8 m/s). Fine-tuned policies trained on each airframe's identified parameters with 0% randomization flew only 5–13 gates in reality despite strong simulation scores, while 10–30% randomization transferred and flew faster than the general policy; the paper reads this as a robustness–speed trade-off controlled by randomization strength.","pith_inferences":["A testable extension the paper does not run: within the sampled parameter envelope, the same general policy should transfer to a third airframe such as a 4-inch racer without retraining; comparing its gate counts to the 3-inch and 5-inch results would directly test the breadth of the generalization.","Because the paper found that adding action history or parameter inputs did not improve reward, the policy likely learns a control law that is insensitive to the randomized parameters rather than an explicit parameter estimate; this suggests future work could investigate what invariant the network encodes.","The time-optimal comparison suggests the main remaining gap to optimal lap times is the reward function, not the generalization mechanism; combining domain randomization with a reward that encourages flying closer to gate edges could plausibly recover speed without losing transferability."],"forward_implications":["A single network can be deployed on multiple physically distinct race quadcopters without retraining or per-platform parameter identification.","With zero randomization, otherwise strong simulation policies fail in reality, so some randomization is necessary for sim-to-real transfer in this racing task.","More randomization buys transferability but costs top speed: the 10% and 20% fine-tuned policies outran the general policy, while 30% randomization ran slower.","The general policy's real-world episode rewards were close to but below the fine-tuned 10% and 20% policies, suggesting the price of one-net-fits-all is moderate rather than prohibitive."],"supporting_citations":[{"why":"Supplies the parametric quadcopter dynamics and observation structure used as the training simulator.","marker":"[14]"},{"why":"Provides the sparse progress-plus-penalty reward function that shapes the racing behavior.","marker":"[5]"},{"why":"Defines the racing gate task and the randomized starting-state setup used in training.","marker":"[3]"},{"why":"Provides the proximal policy optimization algorithm used to train all policies.","marker":"[20]"},{"why":"Provides the reinforcement learning implementation used to run the training.","marker":"[21]"},{"why":"Identifies the partial-observability issue caused by randomized parameters becoming part of the state in the MDP.","marker":"[15]"}],"fun_headline_variants":["One neural network races 3-inch and 5-inch quadcopters","Domain randomization lets one net fly two different drones","Single AI controller pilots 3-inch and 5-inch race drones","One network, two platforms: domain randomization wins","Universal drone racing net adapts across diverse platforms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the parametric simulator, with all platform coefficients estimated from manual flight data and gyroscopic effects ignored, captures the dynamics that matter at racing speeds; if effects such as motor saturation, blade flapping, or asymmetric inertia become significant at 10 m/s and beyond, the trained policies would not transfer to new platforms.","fun_headline_variants_meta":{"raw":{"variants":["One neural network races 3-inch and 5-inch quadcopters","Domain randomization lets one net fly two different drones","Single AI controller pilots 3-inch and 5-inch race drones","One network, two platforms: domain randomization wins","Universal drone racing net adapts across diverse platforms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":2968,"prompt_tokens":990,"completion_tokens":1978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1912}},"tokens_in":606,"tokens_out":1978,"duration_ms":14790,"temperature":1.0,"reasoning_tokens":1912,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:59:01.193624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained general network and fly it on a third quadcopter whose identified parameters fall inside the randomization ranges, such as a 4-inch racer, under the same track and measurement setup; if it passes substantially fewer gates or crashes within the same 12-second episodes, the claimed cross-platform generalization is refuted. In simulation, the equivalent test is to evaluate the general policy on parameter sets drawn from the same distributions but held out from training; if the crash rate rises far above the reported 1.2–5.4%, the policy has memorized the randomization ranges rather than generalizing.","supporting_citations":[{"cited_title":"End-to-end reinforcement learning for time-optimal quadcopter flight,","cited_arxiv_id":null,"evidence_quote":"Supplies the parametric quadcopter dynamics and observation structure used as the training simulator."},{"cited_title":"Au- tonomous drone racing with deep reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Defines the racing gate task and the randomized starting-state setup used in training."}],"review_version":1}