{"id":"566eee5e-f1d2-44ba-8565-67cbb546c94f","arxiv_id":"1908.03478","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A model-based iterative learning control scheme transfers tracking-error corrections across heterogeneous UAVs and scenarios using known scaling and dynamics mapping matrices.","lead":"This paper proposes a learning algorithm that lets a target drone improve its trajectory tracking by learning from the errors of other drones, even when those drones have different dynamics or fly different paths. The idea is to apply each drone's physical model to rescale and transfer the learning signal, and the authors test it in simulations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Equation (11) silently cancels the target's own tracking error T{rd}; the final relation and the convergence conditions are only valid if the robust filters sum to identity, a condition never stated or enforced.","rationale":"The reader's weakest_assumption identifies exactly the same silent cancellation in Eq. (11), and my independent derivation confirms it: the dropped term is (1 - sum alpha_k) T{rd}, not a negligible higher-order quantity. This is an internal algebraic inconsistency, not a disagreement with a modeling consensus. It directly undermines the claimed convergence guarantee because the H-infinity conditions (12)/(20)/(23) bound only the weighted combination of training errors and do not control the residual (1 - sum alpha_k) T{rd}. The paper's simulations use one training UAV with alpha = 1, so the condition holds by construction; the generalization to multiple training UAVs is therefore unproven. The Section 4 approximation F_k approx F near frequencies where ||G|| is not large adds another layer of uncertainty, but the filter-sum condition is the more fundamental gap. This warrants a conditional verdict: the algorithm may work when sum alpha_k = 1 and the model approximations hold, but the paper must state and enforce the filter-sum condition and quantify the approximation error. I do not see grounds to reject the paper outright; the core idea is plausible and the single-UAV demonstrations are consistent with the corrected relationship. The proposed concrete test would settle whether the missing condition is indeed the governing factor.","tokens_in":8215,"tokens_out":5800,"duration_ms":59118,"concrete_test":"Re-derive Eq. (11) from (9)-(10) without discarding T{rd} - sum Lambda_k^{-1} alpha_k T{Lambda_k rd}, obtaining h = sum Lambda_k^{-1}(F L_k + alpha_k I){h_k} + (1 - sum alpha_k) T{rd}. Then reproduce the Section 5 numerical experiment with N = 2 training UAVs, first with alpha_1 = alpha_2 = 0.5 and then with alpha_1 = alpha_2 = 0.7, choosing L_k so that condition (23) holds. If the second case exhibits an additional tracking-error component close to -0.4 times the target's unlearned error T{rd} while the first case does not, the missing condition is confirmed and the convergence guarantee must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central derivation collapses at Eq. (11). Moving from the second to the third line, the term T{rd} - sum_{k in Omega} Lambda_k^{-1} alpha_k T{Lambda_k rd} is discarded. Because T is assumed to be a diagonal LTI system and Lambda_k are diagonal constant scaling matrices, T{Lambda_k rd} = Lambda_k T{rd}, so this term equals (1 - sum alpha_k) T{rd}. The final equality h = sum_k Lambda_k^{-1}(F L_k + alpha_k I){h_k} therefore requires sum alpha_k = 1 (or T{rd} = 0). The paper never states this condition, and the convergence conditions (12), (20), and (23) do not include it. If sum alpha_k != 1, the actual error contains an extra (1 - sum alpha_k) T{rd} term; even when the H-infinity norm condition is satisfied, h need not be small because the unlearned target error T{rd} can be large. The numerical studies in Sections 3-5 use a single training UAV with alpha = 1, so the assumption is satisfied trivially and the flaw is invisible. The Section 4 approximation F_k approx F is a separate source of error near frequencies where ||G|| is not large, but it is secondary; the missing filter-sum condition is the load-bearing issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a feedforward learning algorithm for UAV trajectory tracking in which the learning signal is generated from the tracking errors and control signals of other ('training') UAVs. The claimed contribution is that, by explicitly using physical-model transformations (reference scaling matrices Lambda_k and dynamics-mapping matrices M_k), a target UAV can learn in near real time from training UAVs with different dynamics and different reference trajectories. Sections 3, 4, and 5 present formulations for heterogeneous scenarios, heterogeneous agents, and their unification, each followed by a numerical simulation; convergence is stated in terms of H-infinity norm conditions. The paper concludes that the algorithm demonstrates great learning efficiency and reliability.","tokens_in":8642,"tokens_out":6014,"duration_ms":61223,"significance":"The core idea is potentially useful: if the derivation were correct, it would turn ILC-style error updates into a one-shot, near-real-time transfer between different vehicles by using model knowledge to overcome heterogeneity. Algorithm 1 is concrete enough to implement, and the simulation plots in Figs. 3, 5, and 6 show plausible improvement over tracking without learning. However, the central error-propagation equations contain unstated cancellations and notational errors, and the numerical evidence is entirely qualitative. The contribution is incremental over the authors' prior ILC design [23], and the convergence conditions are not proved in this manuscript; the value of the paper is in the concept and preliminary demonstration rather than in a fully verified algorithm. No machine-checked proofs or reproducible code are provided, so the positive assessment rests on the analytic derivations, which currently require correction.","major_comments":[{"comment":"The step from the fourth displayed line to the fifth displayed line drops the term T{r_d} - sum_{k in Omega} Lambda_k^{-1} alpha_k T{Lambda_k r_d}. Since T is a diagonal LTI system and Lambda_k is a diagonal constant matrix, T{Lambda_k r_d} = Lambda_k T{r_d}, so the dropped term equals (1 - sum_k alpha_k) T{r_d}. Eq. (11) is therefore valid only if sum_k alpha_k = 1 or T{r_d} = 0, and this condition is not stated anywhere in Section 3 and is not included in the convergence condition (12). If it is not enforced, the actual target error contains the additional unlearned term (1 - sum_k alpha_k) T{r_d}, which can be large even when the H-infinity condition (12) holds. The numerical verification in Section 3.2 uses a single training UAV with alpha_1 = 1, so the missing condition is satisfied trivially and the flaw is invisible. This must be fixed by either imposing sum_k alpha_k = 1 as an explicit design constraint or carrying the residual term through the convergence analysis.","section":"Section 3, Eq. (11)"},{"comment":"The derivation of Eq. (19) is not valid as written. The substitution F{s_k} = h_k - T_k{r_d} requires F = F_k, while the text has only argued F_k approx F under ||G|| >> 1. In addition, the simplification of T{r_d} - sum M_k^{-1} alpha_k T_k{r_d} to zero requires both the approximation T_k approx T M_k and the same sum_k alpha_k = 1 condition from the previous comment; neither is stated. The displayed expression M^{-1}_k( sum_{k in Omega} alpha_k I + F L_k ){h_k} also has a free k on M^{-1}_k outside the sum, making the equation ill-formed. Please rewrite Eq. (19) with correct summation indices, state the approximations explicitly, and either impose sum_k alpha_k = 1 or include the residual terms in the convergence bound.","section":"Section 4, Eq. (19)"},{"comment":"From Eq. (5), alpha_k I + F L_k approx 0, the implied learning filter is L_k approx -alpha_k F^{-1}, not L_k approx alpha_k F^{-1} as written in Eq. (6). The missing negative sign would change the sign of the feedforward correction and, if implemented as printed, would amplify rather than cancel the error. Please correct Eq. (6) and state which sign is actually used in the numerical simulations.","section":"Section 2, Eq. (6)"},{"comment":"The convergence conditions are stated by reference to prior work [23] with the phrase 'designed such that the above condition is satisfied [23]', but no proof is given that the stated H-infinity condition bounds the target error in the heterogeneous multi-UAV setup. The conditions are derived from the error-propagation equations (11), (19), and (22), which currently contain the unstated cancellations described above. A self-contained proof, or a precise statement of how [23] applies after the added transformation matrices Lambda_k and M_k, is necessary to support the central convergence claim.","section":"Sections 3-5, Eqs. (12), (20), (23)"},{"comment":"The numerical studies report only qualitative statements such as 'very small overshoot and oscillations' and show trajectory plots without quantitative error metrics. No tracking-error norms, overshoot percentages, or comparisons with a no-learning baseline are provided. Since the central claim is a quantitative error-reduction property, the paper should report quantitative performance measures for the target UAV with and without the proposed learning signal.","section":"Sections 3.2, 4.2, and 5"}],"minor_comments":[{"comment":"Eq. (21) uses the symbol s_k on both sides of the definition; the training UAV's stored learning signal and the contribution sent to the target UAV should be distinguished notationally.","section":"Section 5, Eq. (21)"},{"comment":"The input list 'F_1, F_2, ..., F_k, F' is ambiguous because k is used both as a running index and as the last index; use explicit maximum index N or a different letter.","section":"Algorithm 1"},{"comment":"There are typographical errors including 'scearios' in Section 2, 'feedforward look' in Section 4, and 'quad-robot' in Section 4; these should be corrected.","section":"Throughout"},{"comment":"The Bode plot figure would benefit from clearer labeling of which subplot corresponds to which input-output channel and from a legend identifying the training versus target UAV consistently across all subplots.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a preliminary conference paper, and the core idea is interesting but the analytic derivations need a careful fix. The missing sum-alpha condition and the notational errors in Eq. (19) are fixable, but the convergence conditions should be made self-contained or precisely reduced to [23]. I would not reject outright because the concept is sound enough to repair; however, the revision must address the load-bearing derivation issues and add quantitative validation before the paper can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is modest but real: it takes the authors' earlier model-based learning framework for identical UAVs and adds scaling matrices Λ_k and dynamics mappings M_k so a target UAV can learn from training UAVs with different reference trajectories and different dynamics. That is a natural ILC extension, not a breakthrough, but it is clearly spelled out and the unification in Algorithm 1 is useful.\n\nThe physical-model motivation is sensible, the LTI assumption for the position loop is standard, and the convergence conditions are at least framed in H-infinity terms, citing their prior ILC design. The simulations show qualitative improvement, and the paper is honest about being preliminary.\n\nThe load-bearing issue is Eq. (11). Moving from the second to the third line, the term T{rd} − Σ Λ_k^{-1} α_k T{Λ_k rd} is dropped. Since T is diagonal LTI and Λ_k are diagonal constants, this equals (1 − Σ α_k) T{rd}. The final expression h = Σ Λ_k^{-1}(F L_k + α_k I){h_k} only holds if Σ α_k = 1 or T{rd} = 0. That condition is never stated, and conditions (12), (20), (23) don't include it. In the simulations there is a single training UAV with α_1 = 1, so the issue is invisible. If Σ α_k ≠ 1, the unlearned target error term survives and the convergence guarantee collapses. This is not a minor typo; it is the central derivation.\n\nThe second issue is Eq. (19): the notation is confused (M_k^{-1} pulled inside a sum over k in a way that doesn't parse), and it relies on F_k ≈ F via ‖G‖ ≫ 1, which is a frequency-dependent approximation that can fail. The numerical studies are purely qualitative: no error metrics, no baselines, no comparison to flying without learning. The convergence analysis is also cited from prior work rather than proved here, which is acceptable for a preliminary study but means the reader has to trust that the conditions carry over.\n\nWho is this for? People working on ILC transfer or multi-UAV learning. It deserves a serious referee because the idea is sound enough to fix and the flaw is specific and addressable. I'd want the authors to state the Σ α_k = 1 condition (or redesign the law), clean up Eq. (19), and add quantitative comparisons.\n\nSend it to review, but expect major revision.","headline":"A plausible ILC extension that stumbles on an unstated filter-sum condition in the main derivation; worth a referee but needs a fix.","tokens_in":8979,"tokens_out":1580,"would_cite":false,"duration_ms":15189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a physical-model-based learning algorithm that lets a target UAV improve its trajectory tracking by learning from other UAVs' flight errors in near real time, even when the UAVs differ in dynamics and reference path…","keywords":["physical model oriented learning","iterative learning control","UAV trajectory tracking","feedforward learning signal","heterogeneous UAVs","learning from other UAVs","near real-time learning"],"falsifier":"Simulate the same tracking problem with two training UAVs and choose weights $\\alpha_1=0.3$, $\\alpha_2=0.4$ so that $\\sum\\alpha_k=0.7\\neq1$. If the paper's convergence claim is exact, the target error should still collapse to the weighted combination; the derivation predicts a leftover $0.3\\,T\\{r_d\\}$ term instead. Comparing the simulated error floor with this predicted residual directly settles whether the dropped term matters.","tokens_in":8015,"feed_emoji":"🚁","tokens_out":7591,"duration_ms":74221,"temperature":0.7,"pith_summary":"This paper proposes a physical-model-based learning algorithm that lets a target UAV improve its trajectory tracking by learning from other UAVs' flight errors in near real time. The algorithm is motivated by iterative learning control but replaces repeated trials on the same task with an analytical mapping: scaling matrices convert reference-path differences and transfer-function matrices convert dynamic differences, turning a training UAV's error into a feedforward learning signal. The central claim is that this signal drives the target's tracking error to a weighted combination of the training UAVs' errors, with convergence guaranteed when a frequency-domain norm condition holds. The authors verify the claim in numerical studies where the target UAV tracks an aggressive trajectory while learning from another UAV flying one sampling step ahead. If the claim is right, it would give UAVs a way to share control experience across different airframes and missions without offline training.","feed_headline":"A drone can borrow another drone's tracking skill in near real time","feed_subtitle":"Physics-based filters turn another UAV's error signal into a feedforward correction, no retraining needed.","key_machinery":"The load-bearing object is the unified learning signal of Eq. (21), $s_k=\\Lambda_k^{-1}M_k^{-1}(\\alpha_k\\{s_k\\}+L_k\\{h_k\\})$. Here $\\Lambda_k$ is a diagonal scaling matrix that maps the target's reference trajectory onto the training UAV's reference, $M_k$ is a transfer-function matrix representing the dynamic mapping from the training UAV to the target UAV, $\\alpha_k$ is a robust filter, and $L_k$ is the learning filter. The learning filter is chosen so that $\\alpha_k I + F L_k \\approx 0$, i.e., $L_k\\approx \\alpha_k F^{-1}$; this identity is what converts another UAV's tracking error into an approximately cancelling correction for the target. The convergence proof rests on the same style of condition used in robust iterative learning control: the weighted composite filter must have infinity norm below $1/N$ for $N$ training UAVs. Inserting the correction in the feedforward loop keeps stability of the target's tracking system unchanged.","core_discovery":"The paper's central claim is that the target UAV's tracking error $h$ can be made to follow a weighted combination of training UAVs' errors $h_k$, rather than being driven by its own unlearned error. The unified learning algorithm uses $s_k=\\Lambda_k^{-1}M_k^{-1}(\\alpha_k\\{s_k\\}+L_k\\{h_k\\})$, where $\\Lambda_k$ scales the target's reference path to the training UAV's scenario and $M_k$ maps the training UAV's open-loop dynamics to the target's. The paper derives $h\\approx \\sum_k \\Lambda_k^{-1}M_k^{-1}(\\alpha_k I+FL_k)\\{h_k\\}$ and designs the learning filters $L_k\\approx \\alpha_k F^{-1}$ so that each term $\\alpha_k I+FL_k$ is small. Convergence is stated as an infinity-norm bound on the composite filters ($\\|[\\cdots]\\|_\\infty<1/N$ in the multi-UAV case), and because the learning signal enters only the feedforward path, it does not alter closed-loop stability. The numerical section shows the target tracking a sharp-turn reference with small overshoot while learning from a training UAV with different dynamics and a different reference.","pith_inferences":["The derivation between lines two and three of Eq. (11) drops the target's own baseline-error term $T\\{r_d\\}$; the final expression is exact only when $\\sum\\alpha_k=1$ or $T\\{r_d\\}=0$. The numerical example sets $\\alpha_1=1$, which sidesteps the issue, but the general convergence claim depends on an unstated condition.","The approximation $\\|G\\|\\gg 1$, used to replace each training UAV's dynamics by the target's, becomes doubtful at frequencies where the position-control loop gain is not large; a Bode-magnitude comparison would show where the learning transfer guarantee should be restricted.","A natural strengthening is to require the robust filters to satisfy $\\sum_k\\alpha_k=1$ (or to keep the residual term as an explicit disturbance), which would close the cancellation gap without changing the algorithm's structure."],"forward_implications":["A fleet of UAVs with different airframes and different reference paths can pool their flight errors, so a new mission can be learned one sampling step before it begins instead of after many offline trials.","Because the learning signal goes into the feedforward path, closed-loop stability is untouched; the only design check is the norm bound on the composite filters.","The learning filter design has an explicit target, $L_k\\approx \\alpha_k F^{-1}$, so the algorithm is implementable whenever a model of the position loop's transfer function is available.","The algorithm inherits iterative learning control's ability to reject repetitive tracking errors, but relaxes the requirement that the same task be repeated by the same system."],"supporting_citations":[{"why":"Previous model-based learning algorithm for UAVs with identical dynamics; this paper generalizes it to heterogeneous scenarios and agents.","marker":"[26]"},{"why":"Provides the differential-flatness transformation and LTI position-loop model used to justify the transfer-function formulation.","marker":"[17]"},{"why":"Motivates applying iterative learning control to quadrotor trajectory tracking, the basis for the learning-filter structure.","marker":"[21]"},{"why":"Supplies the robust ILC filter design method the paper uses to satisfy the infinity-norm convergence condition.","marker":"[23]"}],"fun_headline_variants":["UAVs learn from each other's mistakes in near real time","Physics-based transfer lets drones share tracking skills","Drone learns from another drone's errors via physics filters","Cross-UAV learning from error signals in near real time","Borrowing another UAV's tracking error for faster learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The derivation assumes that the target UAV's own tracking error, $T\\{r_d\\}$, cancels out of the final formula; that cancellation requires the learning weights to add to one, or that error to be zero, and the paper never states this requirement.","fun_headline_variants_meta":{"raw":{"variants":["UAVs learn from each other's mistakes in near real time","Physics-based transfer lets drones share tracking skills","Drone learns from another drone's errors via physics filters","Cross-UAV learning from error signals in near real time","Borrowing another UAV's tracking error for faster learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1299,"prompt_tokens":901,"completion_tokens":398,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":517,"tokens_out":398,"duration_ms":4154,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:11:04.299575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the same tracking problem with two training UAVs and choose weights $\\alpha_1=0.3$, $\\alpha_2=0.4$ so that $\\sum\\alpha_k=0.7\\neq1$. If the paper's convergence claim is exact, the target error should still collapse to the weighted combination; the derivation predicts a leftover $0.3\\,T\\{r_d\\}$ term instead. Comparing the simulated error floor with this predicted residual directly settles whether the dropped term matters.","supporting_citations":[{"cited_title":"A scalable model- based learning algorithm with application to uavs,","cited_arxiv_id":null,"evidence_quote":"Previous model-based learning algorithm for UAVs with identical dynamics; this paper generalizes it to heterogeneous scenarios and agents."},{"cited_title":"Minimum snap trajectory generation and control for quadrotors,","cited_arxiv_id":null,"evidence_quote":"Provides the differential-flatness transformation and LTI position-loop model used to justify the transfer-function formulation."},{"cited_title":"Optimization-based iterative learning for precise quadro- copter trajectory tracking,","cited_arxiv_id":null,"evidence_quote":"Motivates applying iterative learning control to quadrotor trajectory tracking, the basis for the learning-filter structure."},{"cited_title":"Design of arbitrary-order robust iterative learning control based on robust control theory,","cited_arxiv_id":null,"evidence_quote":"Supplies the robust ILC filter design method the paper uses to satisfy the infinity-norm convergence condition."}],"review_version":1}