REVIEW 2 major objections 1 minor 43 references
SEArch: Optimistic Policy Selection Between Scene Noise and Drift for UAV Radar Search
T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read An online selector over a library of radar detectors achieves regret that scales with noise level times square root of time plus square root of scene transitions, without oracle knowledge of when scenes change.
desk verdict SEArch gives a clean extension of optimistic FTRL to mixed stochastic-adversarial regret for detector switching, but the UAV radar setting may not supply the per-round loss observations the analysis needs. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Stochastically Extended Adversary (SEA) model that couples stochastic noise inside each scene with adversarial shifts across scenes; SEArch performs policy selection with optimistic Follow-the-Regularized-Leader using an adaptive learning rate.
What would settle it
Collect regret traces from SEArch and W-SEArch on a sequence of radar scenes whose transition count J and per-scene noise levels are known in advance, then check whether the measured cumulative regret stays inside the envelope O(σ̄_T √T + √J) or deviates systematically.
Extended reading notes
Core claim
SEArch instantiates the SEA framework via an optimistic Follow-the-Regularized-Leader selector equipped with an adaptive learning rate and obtains regret O(σ̄_T √T + √J), where σ̄_T captures radar measurement noise and J counts scene transitions over horizon T; the windowed W-SEArch variant that restarts every w rounds obtains O(σ̄_I √w) regret provided there is at most one transition per window.
Load-bearing premise
A fixed library of specialized detectors exists whose per-scene performance can be observed or estimated in real time by the resource-limited UAV.
Editorial extensions
If this is right
- Regret grows only with the square root of the number of scene changes rather than linearly, so longer missions remain feasible.
- W-SEArch keeps regret controlled even under rapid scene changes by resetting every w steps.
- No external forecast of scene boundaries is required; the algorithm adapts using only observed losses.
- Experiments report up to 30 percent lower regret than non-adaptive baselines across varied non-stationary radar settings.
Reading between the lines
- The same selector structure could be applied to other onboard perception tasks where several algorithms compete under shifting conditions, such as camera-based object detection.
- Real-time estimation of local noise level σ̄ becomes the practical bottleneck once the theoretical bound is accepted.
- Dynamically choosing window size w according to recently observed transition frequency could tighten the regret further.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formulates UAV radar target search as an online policy selection problem over a fixed library of detectors. It introduces the Stochastically Extended Adversary (SEA) framework to jointly handle intra-scene stochastic noise and inter-scene adversarial shifts without oracle knowledge of dynamics, instantiates it via the lightweight optimistic FTRL algorithm SEArch (with adaptive rate) that achieves regret O(¯σ_T √T + √J), and proposes the windowed W-SEArch variant achieving O(¯σ_I √w) under at most one transition per window. Experiments are reported to yield up to 30% regret reduction versus non-adaptive baselines.
Significance. If the regret analysis holds under the paper's feedback model, the SEA framework's explicit treatment of mixed stochastic-adversarial non-stationarity, together with the resource-light online selector, would constitute a useful contribution to adaptive perception on constrained platforms. The explicit regret decomposition separating measurement noise from scene transitions is a clear strength, as is the claim of no oracle scene dynamics. The reported experimental gains, if supported by full details, would strengthen the practical case.
major comments (2)
- [Abstract] Abstract: the SEA framework and SEArch regret bound O(¯σ_T √T + √J) are derived under the assumption that the loss of the selected detector is observed (or estimable) after each round to enable the OFTRL update. The target UAV radar application (respiration micro-motion through occlusions) provides no such immediate ground-truth feedback, so the update rule and the derived bound do not apply as stated; this assumption is load-bearing for the central claim that adaptation works without oracle knowledge.
- [Abstract] Abstract: the phrasing 'without requiring oracle knowledge of scene dynamics' addresses only the inter-scene transitions J but leaves unaddressed the per-round loss observability required by the optimistic FTRL selector; the two are distinct and the latter is necessary for the regret guarantee to transfer to the UAV setting.
minor comments (1)
- The abstract states 'Experiments show up to 30% regret reduction' but supplies no information on the number of trials, scene-transition models, detector library size, or how losses were computed in simulation; these details are needed for reproducibility even if the main text contains them.
Simulated Author's Rebuttal
We thank the referee for the careful reading and for highlighting the feedback model assumptions. We agree that the per-round loss observability required by the OFTRL update is a distinct and load-bearing assumption that was not sufficiently distinguished from the lack of oracle scene dynamics in the abstract and introduction. We will revise the manuscript to clarify this point explicitly.
read point-by-point responses
-
Referee: [Abstract] Abstract: the SEA framework and SEArch regret bound O(¯σ_T √T + √J) are derived under the assumption that the loss of the selected detector is observed (or estimable) after each round to enable the OFTRL update. The target UAV radar application (respiration micro-motion through occlusions) provides no such immediate ground-truth feedback, so the update rule and the derived bound do not apply as stated; this assumption is load-bearing for the central claim that adaptation works without oracle knowledge.
Authors: We acknowledge that the regret analysis is derived under the standard full-information feedback model of online learning, in which the loss of the played action is observed after each round. The UAV radar application indeed lacks immediate ground-truth labels. We will revise the abstract, introduction, and a new dedicated paragraph in Section 2 to state the feedback assumption explicitly, to separate it from the 'no oracle scene dynamics' claim, and to discuss practical loss estimation via proxy radar statistics or delayed feedback. The SEA framework itself remains valid under the stated feedback model; the revision will make the scope of the guarantees transparent. revision: yes
-
Referee: [Abstract] Abstract: the phrasing 'without requiring oracle knowledge of scene dynamics' addresses only the inter-scene transitions J but leaves unaddressed the per-round loss observability required by the optimistic FTRL selector; the two are distinct and the latter is necessary for the regret guarantee to transfer to the UAV setting.
Authors: We agree that the current phrasing conflates the two issues. The sentence will be rewritten to read: 'without requiring oracle knowledge of scene transition times or dynamics, under the assumption that per-round losses are observed or estimable.' Corresponding clarifications will appear in the model section and the abstract. revision: yes
Circularity Check
No significant circularity; derivation is self-contained
full rationale
The paper introduces the SEA framework to model the coupled stochastic-adversarial setting and instantiates it via SEArch (OFTRL with adaptive rate) to obtain the stated regret bound. The abstract and provided text present this as a modeling choice followed by standard online-learning analysis, without any quoted reduction of the bound to fitted parameters, self-referential definitions, or load-bearing self-citations that collapse the claim. The central result is a theoretical guarantee under the model's observability assumptions rather than a tautology. No steps meet the criteria for circularity.
Assumptions & free parameters
assumptions (1)
- domain assumption A fixed library of specialized detectors exists with observable per-scene performance
invented entities (1)
-
Stochastically Extended Adversary (SEA) framework
Cite this review
Pith. "Pith review of SEArch: Optimistic Policy Selection Between Scene Noise and Drift for UAV Radar Search." pith.science (2026). https://pith.science/paper/F3FWGIVS
@misc{pith2026260601325,
author = {Pith},
title = {Pith review of: SEArch: Optimistic Policy Selection Between Scene Noise and Drift for UAV Radar Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/F3FWGIVS}},
note = {Machine review of arXiv:2606.01325}
}
abstract
Unmanned Aerial Vehicles (UAVs) equipped with radar sensors are deployed for target search missions in diverse environments, where targets exhibit characteristic signatures (e.g., respiration micro-motion in human search) detectable through occlusions. A fundamental challenge arises from shifts in radar statistics as the UAV moves through a dynamic and potentially non-stationary environment, rendering any fixed signal-processing strategy suboptimal; yet perception and adaptation must run onboard a resource-constrained aerial node in real time. Since no single detector performs well across all conditions, we adopt a multi-policy paradigm and formulate UAV target search as an online policy selection problem over a library of specialized detectors, with performance measured by regret, the cumulative loss gap relative to the best policy in each scene. The setting couples in-scene stochastic noise with inter-scene shifts. Whereas prior methods capture only one regime, we account for both through the Stochastically Extended Adversary (SEA) framework, without requiring oracle knowledge of scene dynamics. Because adaptation must run at the UAV, we instantiate SEA through \textsc{SEArch}, a lightweight optimistic Follow the Regularized Leader (OFTRL) selector with an adaptive learning rate, achieving regret $O(\bar{\sigma}_T \sqrt{T} + \sqrt{J})$, where $\bar{\sigma}_T$ captures radar measurement noise and $J$ is the number of scene transitions over the mission horizon $T$. To enable rapid adaptation under frequent scene changes, we further introduce \textsc{W-SEArch}, a windowed variant that restarts every $w$ rounds and achieves regret $O(\bar{\sigma}_I \sqrt{w})$ under at most one transition per window. Experiments show up to 30\% regret reduction compared to non-adaptive baselines across a range of non-stationary settings.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Through-the-wall human respiration detection using UWB impulse radar on hovering drone,
B. P. A. Rohman, M. B. Andra, and M. Nishimoto, “Through-the-wall human respiration detection using UWB impulse radar on hovering drone,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 6572–6584, 2021. (a) Single-regime dominance (J= 0). (b) Rapid regime switching (J= 3). (c) Gradual regime drift (J= 2). Fig....
2021
-
[2]
Respiration detection of ground injured human target using UWB radar mounted on a hovering UA V,
X. Jing, X. Zhu, Z. Li, F. Luo, and B. Peng, “Respiration detection of ground injured human target using UWB radar mounted on a hovering UA V,”Drones, vol. 6, no. 9, p. 235, 2022
2022
-
[3]
Multi-target path planning with probabilistic detection in cluttered environments,
N. Khial, N. Mhaisen, L. Ismail, M. Mabrok, and A. Mohamed, “Multi-target path planning with probabilistic detection in cluttered environments,” inICC 2025-IEEE International Conference on Com- munications. IEEE, 2025, pp. 1298–1303
2025
-
[4]
Pdsr: Efficient uav deployment for swift and accurate post- disaster search and rescue,
A. A. Abdellatif, A. Elmancy, A. Mohamed, A. Massoud, W. Lebda, and K. K. Naji, “Pdsr: Efficient uav deployment for swift and accurate post- disaster search and rescue,”IEEE Internet of Things Magazine, vol. 8, no. 3, pp. 149–156, 2025
2025
-
[5]
DRONE-RL: Dynamic reinforcement learning for on- line navigation of UA Vs in evolving environments,
N. Khial, M. S. Allahham, N. Mhaisen, L. Ismail, M. Mabrok, and A. Mohamed, “DRONE-RL: Dynamic reinforcement learning for on- line navigation of UA Vs in evolving environments,”Knowledge-Based Systems, p. 115147, 2025
2025
-
[6]
Hazan,Introduction to Online Convex Optimization
E. Hazan,Introduction to Online Convex Optimization. MIT Press, 2016
2016
-
[7]
An online learning framework for UA V search mission in adversarial environ- ments,
N. Khial, N. Mhaisen, M. A. Mabrok, and A. Mohamed, “An online learning framework for UA V search mission in adversarial environ- ments,”Expert Systems with Applications, vol. 267, p. 126136, 2025
2025
-
[8]
Online trajectory optimization using inexact gradient feedback for time-varying environments,
M. K. Nutalapati, A. S. Bedi, K. Rajawat, and M. Coupechoux, “Online trajectory optimization using inexact gradient feedback for time-varying environments,”IEEE Transactions on Signal Processing, vol. 68, pp. 4824–4838, 2020
2020
Show all 43 references
-
[9]
Between stochastic and adversarial online convex optimization: Improved regret bounds via smoothness,
S. Sachs, H. Hadiji, T. van Erven, and C. Guzm ´an, “Between stochastic and adversarial online convex optimization: Improved regret bounds via smoothness,” inAdvances in Neural Information Processing Systems (NeurIPS), vol. 35, 2022, pp. 691–702
2022
-
[10]
Online learning with predictable se- quences,
A. Rakhlin and K. Sridharan, “Online learning with predictable se- quences,” inConference on Learning Theory (COLT), ser. PMLR, vol. 30, 2013, pp. 993–1019
2013
-
[11]
On the dynamic regret of following the regularized leader: Optimism with history pruning,
N. Mhaisen and G. Iosifidis, “On the dynamic regret of following the regularized leader: Optimism with history pruning,” inProceedings of the International Conference on Machine Learning (ICML), ser. PMLR, vol. 267, 2025, pp. 43 990–44 016
2025
-
[12]
Understanding adam optimizer via online learning of updates: Adam is ftrl in disguise,
K. Ahn, Z. Zhang, Y . Kook, and Y . Dai, “Understanding adam optimizer via online learning of updates: Adam is ftrl in disguise,” inInternational Conference on Machine Learning, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267406311
2024
-
[13]
Reinforcement learning framework for UA V-based target localization applications,
M. Shurrab, R. Mizouni, S. Singh, and H. Otrok, “Reinforcement learning framework for UA V-based target localization applications,” Internet of Things, vol. 23, p. 100867, 2023. (a) Single transition (J= 1). (b) Five transitions (J= 5). (c) Five transitions with inconsistent d...
2023
-
[14]
A predictive target tracking framework for IoT using CNN–LSTM,
L. A. Hussain, S. Singh, R. Mizouni, H. Otrok, and E. Damiani, “A predictive target tracking framework for IoT using CNN–LSTM,” Internet of Things, vol. 22, p. 100744, 2023
2023
-
[15]
Human body detection and geolocalization for UA V search and rescue missions using color and thermal imagery,
P. Rudol and P. Doherty, “Human body detection and geolocalization for UA V search and rescue missions using color and thermal imagery,” inIEEE Aerospace Conference, 2008, pp. 1–8
2008
-
[16]
Monocular 3D pose estimation and tracking by detection,
M. Andriluka, S. Roth, and B. Schiele, “Monocular 3D pose estimation and tracking by detection,” inIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2010, pp. 623–630
2010
-
[17]
Obstacle-aware human localization via channel impulse response,
N. Qassmi, M. Hilou, M. S. Allahham, A. Mohamed, and L. Ismail, “Obstacle-aware human localization via channel impulse response,” in 2026 IEEE 23rd Consumer Communications & Networking Conference (CCNC). IEEE, 2026, pp. 1–6
2026
-
[18]
Intelligent uav swarm cooperation for multiple targets tracking,
L. Zhou, S. Leng, Q. Liu, and Q. Wang, “Intelligent uav swarm cooperation for multiple targets tracking,”IEEE Internet of Things Journal, vol. 9, no. 1, pp. 743–754, 2021
2021
-
[19]
A deep learning framework for target localization in error-prone environment,
S. K. Mohammed, S. Singh, R. Mizouni, and H. Otrok, “A deep learning framework for target localization in error-prone environment,”Internet of Things, vol. 22, p. 100713, 2023
2023
-
[20]
Survey on coverage path planning with unmanned aerial vehicles,
T. M. Cabreira, L. B. de Brisolara, and P. R. Ferreira Jr., “Survey on coverage path planning with unmanned aerial vehicles,”Drones, vol. 3, no. 1, p. 4, 2019
2019
-
[21]
Efficient path planning for UA V formation via comprehensively improved particle swarm optimization,
S. Shao, Y . Peng, C. He, and Y . Du, “Efficient path planning for UA V formation via comprehensively improved particle swarm optimization,” ISA Transactions, vol. 97, pp. 415–430, 2020
2020
-
[22]
Distributed robotic sensor networks: An information-theoretic approach,
B. J. Julian, M. Angermann, M. Schwager, and D. Rus, “Distributed robotic sensor networks: An information-theoretic approach,”Interna- tional Journal of Robotics Research, vol. 31, no. 10, pp. 1134–1154, 2012
2012
-
[23]
Uav path planning in a dynamic environment via partially observable markov decision process,
S. Ragi and E. K. Chong, “Uav path planning in a dynamic environment via partially observable markov decision process,”IEEE Transactions on Aerospace and Electronic Systems, vol. 49, no. 4, pp. 2397–2412, 2013
2013
-
[24]
Machine learning-aided operations and communications of unmanned aerial vehicles: A contemporary survey,
H. Kurunathan, H. Huang, K. Li, W. Ni, and E. Hossain, “Machine learning-aided operations and communications of unmanned aerial vehicles: A contemporary survey,”IEEE Communications Surveys & Tutorials, vol. 26, no. 1, pp. 496–533, 2023
2023
-
[25]
Consensus-based decentralized auctions for robust task allocation,
H.-L. Choi, L. Brunet, and J. P. How, “Consensus-based decentralized auctions for robust task allocation,”IEEE Transactions on Robotics, vol. 25, no. 4, pp. 912–926, 2009
2009
-
[26]
Distributed multi-robot coordi- nation in area exploration,
W. Sheng, Q. Yang, J. Tan, and N. Xi, “Distributed multi-robot coordi- nation in area exploration,”Robotics and Autonomous Systems, vol. 54, no. 12, pp. 945–955, 2006
2006
-
[27]
Near-optimal sensor placements in Gaussian processes: Theory, efficient algorithms and empirical stud- ies,
A. Krause, A. Singh, and C. Guestrin, “Near-optimal sensor placements in Gaussian processes: Theory, efficient algorithms and empirical stud- ies,”Journal of Machine Learning Research, vol. 9, pp. 235–284, 2008
2008
-
[28]
Reinforcement learning for agile active target sensing with a UA V,
H. Goel, L. Jarin Lipschitz, S. Agarwal, S. Manjanna, and V . Kumar, “Reinforcement learning for agile active target sensing with a UA V,” 2022
2022
-
[29]
Bandit submodular maximization for multi-robot coordination in unpredictable and partially observable environments,
Z. Xu, X. Lin, and V . Tzoumas, “Bandit submodular maximization for multi-robot coordination in unpredictable and partially observable environments,” 2023
2023
-
[30]
Online caching with optimistic learning,
N. Mhaisen, G. Iosifidis, and D. Leith, “Online caching with optimistic learning,” in2022 IFIP Networking Conference (IFIP Networking). IEEE, 2022, pp. 1–9
2022
-
[31]
A modern introduction to online learning,
F. Orabona, “A modern introduction to online learning,” 2019
2019
-
[32]
Lattimore and C
T. Lattimore and C. Szepesv ´ari,Bandit Algorithms. Cambridge University Press, 2020
2020
-
[33]
Online learning: A compre- hensive survey,
S. C. Hoi, D. Sahoo, J. Lu, and P. Zhao, “Online learning: A compre- hensive survey,”Neurocomputing, vol. 459, pp. 249–289, 2021
2021
-
[34]
Parameter-free algo- rithms for the stochastically extended adversarial model,
S. Wang, A. Barik, P. Zhao, and V . Y . Tan, “Parameter-free algo- rithms for the stochastically extended adversarial model,”arXiv preprint arXiv:2510.04685, 2025
2025
-
[35]
A second-order bound with excess losses,
P. Gaillard, G. Stoltz, and T. V . Erven, “A second-order bound with excess losses,” inConference on Learning Theory (COLT), 2014, pp. 176–196
2014
-
[36]
Achieving all with no parameters: AdaNor- malHedge,
H. Luo and R. E. Schapire, “Achieving all with no parameters: AdaNor- malHedge,” inConference on Learning Theory (COLT), 2015, pp. 1286– 1304
2015
-
[37]
Online optimization: Competing with dynamic comparators,
A. Jadbabaie, A. Rakhlin, S. Shahrampour, and K. Sridharan, “Online optimization: Competing with dynamic comparators,” inInternational Conference on Artificial Intelligence and Statistics (AISTATS), 2015, pp. 398–406
2015
-
[38]
Online meta- learning,
C. Finn, A. Rajeswaran, S. Kakade, and S. Levine, “Online meta- learning,” inInternational Conference on Machine Learning (ICML), 2019, pp. 1920–1930
2019
-
[39]
Continuous adaptation via meta-learning in nonstationary and competitive environments,
M. Al-Shedivat, T. Bansal, Y . Burda, I. Sutskever, I. Mordatch, and P. Abbeel, “Continuous adaptation via meta-learning in nonstationary and competitive environments,” 2017
2017
-
[40]
Adaptive online learning in dynamic environments,
L. Zhang, S. Lu, and Z.-H. Zhou, “Adaptive online learning in dynamic environments,”Advances in neural information processing systems, vol. 31, 2018
2018
-
[41]
Partially lazy gradient descent for smoothed online learning,
N. Mhaisen and G. Iosifidis, “Partially lazy gradient descent for smoothed online learning,” inThe 29th International Conference on Artificial Intelligence and Statistics, 2026. [Online]. Available: https://openreview.net/forum?id=hIAkL2BYCI
2026
-
[42]
Finite-time analysis of the multiarmed bandit problem,
P. Auer, N. Cesa-Bianchi, and P. Fischer, “Finite-time analysis of the multiarmed bandit problem,”Machine Learning, vol. 47, no. 2, pp. 235– 256, 2002
2002
-
[43]
The non- stochastic multiarmed bandit problem,
P. Auer, N. Cesa-Bianchi, Y . Freund, and R. E. Schapire, “The non- stochastic multiarmed bandit problem,”SIAM Journal on Computing, vol. 32, no. 1, pp. 48–77, 2002
2002
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.