REVIEW 3 major objections 5 minor 54 references
UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation
T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper claims that a training-free pipeline—periodic open-vocabulary detection, short-segment mask propagation, and lifecycle-aware identity association—can outperform existing open-vocabulary video instance segmentation methods on UAV v
desk verdict Useful new UAV benchmark and training-free baseline, but the headline margins are vulnerable to threshold tuning on the test set and SAM3-generated GT masks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
AeroTrack itself is the central mechanism: an Open-Vocabulary Recognizer runs text-conditioned detection every Δ frames, a Video Instance Segmenter propagates masks within the resulting segment and resets its memory at each boundary, and LIA (Lifecycle-aware ID Association) maps segment-local IDs to video-global IDs via a lifecycle bank of active and lost trajectories, hard geometric gating (scale ratio and center distance), and a matching score combining IoU, center similarity, and area similarity. This three-way split converts a long, dense, open-set video problem into bounded sub-problems: periodic detection covers newly entering/reappearing targets, short-segment propagation keeps memory
What would settle it
Re-annotate a random subset of AeroVIS (e.g., 20 videos) with independent human annotators who do not use SAM3, then recompute overall HOTA for the best AeroTrack variant and for native SAM3; if the 55.39-vs-25.61 gap shrinks substantially, the central comparison depends on the benchmark's SAM3-generated masks.
Extended reading notes
Core claim
The central claim is that UAV open-vocabulary video instance segmentation can be decomposed into three coordinated, training-free stages—periodic open-vocabulary detection on key frames, short-segment mask propagation with internal state resets, and output-level identity association—and that this decomposition is both effective and resource-efficient. On the authors' AeroVIS benchmark, all five AeroTrack variants beat the best transferred general-scene method; the top variant (Grounding DINO for detection, SAM3 for segmentation) reports 55.39 overall HOTA and higher DetA than AssA, indicating that detection coverage of text-specified targets, not identity association for already-found target
Load-bearing premise
The AeroVIS benchmark's ground-truth masks are generated by SAM3, the same model family used inside the best AeroTrack variants, with only manual reinspection and no reported inter-annotator agreement; if those masks contain systematic errors, the reported HOTA margins are inflated.
Editorial extensions
If this is right
- If the reported results hold, a training-free composition of existing visual foundation models is a strong and immediately usable baseline for UAV open-vocabulary video instance segmentation, removing the need for domain-specific annotations.
- Segmented propagation with state resets cuts peak memory growth so that pipelines can handle more than ten times the target budget of whole-video propagation (at least 80 dense targets in the stress test), without sacrificing speed.
- LIA's consistent +27 to +32 HOTA gain shows that identity fragmentation across segment boundaries—not within-segment mask quality—is the main obstacle to global trajectories in this setting.
- AeroVIS, with 9 UAV categories and 8,279 trajectories, provides a reproducible evaluation protocol for future open-vocabulary instance segmentation research on UAV video.
- Generalization results on YouTube-VIS 2019/2021 and LV-VIS indicate that the UAV-specific design does not degrade open-vocabulary capability on ordinary video.
Reading between the lines
- Because AeroVIS ground-truth masks are generated by SAM3 with manual reinspection, and the best AeroTrack variants also use SAM3 as segmenter, part of the performance gap may reflect shared model bias; an independent human-annotated subset would settle this.
- The recognize-segment-associate decomposition is not UAV-specific; the same pattern could apply to other long-video, dense-scene, open-set tasks (surveillance, wildlife monitoring) where per-frame detection is too expensive.
- LIA operates only on output-level boxes, so it is a plug-and-play identity-repair module that could be paired with any segmenter that resets state, independent of the specific detector and segmenter choices.
- The refresh interval Δ is a manually fixed trade-off; adaptively choosing Δ based on scene motion or target density is a natural next step that the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces UAV-OVVIS, a new task for open-vocabulary video instance segmentation in UAV footage, together with AeroTrack, a training-free framework that composes existing visual foundation models. AeroTrack periodically runs an open-vocabulary detector (YOLO-World, Grounding DINO, or SAM3) on key frames, uses a promptable segmenter (SAM2 or SAM3) to propagate masks within short segments, and employs a Lifecycle-aware ID Association (LIA) module to merge segment-level local IDs into globally consistent video identities. The authors construct AeroVIS, a benchmark with 9 categories and 8,279 trajectories derived from UAV tracking datasets, and report that the best AeroTrack variant (Grounding DINO + SAM3) achieves 55.39 Overall HOTA versus 25.61 for Native SAM3 and 34.98 for DEVA. They also evaluate on YouTube-VIS 2019/2021 and LV-VIS and provide memory/speed analyses showing that segmented propagation scales better than running SAM3 over whole videos.
Significance. If the empirical claims hold, the paper makes a useful contribution: it defines a practically relevant task, provides a modular and training-free strong baseline, and constructs the first UAV-specific benchmark for open-vocabulary video instance segmentation. The five variants are internally consistent, the LIA ablation is large, and the memory/speed analysis in Figure 5 is concrete and informative. The method is not a derivational circularity: no parameters are fitted to the reported numbers, and the underlying VFMs are external. The main value, however, rests on the trustworthiness of the AeroVIS benchmark and the fairness of the comparison protocol, and those are currently not established with sufficient rigor.
major comments (3)
- [Implementation Details/Setup; Table 1] The paper states that 'unified inference settings' are used, 'configuring only the single thresholds of the Recognizer and Segmenter at the category level while keeping all other hyperparameters unchanged.' No AeroVIS validation split is described, and no evidence is given that these per-category thresholds were selected without tuning on the evaluation videos. Baselines are evaluated with 'recommended configurations' and do not receive analogous per-category adaptation. Since HOTA is sensitive to detection precision/recall through DetA, this asymmetry could explain a substantial part of the reported margin (55.39 vs 34.98 DEVA and 25.61 SAM3). Please clarify how thresholds were chosen, add a held-out validation protocol, report threshold values and sensitivity, and/or apply the same category-level threshold adaptation to the baselines.
- [Figure 3; Dataset] AeroVIS ground-truth masks are generated by SAM3 from box/identity priors with 'manual reinspection,' while the strongest AeroTrack variants use SAM3 as the Segmenter (Table 1). This is a partial evaluative circularity. No inter-annotator agreement, number of corrected masks, or per-category audit statistics are reported, so it is impossible to quantify whether the benchmark systematically over-rewards SAM3-style mask predictions. The SAM2-based variants also outperform baselines, which mitigates the concern, but the absolute HOTA values and the comparison with Native SAM3 are not fully interpretable without independent annotation-quality analysis. Please report mask-quality statistics and a subset of masks generated/verified by an independent pipeline or annotators.
- [Table 3; Ablation Study] The +27 to +32 HOTA gain from LIA is large, but the 'w/o LIA' condition leaves local IDs unassociated across segments, so under the global-ID HOTA protocol the low baseline is partly by construction. To substantiate the LIA design as opposed to merely having any cross-segment association, please compare LIA with a simpler association method (e.g., pure IoU/center-distance greedy matching without lifecycle decay) and ablate the lifecycle parameters (rho_min, rho_max, gamma, tau_life, T_lost) individually. This is not central to the headline comparison, but it is load-bearing for the paper's component-level claim.
minor comments (5)
- [Table 2] The SAM3 result on LV-VIS is from the test split, while all other results are from the val split. Direct comparison across splits is misleading; use the same split or clearly separate the footnote.
- [Comparison on AeroVIS] The text says 'which we further analyze in §.' with a missing section number. Please fill in the reference.
- [Throughout] There are widespread OCR-like spacing artifacts ('UA V', 'traffic', 'difficult', 'official'). These should be cleaned before publication.
- [Figure 4] The caption lists queries on the same line ('bicycle car road ...'), which is hard to parse. It would be clearer to format each query as a separate item or use a table.
- [Dataset] AeroVIS is described as containing 117 videos and 8,279 trajectories, but no explicit train/validation/test split is given. Since the paper tunes thresholds on the benchmark, a formal split should be defined and released.
Circularity Check
Partial evaluative circularity: AeroVIS reference masks are SAM3-derived while the top AeroTrack variants use SAM3 as Segmenter; no derivational or self-citation circularity.
-
other
[Dataset (AeroVIS construction, Figure 3); compared with Table 1 and Introduction (Segmenter selection)]
"AeroVIS uses frame-level manual box annotations from the source datasets as positional and identity priors to drive SAM3 mask candidate generation, and obtains YTVIS-style instance-level segmentation trajectories through temporal organization, quality filtering, and manual reinspection."
The reference trajectories that define AeroVIS ground truth are produced by SAM3 (with later human reinspection). AeroTrack's strongest variants in Table 1 ('SAM3 + SAM3' 54.43, 'Grounding DINO + SAM3' 55.39) use the same SAM3 model as their Segmenter. Evaluating a SAM3-based segmenter against SAM3-derived masks means the mask-quality component (Mask J, mAP, and, through DetA, HOTA) partly measures agreement with the method's own mask generator rather than with independently annotated UAV instance masks. The effect is partial: SAM2-based variants also beat the baselines, human reinspection adds some independent grounding, and no derivation reduces to fitted values.
full rationale
No derivational circularity is present: AeroTrack is a training-free composition of externally published VFMs, LIA is a hand-specified geometric matcher with no fitted parameters, and none of the paper's equations re-enter as predictions. The load-bearing issue is evaluative. AeroVIS constructs GT from SAM3 mask candidates; the best AeroTrack rows are SAM3-segmenter rows, so part of the HOTA margin (55.39 vs Native SAM3 25.61 and DEVA 34.98) is self-comparison on mask geometry. Manual reinspection is asserted but unquantified (no inter-annotator agreement or audit statistics), so it only partially breaks the circularity. Independent support comes from YTVIS-19/21 and LV-VIS, where GT is human-annotated and the official protocol is used; there AeroTrack is competitive but not uniformly best, confirming the method has content beyond the AeroVIS construction. A separate fairness risk exists in Setup: 'we adopt unified inference settings, configuring only the single thresholds of the Recognizer and Segmenter at the category level while keeping all other hyperparameters unchanged,' while baselines use 'recommended configurations'; without a described validation split this could inflate margins, but the paper never states thresholds were fit to AeroVIS, so it is not a demonstrated by-construction circularity. No load-bearing self-citations or imported uniqueness theorems were found.
Assumptions & free parameters
free parameters (6)
- Key-frame refresh interval Δ =
Δ = 25 frames
- Category-level Recognizer/Segmenter thresholds =
Not reported (one per category)
- LIA gating parameters (ρ_min, ρ_max, γ) =
Not reported
- LIA matching score weights =
0.45 IoU, 0.30 center, 0.25 area
- LIA acceptance threshold τ_life and lost-track expiry T_lost =
Not reported
- Prompt post-processing target budget =
Not reported
assumptions (4)
- domain assumption Existing visual foundation models transfer sufficiently well to UAV imagery (small, dense, top-down, moving-camera objects) without adaptation.
- domain assumption AeroVIS is a valid independent ground truth: box annotations from source tracking datasets give correct identity priors, and SAM3-generated masks survive quality filtering and manual reinspection.
- domain assumption Category names as text prompts are sufficient open-vocabulary queries for fair evaluation.
- domain assumption HOTA and mAP protocols from prior VIS work transfer to the new UAV-OVVIS task.
Cite this review
Pith. "Pith review of UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation." pith.science (2026). https://pith.science/paper/XE2S4YMZ
@misc{pith2026260708075,
author = {Pith},
title = {Pith review of: UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XE2S4YMZ}},
note = {Machine review of arXiv:2607.08075}
}
read the original abstract
Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box-level detection and tracking over predefined categories, making it difficult to jointly support flexible queries and fine-grained instance-level understanding of temporal dynamics in open scenarios. To this end, we introduce a new task, UAV Open-Vocabulary Video Instance Segmentation (UAV-OVVIS), which aims to discover targets in UAV videos according to open-vocabulary queries and output instance segmentation trajectories with globally consistent identities. Considering the scarcity of instance-level annotations in UAV scenarios, we propose AeroTrack, a training-free framework that coordinates existing visual foundation models to realize UAV-OVVIS. AeroTrack performs target discovery and segmentation through periodic open-vocabulary detection and short-segment mask propagation, and introduces Lifecycle-aware ID Association (LIA) to recover global identities under segment-wise inference. Based on this framework, we instantiate five feasible variants and construct AeroVIS, a UAV-OVVIS evaluation benchmark containing 9 UAV object categories and 8,279 trajectories. Experiments show that AeroTrack achieves better overall performance than the evaluated OV-VIS methods transferred to AeroVIS, while demonstrating good open-vocabulary transferability and dense-target handling capability in long UAV videos. The AeroTrack framework and the AeroVIS dataset will be open-sourced upon acceptance.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Yang, Linjie and Fan, Yuchen and Xu, Ning , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[2]
Proceedings of the European Conference on Computer Vision , pages =
Xu, Ning and Yang, Linjie and Fan, Yuchen and Yang, Jianchao and Yue, Dingcheng and Liang, Yuchen and Price, Brian and Cohen, Scott and Huang, Thomas , title =. Proceedings of the European Conference on Computer Vision , pages =
-
[3]
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
Perazzi, Federico and Pont-Tuset, Jordi and McWilliams, Brian and Van Gool, Luc and Gross, Markus and Sorkine-Hornung, Alexander , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =
-
[4]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Wang, Haochen and Yan, Cilin and Wang, Shuai and Jiang, Xiaolong and Tang, Xu and Hu, Yao and Xie, Weidi and Gavves, Efstratios , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[5]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Guo, Pinxue and Huang, Tony and He, Peiyang and Liu, Xuefeng and Xiao, Tianjun and Chen, Zhaoyu and Zhang, Wenqiang , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2025 , doi =
2025
-
[6]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
Wang, Weiyao and Feiszli, Matt and Wang, Heng and Tran, Du and Torresani, Lorenzo , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =
-
[7]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =
Athar, Ali and Luiten, Jonathon and Voigtlaender, Paul and Khurana, Tarasha and Dave, Achal and Leibe, Bastian and Ramanan, Deva , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =
-
[8]
Proceedings of the European Conference on Computer Vision , pages =
Dave, Achal and Khurana, Tarasha and Tokmakov, Pavel and Schmid, Cordelia and Ramanan, Deva , title =. Proceedings of the European Conference on Computer Vision , pages =
Show all 54 references
-
[9]
Qi, Jiyang and Gao, Yan and Hu, Yao and Wang, Xinggang and Liu, Xiaoyu and Bai, Xiang and Belongie, Serge and Yuille, Alan and Torr, Philip H. S. and Bai, Song , title =. International Journal of Computer Vision , year =
-
[10]
2024 , eprint =
Cheng, Zesen and Li, Kehan and Li, Hao and Jin, Peng and Liu, Chang and Zheng, Xiawu and Ji, Rongrong and Chen, Jie , title =. 2024 , eprint =
2024
-
[11]
Proceedings of the European Conference on Computer Vision , year =
Fang, Hao and Wu, Peng and Li, Yawei and Zhang, Xinxin and Lu, Xiankai , title =. Proceedings of the European Conference on Computer Vision , year =
-
[12]
IEEE Transactions on Circuits and Systems for Video Technology , year =
Zhu, Wenqi and Cao, Jiale and Xie, Jin and Yang, Shuangming and Pang, Yanwei , title =. IEEE Transactions on Circuits and Systems for Video Technology , year =
-
[13]
Proceedings of the IEEE/CVF International Conference on Computer Vision , year =
Cheng, Ho Kei and Oh, Seoung Wug and Price, Brian and Schwing, Alexander and Lee, Joon-Young , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , year =
-
[14]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Cheng, Ho Kei and Oh, Seoung Wug and Price, Brian and Lee, Joon-Young and Schwing, Alexander , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[15]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Mei, Haiyang and Zhang, Pengyu and Shou, Mike Zheng , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[16]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Wu, Junfeng and Jiang, Yi and Liu, Qihao and Yuan, Zehuan and Bai, Xiang and Bai, Song , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[17]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =
Xiao, Bin and Peng, Haiping and Weber, Markus and Azizi, Simon and Bower, Trevor and Shi, Shuo and Wei, Grant and Xie, Yinan and Zhang, Ce and Chen, Yan and Hu, Wen and Yang, Lijuan and others , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
-
[18]
2024 , eprint =
Yan, Bin and Sundermeyer, Martin and Tan, David Joseph and Lu, Huchuan and Tombari, Federico , title =. 2024 , eprint =
2024
-
[19]
Proceedings of the 38th International Conference on Machine Learning , pages =
Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya , title =. Proceedings of the 38th International Co...
-
[20]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Cheng, Tianheng and Song, Lin and Ge, Yixiao and Liu, Wenyu and Wang, Xinggang and Shan, Ying , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =. 2024 , doi =
2024
-
[21]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Li, Liunian Harold and Zhang, Pengchuan and Zhang, Haotian and Yang, Jianwei and Li, Chunyuan and Zhong, Yiwu and Wang, Lijuan and Yuan, Lu and Zhang, Lei and Hwang, Jenq-Neng and Chang, Kai-Wei and Gao, Jianfeng , title =. Proceedings of the IEEE/CVF Conference on Computer Vi...
-
[22]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
Liang, Feng and Wu, Bichen and Dai, Xiaoliang and Li, Kunpeng and Zhao, Yinan and Zhang, Hang and Zhang, Peizhao and Vajda, Peter and Marculescu, Diana , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =
-
[23]
Detecting Twenty-thousand Classes Using Image-level Supervision , booktitle =
Zhou, Xingyi and Girdhar, Rohit and Joulin, Armand and Kr. Detecting Twenty-thousand Classes Using Image-level Supervision , booktitle =
-
[24]
Proceedings of the European Conference on Computer Vision , pages =
Minderer, Matthias and Gritsenko, Alexey and Stone, Austin and Neumann, Maxim and Weissenborn, Dirk and Dosovitskiy, Alexey and Mahendran, Aravindh and Arnab, Anurag and Dehghani, Mostafa and Shen, Zhuoran and Wang, Xiao and Zhai, Xiaohua and Kipf, Thomas and Houlsby, Neil , t...
-
[25]
Proceedings of the European Conference on Computer Vision , year =
Liu, Shilong and Zeng, Zhaoyang and Ren, Tianhe and Li, Feng and Zhang, Hao and Yang, Jie and Jiang, Qing and Li, Chunyuan and Yang, Jianwei and Su, Hang and Zhu, Jun and Zhang, Lei , title =. Proceedings of the European Conference on Computer Vision , year =
-
[26]
Emerging Properties in Self-Supervised Vision Transformers , booktitle =
Caron, Mathilde and Touvron, Hugo and Misra, Ishan and J. Emerging Properties in Self-Supervised Vision Transformers , booktitle =
-
[27]
and Lo, Wan-Yen and Dollar, Piotr and Girshick, Ross , title =
Kirillov, Alexander and Mintun, Eric and Ravi, Nikhila and Mao, Hanzi and Rolland, Chloe and Gustafson, Laura and Xiao, Tete and Whitehead, Spencer and Berg, Alexander C. and Lo, Wan-Yen and Dollar, Piotr and Girshick, Ross , title =. Proceedings of the IEEE/CVF International ...
2023
-
[28]
Advances in Neural Information Processing Systems , year =
Zou, Xueyan and Yang, Jianwei and Zhang, Hao and Li, Feng and Li, Linjie and Wang, Jianfeng and Wang, Lijuan and Gao, Jianfeng and Lee, Yong Jae , title =. Advances in Neural Information Processing Systems , year =
-
[29]
arXiv preprint arXiv:2401.14159 , year =
Ren, Tianhe and Liu, Shilong and Zeng, Ailing and Lin, Jing and Li, Kunchang and Cao, He and Chen, Jiayu and Huang, Xinyu and Chen, Yukang and Yan, Feng and Zeng, Zhaoyang and Zhang, Hao and Li, Feng and Yang, Jie and Li, Hongyang and Jiang, Qing and Zhang, Lei , title =. arXi...
-
[30]
Advances in Neural Information Processing Systems , volume =
Yang, Zongxin and Wei, Yunchao and Yang, Yi , title =. Advances in Neural Information Processing Systems , volume =
-
[31]
, title =
Cheng, Ho Kei and Schwing, Alexander G. , title =. Proceedings of the European Conference on Computer Vision , pages =
-
[32]
Proceedings of the International Conference on Learning Representations , year =
Ravi, Nikhila and Gabeur, Valentin and Hu, Yuan-Ting and Hu, Ronghang and Ryali, Chaitanya and Ma, Tengyu and Khedr, Haitham and Radle, Roman and Rolland, Chloe and Gustafson, Laura and Mintun, Eric and Pan, Junting and Alwala, Kalyan Vasudev and Carion, Nicolas and Wu, Chao-Y...
-
[33]
2025 , eprint =
Carion, Nicolas and Gustafson, Laura and Hu, Yuan-Ting and Debnath, Shoubhik and Hu, Ronghang and Suris, Didac and Ryali, Chaitanya and Alwala, Kalyan Vasudev and Khedr, Haitham and Huang, Andrew and Lei, Jie and Ma, Tengyu and others , title =. 2025 , eprint =
2025
-
[34]
International Journal of Computer Vision , volume =
Luiten, Jonathon and Osep, Aljosa and Dendorfer, Patrick and Torr, Philip and Geiger, Andreas and Leal-Taixe, Laura and Leibe, Bastian , title =. International Journal of Computer Vision , volume =. 2021 , doi =
2021
-
[35]
2018 , eprint =
Zhu, Pengfei and Wen, Longyin and Bian, Xiao and Ling, Haibin and Hu, Qinghua , title =. 2018 , eprint =
2018
-
[36]
2018 , eprint =
Du, Dawei and Qi, Yuankai and Yu, Hongyang and Yang, Yifan and Duan, Kaiwen and Li, Guorong and Zhang, Weigang and Huang, Qingming and Tian, Qi , title =. 2018 , eprint =
2018
-
[37]
2021 , eprint =
Wen, Longyin and Du, Dawei and Zhu, Pengfei and Hu, Qinghua and Wang, Qilong and Bo, Liefeng and Lyu, Siwei , title =. 2021 , eprint =
2021
-
[38]
2022 , eprint =
Zhang, Chunhui and Huang, Guanjie and Liu, Li and Huang, Shan and Yang, Yinan and Wan, Xiang and Ge, Shiming and Tao, Dacheng , title =. 2022 , eprint =
2022
-
[39]
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =
Varga, Leon Amadeus and Kiefer, Benjamin and Messmer, Martin and Zell, Andreas , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2022 , doi =
2022
-
[40]
IEEE Transactions on Geoscience and Remote Sensing , volume =
Liu, Fan and Chen, Delong and Guan, Zhangqingyun and Zhou, Xiaocong and Zhu, Jiale and Ye, Qiaolin and Fu, Liyong and Zhou, Jun , title =. IEEE Transactions on Geoscience and Remote Sensing , volume =. 2024 , doi =
2024
-
[41]
2024 , eprint =
Li, Kaiyu and Liu, Ruixun and Cao, Xiangyong and Bai, Xueru and Zhou, Feng and Meng, Deyu and Wang, Zhi , title =. 2024 , eprint =
2024
-
[42]
2025 , eprint =
Dutta, Saikat and Vasim, Akhil and Gole, Siddhant and Rezatofighi, Hamid and Banerjee, Biplab , title =. 2025 , eprint =
2025
-
[43]
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages =
Zamir, Syed Waqas and Arora, Aditya and Gupta, Akshita and Khan, Salman and Khan, Fahad Shahbaz and Zhu, Fan and Shao, Ling and Xia, Gui-Song and Bai, Xiang , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages =
-
[44]
2018 , eprint =
Lyu, Ye and Vosselman, George and Xia, Guisong and Yilmaz, Alper and Yang, Michael Ying , title =. 2018 , eprint =
2018
-
[45]
2023 , eprint =
Cai, Wenxiao and Jin, Ke and Hou, Jinyan and Guo, Cong and Wu, Letian and Yang, Wankou , title =. 2023 , eprint =
2023
-
[46]
Sensors , volume =
Huang, Ying and Zhang, Yinhui and He, Zifen and Deng, Yunnan , title =. Sensors , volume =. 2025 , doi =
2025
-
[47]
2024 , eprint =
Chen, Yu-Hsi , title =. 2024 , eprint =
2024
-
[48]
2025 , eprint =
Sirma, Aykut and Plastropoulos, Angelos and Tang, Gilbert and Zolotas, Argyrios , title =. 2025 , eprint =
2025
-
[49]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Li, Kaiyu and Cao, Xiangyong and Deng, Yupeng and Pang, Chao and Xin, Zepeng and Qiao, Hui and Gong, Tieliang and Meng, Deyu and Wang, Zhi , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =
2026
-
[50]
2026 , eprint =
Dou, Mingyu and Qiu, Shi and Hu, Ming and Chen, Yifan and Ye, Huping and Liao, Xiaohan and Sun, Zhe , title =. 2026 , eprint =
2026
-
[51]
IEEE International Conference on Robotics and Automation (ICRA) , year =
Blei, Yannik and Krawez, Michael and Nilavadi, Nisarga and Kaiser, Tanja Katharina and Burgard, Wolfram , title =. IEEE International Conference on Robotics and Automation (ICRA) , year =
-
[52]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Ma, Jianbo and Luo, Hui and Chen, Qi and Qi, Yuankai and Sun, Yumei and Beheshti, Amin and Zhang, Jianlin and Yang, Ming-Hsuan , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =
2026
-
[53]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Chen, Chenglizhao and Liang, Shaofeng and Guan, Runwei and Sun, Xiaolou and Zhao, Haocheng and Jiang, Haiyun and Huang, Tao and Ding, Henghui and Han, Qing-Long , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =
2026
-
[54]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Li, Chen and Zhao, Rui and Wang, Zeyu and Xu, Huiying and Zhu, Xinzhong , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2025 , doi =
2025
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.