Pith. sign in

REVIEW 3 major objections 5 minor 63 references

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that long-horizon UAV vision-language navigation should be treated as one coupled policy-state construction problem, and that its semantic-to-decision pipeline achieves top results on the AerialVLN and OpenFly benchmarks.

desk verdict Competent systems integration for UAV-VLN with a plausible but statistically unproven SOTA claim; the ablations are the most convincing part. read the letter →

arxiv 2608.09564 v1 pith:VZBENNHM submitted 2026-08-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords UAVvision-languagenavigationsemanticgroundingdynamictemporalaggregationlocal-optimumcognitionGRPOAerialVLNOpenFlylong-horizon
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the failures of long-horizon UAV vision-language navigation, namely missed landmarks, noisy memory, loops, and false stops, are not three separate problems but one problem of how the policy state is built. It proposes a pipeline that grounds the current view in object-level semantics and relative positions, reweights the full history buffer while adding a few structured landmark prompts, and then applies topology-aware stagnation detection plus reinforcement learning under a composite reward. The intended payoff is an aerial agent that follows natural-language instructions over long routes more reliably in simulation, with lower navigation error and more successful stops. The authors report state-of-the-art results on AerialVLN and OpenFly, with cumulative ablations showing each stage contributes monotonically to the final performance.

What carries the argument

The central mechanism is a decoder context assembled from four streams: the grounded current-view token, the dynamic-temporal-aggregation (DTA) filtered history, the sparse landmark-prompt memory, and the conditional local-optimum-cognition (LOC) frontier cue. This context is what does the work: every module is designed to make the state semantically faithful to the instruction and topologically robust over long horizons. DTA computes instruction-conditioned relevance weights over the full history buffer and keeps a sparse grounding side branch that serializes the top-$K$ frames into a structured prompt; LOC monitors the topology graph for stagnation over a $\Delta$-step window and selects a frontier node balancing semantic agreement and graph distance; and group-relative policy optimization (GRPO) refines the policy with group-normalized advantages and a KL anchor to a behavior-cloned reference.

What would settle it

Run the same evaluation protocol with ten or more seeds for both this method and Fly0 on the AerialVLN-S unseen split and compute a paired significance test under identical conditions; if the success-rate gap of 61.38 versus 60.07 or the navigation-error gap of 50.69 versus 51.23 falls within the standard deviation of either method, the claimed state-of-the-art ordering is not established. A simpler check is to report per-seed unseen-split success rates and see whether any single seed of the proposed method falls below Fly0's reported mean.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that a UAV navigating by language over long horizons is best modeled as a policy-state construction problem rather than as separate perception, memory, and planning subproblems. The framework's instruction-grounded semantic enhancement injects object-level semantics and relative spatial cues into the current observation; dynamic temporal aggregation uses an instruction-conditioned relevance weight to combine the full history while converting up to two high-relevance frames into structured landmark prompts; and local-optimum cognition detects when the topology graph has not expanded and conditionally injects a semantically aligned frontier cue. Training then uses group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. On AerialVLN-S validation and the OpenFly test set, the authors report the best navigation error, success rate, and oracle success rate among compared methods, including improving unseen-split success rate from 60.07 to 61.38 and OpenFly success rate from 64.67 to 67.31 relative to the strongest prior baseline.

Load-bearing premise

The state-of-the-art claim rests on the assumption that three training seeds, averaged without standard deviations, separate this method from baselines whose numbers come from separate papers; on the unseen AerialVLN split the margins over Fly0 are only about 1.3 success-rate points and 0.54 meters of navigation error, so if run-to-run variation is comparable to a few success-rate points, the reported ordering could change.

Editorial extensions

If this is right

  • If the claim is right, long-horizon UAV-VLN should be treated as a single state-construction problem; splitting it into perception, memory, and planning subproblems misses the coupling that causes failures.
  • Instructions can steer history selection: relevance-aware weighting plus sparse landmark prompts should allow the agent to recover delayed landmark cues that global image features overlook.
  • Topology-aware stagnation cues should reduce local loops and false stops, improving success rate and oracle success rate without retraining the planner.
  • The composite reward of progress, goal completion, semantic matching, and path compliance should train more stably than any single reward, as the ablations show each term adds an independent gain.
  • The method keeps competitive performance under simulated image noise and localization drift, suggesting the coupled policy state is reasonably stable under moderate perception corruption.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the same state-construction recipe, grounded current view, weighted history, sparse landmark prompts, and stagnation-aware frontier cue, is not UAV-specific and could transfer to other sparse-landmark navigation domains such as ground or underwater VLN; that is a testable hypothesis the paper does not run.
  • Beyond the paper: because the gains over the strongest baseline are small on the unseen AerialVLN split, a shared-codebase comparison with matched seeds and a paired significance test would settle whether the reported margin is real; the paper's own caveat about non-paired baselines points to the same check.
  • Beyond the paper: the reported robustness to noise and drift is demonstrated in simulation under additive corruption; extrapolating to real closed-loop flight would require testing under sensor-coupled perception and localization failure, which the authors list as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a unified pipeline for long-horizon UAV vision-language navigation, coupling instruction-grounded semantic enhancement of the current observation, relevance-aware dynamic temporal aggregation (DTA) with a sparse landmark-prompt branch, local-optimum cognition (LOC) for topology-aware recovery, and GRPO-based policy refinement under a composite trajectory reward. The method is evaluated on the AerialVLN-S validation splits and the OpenFly test set, reporting state-of-the-art numbers under NE, SR, and OSR. The paper also provides cumulative ablations, component ablations, and robustness checks under image noise and localization drift, and it releases code.

Significance. If the reported results are reliable, the paper makes a useful systems-level contribution to UAV-VLN: it explicitly separates current-view grounding, history selection, and topology-aware decision refinement, and it provides an end-to-end trainable policy rather than relying on external planners. The design is well motivated, the training/evaluation separation is clean (privileged goal information is used only in rewards, not in the policy state), and the use of frozen CLIP models for reward and revisit verification avoids an obvious circularity. The authors also ship code, which supports reproducibility. The main weakness is that the headline SOTA claim is supported by three-seed averages with no variance estimates and by cross-report comparisons against baselines from separate publications; on the AerialVLN-S unseen split the margin over the strongest baseline is only 1.31 SR points and 0.54 NE points. Because the paper's central claim is empirical state-of-the-art performance, this statistical fragility is load-bearing.

major comments (3)
  1. [Section 4.3, Tables 1 and 2] The SOTA claim rests on cross-report comparisons in which all numbers are three-seed averages without standard deviations. On AerialVLN-S validation-unseen, the margin over Fly0 is 1.31 SR (61.38 vs. 60.07) and 0.54 NE (50.69 vs. 51.23), and the paper itself states that the results 'do not constitute a paired significance test against baseline values obtained from separate reports.' Three seeds are too few to establish a 1-point ordering if run-to-run variation is on the order of one or two SR points. The OpenFly margins are larger, but they are still single-point comparisons without shared evaluation harness or per-seed statistics. Please report per-seed results and standard deviations or confidence intervals for all main tables, and either run the strongest baselines under the same evaluation harness or temper the claim from 'state-of-the-art' to 'competitive' unless the uncertainty is quantified.
  2. [Section 4.4, Table 3 and leave-one-out text] The cumulative ablations and the leave-one-out numbers are presented as single values with no variance estimates. While the aggregate improvement from baseline to full model is large (SR 35.85 to 61.38), the per-component gains (e.g., +5.15 SR for LOC, +4.91 SR for sparse grounding) are exactly the kind of differences that can be sensitive to seed-level noise. The sentence in Section 4.3 that 'three-seed results reliably quantify run-to-run variance' is not supported because no variance is reported anywhere in the paper. Please provide seed-level results and error bars for the ablations, or explicitly state that component-level ordering is not statistically established.
  3. [Sections 3.4 and 3.5, Eqs. (15)-(27)] The composite reward and GRPO procedure depend on a large set of hyperparameters (lambda_p, lambda_g, lambda_s, lambda_r, lambda_c, epsilon_g, epsilon_n, w, mu, delta_r, delta_h, rho_v, epsilon_d, Delta, epsilon_loc, H, K_r, epsilon_c, beta, M). The manuscript defers all values to the supplementary material. Since reward-shaping coefficients directly determine the learned policy and the ablations in Section 4.4 are part of the evidence chain, please ensure that the complete configuration, including initialization and sensitivity, is available either in the main text or in a clearly accessible appendix; without these values the reported comparisons cannot be reproduced.
minor comments (5)
  1. [Title and Abstract] The title contains an unintended space in 'UA V Vision-Language Navigation'; please correct it to 'UAV Vision-Language Navigation'.
  2. [Section 3.4, after Eq. (14)] The phrase 'LOC normalize semantic agreement' should be 'LOC normalizes semantic agreement'.
  3. [Section 3.4, paragraph 'GRPO-based policy refinement'] The word 'costrains' is a typo for 'constrains'.
  4. [Figure 4] The legend contains both a general 'Methods' block and a separate 'Fly0'/'Ours' highlight, which is redundant and potentially confusing; please unify the legend entries.
  5. [Tables 1 and 2] The use of underlining for second-best results is not explained in the table captions; please state the convention explicitly.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: training, reward design, and evaluation are cleanly separated, and the SOTA claim rests on held-out benchmark numbers rather than on fitted parameters or self-citation chains.

full rationale

The paper's central derivation chain is self-contained. The policy state is constructed from egocentric RGB-D observations, onboard pose estimates, retained history, a topology graph, and the instruction (Eq. 1), and the decoder context is assembled from these inputs via semantic-spatial enhancement, DTA, and LOC (Eqs. 3-14, 22). Training uses behavior cloning followed by GRPO with a composite reward (Eqs. 15-25); the privileged goal region g and simulator distance d(s,g) enter only the reward and evaluation interfaces, not the decoder context, so the policy is not given the evaluation target at inference. The reported NE, SR, and OSR are computed by the standard benchmark protocol against held-out splits, with ablations on the same unseen split showing monotonic improvement from component additions; no parameter is fitted to the test metric and then reported as a prediction. The explicit disclaimer that three-seed results 'do not constitute a paired significance test against baseline values obtained from separate reports' is a statistical limitation and a correctness risk, not evidence of circularity: the comparison being fragile does not mean the method's output is equivalent to its input by construction. The self-citations in the reference list (e.g., Du et al. 2023, Huang et al. 2022, Li et al. 2026, Wu et al. 2024/2026) are contextual citations for UAV object detection challenges and are not load-bearing for the navigation results. No uniqueness theorem is imported from the authors' prior work, no ansatz is smuggled in via self-citation, and no known empirical pattern is renamed as a new contribution. The SOTA claim may be statistically fragile, but it is not circular.

Assumptions & free parameters 6 free parameters · 7 assumptions · 0 invented entities

The central claim rests on architectural choices, hand-chosen reward thresholds, and benchmark comparability rather than on a derivation from first principles. The paper introduces no new physical entities. The most consequential unstated assumptions are comparability of separate benchmark reports and adequacy of three-seed averages without variance.

free parameters (6)
  • Composite reward weight vector (lambda_p, lambda_g, lambda_s, lambda_r, lambda_c) = Not specified in main text; deferred to supplementary
    Used in Eq. 15-19 to balance progress, goal, semantic, and path-compliance terms; chosen by hand and directly affects training behavior and final performance.
  • Goal success radius epsilon_g and distance normalization epsilon_n = Not specified in main text
    Used in Eq. 16-17 to define goal completion and episode-normalized progress; these thresholds change the reward landscape.
  • Path compliance parameters (w, mu, delta_r, delta_h, rho_v, epsilon_d) = Not specified in main text
    Used in Eq. 19-21 to detect revisits and persistent deviation; chosen by hand and essential to the path penalty.
  • LOC stagnation window Delta and epsilon_loc = Not specified in main text
    Used in Eq. 13-14 to detect local optima and normalize graph distance; controls when the frontier cue is injected.
  • History buffer size H and Top-K K_r = 2 = K_r = 2 stated; H not specified
    Architecture choices in Eq. 8-9 that determine how much history is aggregated and how many frames become sparse prompts.
  • GRPO hyperparameters (clip width epsilon_c, KL scale beta, group size M) = Not specified in main text
    Used in Eq. 23-25 for policy optimization; standard but values are not reported in the main text.
assumptions (7)
  • domain assumption The POMDP formulation with egocentric RGB-D observations and local pose estimates, without goal coordinates, adequately captures the UAV-VLN problem.
    Section 3.1 defines the entire learning problem around this state representation; all training and inference depend on it.
  • domain assumption The simulator provides odometry-equivalent pose estimates; the synthetic drift study is a sufficient proxy for real localization error.
    Section 3.1 and the robustness experiment in Section 4.4 rely on this to justify deployment relevance.
  • domain assumption Benchmark numbers from different papers are comparable under the same standard protocols.
    Tables 1 and 2 compare reported numbers from separate publications without re-running baselines in the same harness.
  • domain assumption Three-seed averages, without standard deviations, quantify run-to-run stability.
    Section 4.2 states 'All the main results are averaged over three random seeds' and Section 4.3 acknowledges this does not constitute a paired significance test.
  • domain assumption Pretrained detectors (Oriented R-CNN, Grounding DINO) and Qwen2.5-VL with LoRA provide reliable visual and language grounding.
    Section 4.2 lists these components as the perception and reasoning backbone; failures in these models would propagate to the policy state.
  • domain assumption Frozen CLIP similarity is a valid proxy for semantic match between observations and instructions.
    Eq. 18 and Eq. 21 use CLIP descriptors for the semantic reward and revisit detection.
  • standard math GRPO with KL anchoring to a behavior-cloned policy is a valid policy optimization method.
    Eq. 23-25 use the standard GRPO surrogate from prior work; the paper does not derive or verify it beyond citation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation." pith.science (2026). https://pith.science/paper/VZBENNHM

@misc{pith2026260809564,
  author       = {Pith},
  title        = {Pith review of: From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VZBENNHM}},
  note         = {Machine review of arXiv:2608.09564}
}
read the original abstract

UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. First, we present an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state. Subsequently, we develop a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frames into structured landmark prompts for the decoder. Finally, we devise a topology-aware decision method that combines local-optimum cognition with group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. Experiments on the widely used AerialVLN and OpenFly benchmarks clearly demonstrate that our method achieves state-of-the-art performance.

Figures

Figures reproduced from arXiv: 2608.09564 by the authors.

Figure 1
Figure 1. Illustration of the long-horizon UAV-VLN task. The example route shows that successful navigation depends on [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of the proposed semantic-to-decision pipeline. The model encodes observation, instruction, and topology [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Overview of dynamic temporal aggregation with [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Qualitative example of a successful navigation [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 47 canonical work pages

  1. [1]

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sün- derhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018. Vision-and- Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3674–3683

  2. [2]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...

  3. [3]

    Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, and Renxin Zhong. 2025. FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 6659–6676

  4. [4]

    Chen, Jo Chuang, Marynel Vázquez, and Silvio Savarese

    Kevin Chen, Junshen K. Chen, Jo Chuang, Marynel Vázquez, and Silvio Savarese

  5. [5]

    Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. 2021. His- tory Aware Multimodal Transformer for Vision-and-Language Navigation. In Proceedings of the Annual Conference on Neural Information Processing Systems, Vol. 34. 5834–5847

  6. [6]

    Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. 2022. Think Global, Act Local: Dual-Scale Graph Transformer for Vision-and-Language Navigation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16516–16526

  7. [7]

    Bowei Du, Yecheng Huang, Jiaxin Chen, and Di Huang. 2023. Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object De- tection on Drone Images. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13435–13444

  8. [8]

    Yue Fan, Winson Chen, Tongzhou Jiang, Chun Zhou, Yi Zhang, and Xin Eric Wang. 2023. Aerial Vision-and-Dialog Navigation. InFindings of the Association for Computational Linguistics: ACL 2023. 3043–3061

Show all 63 references
  1. [9]

    Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis- Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018. Speaker-Follower Models for Vision-and-Language Navigation. In Proceedings of the Annual Conference on Neu...

  2. [10]

    Yunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu, Zhen Li, Pengan Chen, Qizhi Chen, Zhonghan Tang, Liansheng Wang, Penghui Yang, Yiwen Tang, Yuhang Tang, Shuai Liang, Songyi Zhu, Ziqin Xiong, Yifei Su, Xinyi Ye, Jianan Li, Yan Ding, Dong Wang, Xuelong Li, Zhigang Wang, and Bin...

  3. [11]

    Yunpeng Gao, Zhigang Wang, Pengfei Han, Linglin Jing, Dong Wang, and Bin Zhao. 2024. Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation. arXiv preprint arXiv:2410.08500 (2024)

  4. [12]

    Yicong Hong, Cristian Rodriguez-Opazo, Yuankai Qi, Qi Wu, and Stephen Gould

  5. [13]

    Chih Yao Hu, Yang-Sen Lin, Yuna Lee, Chih-Hai Su, Jie-Ying Lee, Shr-Ruei Tsai, Chin-Yang Lin, Kuan-Wen Chen, Tsung-Wei Ke, and Yu-Lun Liu. 2025. See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation. InProceedings of the 9th Conference on Robo...

  6. [14]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InProceedings of the International Conference on Learning Representations

  7. [15]

    Yecheng Huang, Jiaxin Chen, and Di Huang. 2022. UFPMP-Det: Toward Accurate and Efficient Object Detection on Drone Imagery. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 1026–1033

  8. [16]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. InProceedings of the International Conference on Learning Representations

  9. [17]

    Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee

  10. [18]

    Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, and Nakamasa Inoue. 2025. CityNav: A Large-Scale Dataset for Real-World Aerial Navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 5912–5922

  11. [19]

    Tianshun Li, Tianyi Huai, Zhen Li, Yichun Gao, Haoang Li, and Xinhu Zheng

  12. [20]

    InProceedings of the European Conference on Computer Vision

    Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments. InProceedings of the European Conference on Computer Vision. 104–120

  13. [21]

    Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li, Weihong Ren, and Yang Cong

  14. [22]

    Fei Liu, Shichao Xie, Minghua Luo, Zedong Chu, Junjun Hu, Xiaolong Wu, and Mu Xu. 2025. NavForesee: A Unified Vision-Language World Model for Hi- erarchical Planning and Dual-Horizon Navigation Prediction. arXiv preprint arXiv:2512.01550 (2025)

  15. [23]

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. 2024. Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection. InProceedings of the European Conference ...

  16. [24]

    Wenhao Li, Zimeng Wu, Yu Wu, Zehua Fu, and Jiaxin Chen. 2026. Visual Proto- type Conditioned Focal Region Generation for UAV-Based Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. 3772–3782

  17. [25]

    Adam Pardyl, Dominik Matuszek, Mateusz Przebieracz, Marek Cygan, Bartosz Zieliński, and Maciej Wolczyk. 2025. FlySearch: Exploring How Vision-Language Models Explore. InProceedings of the Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Vol. 38

  18. [26]

    arXiv preprint arXiv:2511.06182 (2025)

    OpenVLN: Open-world Aerial Vision-Language Navigation. arXiv preprint arXiv:2511.06182 (2025)

  19. [27]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models from Natural Language Supervision. InProceedings ...

  20. [28]

    Oleg Sautenkov, Yasheerah Yaqoot, Artem Lykov, Muhammad Ahsan Mustafa, Grik Tadevosyan, Aibek Akhmetkazy, Miguel Altamirano Cabrera, Mikhail Mar- tynov, Sausar Karaf, and Dzmitry Tsetserukou. 2025. UAV-VLA: Vision-Language- Action System for Large Scale Aerial Mission Generati...

  21. [29]

    Shubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang, Yanning Zhang, and Qi Wu. 2023. AerialVLN: Vision-and-Language Navigation for UAVs. InProceedings of the IEEE/CVF International Conference on Computer Vision. 15384–15394

  22. [30]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov

  23. [31]

    Yuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang, Anton van den Hengel, and Qi Wu. 2021. The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1655–1664

  24. [32]

    Xiangyu Shi, Zerui Li, Wenqi Lyu, Jiatong Xia, Feras Dayoub, Yanyuan Qiao, and Qi Wu. 2025. SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation. InProceedings of the IEEE/RSJ International Conference on Intelligent Robots and Sy...

  25. [33]

    Yifei Su, Dong An, Kehan Chen, Weichen Yu, Baiyang Ning, Yonggen Ling, Yan Huang, and Liang Wang. 2025. Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 7060–7068

  26. [34]

    Pranav Saxena, Nishant Raghuvanshi, and Neena Goveas. 2025. UAV-VLN: End-to-End Vision Language Guided Navigation for UAVs. arXiv preprint arXiv:2504.21432 (2025)

  27. [35]

    Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan- Fang Wang, William Yang Wang, and Lei Zhang. 2019. Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Naviga- tion. InProceedings of the IEEE/CVF Conference on ...

  28. [36]

    Xiangyu Wang, Donglin Yang, Yue Liao, Wenhao Zheng, Wenjun Wu, Bin Dai, Hongsheng Li, and Si Liu. 2025. UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning. InProceedings of the Annual Conference on Neural Information Processing Systems Datase...

  29. [37]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300 (2024)

  30. [38]

    Ke Wu, Jiaxin Chen, and Miao Wang. 2024. Domain Adaptive Object Detection for UAV-based Images by Robust Representation Learning and Multiple Pseudo- label Aggregation. InProceedings of the 1st International Workshop on Efficient Multimedia Computing under Limited. 59–67

  31. [39]

    Ke Wu, Yanan Zhang, Yingjie Gao, Wenhao Li, Chenyu Zhou, XinZhu Ma, Jiaxin Chen, and Di Huang. 2026. DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images. InProceedings of the European Conference on Computer Vision. To appear

  32. [40]

    Yonglin Tian, Fei Lin, Yiduo Li, Tengchao Zhang, Qiyao Zhang, Xuan Fu, Jun Huang, Xingyuan Dai, Yutong Wang, Chunwei Tian, Bai Li, Yisheng Lv, Levente Kovács, and Fei-Yue Wang. 2025. UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility. arXiv preprin...

  33. [41]

    Xingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao, and Junwei Han. 2021. Oriented R-CNN for Object Detection. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3500–3509

  34. [42]

    Haotian Xu, Yue Hu, Chen Gao, Zhengqiu Zhu, Yong Zhao, and Quanjun Yin

  35. [43]

    Xiangyu Wang, Donglin Yang, Ziqin Wang, Hohin Kwan, Jinyu Chen, Wenjun Wu, Hongsheng Li, Yue Liao, and Si Liu. 2025. Towards Realistic UAV Vision- Language Navigation: Platform, Benchmark, and Methodology. InProceedings of the International Conference on Learning Representatio...

  36. [44]

    Fanglong Yao, Youzhi Liu, Wenyi Zhang, Zhengqiu Zhu, Chenglong Li, Nayu Liu, Peng Hu, Yuanchang Yue, Kaiwen Wei, Xin He, Xudong Zhao, Zihan Wei, Haotian Xu, Zhiyuan Wang, Gujie Shao, Liu Yang, Dan Zhao, and Yong Yang

  37. [45]

    Xuan Yao, Junyu Gao, and Changsheng Xu. 2025. NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments. InProceedings of the IEEE/CVF International Conference on Computer Vision. 5536– 5546

  38. [46]

    Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, and Si Liu. 2025. AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation. InProceedings of the 33rd ACM International Conference on Multimedia. 2576–2585

  39. [47]

    Siqi Zhang, Yanyuan Qiao, Qunbo Wang, Longteng Guo, Zhihua Wei, and Jing Liu

  40. [48]

    Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng, Baining Zhao, Qian Zhang, Jinqiang Cui, Xinlei Chen, and Yong Li. 2025. CityNavAgent: Aerial Vision-and- Language Navigation with Hierarchical Semantic Planning and Global Memory. InProceedings of the 63rd Annual Meeting of th...

  41. [49]

    Xinyuan Zhang, Yonglin Tian, Fei Lin, Yue Liu, Jing Ma, Kornélia Sára Szat- máry, and Fei-Yue Wang. 2025. LogisticsVLN: Vision-Language Navigation for Low-Altitude Terminal Delivery Based on Agentic UAVs. arXiv preprint arXiv:2505.03460 (2025)

  42. [50]

    Zhenxing Xu, Yihong Lu, Weidong Bao, Zhengqiu Zhu, Jingxuan Zhou, Zhichuang Wang, Ji Wang, Lihua Liu, and Wei He. 2026. Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation. arXiv preprint arXiv:2602.15875 (2026)

  43. [51]

    Ji Zhao and Xiao Lin. 2025. General-Purpose Aerial Intelligent Agents Empowered by Large Language Models. arXiv preprint arXiv:2503.08302 (2025)

  44. [52]

    AeroVerse-Review: Comprehensive Survey on Aerial Embodied Vision- and-Language Navigation.The Innovation Informatics1, 1 (2025), 100015

  45. [53]

    Yue Zhou, Xue Yang, Gefan Zhang, Jiabao Wang, Yanyi Liu, Liping Hou, Xue Jiang, Xingzhao Liu, Junchi Yan, Chengqi Lyu, Wenwei Zhang, and Kai Chen

  46. [54]

    Jiazhao Zhang, Kunyu Wang, Rongtao Xu, Gengze Zhou, Yicong Hong, Xiaomeng Fang, Qi Wu, Zhizheng Zhang, and He Wang. 2024. NaVid: Video-Based VLM Plans the Next Step for Vision-and-Language Navigation. InProceedings of Ro- botics: Science and Systems, Vol. 20. 16 pages

  47. [56]

    arXiv preprint arXiv:2503.13966 (2025)

    FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks. arXiv preprint arXiv:2503.13966 (2025)

  48. [59]

    Ganlong Zhao, Guanbin Li, Jia Pan, and Yizhou Yu. 2025. Aerial Vision-and- Language Navigation with Grid-based View Selection and Map Construction. arXiv preprint arXiv:2503.11091 (2025)

  49. [61]

    Gengze Zhou, Yicong Hong, and Qi Wu. 2024. NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 7641–7649

  50. [2017]

    arXiv preprint arXiv:1707.06347 (2017)

    Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347 (2017)

  51. [2020]

    InProceedings of the Annual Conference on Neural Information Processing Systems, Vol

    A Language and Visual Entity Relationship Graph for Agent Navigation. InProceedings of the Annual Conference on Neural Information Processing Systems, Vol. 33. 7685–7696

  52. [2021]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Topological Planning with Transformers for Vision-and-Language Naviga- tion. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11271–11281

  53. [2022]

    In Proceedings of the 30th ACM International Conference on Multimedia

    MMRotate: A Rotated Object Detection Benchmark Using PyTorch. In Proceedings of the 30th ACM International Conference on Multimedia. 7331–7334

  54. [2025]

    InProceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems

    SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments. InProceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. 17199–17206

  55. [2026]

    GeoNav: Empowering MLLMs with Dual-Scale Geospatial Reasoning for Language-Goal Aerial Navigation.Pattern Recognition177 (2026), 113365

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.