{"total":10,"items":[{"citing_arxiv_id":"2607.06706","ref_index":134,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review","primary_cat":"cs.RO","submitted_at":"2026-07-07T18:24:09+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.5,"formal_verification":"none","one_line_summary":"Bimanual VLA coordination strategies, training recipes, and continuous action chunking transfer to unmanned aerial systems; the survey maps 183 works and lists fourteen shared research directions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12859","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots","primary_cat":"cs.RO","submitted_at":"2026-06-11T03:42:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AIR-VLA+ introduces cascaded manipulation and movement decoders plus asymmetric MoE to decouple action scales in aerial manipulation, reporting 48.0 average score and 80.2% task completion gain over single-head baseline on AIR-VLA benchmark.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.11618","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Vision-Language-Action Models Meet World Models: Embodied Agentic AI for Low-Altitude Wireless Networks","primary_cat":"cs.IT","submitted_at":"2026-06-10T03:30:05+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"Proposes an Embodied Agentic UAV framework using VLA models and World Models with closed-loop memory mechanisms for autonomous control in low-altitude wireless networks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.06147","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation","primary_cat":"cs.AI","submitted_at":"2026-06-04T13:23:05+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"WorldFly integrates a world model into a VLA framework via dual-branch coupled flow matching to jointly generate future videos and actions, outperforming baselines on an urban canyon traversal benchmark especially in unseen environments.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.00104","ref_index":13,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs","primary_cat":"cs.RO","submitted_at":"2026-05-26T10:03:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"PEACE decouples single-pass LLM planning from PX4 execution via ROS 2 and a constraint layer, with modular 3D perception, and shows feasibility in Gazebo SITL with improved explainability and fewer LLM calls.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.19728","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls","primary_cat":"cs.CV","submitted_at":"2026-05-19T12:02:17+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Aero-World adapts a pretrained latent diffusion transformer for action-conditioned aerial video generation by injecting inertial action tokens and using a frozen latent-space Physics Probe for inertial consistency supervision during LoRA finetuning, with a new AeroBench benchmark showing improved AA","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.12735","ref_index":272,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The Unified Autonomy Stack: Toward a Blueprint for Generalizable Robot Autonomy","primary_cat":"cs.RO","submitted_at":"2026-05-12T20:39:35+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"An open-sourced Unified Autonomy Stack fuses LiDAR, radar, vision and inertial data with sampling-based planning and control barrier functions to deliver resilient autonomy on aerial and ground robots in challenging real-world settings.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.13654","ref_index":85,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap","primary_cat":"cs.RO","submitted_at":"2026-04-15T09:20:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"A survey of UAV vision-and-language navigation that establishes a methodological taxonomy, reviews resources and challenges, and proposes a forward-looking research roadmap.","context_count":1,"top_context_role":"method","top_context_polarity":"use_method","context_text":"GR00T N1 [25] 2025 VLA Dual-system architectureOpen foundation model for generalist humanoid robots NWM [83] 2025 World Model Controllable video generation Navigation World Models predicting future visual observations Cosmos-Reason1 [26] 2025 WM + Reasoning Physical common senseCombining world models with embodied reasoning GRaD-Nav++ [84] 2025 VLA MoE + Diff. RL in 3DGSLightweight onboard VLA at 25 Hz real-time control RaceVLA [85] 2025 VLA End-to-end FPV controlFirst VLA for high-speed autonomous drone racing TrackVLA [86] 2025 VLA Unified LLM backbone Integrated target recognition and trajectory planning NaVILA [87] 2025 Hierarchical VLM planner + RL executor Mid-level language actions bridging reasoning and control Swarm-GPT [88] 2023 LLM + Safety LLM planner + safety filter Language-based multi-UA V swarm coordination"},{"citing_arxiv_id":"2507.01925","ref_index":28,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Survey on Vision-Language-Action Models: An Action Tokenization Perspective","primary_cat":"cs.RO","submitted_at":"2025-07-02T17:34:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"The survey frames VLA models as pipelines that generate progressively grounded action tokens and classifies those tokens into eight types to guide future development.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"extend the scaling laws [19, 20] observed in vision and language domains to the embodied setting, collecting large-scale embodied datasets and training generalist agents end-to-end on top of vision-language foundation models [21, 22, 23]. These diverse approaches have led to a rapid proliferation of VLA models in robotic manipulation [24, 25], navigation [26, 27], and autonomous driving [28, 29, 30], demonstrating promising capabilities in multitask learning [31], long-horizon task completion [22], and strong generalization [32]. By leveraging foundation model intelligence, they offer new directions for addressing long-standing challenges in embodied AI, such as data scarcity and poor cross-embodiment transferability, and pave the way foragents"},{"citing_arxiv_id":"2506.14009","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics","primary_cat":"cs.RO","submitted_at":"2025-06-16T21:12:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"GRaD-Nav++ combines 3D Gaussian Splatting simulation and differentiable RL to train an onboard VLA policy that achieves 50-83% success on language-guided drone navigation tasks in simulation and real hardware.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}