ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.
NavGPT: Explicit Reasoning in Vision- and-Language Navigation with Large Language Models
3 Pith papers cite this work, alongside 139 external citations. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
SpaceVLN proposes a stagewise closed-loop framework using Spatial Cognitive Memory and Spatial-CoT for zero-shot vision-and-language navigation and object-goal navigation, reporting SOTA results on R2R-CE, RxR-CE, GN-Bench, and HM3D-OVON plus real-robot tests.
Instruction understanding is reframed as an evolving Instruction-as-State variable conditioned on perceptual state and realized via the S-EGIU coarse-to-fine framework, reporting a +2.68% SPL gain on REVERIE Test Unseen.
citing papers explorer
-
ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue
ESARBench is the first unified benchmark for MLLM-driven UAV agents that must explore, locate clues, and decide on victim positions in photorealistic simulated SAR environments.
-
SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning
SpaceVLN proposes a stagewise closed-loop framework using Spatial Cognitive Memory and Spatial-CoT for zero-shot vision-and-language navigation and object-goal navigation, reporting SOTA results on R2R-CE, RxR-CE, GN-Bench, and HM3D-OVON plus real-robot tests.
-
Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation
Instruction understanding is reframed as an evolving Instruction-as-State variable conditioned on perceptual state and realized via the S-EGIU coarse-to-fine framework, reporting a +2.68% SPL gain on REVERIE Test Unseen.