BridgeEQA creates a new benchmark and EMVR method for embodied agents to perform question answering on real-world bridge inspections using egocentric images and professional reports.
Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
HUMEMBR introduces a continuous memory system with retrieval for embodied robots that learns human routines to support long-horizon question answering and navigation while using fewer tokens than full-context LLM baselines.
A monocular RGB-only aerial VLN framework outperforms baselines via prompt-guided multi-task learning, keyframe selection, and label reweighting on AerialVLN and OpenFly benchmarks.
AirVista-II integrates agent-based task identification and scheduling, multimodal perception, and scenario-tailored keyframe extraction to deliver high-quality zero-shot semantic understanding for embodied UAVs in dynamic environments.
citing papers explorer
-
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
BridgeEQA creates a new benchmark and EMVR method for embodied agents to perform question answering on real-world bridge inspections using egocentric images and professional reports.
-
HUMEMBR: Learning Human Routines for Predictive Embodied Navigation
HUMEMBR introduces a continuous memory system with retrieval for embodied robots that learns human routines to support long-horizon question answering and navigation while using fewer tokens than full-context LLM baselines.
-
Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning
A monocular RGB-only aerial VLN framework outperforms baselines via prompt-guided multi-task learning, keyframe selection, and label reweighting on AerialVLN and OpenFly benchmarks.
-
AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding
AirVista-II integrates agent-based task identification and scheduling, multimodal perception, and scenario-tailored keyframe extraction to deliver high-quality zero-shot semantic understanding for embodied UAVs in dynamic environments.