Pith. sign in

REVIEW 6 cited by

CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12532 v3 pith:6ZWJS2ZC submitted 2025-02-18 cs.AI

classification cs.AI
keywords cityeqaansweringagentembodiedquestiontaskurbanbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings-spanning environment, action, and perception-largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions through active exploration in dynamic city spaces. To support this task, we present CityEQA-EC, the first benchmark dataset featuring 1,412 human-annotated tasks across six categories, grounded in a realistic 3D urban simulator. Moreover, we propose Planner-Manager-Actor (PMA), a novel agent tailored for CityEQA. PMA enables long-horizon planning and hierarchical task execution: the Planner breaks down the question answering into sub-tasks, the Manager maintains an object-centric cognitive map for spatial reasoning during the process control, and the specialized Actors handle navigation, exploration, and collection sub-tasks. Experiments demonstrate that PMA achieves 60.7% of human-level answering accuracy, significantly outperforming competitive baselines. While promising, the performance gap compared to humans highlights the need for enhanced visual reasoning in CityEQA. This work paves the way for future advancements in urban spatial intelligence. Dataset and code are available at https://github.com/BiluYong/CityEQA.git.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

    cs.RO 2026-07 conditional novelty 7.0 of 10

    ActiveFly-Bench defines Air-EQA, Observation Behavior Planning, and 7-DoF FLUC tasks on 10k real/sim trajectories so UAV agents must plan, fly, and answer questions they cannot solve from the start view.

  2. IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios

    cs.CV 2025-05 conditional novelty 7.0 of 10

    IndustryEQA offers 1,344 video-based question-answer pairs across six categories, with a focus on equipment and human safety, plus evaluations of several vision-language models.

  3. USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents

    cs.AI 2025-05 conditional novelty 6.0 of 10

    USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.

  4. IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

    cs.RO 2025-11 conditional novelty 5.0 of 10

    On IndustryNav, a dynamic Unity warehouse navigation benchmark, nine VLLMs earned only 4.9–65.3% success and high collision/warning rates, with closed-source models ahead.

  5. Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation

    cs.RO 2026-02 conditional novelty 4.0 of 10

    A zero-shot aerial VLN system that has an MLLM output only 2D image coordinates, then uses depth unprojection and Ego-Planner to navigate, reporting >20 percentage-point SR gains and 31–37% NE reductions over baselines.

  6. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools