Pith. sign in

REVIEW 40 cited by

ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.13171 v2 pith:YDYQHX3J submitted 2020-06-23 cs.CV cs.RO

classification cs.CVcs.RO
keywords navigatingobjectnavobjectrecommendationstaskagentembodiedenvironment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We revisit the problem of Object-Goal Navigation (ObjectNav). In its simplest form, ObjectNav is defined as the task of navigating to an object, specified by its label, in an unexplored environment. In particular, the agent is initialized at a random location and pose in an environment and asked to find an instance of an object category, e.g., find a chair, by navigating to it. As the community begins to show increased interest in semantic goal specification for navigation tasks, a number of different often-inconsistent interpretations of this task are emerging. This document summarizes the consensus recommendations of this working group on ObjectNav. In particular, we make recommendations on subtle but important details of evaluation criteria (for measuring success when navigating towards a target object), the agent's embodiment parameters, and the characteristics of the environments within which the task is carried out. Finally, we provide a detailed description of the instantiation of these recommendations in challenges organized at the Embodied AI workshop at CVPR 2020 http://embodied-ai.org .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 40 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition

    cs.RO 2026-08 conditional novelty 7.0 of 10

    FUSE, an entropy-gated planner that combines amortized viewpoint prediction with explicit semantic-geometric exploration, achieves the best non-oracle active functional grounding results on a new Habitat benchmark whi...

  2. SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

    cs.RO 2026-08 conditional novelty 7.0 of 10

    SAIN compiles dialogue answers into persistent value, room, graph, and object memories, raising SR from 20.2 to 25.4 on VL-LN IIGN without task-specific policy training.

  3. SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A training-free skill layer that modifies the VLM's value map improves zero-shot object-goal navigation SPL by up to 6.0 points on MP3D and HM3D.

  4. REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation

    cs.RO 2026-03 conditional novelty 7.0 of 10

    REST replaces isolated waypoint subgoals with a Steiner-tree-compacted option space of full paths and lets an LLM pick among textualized branches, ranking among the top zero-shot ObjectNav methods on Gibson, HM3D, and HSSD.

  5. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.

  6. FrontierNet: Learning Visual Cues to Explore

    cs.RO 2025-01 conditional novelty 7.0 of 10

    FrontierNet learns to propose frontier exploration targets and their information gain from RGB images plus monocular depth, improving early-stage mapped volume in simulation and on a real robot.

  7. The One RING: a Robotic Indoor Navigation Generalist

    cs.RO 2024-12 conditional novelty 7.0 of 10

    A simulation-trained policy that randomizes robot body and camera configurations generalizes zero-shot to real robots it has never seen.

  8. From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

    cs.RO 2026-07 conditional novelty 6.5 of 10

    REALM, a visibility-aware plug-and-play last-meters module trained on the new REVERIE-AIM dataset, consistently raises instance proximity and grounding success on four VLN backbones.

  9. SpikingNav: Robust Embodied Navigation with Spiking Neural Policies

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A spiking sensing encoder and spiking policy network improve ObjectNav success under visual corruptions (8.45% to 13.71%) while using fewer parameters and fewer FLOPs than a matched ANN baseline.

  10. SSTG-Nav: Metric-Grounded Spatial-Semantic Topological Graphs for Reusable Object Navigation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A pre-explored metric-semantic topology with depth-grounded standoffs, multi-view fusion, and sequential verification achieves high success in repeated object navigation.

  11. BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

    cs.RO 2026-07 conditional novelty 6.0 of 10

    BioVLN introduces a three-zone operational envelope for biomedical lab navigation, with 47 scenes and benchmarks showing multi-point operation-area goals raise success to 83–92% while cutting unsafe proximity.

  12. Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

    cs.RO 2026-07 conditional novelty 6.0 of 10

    An object-centric, training-free pipeline using CLIP-derived room-probability vectors to score frontiers improves zero-shot ObjectNav success by a relative 3% over an image-based baseline on HM3D.

  13. Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A new 120-mission drone benchmark finds the best off-the-shelf multimodal AI completes 34.8% of missions versus 84.4% for humans, with scaling helping but not closing the gap.

  14. ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments

    cs.RO 2026-07 conditional novelty 6.0 of 10

    ZONDA combines height-difference stair traversal, multi-view VLM target verification, and pedestrian tracking to achieve SOTA zero-shot ObjectNav on MP3D and robust results on the new HM3D-DYNA benchmark.

  15. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.

  16. ABot-N1: Toward a General Visual Language Navigation Foundation Model

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A slow–fast VLN model that routes five navigation tasks through CoT plus image-space pixel goals reaches SOTA on established and new urban benchmarks, including 77.3% POI arrival.

  17. Semantic Evidence Regulation via Relational Bias for Zero-Shot Object Navigation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    DB-Nav/SER-Nav improves zero-shot object navigation by reranking frontier goals using activation from object co-occurrence and inhibition from similar distractors and failed visits.

  18. DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Organizing zero-shot object navigation around tracked directional exits with 240° inspection and VLM verification yields 50.2% SR / 32.6% SPL on HM3D-OVON and best SPL on HM3Dv2 and MP3D.

  19. Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A new step-level measure of visual grounding in long-horizon vision-language models predicts out-of-distribution generalization (r=0.83), and varies independently of model scale and in-distribution accuracy.

  20. DSCD-Nav: Dual-Stance Cooperative Debate for Object Navigation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    A dual-stance debate between a goal-focused and a safety-focused VLM, plus arbitration and optional micro-probing, improves zero-shot object navigation success and path efficiency on HM3Dv1, HM3Dv2, MP3D, and GOAT.

  21. Visual-Language-Guided Task Planning for Horticultural Robots

    cs.RO 2026-01 conditional novelty 6.0 of 10

    A vision-language model drives a simulated greenhouse robot through simple crop-inspection tasks with ~87% success, but long multi-target tasks collapse to under 10% success.

  22. VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs

    cs.RO 2025-12 conditional novelty 6.0 of 10

    VL-LN Bench turns instance-goal navigation into an interactive dialog task, contributes a 41k-trajectory house-scale benchmark with a GPT-4o oracle, and shows active questioning improves embodied agents' success.

  23. From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A brain-inspired navigation system stores landmarks, routes, and map-like voxel features in structured spatial memory and uses MLLM-powered retrieval to achieve strong results across object, instance, instruction, and...

  24. MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

    cs.RO 2025-08 conditional novelty 6.0 of 10

    MAG-Nav uses active viewpoint selection and memory replay with GPT-4o to achieve state-of-the-art 40.8% success in zero-shot language-driven object navigation on GOAT-Bench/HM3D.

  25. OctoNav: Towards Generalist Embodied Navigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.

  26. RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RATE-Nav reduces redundant exploration in zero-shot object navigation by segmenting the map into regions and using VLM judgments to terminate unproductive region searches.

  27. ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A GRPO-based reinforcement learning framework teaches an MLLM to propose zoom-in regions, improving small-object detection and interactive segmentation under a fixed sensing budget.

  28. SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    SORT3D is a zero-shot 3D grounding system where an LLM calls hand-built spatial functions and uses 2D captions, matching or beating prior zero-shot methods on several view-dependent subsets while running on real robots.

  29. ApexNav: An Adaptive Exploration Strategy for Zero-Shot Object Navigation with Target-centric Semantic Fusion

    cs.RO 2025-04 conditional novelty 6.0 of 10

    ApexNav adaptively switches between semantic-guided and geometry-based exploration, and fuses multi-frame target-centric evidence, to set a new state of the art in zero-shot object navigation.

  30. Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A monocular video is converted into a Gaussian-splatting plus mesh simulation world, and navigation policies trained there transfer to a real robot better than mesh-only training does.

  31. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    One navigation model with state-adaptive mixture-of-experts routing matches or exceeds task-specific agents on several of seven navigation benchmarks.

  32. Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues

    cs.AI 2024-12 conditional novelty 6.0 of 10

    In collaborative instance navigation, an uncertainty-aware training-free method (AIUTA) lets an agent ask concise questions to find a specific object instance with minimal human input.

  33. FIction: 4D Future Interaction Prediction from Video

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FICTION predicts future 3D interaction locations and body poses up to three minutes ahead from egocentric video and a 3D scene map, and claims substantial gains over prior methods on a new Ego-Exo4D benchmark.

  34. Discounted Beta-Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable Rewards

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    Discounted Beta-Bernoulli reward estimation reduces variance and variance collapse in group RLVR, improving GRPO Acc@8 on reasoning benchmarks at no extra cost.

  35. TopoNav: Topological Graphs as a Key Enabler for Advanced Object Navigation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A zero-shot object navigation system that builds a text-based topological memory graph, queried by GPT-4o, reports state-of-the-art success rates of 60.1% on HM3D and 45.5% on MP3D.

  36. DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A zero-shot VLM navigation policy with dynamic waypoint selection and graph memory reports state-of-the-art results among VLM-based methods on ObjectNav and GOAT-Bench, plus real-world tests on a quadruped.

  37. ForesightNav: Learning Scene Imagination for Efficient Exploration

    cs.RO 2025-04 conditional novelty 5.0 of 10

    ForesightNav uses a learned imagination module that predicts unseen occupancy and CLIP semantic features from partial maps, then selects navigation goals from the imagined map, improving simulated PointNav and ObjectN...

  38. Exploring the Generalizability of Geomagnetic Navigation: A Deep Reinforcement Learning approach with Policy Distillation

    cs.RO 2025-02 conditional novelty 5.0 of 10

    Proposes TD3-STEPD, a deep reinforcement learning method that distills several region-specific geomagnetic navigation policies into one student policy that generalizes to unseen simulated areas.

  39. Multimodal Perception for Goal-oriented Navigation: A Survey

    cs.RO 2025-04 conditional novelty 2.0 of 10

    A literature survey that categorizes multimodal goal-oriented navigation methods into six inference domains and claims this taxonomy reveals cross-task computational patterns.

  40. TANGO: Training-free Embodied AI Agents for Open-world Tasks

    cs.AI 2024-12

Pith tools