Pith. sign in

REVIEW 9 cited by

SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08189 v1 pith:JVEOKRXK submitted 2024-10-10 cs.CV cs.RO

classification cs.CVcs.RO
keywords scenenavigationobjectzero-shotgraphmethodssg-navability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a new framework for zero-shot object navigation. Existing zero-shot object navigation methods prompt LLM with the text of spatially closed objects, which lacks enough scene context for in-depth reasoning. To better preserve the information of environment and fully exploit the reasoning ability of LLM, we propose to represent the observed scene with 3D scene graph. The scene graph encodes the relationships between objects, groups and rooms with a LLM-friendly structure, for which we design a hierarchical chain-of-thought prompt to help LLM reason the goal location according to scene context by traversing the nodes and edges. Moreover, benefit from the scene graph representation, we further design a re-perception mechanism to empower the object navigation framework with the ability to correct perception error. We conduct extensive experiments on MP3D, HM3D and RoboTHOR environments, where SG-Nav surpasses previous state-of-the-art zero-shot methods by more than 10% SR on all benchmarks, while the decision process is explainable. To the best of our knowledge, SG-Nav is the first zero-shot method that achieves even higher performance than supervised object navigation methods on the challenging MP3D benchmark.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical room-and-object memory that persists across independent ObjectNav episodes yields small success-rate gains, but most of the gain comes from within-episode memory rather than the cross-episode component.

  2. Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    SGImagineNav uses an imagined hierarchical scene graph, filled in by an LLM, that guides a robot to unseen objects and achieves 65.4% and 66.8% success on HM3D and HSSD.

  3. Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A VLN agent that recursively imagines future views and layouts in a fixed-size neural grid, and adaptively aligns instruction parts to grid cells, achieves state-of-the-art success rates on R2R-CE and ObjectNav.

  4. Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Qwen-RobotNav introduces a parameterized navigation model supporting multiple task modes and controllable observation parameters, trained on 15.6M samples with vision-language co-training to achieve SOTA results on be...

  5. Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Qwen-RobotNav provides a parameterized navigation model trained on 15.6M samples with vision-language co-training that achieves SOTA results on benchmarks and zero-shot transfer to real robots.

  6. G-DRAGON: Geospatial Reasoning and Dynamic Planning for Retrieval-Augmented Outdoor Navigation

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    G-DRAGON framework maps language commands to OSM coordinates via lightweight LLM for global planning and uses frontier exploration for local targets, outperforming baselines in simulation and completing real UGV perso...

  7. MORN: Metacognitive Object-Goal Regulation for Resource-Rational Long-Horizon Navigation

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    MORN augments frozen VLM-based object navigation agents with a System 2 meta-controller using Potentiality Index, Persistence Gating, and Evidence Accumulation to improve goal completion rate from 0.23 to 0.30 and red...

  8. SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations

    cs.RO 2026-03 reject novelty 5.0 of 10

    SG-CoT grounds an LLM planner's chain-of-thought in a scene graph via iterative API queries, improving ambiguity detection and clarification in simulated manipulation, though its success metric credits any clarifying ...

  9. MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation

    cs.RO 2025-09 conditional novelty 4.0 of 10

    MoTo turns existing fixed-base manipulation models into mobile manipulators by using VLM-picked contact keypoints and trajectory optimization to find docking points, with no training of MoTo itself.

Pith tools