Pith. sign in

REVIEW 2 cited by

Navigating to Objects Specified by Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01192 v1 pith:FNV265ZG submitted 2023-04-03 cs.CV cs.RO

classification cs.CVcs.RO
keywords goalinstancesystemexplorationimagessuccesstaskachieving
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Images are a convenient way to specify which particular object instance an embodied agent should navigate to. Solving this task requires semantic visual reasoning and exploration of unknown environments. We present a system that can perform this task in both simulation and the real world. Our modular method solves sub-tasks of exploration, goal instance re-identification, goal localization, and local navigation. We re-identify the goal instance in egocentric vision using feature-matching and localize the goal instance by projecting matched features to a map. Each sub-task is solved using off-the-shelf components requiring zero fine-tuning. On the HM3D InstanceImageNav benchmark, this system outperforms a baseline end-to-end RL policy 7x and a state-of-the-art ImageNav model 2.3x (56% vs 25% success). We deploy this system to a mobile robot platform and demonstrate effective real-world performance, achieving an 88% success rate across a home and an office environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    GC-VLN decomposes a navigation instruction into a graph of spatial constraints, solves the constraints with an optimizer, and beats prior zero-shot methods on VLN-CE benchmarks without any training.

  2. RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation

    cs.CV 2025-04 conditional novelty 6.0 of 10

    Correlation between goal and current observation, refined by a direction-aware pyramid, yields state-of-the-art image-goal navigation on Gibson, MP3D, and HM3D, with the largest gains under user-matched goal viewpoints.

Pith tools