REVIEW 2 cited by
End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent Phenomenon
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Most recent work in goal oriented visual navigation resorts to large-scale machine learning in simulated environments. The main challenge lies in learning compact representations generalizable to unseen environments and in learning high-capacity perception modules capable of reasoning on high-dimensional input. The latter is particularly difficult when the goal is not given as a category ("ObjectNav") but as an exemplar image ("ImageNav"), as the perception module needs to learn a comparison strategy requiring to solve an underlying visual correspondence problem. This has been shown to be difficult from reward alone or with standard auxiliary tasks. We address this problem through a sequence of two pretext tasks, which serve as a prior for what we argue is one of the main bottleneck in perception, extremely wide-baseline relative pose estimation and visibility prediction in complex scenes. The first pretext task, cross-view completion is a proxy for the underlying visual correspondence problem, while the second task addresses goal detection and finding directly. We propose a new dual encoder with a large-capacity binocular ViT model and show that correspondence solutions naturally emerge from the training signals. Experiments show significant improvements and SOTA performance on the two benchmarks, ImageNav and the Instance-ImageNav variant, where camera intrinsics and height differ between observation and goal.
Forward citations
Cited by 2 Pith papers
-
Memory Proxy Maps for Visual Navigation
A three-tier feudal navigation agent with a self-supervised memory proxy map, a human-imitation waypoint network, and a low-level action classifier achieves state-of-the-art image-goal navigation in unseen Gibson envi...
-
Multimodal Perception for Goal-oriented Navigation: A Survey
A literature survey that categorizes multimodal goal-oriented navigation methods into six inference domains and claims this taxonomy reveals cross-task computational patterns.
Discussion (0). Continue with ORCID to comment.