REVIEW 4 cited by
Auxiliary Tasks and Exploration Enable ObjectNav
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
ObjectGoal Navigation (ObjectNav) is an embodied task wherein agents are to navigate to an object instance in an unseen environment. Prior works have shown that end-to-end ObjectNav agents that use vanilla visual and recurrent modules, e.g. a CNN+RNN, perform poorly due to overfitting and sample inefficiency. This has motivated current state-of-the-art methods to mix analytic and learned components and operate on explicit spatial maps of the environment. We instead re-enable a generic learned agent by adding auxiliary learning tasks and an exploration reward. Our agents achieve 24.5% success and 8.1% SPL, a 37% and 8% relative improvement over prior state-of-the-art, respectively, on the Habitat ObjectNav Challenge. From our analysis, we propose that agents will act to simplify their visual inputs so as to smooth their RNN dynamics, and that auxiliary tasks reduce overfitting by minimizing effective RNN dimensionality; i.e. a performant ObjectNav agent that must maintain coherent plans over long horizons does so by learning smooth, low-dimensional recurrent dynamics. Site: https://joel99.github.io/objectnav/
Forward citations
Cited by 4 Pith papers
-
The One RING: a Robotic Indoor Navigation Generalist
A simulation-trained policy that randomizes robot body and camera configurations generalizes zero-shot to real robots it has never seen.
-
Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
An object-centric, training-free pipeline using CLIP-derived room-probability vectors to score frontiers improves zero-shot ObjectNav success by a relative 3% over an image-based baseline on HM3D.
-
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
In modular RL-based object-goal navigation, perception quality and test-time strategies dominate performance; policy architecture and observation-space choices contribute little under the tested settings.
-
OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments
OpenIN uses a dynamically updated scene graph of carried-by relationships, plus LLM and VLM guidance, to navigate to specific moved objects in homes, reporting higher success than two open-vocabulary baselines.
Discussion (0). Continue with ORCID to comment.