Pith. sign in

REVIEW 25 cited by

GOAT: GO to Any Thing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06430 v1 pith:76DRI4MQ submitted 2023-11-10 cs.RO

GOAT: GO to Any Thing

classification cs.RO
keywords goatdifferentnavigationsuccessadditioncategorydescriptionsenvironment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In deployment scenarios such as homes and warehouses, mobile robots are expected to autonomously navigate for extended periods, seamlessly executing tasks articulated in terms that are intuitively understandable by human operators. We present GO To Any Thing (GOAT), a universal navigation system capable of tackling these requirements with three key features: a) Multimodal: it can tackle goals specified via category labels, target images, and language descriptions, b) Lifelong: it benefits from its past experience in the same environment, and c) Platform Agnostic: it can be quickly deployed on robots with different embodiments. GOAT is made possible through a modular system design and a continually augmented instance-aware semantic memory that keeps track of the appearance of objects from different viewpoints in addition to category-level semantics. This enables GOAT to distinguish between different instances of the same category to enable navigation to targets specified by images and language descriptions. In experimental comparisons spanning over 90 hours in 9 different homes consisting of 675 goals selected across 200+ different object instances, we find GOAT achieves an overall success rate of 83%, surpassing previous methods and ablations by 32% (absolute improvement). GOAT improves with experience in the environment, from a 60% success rate at the first goal to a 90% success after exploration. In addition, we demonstrate that GOAT can readily be applied to downstream tasks such as pick and place and social navigation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MCNav: Memory-Aware Dynamic Cognitive Map for Zero-shot Goal-oriented Navigation

    cs.RO 2026-05 unverdicted novelty 7.0

    MCNav builds a dynamic cognitive map with goal re-validation and missed-goal re-exploration to reach state-of-the-art results on instance-level zero-shot navigation in HM3D environments.

  2. Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

    cs.LG 2026-05 unverdicted novelty 7.0

    Embedding Temporal Logic (ETL) performs runtime monitoring directly in learned embedding spaces using distance-based predicates composed with temporal operators, supported by conformal calibration for reliable predica...

  3. Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

    cs.LG 2026-05 unverdicted novelty 7.0

    Embedding Temporal Logic enables runtime monitoring of temporally extended perceptual behaviors by defining predicates via distances between observed and reference embeddings in learned spaces, with conformal calibrat...

  4. ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

    cs.AI 2026-05 unverdicted novelty 7.0

    ProCompNav disambiguates ambiguous instance navigation queries via candidate-pool construction followed by attribute-based comparative binary questions that prune distractors, yielding higher success rates and shorter...

  5. ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

    cs.AI 2026-05 unverdicted novelty 7.0

    ProCompNav improves success rate and shortens user responses in ambiguous instance navigation by using comparative binary questions that prune a candidate pool rather than requesting detailed descriptions.

  6. Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation

    cs.CV 2026-06 unverdicted novelty 6.0

    Proposes online hierarchical 3D scene graph construction paired with belief-based planning to improve zero-shot semantic navigation performance in unseen environments.

  7. SurveilNav: Collaborative Object Goal Navigation with Robot and Surveillance System

    cs.RO 2026-06 unverdicted novelty 6.0

    SurveilNav integrates robot local perception with multi-view surveillance for improved collaborative object goal navigation and reports SOTA results on HM3D.

  8. Autonomous Frontier-Based Exploration with VLM Guidance

    cs.RO 2026-05 unverdicted novelty 6.0

    A VLM-based method for selecting exploration frontiers in robotics achieves up to 24% better map coverage than standard geometric heuristics in simulated indoor environments.

  9. Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation

    cs.RO 2026-05 unverdicted novelty 6.0

    SAGE trains agents in physics-grounded semantic abstractions via RL with asymmetric clipping, achieving 53.21% LLM-Match Success on A-EQA (+9.7% over baseline) and encouraging physical robot transfer.

  10. ProCompNav: Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries

    cs.AI 2026-05 unverdicted novelty 6.0

    ProCompNav builds a candidate pool from ambiguous queries then uses pool-splitting binary questions for disambiguation, improving success rate and shortening responses on CoIN-Bench and TextNav.

  11. DM$^3$-Nav: Decentralized Multi-Agent Multimodal Multi-Object Semantic Navigation

    cs.MA 2026-04 unverdicted novelty 6.0

    DM³-Nav delivers decentralized multi-agent semantic navigation for multimodal open-vocabulary multi-object tasks that matches centralized baselines in simulation and succeeds in real-world robot deployments.

  12. Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies

    cs.RO 2026-04 unverdicted novelty 6.0

    VLMs generalize affordance inference to non-humanoid robots but produce inconsistent results with a conservative bias of low false positives and high false negatives, especially for novel object manipulations.

  13. FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation

    cs.RO 2026-04 unverdicted novelty 6.0

    FSUNav's dual brain-inspired modules achieve state-of-the-art zero-shot goal navigation across heterogeneous robots with improved speed, safety, and generalization.

  14. Memory Over Maps: 3D Object Localization Without Reconstruction

    cs.RO 2026-03 unverdicted novelty 6.0

    A map-free localization method stores posed RGB-D keyframes, retrieves and re-ranks them with a VLM, then fuses sparse depth for on-demand 3D target estimates, matching reconstruction-based performance on navigation b...

  15. FeudalNav: A Simple Framework for Visual Navigation

    cs.RO 2026-01 unverdicted novelty 6.0

    FeudalNav decomposes visual navigation into hierarchical levels with a visual-similarity latent memory, delivering competitive Habitat AI results without any odometry.

  16. SplatSearch: Instance Image Goal Navigation for Mobile Robots using 3D Gaussian Splatting and Diffusion Models

    cs.RO 2025-11 conditional novelty 6.0

    SplatSearch combines sparse-view 3D Gaussian Splatting, multi-view diffusion inpainting, and semantic/visual frontier scoring to achieve viewpoint-invariant instance image-goal navigation in unknown environments.

  17. TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals

    cs.RO 2025-09 conditional novelty 6.0

    A navigation pipeline that bridges object-level global planning with traversability-aware local control, using only RGB images and pretrained models, improves success over prior zero-shot and learned baselines in simulation.

  18. Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

    cs.RO 2025-07 unverdicted novelty 6.0

    RIGVid shows that filtered AI-generated videos can serve as effective supervision for complex robotic manipulation tasks without any real demonstrations.

  19. Learning Category-level Last-meter Navigation from RGB Demonstrations of a Single-instance

    cs.RO 2025-12 conditional novelty 5.0

    An RGB-only imitation-learning policy trained on one chair generalizes to unseen chairs and environments for last-meter base positioning, though reported success depends on an added heuristic stopping rule and a 0.3 m...

  20. N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout

    cs.RO 2025-09 conditional novelty 5.0

    N2M predicts preferable base poses for manipulation policies from ego-centric point clouds, learned from rollouts, lifting success from 3% to 54% in the PnPCounterToCab task.

  21. Collision-Aware Object-Goal Visual Navigation via Two-Stage Deep Reinforcement Learning

    cs.RO 2025-02 unverdicted novelty 5.0

    A two-stage DRL method with a supervised collision prediction stage improves collision-free success rates on object-goal visual navigation tasks in simulation and real rooms.

  22. A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration

    cs.RO 2026-04 unverdicted novelty 4.0

    Introduces a hierarchical VLN architecture with asynchronous layers, incremental memory graph, and WTRP-based exploration that improves success and efficiency on resource-constrained robots.

  23. A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration

    cs.RO 2026-04 unverdicted novelty 4.0

    A modular VLN architecture builds a cognitive memory graph, decomposes it for VLM reasoning, and solves a weighted traveling repairman problem for context-aware exploration to achieve real-time performance and higher ...

  24. DELIVER: A System for LLM-Guided Coordinated Multi-Robot Pickup and Delivery using Voronoi-Based Relay Planning

    cs.RO 2025-08 conditional novelty 4.0

    DELIVER combines an LLM parser, Voronoi territory division, and relay handoffs so multiple robots can cooperatively deliver an item from a single spoken instruction.

  25. SPG: Style-Prompting Guidance for Style-Specific Content Creation

    cs.GR 2025-08 unverdicted novelty 4.0

    SPG is not described anywhere in the supplied text; the body is a different paper (OVSegDT) about robot navigation.