Pith. sign in

REVIEW 23 cited by

HomeRobot: Open-Vocabulary Mobile Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11565 v2 pith:QHQGMDNY submitted 2023-06-20 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords homerobotovmmenvironmentsmanipulationacrossbaselinescomponentexperiments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

HomeRobot (noun): An affordable compliant robot that navigates homes and manipulates a wide range of objects in order to complete everyday tasks. Open-Vocabulary Mobile Manipulation (OVMM) is the problem of picking any object in any unseen environment, and placing it in a commanded location. This is a foundational challenge for robots to be useful assistants in human environments, because it involves tackling sub-problems from across robotics: perception, language understanding, navigation, and manipulation are all essential to OVMM. In addition, integration of the solutions to these sub-problems poses its own substantial challenges. To drive research in this area, we introduce the HomeRobot OVMM benchmark, where an agent navigates household environments to grasp novel objects and place them on target receptacles. HomeRobot has two components: a simulation component, which uses a large and diverse curated object set in new, high-quality multi-room home environments; and a real-world component, providing a software stack for the low-cost Hello Robot Stretch to encourage replication of real-world experiments across labs. We implement both reinforcement learning and heuristic (model-based) baselines and show evidence of sim-to-real transfer. Our baselines achieve a 20% success rate in the real world; our experiments identify ways future research work improve performance. See videos on our website: https://ovmm.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.

  2. DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A new outdoor vision-and-language navigation dataset that pairs driving instructions with continuous, partially dynamic simulated streets, plus a reinforcement learning baseline.

  3. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Combodied Agents are defined as human-centered AI systems that model a person's ongoing state and agency, predict outcomes of possible interventions, and choose proportionate, consent-aware support.

  4. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.

  5. TypeGo: An OS Runtime for Embodied Agents

    cs.SE 2026-07 conditional novelty 6.0 of 10

    An OS-style multi-cadence runtime for LLM-driven robots hides planning latency and schedules concurrent tasks, cutting TTFA by 73% and per-step delay by 50% on a Go2 prototype suite.

  6. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Aligning temporal granularity, action subspaces, and train-test conditioning yields SOTA long-horizon mobile and fine-grained manipulation success for a unified world-action model.

  7. Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot

    cs.RO 2026-01 unverdicted novelty 6.0 of 10

    Genie Sim 3.0 introduces an LLM-powered scene generator, the first LLM-based automated evaluation benchmark, and a large open synthetic dataset that demonstrates zero-shot sim-to-real transfer for robotic manipulation...

  8. LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes

    cs.RO 2025-08 reject novelty 6.0 of 10

    A staged teacher-student pipeline lets a simulated humanoid relocate two objects in sequence without resets, from egocentric RGB and language, beating the single-task baseline on 350 training and 66 unseen layouts.

  9. Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Inter-LLM interleaves LLM task selection with motion-planning cost feedback through a multimodal similarity estimator, reporting 30% lower mission cost than SayPlan and MoMa-LLM in a simulated household setting.

  10. A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

    cs.RO 2025-02 conditional novelty 6.0 of 10

    IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.

  11. Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PartCATSeg improves open-vocabulary part segmentation by separating object- and part-level cost volumes, adding a compositional loss, and injecting DINO structural guidance, achieving over 10% h-IoU gains on three benchmarks.

  12. VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

    cs.RO 2024-12 conditional novelty 6.0 of 10

    VLABench introduces 100 manipulation task categories with long-horizon reasoning, and shows that current VLAs and VLMs fail most of them.

  13. From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A single MLLM-based agent, finetuned with cross-domain supervision and online RL, achieves strong zero-shot generalization across manipulation, navigation, games, UI control, and planning.

  14. SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    One navigation model with state-adaptive mixture-of-experts routing matches or exceeds task-specific agents on several of seven navigation benchmarks.

  15. WildLMa: Long Horizon Loco-Manipulation in the Wild

    cs.RO 2024-11 conditional novelty 6.0 of 10

    WildLMa combines VR teleoperation with whole-body control, CLIP-based language-conditioned imitation learning, and an LLM planner to give a quadruped robot reusable manipulation skills that generalize to unseen object...

  16. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...

  17. RobotMover: Learning to Move Large Objects From Human Demonstrations

    cs.RO 2025-02 conditional novelty 5.0 of 10

    RobotMover trains a Spot robot to move large objects in the real world by imitating human-object interaction demonstrations through a compact Interaction Chain reward, without fine-tuning on hardware.

  18. InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction

    cs.RO 2024-12 conditional novelty 5.0 of 10

    InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.

  19. ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

    cs.AI 2026-07 conditional novelty 4.0 of 10

    An episode-level variant of GRPO (ESPO) improves personalized GUI-agent reasoning on the 102-episode SmartSpot benchmark, outperforming step-wise and outcome-only training baselines.

  20. MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation

    cs.RO 2025-09 conditional novelty 4.0 of 10

    MoTo turns existing fixed-base manipulation models into mobile manipulators by using VLM-picked contact keypoints and trajectory optimization to find docking points, with no training of MoTo itself.

  21. Embodied AI Agents: Modeling the World

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Embodied AI agents should be built around physical world models plus a mental world model of the user, with virtual, wearable, and robotic agents sharing this core.

  22. ManiSkill-ViTac 2025: Challenge on Manipulation Skill Learning With Vision and Tactile Sensing

    cs.RO 2024-11 unverdicted novelty 4.0 of 10

    The paper describes the tasks, simulation, evaluation metrics, reference policy, and prizes of a 2025 robot manipulation challenge combining touch and vision.

  23. A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI

    cs.RO 2025-05 conditional novelty 2.0 of 10

    A review of navigation and manipulation simulators, datasets, and methods, framed around the sim-to-real gap.

Pith tools