Pith. sign in

REVIEW 29 cited by

Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.13724 v1 pith:XGQPD2A2 submitted 2023-10-19 cs.HC cs.AIcs.CVcs.GRcs.MAcs.RO

classification cs.HCcs.AIcs.CVcs.GRcs.MAcs.RO
keywords humanoidrobotcollaborativehabitathumansocialpoliciessimulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and motion, all while ensuring high simulation speed. (2) Human-in-the-loop infrastructure: enabling real human interaction with simulated robots via mouse/keyboard or a VR interface, facilitating evaluation of robot policies with human input. (3) Collaborative tasks: studying two collaborative tasks, Social Navigation and Social Rearrangement. Social Navigation investigates a robot's ability to locate and follow humanoid avatars in unseen environments, whereas Social Rearrangement addresses collaboration between a humanoid and robot while rearranging a scene. These contributions allow us to study end-to-end learned and heuristic baselines for human-robot collaboration in-depth, as well as evaluate them with humans in the loop. Our experiments demonstrate that learned robot policies lead to efficient task completion when collaborating with unseen humanoid agents and human partners that might exhibit behaviors that the robot has not seen before. Additionally, we observe emergent behaviors during collaborative task execution, such as the robot yielding space when obstructing a humanoid agent, thereby allowing the effective completion of the task by the humanoid agent. Furthermore, our experiments using the human-in-the-loop tool demonstrate that our automated evaluation with humanoids can provide an indication of the relative ordering of different policies when evaluated with real human collaborators. Habitat 3.0 unlocks interesting new features in simulators for Embodied AI, and we hope it paves the way for a new frontier of embodied human-AI interaction capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A new physics-enabled benchmark with 10,000 indoor, outdoor, and indoor-to-outdoor navigation episodes shows current zero-shot agents fail most when crossing the indoor-outdoor boundary, especially on the new PlaceNav task.

  2. Towards Generalizable Robotic Manipulation in Dynamic Environments

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    DOMINO supplies 110K+ dynamic dual-arm trajectories across 35 tasks, and PUMA’s optical-flow history plus object-centric future queries raise dynamic success rate by 6.3 points over strong VLA baselines.

  3. The One RING: a Robotic Indoor Navigation Generalist

    cs.RO 2024-12 conditional novelty 7.0 of 10

    A simulation-trained policy that randomizes robot body and camera configurations generalizes zero-shot to real robots it has never seen.

  4. ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A new human-centric agentic AI paradigm, Combodied Agents, organizes perception, memory, prediction, and intervention around the evolving human state and agency over time.

  5. SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    SONG, a benchmark with 1,000 photorealistic Gaussian scenes, 500 animated human avatars, and 500 difficulty-graded episodes, finds current vision-based social navigation policies succeed below 22% in easy episodes and...

  6. ReferTrack: Referring Then Tracking for Embodied Visual Tracking

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A refer-then-track policy picks the target from indexed detections before planning waypoints, achieving state-of-the-art single-view results on EVT-Bench and approaching multi-camera performance.

  7. Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Streaming intermediate hidden states from a slow VLM to a fast flow-matching planner, token by token, improves dynamic social VLN success and reduces observation staleness.

  8. UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

    cs.AI 2026-07 conditional novelty 6.0 of 10

    UESF-Bench is a 1.43M-sample simulated benchmark for embodied agents that must first find a language-described person and then follow them; SeekFollow-VLA with task-driven routing outperforms the paper's internal baselines.

  9. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.

  10. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  11. Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

    cs.CR 2026-03 unverdicted novelty 6.0 of 10

    The survey organizes over 400 papers on embodied AI safety into a multi-level taxonomy and flags overlooked issues such as fragile multimodal fusion and unstable planning under jailbreaks.

  12. HIMM: Human-Inspired Long-Term Memory Modeling for Embodied Exploration and Question Answering

    cs.RO 2026-02 conditional novelty 6.0 of 10

    An embodied-agent memory framework that disentangles episodic and semantic memories, retrieves past experiences via visual reasoning, and distills program-style rules achieves new state-of-the-art results on A-EQA and...

  13. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

    cs.CV 2025-11 conditional novelty 6.0 of 10

    AREA3D fuses feed-forward 3D confidence and vision-language region reasoning to select informative viewpoints, improving sparse-view 3D reconstruction quality.

  14. What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework

    cs.RO 2025-10 conditional novelty 6.0 of 10

    In modular RL-based object-goal navigation, perception quality and test-time strategies dominate performance; policy architecture and observation-space choices contribute little under the tested settings.

  15. StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

    cs.AI 2025-07 conditional novelty 6.0 of 10

    StarDojo is a 1,000-task benchmark in Stardew Valley combining production and social activities, and the best tested MLLM (GPT-4.1) achieves only 12.7% success on its 100-task subset.

  16. Ella: Embodied Social Agents with Lifelong Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.

  17. OctoNav: Towards Generalist Embodied Navigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.

  18. TrackVLA: Embodied Visual Tracking in the Wild

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...

  19. DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DIPO generates articulated 3D objects from a closed and an open image, and the new PM-X dataset improves generalization to complex objects.

  20. When Incentives Backfire, Data Stops Being Human

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.

  21. SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A two-stage VLM fine-tuning approach, coordinate alignment plus chain-of-thought grounding, improves closed-loop navigation and manipulation success rates over prior point-based spatial reasoning methods.

  22. FIction: 4D Future Interaction Prediction from Video

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FICTION predicts future 3D interaction locations and body poses up to three minutes ahead from egocentric video and a 3D scene map, and claims substantial gains over prior methods on a new Ego-Exo4D benchmark.

  23. ViSTa Dataset: Do vision-language models understand sequential tasks?

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ViSTa is a new hierarchical video benchmark showing that vision-language models recognize objects well but fail to understand action order in sequential tasks.

  24. Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Half physics converts kinematic SMPL-X poses into velocities that drive a physics engine, preserving the original motion when contact-free and giving physically correct responses when collisions occur.

  25. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...

  26. InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction

    cs.RO 2024-12 conditional novelty 5.0 of 10

    InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.

  27. Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

    cs.AI 2025-09 conditional novelty 4.0 of 10

    Agent spatial intelligence is organized into six neuroscience-inspired modules, and the field is reviewed through that lens without any experimental validation.

  28. Generating Actionable Robot Knowledge Bases by Combining 3D Scene Graphs with Robot Ontologies

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A pipeline standardizes heterogeneous robot scene formats into USD, maps them through semantic reporting into an ontology-backed knowledge graph, and a robot uses the graph to answer task queries for breakfast table-setting.

  29. Learning Implicit Social Navigation Behavior using Deep Inverse Reinforcement Learning

    cs.RO 2025-01 conditional novelty 4.0 of 10

    S-MEDIRL, a deep inverse RL method with a bilateral filtering smoothing loss and demonstration extrapolation, learns to yield and avoid deadlock in a narrow crossing, reaching about 92% success.

Pith tools