REVIEW 29 cited by
Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Habitat 3.0: a simulation platform for studying collaborative human-robot tasks in home environments. Habitat 3.0 offers contributions across three dimensions: (1) Accurate humanoid simulation: addressing challenges in modeling complex deformable bodies and diversity in appearance and motion, all while ensuring high simulation speed. (2) Human-in-the-loop infrastructure: enabling real human interaction with simulated robots via mouse/keyboard or a VR interface, facilitating evaluation of robot policies with human input. (3) Collaborative tasks: studying two collaborative tasks, Social Navigation and Social Rearrangement. Social Navigation investigates a robot's ability to locate and follow humanoid avatars in unseen environments, whereas Social Rearrangement addresses collaboration between a humanoid and robot while rearranging a scene. These contributions allow us to study end-to-end learned and heuristic baselines for human-robot collaboration in-depth, as well as evaluate them with humans in the loop. Our experiments demonstrate that learned robot policies lead to efficient task completion when collaborating with unseen humanoid agents and human partners that might exhibit behaviors that the robot has not seen before. Additionally, we observe emergent behaviors during collaborative task execution, such as the robot yielding space when obstructing a humanoid agent, thereby allowing the effective completion of the task by the humanoid agent. Furthermore, our experiments using the human-in-the-loop tool demonstrate that our automated evaluation with humanoids can provide an indication of the relative ordering of different policies when evaluated with real human collaborators. Habitat 3.0 unlocks interesting new features in simulators for Embodied AI, and we hope it paves the way for a new frontier of embodied human-AI interaction capabilities.
Forward citations
Cited by 29 Pith papers
-
NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation
A new physics-enabled benchmark with 10,000 indoor, outdoor, and indoor-to-outdoor navigation episodes shows current zero-shot agents fail most when crossing the indoor-outdoor boundary, especially on the new PlaceNav task.
-
Towards Generalizable Robotic Manipulation in Dynamic Environments
DOMINO supplies 110K+ dynamic dual-arm trajectories across 35 tasks, and PUMA’s optical-flow history plus object-centric future queries raise dynamic success rate by 6.3 points over strong VLA baselines.
-
The One RING: a Robotic Indoor Navigation Generalist
A simulation-trained policy that randomizes robot body and camera configurations generalizes zero-shot to real robots it has never seen.
-
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
A new human-centric agentic AI paradigm, Combodied Agents, organizes perception, memory, prediction, and intervention around the evolving human state and agency over time.
-
SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation
SONG, a benchmark with 1,000 photorealistic Gaussian scenes, 500 animated human avatars, and 500 difficulty-graded episodes, finds current vision-based social navigation policies succeed below 22% in easy episodes and...
-
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
A refer-then-track policy picks the target from indexed detections before planning waypoints, achieving state-of-the-art single-view results on EVT-Bench and approaching multi-camera performance.
-
Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation
Streaming intermediate hidden states from a slow VLM to a fast flow-matching planner, token by token, improves dynamic social VLN success and reduces observation staleness.
-
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
UESF-Bench is a 1.43M-sample simulated benchmark for embodied agents that must first find a language-described person and then follow them; SeekFollow-VLA with task-driven routing outperforms the paper's internal baselines.
-
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.
-
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.
-
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
The survey organizes over 400 papers on embodied AI safety into a multi-level taxonomy and flags overlooked issues such as fragile multimodal fusion and unstable planning under jailbreaks.
-
HIMM: Human-Inspired Long-Term Memory Modeling for Embodied Exploration and Question Answering
An embodied-agent memory framework that disentangles episodic and semantic memories, retrieves past experiences via visual reasoning, and distills program-style rules achieves new state-of-the-art results on A-EQA and...
-
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
AREA3D fuses feed-forward 3D confidence and vision-language region reasoning to select informative viewpoints, improving sparse-view 3D reconstruction quality.
-
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
In modular RL-based object-goal navigation, perception quality and test-time strategies dominate performance; policy architecture and observation-space choices contribute little under the tested settings.
-
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
StarDojo is a 1,000-task benchmark in Stardew Valley combining production and social activities, and the best tested MLLM (GPT-4.1) achieves only 12.7% success on its 100-task subset.
-
Ella: Embodied Social Agents with Lifelong Memory
Ella, an embodied social agent with a name-centric semantic memory and a spatiotemporal episodic memory, outperformed two re-implemented baselines in social influence and leadership tasks in a 3D simulation.
-
OctoNav: Towards Generalist Embodied Navigation
OctoNav-R1, trained with SFT, GRPO, and online RL on the new OctoNav-Bench, achieves 19.4% overall success on mixed-instruction navigation, more than double the best baseline.
-
TrackVLA: Embodied Visual Tracking in the Wild
A single vision-language-action model jointly trained on recognition and tracking data follows described targets at the best reported levels on a public benchmark and transfers zero-shot from simulation to a real quad...
-
DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
DIPO generates articulated 3D objects from a closed and an open image, and the new PM-X dataset improves generalization to complex objects.
-
When Incentives Backfire, Data Stops Being Human
Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.
-
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
A two-stage VLM fine-tuning approach, coordinate alignment plus chain-of-thought grounding, improves closed-loop navigation and manipulation success rates over prior point-based spatial reasoning methods.
-
FIction: 4D Future Interaction Prediction from Video
FICTION predicts future 3D interaction locations and body poses up to three minutes ahead from egocentric video and a 3D scene map, and claims substantial gains over prior methods on a new Ego-Exo4D benchmark.
-
ViSTa Dataset: Do vision-language models understand sequential tasks?
ViSTa is a new hierarchical video benchmark showing that vision-language models recognize objects well but fail to understand action order in sequential tasks.
-
Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions
Half physics converts kinematic SMPL-X poses into velocities that drive a physics engine, preserving the original motion when contact-free and giving physically correct responses when collisions occur.
-
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...
-
InfiniteWorld: A Unified Scalable Simulation Framework for General Visual-Language Robot Interaction
InfiniteWorld presents an Isaac Sim based simulator with unified assets and four benchmarks, including scene graph exploration and social mobile manipulation, but reports zero success on the main social task.
-
Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective
Agent spatial intelligence is organized into six neuroscience-inspired modules, and the field is reviewed through that lens without any experimental validation.
-
Generating Actionable Robot Knowledge Bases by Combining 3D Scene Graphs with Robot Ontologies
A pipeline standardizes heterogeneous robot scene formats into USD, maps them through semantic reporting into an ontology-backed knowledge graph, and a robot uses the graph to answer task queries for breakfast table-setting.
-
Learning Implicit Social Navigation Behavior using Deep Inverse Reinforcement Learning
S-MEDIRL, a deep inverse RL method with a bilateral filtering smoothing loss and demonstration extrapolation, learns to yield and avoid deadlock in a narrow crossing, reaching about 92% success.
Discussion (0). Continue with ORCID to comment.