A hierarchical pipeline generates controllable whole-home 3D scenes from floorplans via LLMs, image models, and VLMs, releasing 300K floorplans and 5K scenes for embodied AI use.
MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments
5 Pith papers cite this work. Polarity classification is still indexing.
abstract
We present MINOS, a simulator designed to support the development of multisensory models for goal-directed navigation in complex indoor environments. The simulator leverages large datasets of complex 3D environments and supports flexible configuration of multimodal sensor suites. We use MINOS to benchmark deep-learning-based navigation methods, to analyze the influence of environmental complexity on navigation performance, and to carry out a controlled study of multimodality in sensorimotor learning. The experiments show that current deep reinforcement learning approaches fail in large realistic environments. The experiments also indicate that multimodality is beneficial in learning to navigate cluttered scenes. MINOS is released open-source to the research community at http://minosworld.org . A video that shows MINOS can be found at https://youtu.be/c0mL9K64q84
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
A collision-triggered visual memory module steering a central-complex compass navigates point goals with no pretraining, but its path-optimization claim depends on an oracle SPL signal.
Consensus recommendations for standardized evaluation measures, problem statements, and benchmarking scenarios in embodied navigation research.
Classical agents outperform learning-based ones on MINOS and Stanford 3D Indoor Spaces, with learned agents weaker at collision avoidance and memory but stronger at handling ambiguity and noise.
A rationale is presented for developing an assistant in Minecraft to advance natural language understanding and dialogue learning.
citing papers explorer
-
HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes
A hierarchical pipeline generates controllable whole-home 3D scenes from floorplans via LLMs, image models, and VLMs, releasing 300K floorplans and 5K scenes for embodied AI use.
-
Insect-inspired Visual Point-goal Navigation
A collision-triggered visual memory module steering a central-complex compass navigates point goals with no pretraining, but its path-optimization claim depends on an oracle SPL signal.
-
On Evaluation of Embodied Navigation Agents
Consensus recommendations for standardized evaluation measures, problem statements, and benchmarking scenarios in embodied navigation research.
-
To Learn or Not to Learn: Analyzing the Role of Learning for Navigation in Virtual Environments
Classical agents outperform learning-based ones on MINOS and Stanford 3D Indoor Spaces, with learned agents weaker at collision avoidance and memory but stronger at handling ambiguity and noise.
-
Why Build an Assistant in Minecraft?
A rationale is presented for developing an assistant in Minecraft to advance natural language understanding and dialogue learning.