Pith. sign in

hub Baseline reference

Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Baseline reference. 89% of citing Pith papers use this work as a benchmark or comparison.

71 Pith papers citing it
Baseline 89% of classified citations
abstract

We present the Habitat-Matterport 3D (HM3D) dataset. HM3D is a large-scale dataset of 1,000 building-scale 3D reconstructions from a diverse set of real-world locations. Each scene in the dataset consists of a textured 3D mesh reconstruction of interiors such as multi-floor residences, stores, and other private indoor spaces. HM3D surpasses existing datasets available for academic research in terms of physical scale, completeness of the reconstruction, and visual fidelity. HM3D contains 112.5k m^2 of navigable space, which is 1.4 - 3.7x larger than other building-scale datasets such as MP3D and Gibson. When compared to existing photorealistic 3D datasets such as Replica, MP3D, Gibson, and ScanNet, images rendered from HM3D have 20 - 85% higher visual fidelity w.r.t. counterpart images captured with real cameras, and HM3D meshes have 34 - 91% fewer artifacts due to incomplete surface reconstruction. The increased scale, fidelity, and diversity of HM3D directly impacts the performance of embodied AI agents trained using it. In fact, we find that HM3D is `pareto optimal' in the following sense -- agents trained to perform PointGoal navigation on HM3D achieve the highest performance regardless of whether they are evaluated on HM3D, Gibson, or MP3D. No similar claim can be made about training on other datasets. HM3D-trained PointNav agents achieve 100% performance on Gibson-test dataset, suggesting that it might be time to retire that episode dataset.

hub tools

citation-role summary

dataset 7 baseline 2

citation-polarity summary

representative citing papers

EAGOR: Embodied Reasoning in Omni-direction

cs.RO · 2026-07-07 · conditional · novelty 7.0

EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

SE(2) Navigation Mesh

cs.RO · 2026-07-01 · unverdicted · novelty 7.0

SE(2) NavMesh adds yaw-dependent traversability layers to polygonal navigation meshes, paired with ASA hierarchical pathfinding and incremental point-cloud updates, claiming over 50% more traversable area than standard navmeshes.

Scaling Diverse Language Generation for 3D Visual Grounding

cs.CL · 2026-06-18 · unverdicted · novelty 7.0

ViGiL3D++ generates diverse 3D visual grounding queries via scene-graph constraint sampling plus LLM language generation, yielding higher diversity than prior datasets and improved benchmark performance.

Open-World Video Segmentation

cs.CV · 2026-06-14 · unverdicted · novelty 7.0

Savvy is a practical system for zero-shot open-world long-horizon video segmentation paired with the OGA granularity-aware evaluation protocol that uses n:1 matching and sever-point detection.

PInVerify: An Offline Embodied Benchmark for Active Instance Verification

cs.CV · 2026-05-28 · unverdicted · novelty 7.0

PInVerify is a new offline embodied benchmark for active instance verification that supplies multi-view captures and 6-sector navigation topology, with MLLM baselines reaching 85.6% after fine-tuning but showing no reliable benefit from tested next-best-view strategies.

Semantic Area Graph Reasoning for Multi-Robot Language-Guided Search

cs.RO · 2026-04-17 · unverdicted · novelty 7.0

SAGR builds a semantic area graph from occupancy maps so LLMs can assign rooms to robots for language-guided search, staying competitive with standard exploration while improving semantic target finding by up to 18.8% in large environments.

UniDAC: Universal Metric Depth Estimation for Any Camera

cs.CV · 2026-03-28 · unverdicted · novelty 7.0

UniDAC achieves universal metric depth estimation across camera types by decoupling relative depth prediction from spatially varying scale estimation using a depth-guided module and distortion-aware positional embedding.

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

cs.CV · 2026-03-19 · unverdicted · novelty 7.0

VGGT-360 delivers geometry-consistent zero-shot panoramic depth by converting panoramas into multi-view 3D reconstructions via VGGT models and three plug-and-play correction modules, then reprojecting the result.

POMA-3D: The Point Map Way to 3D Scene Understanding

cs.CV · 2025-11-20 · unverdicted · novelty 7.0

POMA-3D learns self-supervised 3D scene representations from point maps and improves performance on geometric 3D tasks including navigation and scene retrieval.

Learning Interactive Real-World Simulators

cs.AI · 2023-10-09 · conditional · novelty 7.0

UniSim learns a universal real-world simulator from orchestrated diverse datasets, enabling zero-shot deployment of policies trained purely in simulation.

SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion

cs.CV · 2026-07-06 · conditional · novelty 6.0

SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.

Robots Ask the Way: Communication-Enabled Social Navigation

cs.RO · 2026-07-01 · unverdicted · novelty 6.0

Adding a communication module to social navigation models improves episode success by 10 percentage points in simulated multi-human environments and remains robust to natural language inputs.

Agentic Collaborative Cognition for Zero-Shot 3D Understanding

cs.CV · 2026-06-23 · unverdicted · novelty 6.0 · 2 refs

A collaborative Planning-Perception agent framework using MLLMs constructs a holistic cognitive map through iterative viewpoint supplementation and achieves reported SOTA gains on six 3D benchmarks.

Vesta: A Generalist Embodied Reasoning Model

cs.RO · 2026-06-18 · unverdicted · novelty 6.0

Vesta is a unified embodied generalist model that outperforms specialist baselines by over 20% on average and improves real-world robotic task success by over 35%.

citing papers explorer

Showing 50 of 71 citing papers.