Pith. sign in

hub

Joint 2D-3D-Semantic Data for Indoor Scene Understanding

21 Pith papers cite this work, alongside 686 external citations. Polarity classification is still indexing.

21 Pith papers citing it
686 external citations · Pith
abstract

We present a dataset of large-scale indoor spaces that provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations. The dataset covers over 6,000m2 and contains over 70,000 RGB images, along with the corresponding depths, surface normals, semantic annotations, global XYZ images (all in forms of both regular and 360{\deg} equirectangular images) as well as camera information. It also includes registered raw and semantically annotated 3D meshes and point clouds. The dataset enables development of joint and cross-modal learning models and potentially unsupervised approaches utilizing the regularities present in large-scale indoor spaces. The dataset is available here: http://3Dsemantics.stanford.edu/

hub tools

citation-role summary

dataset 3 background 1

citation-polarity summary

representative citing papers

VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation

cs.CV · 2026-03-19 · unverdicted · novelty 7.0

VGGT-360 delivers geometry-consistent zero-shot panoramic depth by converting panoramas into multi-view 3D reconstructions via VGGT models and three plug-and-play correction modules, then reprojecting the result.

Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

cs.CV · 2026-06-29 · accept · novelty 6.0 · 2 refs

Argus plus Realsee3D deliver state-of-the-art metric camera pose, depth, and point-cloud reconstruction from unordered indoor panoramas via learned covisibility anchoring and geometric factorization.

PointCaM: Cut-and-Mix for Open-Set Point Cloud Learning

cs.CV · 2022-12-05 · unverdicted · novelty 6.0

PointCaM proposes a cut-and-mix mechanism with an Unknown-Point Simulator and Estimator to improve open-set recognition on point clouds by simulating out-of-distribution data and using multi-level features.

Ordinal Neural Collapse as a Representation Prior for Visual Navigation

cs.RO · 2026-06-25 · unverdicted · novelty 5.0

ORION applies ordinal neural collapse to organize visual encoder features along an action-ordinal axis, then integrates the encoder into a diffusion navigation policy, yielding higher success rates than end-to-end and standard neural collapse baselines in simulation and real-world tests.

citing papers explorer

Showing 21 of 21 citing papers.