REVIEW 5 cited by
ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present ScanNet++, a large-scale dataset that couples together capture of high-quality and commodity-level geometry and color of indoor scenes. Each scene is captured with a high-end laser scanner at sub-millimeter resolution, along with registered 33-megapixel images from a DSLR camera, and RGB-D streams from an iPhone. Scene reconstructions are further annotated with an open vocabulary of semantics, with label-ambiguous scenarios explicitly annotated for comprehensive semantic understanding. ScanNet++ enables a new real-world benchmark for novel view synthesis, both from high-quality RGB capture, and importantly also from commodity-level images, in addition to a new benchmark for 3D semantic scene understanding that comprehensively encapsulates diverse and ambiguous semantic labeling scenarios. Currently, ScanNet++ contains 460 scenes, 280,000 captured DSLR images, and over 3.7M iPhone RGBD frames.
Forward citations
Cited by 5 Pith papers
-
Vision as Unified Multimodal Generation
A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
-
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
Multi-view relational distillation transfers geometric knowledge to vision-language models by matching cross-view patch similarity matrices, improving spatial reasoning with minimal overhead.
-
Discovering and using Spelke segments
SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.
-
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding
ViGiL3D is a 350-prompt diagnostic dataset showing that existing 3D visual grounding models lose 20 or more points on linguistically diverse prompts compared to ScanRefer.
-
THUD++: Large-Scale Dynamic Indoor Scene Dataset and Benchmark for Mobile Robots
THUD++ is a 13-scene dynamic indoor RGB-D and trajectory dataset with benchmarks showing existing algorithms struggle in crowded mobile-robot environments.
Discussion (0). Continue with ORCID to comment.