UnderOneFacade is a large-scale cross-continent 3D facade point cloud benchmark with harmonized labels that reveals existing segmentation models achieve at most 33 IoU on fine-grained architectural elements and degrade across geographic domains.
In: Proceedings of the IEEE conference on computer vision and pattern recognition
8 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 8roles
dataset 2polarities
use dataset 2representative citing papers
Adaptive manifold keyframe sampling plus an instruction-pose-aware geometry MoE raises sparse-RGB 3D spatial reasoning to 63.5 average on VSI-Bench, beating strong baselines by 7.8 points.
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.
A query-based attack reconstructs substantial 3D geometry and approximate appearance from scene coordinate regression models, contradicting prior privacy-preserving claims.
FluSplat trains a model with geometric alignment constraints on multi-view edits to produce consistent 3D scene edits from sparse views in a single forward pass without test-time optimization.
ReplicateAnyScene performs fully automated zero-shot video-to-compositional-3D reconstruction by cascading alignments of generic priors from vision foundation models across textual, visual, and spatial dimensions.
Motion-MLLM integrates IMU egomotion data into MLLMs using cascaded filtering and asymmetric fusion to ground visual content in physical trajectories for scale-aware 3D understanding, achieving competitive accuracy at higher speed.
OpenSpatial supplies a principled open-source data engine and 3-million-sample dataset that raises spatial-reasoning model performance by an average of 19 percent on benchmarks.
citing papers explorer
-
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
UnderOneFacade is a large-scale cross-continent 3D facade point cloud benchmark with harmonized labels that reveals existing segmentation models achieve at most 33 IoU on fine-grained architectural elements and degrade across geographic domains.
-
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Adaptive manifold keyframe sampling plus an instruction-pose-aware geometry MoE raises sparse-RGB 3D spatial reasoning to 63.5 average on VSI-Bench, beating strong baselines by 7.8 points.
-
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
ICDepth adapts text-to-video diffusion transformers for video depth estimation via in-context conditioning, achieving SOTA results on benchmarks with 6-13x less training data than prior generative methods.
-
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
A query-based attack reconstructs substantial 3D geometry and approximate appearance from scene coordinate regression models, contradicting prior privacy-preserving claims.
-
FluSplat: Sparse-View 3D Editing without Test-Time Optimization
FluSplat trains a model with geometric alignment constraints on multi-view edits to produce consistent 3D scene edits from sparse views in a single forward pass without test-time optimization.
-
ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment
ReplicateAnyScene performs fully automated zero-shot video-to-compositional-3D reconstruction by cascading alignments of generic priors from vision foundation models across textual, visual, and spatial dimensions.
-
Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding
Motion-MLLM integrates IMU egomotion data into MLLMs using cascaded filtering and asymmetric fusion to ground visual content in physical trajectories for scale-aware 3D understanding, achieving competitive accuracy at higher speed.
-
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
OpenSpatial supplies a principled open-source data engine and 3-million-sample dataset that raises spatial-reasoning model performance by an average of 19 percent on benchmarks.