MonoMobility recovers motion parts and their 3D axes from a monocular video by initializing with depth, flow, and point-cloud registration, then refining via 2D Gaussian splatting optimization.
CAPT: Category-level Articulation Estimation from a Single Point Cloud Using Transformer
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The ability to estimate joint parameters is essential for various applications in robotics and computer vision. In this paper, we propose CAPT: category-level articulation estimation from a point cloud using Transformer. CAPT uses an end-to-end transformer-based architecture for joint parameter and state estimation of articulated objects from a single point cloud. The proposed CAPT methods accurately estimate joint parameters and states for various articulated objects with high precision and robustness. The paper also introduces a motion loss approach, which improves articulation estimation performance by emphasizing the dynamic features of articulated objects. Additionally, the paper presents a double voting strategy to provide the framework with coarse-to-fine parameter estimation. Experimental results on several category datasets demonstrate that our methods outperform existing alternatives for articulation estimation. Our research provides a promising solution for applying Transformer-based architectures in articulated object analysis.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
MonoMobility: Zero-Shot 3D Mobility Analysis from Monocular Videos
MonoMobility recovers motion parts and their 3D axes from a monocular video by initializing with depth, flow, and point-cloud registration, then refining via 2D Gaussian splatting optimization.