Encoding cameras as pixel-aligned raxels lets one video diffusion model jointly denoise video and trajectories, supporting pose estimation, controlled generation, and joint synthesis.
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation
5 Pith papers cite this work, alongside 160 external citations. Polarity classification is still indexing.
years
2026 5representative citing papers
Pixal3D performs pixel-aligned 3D generation from images via back-projected multi-scale feature volumes, achieving fidelity close to reconstruction while supporting multi-view and scene synthesis.
Dense initialization of 3DGS does not consistently beat sparse SfM initialization for standard novel views, but improves off-trajectory generalization; no densification method wins everywhere.
A distilled student policy using monocular depth estimation from cameras outperforms a 2D LiDAR teacher policy in navigating complex 3D obstacles while running fully onboard a Jetson Orin.
LinStereo uses Position-Aware Linear Attention, Hierarchical Semantic Cost Volumes, and Depth Prior Initialization to enable global aggregation in iterative stereo matching at linear complexity, showing improved performance on standard and underwater benchmarks.
citing papers explorer
-
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
Encoding cameras as pixel-aligned raxels lets one video diffusion model jointly denoise video and trajectories, supporting pose estimation, controlled generation, and joint synthesis.
-
Pixal3D: Pixel-Aligned 3D Generation from Images
Pixal3D performs pixel-aligned 3D generation from images via back-projected multi-scale feature volumes, achieving fidelity close to reconstruction while supporting multi-view and scene synthesis.
-
The Role of Initialization in 3D Gaussian Splatting
Dense initialization of 3DGS does not consistently beat sparse SfM initialization for standard novel views, but improves off-trajectory generalization; no densification method wins everywhere.
-
Learning Vision-Based Omnidirectional Navigation: A Teacher-Student Approach Using Monocular Depth Estimation
A distilled student policy using monocular depth estimation from cameras outperforms a 2D LiDAR teacher policy in navigating complex 3D obstacles while running fully onboard a Jetson Orin.
-
LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching
LinStereo uses Position-Aware Linear Attention, Hierarchical Semantic Cost Volumes, and Depth Prior Initialization to enable global aggregation in iterative stereo matching at linear complexity, showing improved performance on standard and underwater benchmarks.