SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior self-supervised baselines.
In: European conference on computer vision
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3years
2026 3roles
dataset 1polarities
use dataset 1representative citing papers
A 6.1M-parameter monocular depth network, distilled from Depth Anything v2-Large over 14.1M multi-domain images, achieves the best zero-shot accuracy–efficiency trade-off among lightweight models across five benchmarks while running in real time on devices from server GPUs to smartphones.
FoundDP integrates DP-derived metric depth with ViT-based structural priors from monocular models, using feature alignment to mitigate defocus blur and improve depth in low-observability areas.
citing papers explorer
-
SS3D: End2End Self-Supervised 3D from Web Videos
SS3D pretrains an end-to-end feed-forward 3D estimator on filtered YouTube-8M videos via SfM self-supervision, MVS filtering, and expert distillation, delivering stronger zero-shot transfer and fine-tuning than prior self-supervised baselines.
-
ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
A 6.1M-parameter monocular depth network, distilled from Depth Anything v2-Large over 14.1M multi-domain images, achieves the best zero-shot accuracy–efficiency trade-off among lightweight models across five benchmarks while running in real time on devices from server GPUs to smartphones.
-
FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
FoundDP integrates DP-derived metric depth with ViT-based structural priors from monocular models, using feature alignment to mitigate defocus blur and improve depth in low-observability areas.