Using a pretrained metric reconstruction model as the detector encoder, with an up-to-scale 3D box head scaled by the model's predicted scale factor, gives stronger online monocular 3D detection and transfer than 2D-to-3D lifting baselines.
IEEE TPAMI45(2), 1992–2008 (2022)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
Using a pretrained metric reconstruction model as the detector encoder, with an up-to-scale 3D box head scaled by the model's predicted scale factor, gives stronger online monocular 3D detection and transfer than 2D-to-3D lifting baselines.