Pith. sign in

REVIEW 1 cited by

STS: Surround-view Temporal Stereo for Multi-view 3D Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10145 v1 pith:OUXSHT4E submitted 2022-08-22 cs.CV

classification cs.CV
keywords depthmonoculardetectionlearningstereotemporalaccuratebackbone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning accurate depth is essential to multi-view 3D object detection. Recent approaches mainly learn depth from monocular images, which confront inherent difficulties due to the ill-posed nature of monocular depth learning. Instead of using a sole monocular depth method, in this work, we propose a novel Surround-view Temporal Stereo (STS) technique that leverages the geometry correspondence between frames across time to facilitate accurate depth learning. Specifically, we regard the field of views from all cameras around the ego vehicle as a unified view, namely surroundview, and conduct temporal stereo matching on it. The resulting geometrical correspondence between different frames from STS is utilized and combined with the monocular depth to yield final depth prediction. Comprehensive experiments on nuScenes show that STS greatly boosts 3D detection ability, notably for medium and long distance objects. On BEVDepth with ResNet-50 backbone, STS improves mAP and NDS by 2.6% and 1.4%, respectively. Consistent improvements are observed when using a larger backbone and a larger image resolution, demonstrating its effectiveness

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MambaDETR: Query-based Temporal Modeling using State Space Model for Multi-View 3D Object Detection

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MambaDETR applies a Mamba/SSM sequence model to temporal fusion of 3D detection queries, achieving 50.8 mAP on nuScenes val with 8 frames and linear memory scaling.

Pith tools