Pith. sign in

AeDet: Azimuth-invariant Multi-view 3D Object Detection

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Recent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization. To preserve the inherent property of the BEV features and ease the optimization, we propose an azimuth-equivariant convolution (AeConv) and an azimuth-equivariant anchor. The sampling grid of AeConv is always in the radial direction, thus it can learn azimuth-invariant BEV features. The proposed anchor enables the detection head to learn predicting azimuth-irrelevant targets. In addition, we introduce a camera-decoupled virtual depth to unify the depth prediction for the images with different camera intrinsic parameters. The resultant detector is dubbed Azimuth-equivariant Detector (AeDet). Extensive experiments are conducted on nuScenes, and AeDet achieves a 62.0% NDS, surpassing the recent multi-view 3D object detectors such as PETRv2 and BEVDepth by a large margin. Project page: https://fcjian.github.io/aedet.

citation-role summary

other 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

other 1

polarities

unclear 1

representative citing papers

Generalizing Monocular 3D Object Detection

cs.CV · 2025-08-27 · conditional · novelty 7.0

A dissertation that improves monocular 3D object detection across occlusions, datasets, object sizes, and camera heights via four complementary techniques, validated on KITTI, Waymo, nuScenes, and CARLA.

citing papers explorer

Showing 1 of 1 citing paper.

  • Generalizing Monocular 3D Object Detection cs.CV · 2025-08-27 · conditional · none · ref 59 · internal anchor

    A dissertation that improves monocular 3D object detection across occlusions, datasets, object sizes, and camera heights via four complementary techniques, validated on KITTI, Waymo, nuScenes, and CARLA.