A BEVFusion-based offboard tracker with density-aware loss weighting, nearest-neighbor relationship targets, and high-resolution sparse features doubles MOTA on a new crowded-pedestrian benchmark (PCP-MV) from 0.172 to 0.353.
Auto4D: Learning to Label 4D Objects from Sequential Point Clouds
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In the past few years we have seen great advances in object perception (particularly in 4D space-time dimensions) thanks to deep learning methods. However, they typically rely on large amounts of high-quality labels to achieve good performance, which often require time-consuming and expensive work by human annotators. To address this we propose an automatic annotation pipeline that generates accurate object trajectories in 3D space (i.e., 4D labels) from LiDAR point clouds. The key idea is to decompose the 4D object label into two parts: the object size in 3D that's fixed through time for rigid objects, and the motion path describing the evolution of the object's pose through time. Instead of generating a series of labels in one shot, we adopt an iterative refinement process where online generated object detections are tracked through time as the initialization. Given the cheap but noisy input, our model produces higher quality 4D labels by re-estimating the object size and smoothing the motion path, where the improvement is achieved by exploiting aggregated observations and motion cues over the entire trajectory. We validate the proposed method on a large-scale driving dataset and show a 25% reduction of human annotation efforts. We also showcase the benefits of our approach in the annotator-in-the-loop setting.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
A BEVFusion-based offboard tracker with density-aware loss weighting, nearest-neighbor relationship targets, and high-resolution sparse features doubles MOTA on a new crowded-pedestrian benchmark (PCP-MV) from 0.172 to 0.353.