VESPA fuses LiDAR geometry with vision-language model semantics to generate open-vocabulary 3D pseudolabels, achieving 52.95% class-agnostic AP and 46.54% 3-class mAP on nuScenes without human supervision.
Liso: Lidar-only self-supervised 3d object detec- tion
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
VESPA: Towards un(Human)supervised Open-World Pointcloud Labeling for Autonomous Driving
VESPA fuses LiDAR geometry with vision-language model semantics to generate open-vocabulary 3D pseudolabels, achieving 52.95% class-agnostic AP and 46.54% 3-class mAP on nuScenes without human supervision.