PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.
Second: Sparsely embed- ded convolutional detection.Sensors, 18(10):3337
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3years
2026 3roles
background 2polarities
background 2representative citing papers
UniD-Shift decomposes 2D and 3D features into shared semantic and private modality-specific subspaces to enable unified semantic segmentation with improved accuracy and cross-domain generalization on SemanticKITTI and nuScenes.
GEM is a new LiDAR world model using deformable Mamba that disentangles dynamic and static features to generate high-fidelity simulations and achieve state-of-the-art results on autonomous driving benchmarks.
citing papers explorer
-
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.
-
UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition
UniD-Shift decomposes 2D and 3D features into shared semantic and private modality-specific subspaces to enable unified semantic segmentation with improved accuracy and cross-domain generalization on SemanticKITTI and nuScenes.
-
GEM: Generating LiDAR World Model via Deformable Mamba
GEM is a new LiDAR world model using deformable Mamba that disentangles dynamic and static features to generate high-fidelity simulations and achieve state-of-the-art results on autonomous driving benchmarks.