PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.
MoE3D: Mixture of Experts Meets Multi-Modal 3D Understanding.arXiv 2025
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7representative citing papers
SEGA3D improves 3D vision-language segmentation on ScanNet and Matterport3D by operating on fine-grained masks with LLM-assisted selection, claiming gains of 8.3 and 5.3 mIoU over prior top methods.
MATE is a multi-modal MoE trajectory policy using a cosine router and stochastic noise to improve expert balance, reporting 4.75% higher average success rate than prior methods on LIBERO under data scarcity.
LER-YOLO reports 89.7% AP50 on the MBU benchmark for misaligned RGB-IR UAV detection by routing among RGB-dominant, IR-dominant, and fusion experts using a spatial reliability map.
A routing framework maintains three parallel 3D feature streams for LiDAR, 4D radar, and fusion, with a lightweight router using weather prompts to dynamically weight them and auxiliary supervision to keep branches distinct, achieving SOTA on K-Radar.
CMTFormer proposes hierarchical cross-modal modules (SAM, CEM, LDFM) plus a spatial prior to fuse RGB and event streams, outperforming prior detectors on DSEC-Detection and PKU-DAVIS-SOD.
HAS-KD combines information-oriented heterogeneous distillation from multi-modal models with adept snapshot distillation from training checkpoints to reach SOTA 3D semantic segmentation on ScanNetV2 and S3DIS without added inference burden.
citing papers explorer
-
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.
-
Segment and Select: Vision-Language Segmentation in 3D Scenarios
SEGA3D improves 3D vision-language segmentation on ScanNet and Matterport3D by operating on fine-grained masks with LLM-assisted selection, claiming gains of 8.3 and 5.3 mIoU over prior top methods.
-
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
MATE is a multi-modal MoE trajectory policy using a cosine router and stochastic noise to improve expert balance, reporting 4.75% higher average success rate than prior methods on LIBERO under data scarcity.
-
LER-YOLO: Reliability-Aware Expert Routing for Misaligned RGB-Infrared UAV Detection
LER-YOLO reports 89.7% AP50 on the MBU benchmark for misaligned RGB-IR UAV detection by routing among RGB-dominant, IR-dominant, and fusion experts using a spatial reliability map.
-
Weather-Conditioned Branch Routing for Robust LiDAR-Radar 3D Object Detection
A routing framework maintains three parallel 3D feature streams for LiDAR, 4D radar, and fusion, with a lightweight router using weather prompts to dynamically weight them and auxiliary supervision to keep branches distinct, achieving SOTA on K-Radar.
-
CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection
CMTFormer proposes hierarchical cross-modal modules (SAM, CEM, LDFM) plus a spatial prior to fuse RGB and event streams, outperforming prior detectors on DSEC-Detection and PKU-DAVIS-SOD.
-
Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation
HAS-KD combines information-oriented heterogeneous distillation from multi-modal models with adept snapshot distillation from training checkpoints to reach SOTA 3D semantic segmentation on ScanNetV2 and S3DIS without added inference burden.