Fusing depth features into Mask2Former with dynamic weighting, plus location-aware and time-aware queries, improves image and video panoptic segmentation on Cityscapes while avoiding video-specific losses.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LiDAR-Camera Fusion for Video Panoptic Segmentation without Video Training
Fusing depth features into Mask2Former with dynamic weighting, plus location-aware and time-aware queries, improves image and video panoptic segmentation on Cityscapes while avoiding video-specific losses.