Y-MAP-Net distills several large vision models into a single real-time convolutional network that predicts depth, normals, pose, segmentation, and captions from RGB.
Multimodal fu- sion via teacher-student network for indoor action recogni- tion
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
Y-MAP-Net distills several large vision models into a single real-time convolutional network that predicts depth, normals, pose, segmentation, and captions from RGB.