A multimodal multi-task architecture combining Mamba-based temporal-spatial features and task-specific gating achieves state-of-the-art accuracy on the AIDE assistive-driving benchmark at real-time speed.
Is space-time attention all you need for video understanding? InInternational Conference on Machine Learnin (ICML), page 4, 2021
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
TEM^3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving
A multimodal multi-task architecture combining Mamba-based temporal-spatial features and task-specific gating achieves state-of-the-art accuracy on the AIDE assistive-driving benchmark at real-time speed.