MMHNet enables video-to-audio models trained on short clips to generalize and generate audio for videos over 5 minutes long.
Vssd: Vision mamba with non-causal state space duality
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
RegNetMamba-2 integrates SSD into CMI and MSF modules for shared structural feature extraction and fusion across scales, reporting improved performance and efficiency versus prior deep learning methods on VIS-SAR, VIS-IR, and VIS-NIR datasets.
citing papers explorer
-
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
MMHNet enables video-to-audio models trained on short clips to generalize and generate audio for videos over 5 minutes long.
-
Cross-Modality Feature Fusion Based on Structured State Space Duality for Multimodal Image Registration Network
RegNetMamba-2 integrates SSD into CMI and MSF modules for shared structural feature extraction and fusion across scales, reporting improved performance and efficiency versus prior deep learning methods on VIS-SAR, VIS-IR, and VIS-NIR datasets.