A fully state-space audio-language model matches transformer counterparts on audio classification and captioning using a 2.8B parameter backbone.
Pengi: An audio language model for audio tasks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
State-Space Large Audio Language Models
A fully state-space audio-language model matches transformer counterparts on audio classification and captioning using a 2.8B parameter backbone.