A supervised audio tokenizer trained with audio tagging labels nearly matches continuous BEATs features for automated audio captioning on Clotho, while conventional discrete tokens degrade performance.
Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Discrete Audio Representations for Automated Audio Captioning
A supervised audio tokenizer trained with audio tagging labels nearly matches continuous BEATs features for automated audio captioning on Clotho, while conventional discrete tokens degrade performance.