Averaging a small number of audio embeddings per class outperforms zero-shot text-embedding classification for CLAP-based audio classification.
For ESC-50 and FSD50K, results are obtained from the respective papers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improving Audio Classification by Transitioning from Zero- to Few-Shot
Averaging a small number of audio embeddings per class outperforms zero-shot text-embedding classification for CLAP-based audio classification.