REVIEW 3 cited by
Adapting Language-Audio Models as Few-Shot Audio Learners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We presented the Treff adapter, a training-efficient adapter for CLAP, to boost zero-shot classification performance by making use of a small set of labelled data. Specifically, we designed CALM to retrieve the probability distribution of text-audio clips over classes using a set of audio-label pairs and combined it with CLAP's zero-shot classification results. Furthermore, we designed a training-free version of the Treff adapter by using CALM as a cosine similarity measure. Experiments showed that the proposed Treff adapter is comparable and even better than fully-supervised methods and adaptation methods in low-shot and data-abundant scenarios. While the Treff adapter shows that combining large-scale pretraining and rapid learning of domain-specific knowledge is non-trivial for obtaining generic representations for few-shot learning, it is still limited to audio classification tasks. In the future, we will explore how to use audio-language models in diverse audio domains.
Forward citations
Cited by 3 Pith papers
-
CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
CLAP-S and CLAP-S+ adapt CLAP models to fiber-optic acoustic recognition by combining support-set retrieval with a fine-tuned adapter, reporting improved few-shot classification accuracy over existing methods.
-
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
Task-specific prompt ensembling with GPT-4-generated attributes and sources improves some zero-shot audio classification datasets while degrading others.
-
Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio
A consistency-guided test-time prompt adaptation method improves CLAP zero-shot audio classification by 4.41% relative on average over DA CLAP across 12 datasets.
Discussion (0). Continue with ORCID to comment.