Pith. sign in

REVIEW

Unsupervised Improvement of Audio-Text Cross-Modal Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.01864 v3 pith:SLR5KDPV submitted 2023-05-03 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audio-textclassificationrepresentationsapproachescross-modalcurationdomain-specificimprove
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in using language models to obtain cross-modal audio-text representations have overcome the limitations of conventional training approaches that use predefined labels. This has allowed the community to make progress in tasks like zero-shot classification, which would otherwise not be possible. However, learning such representations requires a large amount of human-annotated audio-text pairs. In this paper, we study unsupervised approaches to improve the learning framework of such representations with unpaired text and audio. We explore domain-unspecific and domain-specific curation methods to create audio-text pairs that we use to further improve the model. We also show that when domain-specific curation is used in conjunction with a soft-labeled contrastive loss, we are able to obtain significant improvement in terms of zero-shot classification performance on downstream sound event classification or acoustic scene classification tasks.

Discussion (0). Continue with ORCID to comment.

Pith tools