Pith. sign in

REVIEW 1 cited by

Weakly-supervised Automated Audio Captioning via text only training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12242 v1 pith:QVQDJOKU submitted 2023-09-21 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiotextclappaireddataduringembeddingstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (AAC). However, it is labor-intensive and time-consuming to collect a sufficient number of paired audio and captions. Motivated by the recent advances in Contrastive Language-Audio Pretraining (CLAP), we propose a weakly-supervised approach to train an AAC model assuming only text data and a pre-trained CLAP model, alleviating the need for paired target data. Our approach leverages the similarity between audio and text embeddings in CLAP. During training, we learn to reconstruct the text from the CLAP text embedding, and during inference, we decode using the audio embeddings. To mitigate the modality gap between the audio and text embeddings we employ strategies to bridge the gap during training and inference stages. We evaluate our proposed method on Clotho and AudioCaps datasets demonstrating its ability to achieve a relative performance of up to ~$83\%$ compared to fully supervised approaches trained with paired target data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A background-profile subtraction method improves zero-shot sound classification accuracy under noisy conditions, and narrowing the audio-text modality gap further boosts performance.

Pith tools