Pith. sign in

Discrete Audio Representations for Automated Audio Captioning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Discrete Audio Representations for Automated Audio Captioning

cs.SD · 2025-05-21 · conditional · novelty 5.0

A supervised audio tokenizer trained with audio tagging labels nearly matches continuous BEATs features for automated audio captioning on Clotho, while conventional discrete tokens degrade performance.

citing papers explorer

Showing 1 of 1 citing paper.

  • Discrete Audio Representations for Automated Audio Captioning cs.SD · 2025-05-21 · conditional · none · ref 1 · internal anchor

    A supervised audio tokenizer trained with audio tagging labels nearly matches continuous BEATs features for automated audio captioning on Clotho, while conventional discrete tokens degrade performance.