Pith. sign in

REVIEW 6 cited by

Natural Language Supervision for General-Purpose Audio Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05767 v2 pith:NYGTGKWJ submitted 2023-09-11 cs.SD eess.AS

classification cs.SDeess.AS
keywords representationsaudiomodelstasksencoderslanguagelearnmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio-Language models jointly learn multimodal text and audio representations that enable Zero-Shot inference. Models rely on the encoders to create powerful representations of the input and generalize to multiple tasks ranging from sounds, music, and speech. Although models have achieved remarkable performance, there is still a performance gap with task-specific models. In this paper, we propose a Contrastive Language-Audio Pretraining model that is pretrained with a diverse collection of 4.6M audio-text pairs employing two innovative encoders for Zero-Shot inference. To learn audio representations, we trained an audio encoder on 22 audio tasks, instead of the standard training of sound event classification. To learn language representations, we trained an autoregressive decoder-only model instead of the standard encoder-only models. Then, the audio and language representations are brought into a joint multimodal space using Contrastive Learning. We used our encoders to improve the downstream performance by a margin. We extensively evaluated the generalization of our representations on 26 downstream tasks, the largest in the literature. Our model achieves state of the art results in several tasks leading the way towards general-purpose audio representations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reasoning-Aware Multimodal Fusion for Hateful Video Detection

    cs.CV 2025-12 conditional novelty 6.0 of 10

    RAMF's three-stage adversarial VLM reasoning plus local-global/cross-head attention fusion improves hateful video classification on HateMM and MultiHateClip.

  2. TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    TTA-Bench offers a seven-dimension, 2,999-prompt evaluation of ten text-to-audio models with 118,000 human ratings, covering quality, robustness, fairness, bias, and toxicity.

  3. LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.

  4. Assessing the Alignment of Audio Representations with Timbre Similarity Ratings

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Style embeddings computed from CLAP and a sound matching model align with human timbre similarity ratings better than plain embeddings, MFCC, and other pre-trained representations.

  5. Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

    cs.IR 2025-06 conditional novelty 6.0 of 10

    Vela adapts an audio MLLM into a universal text-audio embedding model using 'in one word' prompts, in-context examples, and text-only contrastive training, outperforming CLAP-style models on retrieval benchmarks.

  6. Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual ...

Pith tools