Pith. sign in

REVIEW 4 cited by

Cacophony: An Improved Contrastive Audio-Text Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06986 v3 pith:TWURRAHS submitted 2024-02-10 cs.SD eess.AS

classification cs.SDeess.AS
keywords audio-textaudiomodelcontrastivemodelscacophonycaptioningdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite recent advancements, audio-text models still lag behind their image-text counterparts in scale and performance. In this paper, we propose to improve both the data scale and the training procedure of audio-text contrastive models. Specifically, we craft a large-scale audio-text dataset containing 13,000 hours of text-labeled audio, using pretrained language models to process noisy text descriptions and automatic captioning to obtain text descriptions for unlabeled audio samples. We first train on audio-only data with a masked autoencoder (MAE) objective, which allows us to benefit from the scalability of unlabeled audio datasets. We then train a contrastive model with an auxiliary captioning objective with the audio encoder initialized from the MAE model. Our final model, which we name Cacophony, achieves state-of-the-art performance on audio-text retrieval tasks, and exhibits competitive results on the HEAR benchmark and other downstream tasks such as zero-shot classification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Trained connectors and audio-only gated adapters integrate audio into a frozen vision-language embedding space, preserving base outputs bit-exactly and yielding emergent audio-image retrieval.

  2. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  3. Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders

    cs.SD 2025-01 conditional novelty 6.0 of 10

    Decoding pretrained piano-roll encoder embeddings with three hierarchical language models improves onset-offset-velocity F1 by 0.010 to 0.022 over roll outputs on Maestro.

  4. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

Pith tools