Pith. sign in

REVIEW 12 cited by

MuLan: A Joint Embedding of Music Audio and Natural Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12415 v1 pith:MBSTGMMV submitted 2022-08-26 eess.AS cs.CLcs.SDstat.ML

classification eess.AScs.CLcs.SDstat.ML
keywords musicmulanlanguagetextaudioaudio-textembeddingincluding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of acoustic models that link music audio directly to unconstrained natural language music descriptions. MuLan takes the form of a two-tower, joint audio-text embedding model trained using 44 million music recordings (370K hours) and weakly-associated, free-form text annotations. Through its compatibility with a wide range of music genres and text styles (including conventional music tags), the resulting audio-text representation subsumes existing ontologies while graduating to true zero-shot functionalities. We demonstrate the versatility of the MuLan embeddings with a range of experiments including transfer learning, zero-shot music tagging, language understanding in the music domain, and cross-modal retrieval applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

    cs.SD 2026-08 conditional novelty 6.0 of 10

    RAG-Audio starts frozen audio generators from a retrieved exemplar of the fMRI-decoded CLAP embedding, raising 10-way stimulus identification from 0.14-0.18 to 0.40-0.43 on Brain2Music and cutting FAD by about 10x.

  2. Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

    cs.SD 2026-08 conditional novelty 6.0 of 10

    Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.

  3. Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A zero-shot instrument cloning system feeds raw reference audio into a flow-matching DiT and uses asymmetric CFG to keep melody and timbre control separate.

  4. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

  5. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  6. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

    eess.AS 2025-06 conditional novelty 6.0 of 10

    CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.

  7. GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    GVMGen generates background music from video using spatial and temporal cross-attention to condition a MusicGen decoder, reporting state-of-the-art correspondence and diversity.

  8. MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

    cs.SD 2025-01 conditional novelty 6.0 of 10

    MuQ, trained with masked prediction of Mel-RVQ tokens, beats MERT and MusicFM on the MARBLE average despite a much smaller pre-training set.

  9. An introduction to pitch strength in contemporary popular music analysis and production

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Pitch strength is proposed as a variable, structurally relevant perceptual parameter of contemporary popular music that generative music models should expose as a low-level control.

  10. PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

    cs.SD 2025-09 conditional novelty 4.0 of 10

    PianoBind, a trimodal audio-MIDI-text embedding model trained on piano data, beats general-purpose music embedding models on pop-piano text-to-music retrieval benchmarks.

  11. Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

    cs.CV 2024-11 unverdicted novelty 3.0 of 10

    A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.

  12. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools