Pith. sign in

REVIEW 6 cited by

MuLan: A Joint Embedding of Music Audio and Natural Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12415 v1 pith:MBSTGMMV submitted 2022-08-26 eess.AS cs.CLcs.SDstat.ML

classification eess.AScs.CLcs.SDstat.ML
keywords musicmulanlanguagetextaudioaudio-textembeddingincluding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of acoustic models that link music audio directly to unconstrained natural language music descriptions. MuLan takes the form of a two-tower, joint audio-text embedding model trained using 44 million music recordings (370K hours) and weakly-associated, free-form text annotations. Through its compatibility with a wide range of music genres and text styles (including conventional music tags), the resulting audio-text representation subsumes existing ontologies while graduating to true zero-shot functionalities. We demonstrate the versatility of the MuLan embeddings with a range of experiments including transfer learning, zero-shot music tagging, language understanding in the music domain, and cross-modal retrieval applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A zero-shot instrument cloning system feeds raw reference audio into a flow-matching DiT and uses asymmetric CFG to keep melody and timbre control separate.

  2. DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

    eess.AS 2025-07 conditional novelty 6.0 of 10

    DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.

  3. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  4. CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

    eess.AS 2025-06 conditional novelty 6.0 of 10

    CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.

  5. An introduction to pitch strength in contemporary popular music analysis and production

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Pitch strength is proposed as a variable, structurally relevant perceptual parameter of contemporary popular music that generative music models should expose as a low-level control.

  6. PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

    cs.SD 2025-09 conditional novelty 4.0 of 10

    PianoBind, a trimodal audio-MIDI-text embedding model trained on piano data, beats general-purpose music embedding models on pop-piano text-to-music retrieval benchmarks.

Pith tools