REVIEW 12 cited by
MuLan: A Joint Embedding of Music Audio and Natural Language
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of acoustic models that link music audio directly to unconstrained natural language music descriptions. MuLan takes the form of a two-tower, joint audio-text embedding model trained using 44 million music recordings (370K hours) and weakly-associated, free-form text annotations. Through its compatibility with a wide range of music genres and text styles (including conventional music tags), the resulting audio-text representation subsumes existing ontologies while graduating to true zero-shot functionalities. We demonstrate the versatility of the MuLan embeddings with a range of experiments including transfer learning, zero-shot music tagging, language understanding in the music domain, and cross-modal retrieval applications.
Forward citations
Cited by 12 Pith papers
-
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
RAG-Audio starts frozen audio generators from a retrieved exemplar of the fMRI-decoded CLAP embedding, raising 10-way stimulus identification from 0.14-0.18 to 0.40-0.43 on Brain2Music and cutting FAD by about 10x.
-
Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model
Diff-Symbo generates long, text-controlled symbolic music by autoregressively extending 8-bar latent diffusion segments conditioned on the previous segment's latent.
-
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
A zero-shot instrument cloning system feeds raw reference audio into a flow-matching DiT and uses asymmetric CFG to keep melody and timbre control separate.
-
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.
-
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
GVMGen generates background music from video using spatial and temporal cross-attention to condition a MusicGen decoder, reporting state-of-the-art correspondence and diversity.
-
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
MuQ, trained with masked prediction of Mel-RVQ tokens, beats MERT and MusicFM on the MARBLE average despite a much smaller pre-training set.
-
An introduction to pitch strength in contemporary popular music analysis and production
Pitch strength is proposed as a variable, structurally relevant perceptual parameter of contemporary popular music that generative music models should expose as a low-level control.
-
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
PianoBind, a trimodal audio-MIDI-text embedding model trained on piano data, beats general-purpose music embedding models on pop-piano text-to-music retrieval benchmarks.
-
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.
-
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.
Discussion (0). Continue with ORCID to comment.