Pith. sign in

REVIEW 3 cited by

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.10362 v3 pith:ZGGYMS3G submitted 2025-02-14 cs.SD eess.AS

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

classification cs.SD eess.AS
keywords musictextclampgeneralizationmultilingualretrievalacrossaudio
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities--including sheet music, performance signals, and audio recordings--with multilingual text in a shared representation space, enabling retrieval across unaligned modalities with text as a bridge. It features a multilingual text encoder adaptable to unseen languages, exhibiting strong cross-lingual generalization. Leveraging retrieval-augmented generation, we curated M4-RAG, a web-scale dataset consisting of 2.31 million music-text pairs. This dataset is enriched with detailed metadata that represents a wide array of global musical traditions. To advance future research, we release WikiMT-X, a benchmark comprising 1,000 triplets of sheet music, audio, and richly varied text descriptions. Experiments show that CLaMP 3 achieves state-of-the-art performance on multiple MIR tasks, significantly surpassing previous strong baselines and demonstrating excellent generalization in multimodal and multilingual music contexts.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Text2Score: Generating Sheet Music From Textual Prompts

    cs.SD 2026-05 unverdicted novelty 7.0

    A two-stage framework uses an LLM to plan musical structures from text and then generates conditioned ABC notation sheet music, outperforming baselines in expert-validated evaluations.

  2. Text2Score: Generating Sheet Music From Textual Prompts

    cs.SD 2026-05 conditional novelty 6.0

    Text2Score turns text prompts into sheet music by having an LLM produce a bar-wise structural plan and a hierarchical decoder write ABC notation from that plan.

  3. Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking

    cs.MM 2026-07 unverdicted novelty 5.0

    VTMR is a two-stage video-to-music recommender: joint audio-visual-text retrieval of candidates, then temporal-sequence reranking, lifting R@10 to 18.3 and matching commercial preference.