REVIEW 7 cited by
The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural language descriptions of 706 music recordings, all publicly accessible and released under Creative Common licenses. To showcase the use of our dataset, we benchmark popular models on three key music-and-language tasks (music captioning, text-to-music generation and music-language retrieval). Our experiments highlight the importance of cross-dataset evaluation and offer insights into how researchers can use SDD to gain a broader understanding of model performance.
Forward citations
Cited by 7 Pith papers
-
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
SonicWeave routes chunks of audio through specialized experts, using a learned gate between text-derived prior and local evidence, improving compositional fidelity in unified text-to-audio scene generation.
-
From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music
Prompts for AI music are dominated by genre and story terms, but genre words survive into perception while story-heavy prompts produce the largest semantic mismatch.
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
CMI-Bench converts standard MIR annotations into instruction-following tasks and shows current audio-text LLMs underperform supervised MIR systems across nearly all 14 tasks.
-
Qwen-Audio-3.0-Gen-Preview Technical Report
One diffusion-transformer system with a shared audio codec generates standalone audio, multi-speaker dialogue, and time-structured mixed scenes, with strongest measured advantages in speaker similarity, cross-turn con...
-
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
A contrastive learning framework (CLaMP 3) aligns three music modalities with multilingual text, enabling text-to-music retrieval, cross-lingual retrieval for unseen languages, and emergent cross-modal retrieval.
-
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
A dataset of 362,000 Jamendo instrumental tracks pairs each song with a generated caption and imputed metadata fields, created with a retrieval-based local-LLM pipeline.
-
Improving Controllability and Editability for Pretrained Text-to-Music Generation Models
A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.
Discussion (0). Continue with ORCID to comment.