MusFlow generates music from images, story texts, or captions by aligning all inputs into the CLAP audio embedding space and sampling with conditional flow matching.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MusFlow: Multimodal Music Generation via Conditional Flow Matching
MusFlow generates music from images, story texts, or captions by aligning all inputs into the CLAP audio embedding space and sampling with conditional flow matching.