Prompts for AI music are dominated by genre and story terms, but genre words survive into perception while story-heavy prompts produce the largest semantic mismatch.
From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music
Prompts for AI music are dominated by genre and story terms, but genre words survive into perception while story-heavy prompts produce the largest semantic mismatch.