A cross-modal model extracts performance style from audio and uses it to generate matching piano arrangements from lead sheets, with style transfer and audio-to-MIDI retrieval.
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
A cross-modal model extracts performance style from audio and uses it to generate matching piano arrangements from lead sheets, with style transfer and audio-to-MIDI retrieval.