Automatic music mixing can be reformulated as sequential stem blending using a flow matching model conditioned on the growing submix, with strong in-distribution blending scores and competitive full-mix results.
Rethinking Automatic Music Mixing as Sequential Stem Blending
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Automatic music mixing, the task of automatically combining individual audio tracks into a cohesive mixture, is typically addressed by parallelized architectures that process all input tracks in a single pass. In this work, inspired by how human mix engineers process stems one at a time, we propose a paradigm shift and ask whether automatic music mixing can be reformulated as a sequential stem blending task, where each stem is blended into a growing submix. Specifically, we train a latent flow matching model conditioned on the submix context, enabling sequential processing of an arbitrary number of input tracks. To train the model, we introduce a degradation-based data synthesis strategy that simulates realistic stem blending scenarios from existing multitrack and source separation datasets. Experimental results on both stem blending and automatic music mixing benchmarks demonstrate the effectiveness of the proposed approach. We provide audio examples on the accompanying demo page\footnote{https://sequential-mixing-demo.vercel.app/}.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Rethinking Automatic Music Mixing as Sequential Stem Blending
Automatic music mixing can be reformulated as sequential stem blending using a flow matching model conditioned on the growing submix, with strong in-distribution blending scores and competitive full-mix results.