REVIEW 5 cited by
Style Transfer of Audio Effects with Differentiable Signal Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Style Transfer of Audio Effects with Differentiable Signal Processing
read the original abstract
We present a framework that can impose the audio effects and production style from one recording to another by example with the goal of simplifying the audio production process. We train a deep neural network to analyze an input recording and a style reference recording, and predict the control parameters of audio effects used to render the output. In contrast to past work, we integrate audio effects as differentiable operators in our framework, perform backpropagation through audio effects, and optimize end-to-end using an audio-domain loss. We use a self-supervised training strategy enabling automatic control of audio effects without the use of any labeled or paired training data. We survey a range of existing and new approaches for differentiable signal processing, showing how each can be integrated into our framework while discussing their trade-offs. We evaluate our approach on both speech and music tasks, demonstrating that our approach generalizes both to unseen recordings and even to sample rates different than those seen during training. Our approach produces convincing production style transfer results with the ability to transform input recordings to produced recordings, yielding audio effect control parameters that enable interpretability and user interaction.
Forward citations
Cited by 5 Pith papers
-
Finding Fast Filters
A differentiable DSL combines multi-rate, recurrent, and cascade filter tricks; automated search and compilation produce Pareto-optimal fast filter programs that dominate prior baselines across six filter families.
-
Compiling Differentiable Audio Graphs to Real-Time DSP
ADAC compiles differentiable audio processors to stable FAUST plugins whose behavior matches the source model within floating-point noise.
-
Black-Box Optimization for Identifying and Inverting Audio Dynamic Range Control Effects
Blind DRC parameter estimation and inversion can be framed as derivative-free optimization in a dynamic-histogram feature space, yielding competitive reconstructions against neural baselines.
-
StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems
StemFX predicts tokenized per-stem audio-effect chains with a jointly-trained Transformer encoder-decoder, beating contrastive and prior FX-encoding methods on effect-chain retrieval and real-mix style transfer.
-
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
Conditional flow matching with a rollout loss and early audio-visual fusion achieves state-of-the-art results on visually-guided acoustic highlighting.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.