REVIEW 6 cited by
DDSP: Differentiable Digital Signal Processing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge of how sound is generated and perceived. A third approach (vocoders/synthesizers) successfully incorporates strong domain knowledge of signal processing and perception, but has been less actively researched due to limited expressivity and difficulty integrating with modern auto-differentiation-based machine learning methods. In this paper, we introduce the Differentiable Digital Signal Processing (DDSP) library, which enables direct integration of classic signal processing elements with deep learning methods. Focusing on audio synthesis, we achieve high-fidelity generation without the need for large autoregressive models or adversarial losses, demonstrating that DDSP enables utilizing strong inductive biases without losing the expressive power of neural networks. Further, we show that combining interpretable modules permits manipulation of each separate model component, with applications such as independent control of pitch and loudness, realistic extrapolation to pitches not seen during training, blind dereverberation of room acoustics, transfer of extracted room acoustics to new environments, and transformation of timbre between disparate sources. In short, DDSP enables an interpretable and modular approach to generative modeling, without sacrificing the benefits of deep learning. The library is publicly available at https://github.com/magenta/ddsp and we welcome further contributions from the community and domain experts.
Forward citations
Cited by 6 Pith papers
-
Black-Box Optimization for Identifying and Inverting Audio Dynamic Range Control Effects
Blind DRC parameter estimation and inversion can be framed as derivative-free optimization in a dynamic-histogram feature space, yielding competitive reconstructions against neural baselines.
-
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
SLASH adds DSP-derived absolute pitch objectives, including direct spectrogram generation from F0, to self-supervised pitch estimation and beats DSP and SSL baselines on MIR-1K.
-
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
Anira, a new library for real-time audio neural network inference, is benchmarked across three engines, finding ONNX Runtime fastest for stateless models and LibTorch fastest for stateful models.
-
Go witheFlow: Real-time Emotion Driven Audio Effects Modulation
witheFlow is a lightweight open-source proof-of-concept system for real-time emotion-driven modulation of audio effects in music performance by combining biosignals and audio features.
-
Efficient and Distortion-less Spectrum Multiplexer via Neural Network-based Filter Banks
A neural network built to mirror an oversampled polyphase filter bank multiplexes multiple IoT signals into one wideband stream, learning its filter coefficients by training and reaching about -39 dB NMSE with GPU-acc...
-
Workflow-Based Evaluation of Music Generation Systems
A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.
Discussion (0). Continue with ORCID to comment.