Pith. sign in

Rave:Avariationalautoencoderforfastandhigh-qualityneuralaudiosynthesis

9 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.

9 Pith papers citing it
5 external citations · Pith
abstract

Deep generative models applied to audio have improved by a large margin the state-of-the-art in many speech and music related tasks. However, as raw waveform modelling remains an inherently difficult task, audio generative models are either computationally intensive, rely on low sampling rates, are complicated to control or restrict the nature of possible signals. Among those models, Variational AutoEncoders (VAE) give control over the generation by exposing latent variables, although they usually suffer from low synthesis quality. In this paper, we introduce a Realtime Audio Variational autoEncoder (RAVE) allowing both fast and high-quality audio waveform synthesis. We introduce a novel two-stage training procedure, namely representation learning and adversarial fine-tuning. We show that using a post-training analysis of the latent space allows a direct control between the reconstruction fidelity and the representation compactness. By leveraging a multi-band decomposition of the raw waveform, we show that our model is the first able to generate 48kHz audio signals, while simultaneously running 20 times faster than real-time on a standard laptop CPU. We evaluate synthesis quality using both quantitative and qualitative subjective experiments and show the superiority of our approach compared to existing models. Finally, we present applications of our model for timbre transfer and signal compression. All of our source code and audio examples are publicly available.

citation-role summary

background 1 method 1

citation-polarity summary

fields

cs.SD 8 cs.LG 1

years

2026 8 2025 1

representative citing papers

DEMON: Diffusion Engine for Musical Orchestrated Noise

cs.SD · 2026-05-27 · unverdicted · novelty 7.0

DEMON is a streaming diffusion engine that exposes denoising parameters as playable controls at up to 12.3 decoder completions per second via per-slot scheduling, shared state, source blending, and accelerated decoding.

Latent Fourier Transform

cs.SD · 2026-04-20 · unverdicted · novelty 7.0

LatentFT uses latent-space Fourier transforms and frequency masking in diffusion autoencoders to enable timescale-specific manipulation of musical structure in generative models.

FXplorer: A Map-Based Interface for Exploratory Audio Effect Design

cs.SD · 2026-06-06 · unverdicted · novelty 5.0

FXplorer organizes audio effects in a perceptually informed 2D space with ML embeddings for similarity and semantic search, combining spatial browsing with DAW-style controls for interactive editing and interpolation.

Drivetrain simulation using variational autoencoders

cs.LG · 2025-01-29 · unverdicted · novelty 5.0

Variational autoencoders generate jerk signals from torque inputs in electric drivetrains and outperform physics-based baselines without detailed parametrization.

Hu\'i S\`u: Co-constructing a Dual Feedback Apparatus

cs.SD · 2026-04-28 · unverdicted · novelty 3.0

A musical performance co-produces sound through dual feedback loops between a RAVE-based neural audio instrument and a recurrent neural control system, exploring shared agency with human performers.

citing papers explorer

Showing 9 of 9 citing papers.

  • Structural Bottlenecks on Frequency Representation in End-to-End Audio Models cs.SD · 2026-07-09 · conditional · none · ref 2 · internal anchor

    State-of-the-art strided audio encoders impose predictable alias-collapse and resolution bottlenecks on frequency primitives; Gabor Latent Refactorization recovers much of the lost separability post-hoc.

  • DEMON: Diffusion Engine for Musical Orchestrated Noise cs.SD · 2026-05-27 · unverdicted · none · ref 19

    DEMON is a streaming diffusion engine that exposes denoising parameters as playable controls at up to 12.3 decoder completions per second via per-slot scheduling, shared state, source blending, and accelerated decoding.

  • Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators cs.SD · 2026-05-21 · unverdicted · none · ref 4

    Live Music Diffusion Models adapt bidirectional diffusion for interactive music generation via KV caching and ARC-Forcing, recovering and exceeding discrete autoregressive efficiency while enabling post-training alignment without RL.

  • Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cs.SD · 2026-05-10 · unverdicted · none · ref 35

    MixtureTT performs direct per-stem timbre transfer on polyphonic mixtures via a shared diffusion transformer, outperforming single-stem baselines on SATB choral data while eliminating cascaded separation errors.

  • Latent Fourier Transform cs.SD · 2026-04-20 · unverdicted · none · ref 6

    LatentFT uses latent-space Fourier transforms and frequency masking in diffusion autoencoders to enable timescale-specific manipulation of musical structure in generative models.

  • Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments cs.SD · 2026-04-26 · conditional · none · ref 6

    A portable single-board-computer AI music platform and five case studies demonstrate that remapping inputs, interleaving fast and slow controls, small artist datasets, and cheap hardware can open new artist-centered design spaces for intelligent instruments.

  • FXplorer: A Map-Based Interface for Exploratory Audio Effect Design cs.SD · 2026-06-06 · unverdicted · none · ref 1

    FXplorer organizes audio effects in a perceptually informed 2D space with ML embeddings for similarity and semantic search, combining spatial browsing with DAW-style controls for interactive editing and interpolation.

  • Drivetrain simulation using variational autoencoders cs.LG · 2025-01-29 · unverdicted · none · ref 7

    Variational autoencoders generate jerk signals from torque inputs in electric drivetrains and outperform physics-based baselines without detailed parametrization.

  • Hu\'i S\`u: Co-constructing a Dual Feedback Apparatus cs.SD · 2026-04-28 · unverdicted · none · ref 2

    A musical performance co-produces sound through dual feedback loops between a RAVE-based neural audio instrument and a recurrent neural control system, exploring shared agency with human performers.