Pith. sign in

REVIEW 1 cited by

DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20289 v1 pith:74JRSHOC submitted 2024-05-30 cs.SD cs.AIcs.LG

classification cs.SDcs.AIcs.LG
keywords generationcontroldiffusioninference-timemodelmusicdistilledmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Controllable music generation methods are critical for human-centered AI-based music creation, but are currently limited by speed, quality, and control design trade-offs. Diffusion Inference-Time T-optimization (DITTO), in particular, offers state-of-the-art results, but is over 10x slower than real-time, limiting practical use. We propose Distilled Diffusion Inference-Time T -Optimization (or DITTO-2), a new method to speed up inference-time optimization-based control and unlock faster-than-real-time generation for a wide-variety of applications such as music inpainting, outpainting, intensity, melody, and musical structure control. Our method works by (1) distilling a pre-trained diffusion model for fast sampling via an efficient, modified consistency or consistency trajectory distillation process (2) performing inference-time optimization using our distilled model with one-step sampling as an efficient surrogate optimization task and (3) running a final multi-step sampling generation (decoding) using our estimated noise latents for best-quality, fast, controllable generation. Through thorough evaluation, we find our method not only speeds up generation over 10-20x, but simultaneously improves control adherence and generation quality all at once. Furthermore, we apply our approach to a new application of maximizing text adherence (CLAP score) and show we can convert an unconditional diffusion model without text inputs into a model that yields state-of-the-art text control. Sound examples can be found at https://ditto-music.github.io/ditto2/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A lightweight adapter that adds rotary position embeddings to decoupled cross-attention enables efficient time-varying style control and audio inpainting/outpainting for text-to-music diffusion Transformers.

Pith tools