REVIEW 21 cited by
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in music generation have garnered significant attention, yet existing approaches face critical limitations. Some current generative models can only synthesize either the vocal track or the accompaniment track. While some models can generate combined vocal and accompaniment, they typically rely on meticulously designed multi-stage cascading architectures and intricate data pipelines, hindering scalability. Additionally, most systems are restricted to generating short musical segments rather than full-length songs. Furthermore, widely used language model-based methods suffer from slow inference speeds. To address these challenges, we propose DiffRhythm, the first latent diffusion-based song generation model capable of synthesizing complete songs with both vocal and accompaniment for durations of up to 4m45s in only ten seconds, maintaining high musicality and intelligibility. Despite its remarkable capabilities, DiffRhythm is designed to be simple and elegant: it eliminates the need for complex data preparation, employs a straightforward model structure, and requires only lyrics and a style prompt during inference. Additionally, its non-autoregressive structure ensures fast inference speeds. This simplicity guarantees the scalability of DiffRhythm. Moreover, we release the complete training code along with the pre-trained model on large-scale data to promote reproducibility and further research.
Forward citations
Cited by 21 Pith papers
-
YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
YingMusic-Singer-Plus is a diffusion model for singing voice synthesis that preserves melody from a reference clip while allowing flexible lyric changes without manual alignment, outperforming Vevo2 and introducing th...
-
MusicMark: A Robust Generative Watermarking Framework for Music Generation
Embedding watermark bits into diffusion semantic latents via a frozen-backbone adapter yields far more robust music provenance than post-hoc watermarking under codecs and cover-song attacks.
-
Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Training an acoustic renderer with error-matched, near-miss codec corruption improves music quality when the upstream language model plan is imperfect.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Adding SVS-style phoneme conditioning and a length regulator to melody-guided cover generation cuts phoneme error rate from 45.6% to 18.7% while holding melody metrics.
-
Qwen-Music Technical Report
Qwen-Music generates high-fidelity vocal songs via 25 Hz semantic tokens, Melody-CoT planning, and DiT rendering, claiming SOTA on 13/16 metrics and expert preference over proprietary systems.
-
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Current singing voice synthesis models fail to differentiate musical genres, defaulting to pop-like output regardless of input genre, unless given genre-specific fine-tuning data.
-
Echoes: A semantically-aligned music deepfake detection dataset
A semantically aligned, multi-provider music deepfake dataset is harder for detectors and trains models that transfer better than prior AI-music datasets.
-
SegTune: Structured and Fine-Grained Control for Song Generation
SegTune generates songs where each musical section follows its own text description, using a fine-tuned LLM to predict lyric timings so per-section instructions land in the correct audio window.
-
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.
-
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
JAM is a 530M-parameter flow-matching song generator that adds word- and phoneme-level timing control and duration control, achieving strong lyric fidelity and musicality scores when ground-truth timings are provided.
-
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.
-
AI-Generated Song Detection via Lyrics Transcripts
Transcribing audio with Whisper and classifying the transcript with LLM2Vec detects AI-generated songs from audio alone, nearly matching clean-lyrics accuracy and beating audio-based detectors under perturbations and ...
-
AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
A training-free multi-agent framework that decomposes multimodal inputs into audio events, selects specialized generators, and self-corrects outputs to produce multiple audio types.
-
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
SongEval is a 140-hour benchmark of full-length generated songs rated by expert musicians on five aesthetic dimensions, and trained predictors outperform objective metrics at matching human ratings.
-
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
A unified hierarchical-LM-plus-flow-matching system reports top-tier full-song vocal generation, ranking 2–3 on an external blind leaderboard, but releases neither code nor evaluation data.
-
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation
PER-based preference optimization (DPO, PPO, GRPO) reduces lyric-to-song hallucination in an audio language model, with the largest gains from DPO plus reject sampling.
-
ACE-Step: A Step Towards Music Generation Foundation Model
ACE-Step is a fast, controllable open-source music generation model built from a mel-spectrogram DCAE, a linear DiT, and REPA-style semantic alignment.
-
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.
-
TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
TinyMusician distills MusicGen and applies hand-picked mixed-precision quantization to make a 1.04 GB on-device music generator, but the headline '93% quality, 55% smaller' claims conflict with the paper's own tables.
-
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Across five text-to-music models, aesthetic predictor scores, pairwise human preferences, and reference-based distribution metrics produce inconsistent rankings, so the choice of evaluation metric changes the winner.
Discussion (0). Continue with ORCID to comment.