{"id":"b30366cb-32fa-400e-9fb9-befccf1d9e9f","arxiv_id":"2507.12890","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"DiffRhythm+ improves full-length lyric-to-song generation via balanced data scaling, MuLan-based multimodal style control, and DPO fine-tuning guided by automated aesthetic scorers.","lead":"DiffRhythm+ upgrades an open-source diffusion song generator with a larger balanced training set, text-and-audio style conditioning, and preference optimization driven by automatic music-quality scores. The authors report better intelligibility, musicality, and speed than the original system, while acknowledging a remaining gap to the larger YuE model.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests on treating SongEval/Audiobox scores as proxies for human preference; because those same scorers drive DPO, data filtering, and Tables II–III, the only independent evidence is an unquantified MOS plot, so the claimed improvement needs external validation.","rationale":"The reader's weakest-assumption analysis is sound. I independently looked for an even more foundational flaw and did not find one: the flow-matching objective and DPO loss derivation (Eqs. 1–6) follow standard practice; the efficiency claim (RTF 0.036–0.039 on an RTX 4090) is concrete and plausible; and the release of code and audio samples is genuine supporting evidence. The soft spot is the evidential chain linking the optimization objective to the claimed outcome. Preference optimization is only as good as the reward signal; using the same self-authored aesthetic scorer as reward, data filter, and headline metric makes Tables II–III an in-domain fit rather than an independent evaluation. The paper's own caveats about Audiobox, and the observation that systems exceed ground-truth AudioBox scores, are acknowledgements that the proxies are imperfect, yet no external validation is supplied. The human MOS is the only external anchor, and it is reported only as violin plots without inferential statistics; moreover, it places YuE above DiffRhythm+ on every axis, so the abstract's unqualified \"significant improvements over previous systems\" overclaims relative to the strongest open-source baseline. This does not invalidate the system contributions, but it does mean the central claim is not yet established beyond the authors' own metrics. The conditional verdict is therefore appropriate, and my stress-test does not move it.","tokens_in":11354,"tokens_out":4395,"duration_ms":50635,"concrete_test":"Run a preregistered, blinded pairwise-preference listening test with at least 50 independent listeners (none from the author group) and at least 20 matched lyrics/style prompts per system, comparing DiffRhythm+ full, DiffRhythm full, and YuE; report pairwise preference proportions, effect sizes, and 95% confidence intervals. If DiffRhythm+ is not preferred over YuE at p < 0.05, or if its advantage over DiffRhythm has effect size below 0.5, the abstract and conclusion should be narrowed to \"improvement over DiffRhythm, comparable but not superior to YuE\".","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III-D constructs DPO win-lose pairs from SongEval and Audiobox-aesthetic scores (winner > 3.0, gap > 0.4); Section IV-A uses the same scorers to filter the 25,000-hour SFT subset; and Table II/Table III report those same scorers as the primary evidence that DPO, data scale, and style conditioning improve quality. SongEval is not an independent arbiter: five of its listed authors overlap with the DiffRhythm+ author list, and the paper itself concedes Audiobox \"may not always align with human ratings\" and that several systems \"exceed ground-truth scores\" on AudioBox metrics (Sec. V-B). If SongEval/Audiobox are optimized targets rather than faithful proxies for listener taste, the DPO gain in Table III (mean 2.86 to 3.19) and the \"closing the gap\" conclusion lose external validity. The only independent check, the MOS violin plots in Figure 2, is reported without means, confidence intervals, per-condition sample counts, or significance tests; on all three axes YuE is highest, which is inconsistent with the abstract's unqualified \"significant improvements ... over previous systems\" and supports only a narrower improvement over DiffRhythm.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents DiffRhythm+, an extension of the diffusion-based full-length song generation system DiffRhythm, and claims improvements along three axes: a larger and rebalanced training corpus, multi-modal style conditioning via MuLan embeddings, and a DPO preference-optimization stage whose win/lose pairs are scored by SongEval and Audiobox-aesthetic. The paper reports objective metrics (KL, FAD, CLaMP 3, PER, RTF), aesthetic-model scores (Audiobox and SongEval), ablation results, and MOS violin plots, concluding that DiffRhythm+ substantially closes the gap with state-of-the-art open-source systems such as YuE while retaining a very low real-time factor.","tokens_in":11794,"tokens_out":4603,"duration_ms":59166,"significance":"If the central claims held, DiffRhythm+ would be a practically valuable contribution: an open, fast, full-length lyric-to-song model with flexible text/audio style conditioning, a large and balanced training set, and a concrete preference-alignment recipe. The paper explicitly ships audio samples and code, and the engineering effort behind the 120,000-hour data pipeline and the 1.1B-parameter DiT system is substantial. However, the significance assessment hinges entirely on whether the reported gains in SongEval/Audiobox scores and the MOS violin plots actually reflect listener satisfaction; the current evaluation design does not establish that, so the contribution is potentially useful but not yet convincingly validated.","major_comments":[{"comment":"The primary objective evidence for preference optimization is circular. SongEval and Audiobox-aesthetic scores are used to construct DPO win-lose pairs (Sec. III-D: winner score >3, gap >0.4), to filter the 25,000-hour SFT subset (Sec. IV-A), and then as the main outcome metrics in Tables II and III. Optimizing a model against a proxy and then reporting gains on that same proxy is expected behavior and does not constitute evidence of improved listener satisfaction. This concern is compounded by the fact that several SongEval authors overlap with the DiffRhythm+ author list, and by the ad-hoc linear mapping of Audiobox scores to the 1-5 scale without validation. The paper should either validate SongEval/Audiobox against human ratings in the target domain or provide independent, non-optimized evaluation metrics and human listening results with appropriate statistics.","section":"Sec. III-D, Sec. IV-A, Sec. IV-D, Table II, Table III"},{"comment":"The subjective evaluation is too thin to support the abstract's unqualified claim of \"significant improvements in naturalness, arrangement complexity, and listener satisfaction over previous systems.\" Figure 2 is presented only as violin plots without means, confidence intervals, per-condition sample counts, or significance tests, and the text in Sec. V-A explicitly states that YuE achieves the highest MOS on all aspects. This directly contradicts the abstract's unqualified wording if YuE is counted among \"previous systems.\" At minimum, the authors should report numeric MOS means and dispersion, run statistical significance tests, specify the number of songs and listeners per condition, and either weaken the abstract or restrict the claim to improvements over DiffRhythm.","section":"Sec. V-A, Fig. 2, Abstract"},{"comment":"The ablation labeled \"DPO Winner: GT Winner vs Generated Winner\" shows that training with self-generated, SongEval-scored winners gives a higher SongEval mean (3.19) than training with ground-truth winners (2.94). This is precisely the outcome expected if the model is optimizing the scoring function rather than human taste; it does not demonstrate that generated winners are preferable to ground-truth recordings. The result should be reframed as evidence of proxy overfitting unless accompanied by human evaluation showing that the self-generated-winner DPO model is also preferred by listeners.","section":"Table III, DPO Winner block"}],"minor_comments":[{"comment":"Equation (6) is under-specified: the quantities N, ω(λ_n), and λ_n are not defined, and the sign of the argument inside the log-sigmoid is not derived from the preceding DPO objective. Please provide the full diffusion-DPO derivation or a precise reference to the exact form used.","section":"Sec. III-D, Eq. (6)"},{"comment":"No statistical significance tests or confidence intervals are reported for any of the objective metrics or the MOS results, despite multiple comparisons across models and settings; this makes it impossible to assess whether the reported differences are reliable.","section":"Sec. IV-C, Sec. V-A"},{"comment":"The phrase \"outperforms other baseline models in most aspects\" in Sec. V-A is hard to reconcile with the same section's statement that YuE is highest on all aspects; please clarify the intended comparison set (e.g., excluding YuE) or rephrase.","section":"Abstract and Sec. V-A"},{"comment":"The sample-count notation \"33 / 32\" for Chinese/English is not accompanied by the total number of songs evaluated per condition; please report the N for each block so the reader can judge the stability of the score distributions.","section":"Table III"},{"comment":"The abstract states that code is available at github.com/ASLP-lab/DiffRhythm, but that URL points to the original DiffRhythm repository rather than a DiffRhythm+ repository; please provide the correct link or clarify that only the base model's code is released.","section":"Abstract/Code availability"},{"comment":"The real-time-factor comparisons would be more interpretable if the paper specified the exact hardware, inference batch size, and decoding setup used for each baseline; YuE's reported RTF of 10.385 is much higher than DiffRhythm+'s 0.036, and it should be verified that the measurement protocol is identical across systems.","section":"Sec. IV-B/Sec. IV-D"}],"recommendation":"major_revision","confidential_remarks":"The overlap between the authors of SongEval and the authors of this manuscript, combined with the use of SongEval as both the DPO reward and the primary evaluation metric, makes the external validity of the main claims fragile. I would encourage the editor to require independent human evaluation or an evaluation by a third-party metric set not used during training before accepting. The manuscript is technically interesting and the engineering contribution is real, but the current evidence does not support the abstract's unqualified claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real system with plausible engineering gains, but the headline claim is broader than the evidence. The preference-optimization story is partly circular, because the same self-authored scorer that builds the DPO pairs is the main metric used to show DPO works.\n\nWhat's actually new and good: the combination of a large balanced 120k-hour dataset, MuLan text-plus-audio style conditioning, DPO with automated aesthetic scorers, and multi-frame temporal perturbation is genuinely new as a package. The system is open-source, runs at RTF ~0.04, and beats YuE on PER while coming close on KL/FAD. Those distribution-matching and intelligibility metrics are independent of the aesthetic scorers and show consistent gains over DiffRhythm. That part of the contribution holds up.\n\nWhere it gets soft: SongEval is used to construct DPO winners/losers (Sec. III-D), to filter the SFT data (Sec. IV-A), and then as the primary evidence in Table II and the DPO ablation in Table III. Five SongEval authors overlap with the DiffRhythm+ author list. The ablation row \"Generated Winner vs GT Winner\" showing higher same-metric scores after training on machine-scored generated winners is exactly the expected outcome if you optimize the metric you measure with. The paper itself concedes Audiobox \"may not always align with human ratings\" and that several systems beat ground truth on some AudioBox metrics, so the objective metrics only weakly support the preference-alignment claim. The only independent human evidence, the MOS violin plots in Fig. 2, shows YuE highest on all three axes with no means, confidence intervals, or significance tests. That directly undercuts the abstract's unqualified \"significant improvements over previous systems.\"\n\nStill, the central engineering claim—fast, controllable, open full-length song generation that improves over the authors' own DiffRhythm and approaches YuE—is credible. The circularity is fixable with an independent human preference test and significance reporting, and the abstract should be reworded to say \"improvements over DiffRhythm\" rather than \"previous systems.\"\n\nFor whom: people building fast open song generators and anyone studying preference optimization when reward and metric are the same object. It deserves serious peer review, but the evaluation needs an overhaul before it can be accepted as-is.","headline":"Solid engineering with a partially circular preference-optimization evaluation; the claim of progress over previous systems outruns the paper's own human data.","tokens_in":12189,"tokens_out":2149,"would_cite":true,"duration_ms":27782,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DiffRhythm+ claims full-length, style-controllable songs from lyrics in about ten seconds, at near-YuE quality with far less compute.","keywords":["lyrics-to-song generation","full-length song generation","diffusion model","flow matching","preference optimization","multimodal style conditioning","MuLan embeddings","music aesthetics evaluation"],"falsifier":"Play the exact winner-loser pairs that SongEval selected (e.g., 'DiffR.+ DPO' vs. 'DiffR.+ Pretrain' outputs it rated apart by more than 0.4) to a group of listeners and ask them which they prefer; if their choices agree with SongEval no more than chance, the preference-optimization gains are an artifact of the automated scorer rather than evidence of alignment with human taste.","tokens_in":11159,"feed_emoji":"🎵","tokens_out":7259,"duration_ms":73805,"temperature":0.7,"pith_summary":"DiffRhythm+ is a diffusion-based system that generates full songs — vocals and accompaniment together — directly from lyrics and a style prompt given as text, a reference audio clip, or both. The paper claims that three fixes — rebalancing and scaling up training data to 120,000 hours, replacing raw audio latents with MuLan style embeddings, and adding a preference-optimization stage — jointly raise intelligibility, musicality, arrangement complexity, and listener satisfaction over the predecessor DiffRhythm. If the claim is right, a user can go from raw lyrics to a complete, style-controlled song in about ten seconds on a single GPU, while substantially closing the quality gap with much larger autoregressive systems such as YuE.","feed_headline":"Full songs from lyrics in 10 seconds with near-YuE quality","feed_subtitle":"Style control by text or reference audio, with generation roughly 260 times faster than YuE.","key_machinery":"The engine is a latent diffusion transformer trained with conditional flow matching: a denoiser $v_\\theta(t,y,c)$ moves samples from noise to song latents conditioned on phoneme-tokenized lyrics, a MuLan style embedding, and the diffusion timestep. The new mechanism is the preference stage, which fine-tunes this denoiser with the diffusion DPO loss of Eq. (6): for each prompt, a SongEval- or Audiobox-scored winner and loser are drawn from generated samples, filtered by a score gap of at least 0.4 and a winner score above 3, and the loss pushes the model's noise estimates closer to the winner's trajectory and away from the loser's. Multi-frame temporal perturbation also relaxes lyric-to-audio alignment so that imprecise timestamps do not force rigid phrasing.","core_discovery":"The central claim is that the main bottlenecks of full-length lyric-to-song generation — data imbalance, rigid style conditioning, and the absence of explicit preference alignment — can be fixed inside a non-autoregressive diffusion framework without sacrificing its speed advantage. Concretely, DiffRhythm+ reports a balanced 2:2:1 Chinese/English/instrumental corpus of about 120,000 hours; MuLan-based style embeddings that accept both text and reference audio; and a DPO stage whose winner-loser pairs are built from SongEval and Audiobox-aesthetic scores with a gap threshold of 0.4 and a winner floor of 3.0. The resulting model achieves a phoneme error rate of 14.85%, a KL divergence of 0.488, and an FAD of 1.835 while keeping the real-time factor at 0.036, with human listening scores above DiffRhythm and close to YuE on intelligibility, musicality, and quality.","pith_inferences":["Because the reward signal is automated, a natural next step would be to replace or augment SongEval/Audiobox with human pairwise judgments; a small human preference set might transfer the DPO gains to dimensions these scorers underweight, such as emotional expression.","The same MuLan conditioning could make style transfer a direct operation — swap the reference audio while keeping the lyrics — though the paper does not test cover-song behavior explicitly.","The roughly 260x speed margin over YuE suggests that, if independent listening tests confirm parity, diffusion-based song generation could become the default engine for real-time interactive music tools, with autoregressive models reserved for offline high-budget production."],"forward_implications":["Songwriters and content creators can specify genre, emotion, and instrumentation in natural language or by providing a reference track, and receive a complete song with both vocals and accompaniment in a single forward pass.","At a real-time factor of about 0.04, full-length songs can be generated on an RTX 4090 in around ten seconds, making interactive, iterative song creation practical.","The balanced training data lifts Chinese-language song quality, reducing the repetition and omission-of-lyrics failure mode that plagued the predecessor.","The DPO stage increases the fraction of samples rated 3–5 on SongEval from 59.5% after SFT to 81.5%, with 8 DPO epochs and a 0.4 score gap as the reported best configuration.","The model approaches YuE's KL/FAD numbers at a real-time factor of 0.036–0.039 versus 10.385, showing that near-YuE quality is obtainable without autoregressive decoding."],"supporting_citations":[{"why":"The predecessor DiffRhythm supplies the VAE, DiT backbone, and multi-frame temporal perturbation that DiffRhythm+ extends.","marker":"[25]"},{"why":"YuE is the large autoregressive baseline DiffRhythm+ compares against and whose data-scaling results motivate the dataset expansion.","marker":"[21]"},{"why":"MuLan provides the joint audio-text embedding space that enables multimodal style conditioning.","marker":"[26]"},{"why":"MuQ-MuLan is the specific instantiation used to extract expressive style embeddings from audio prompts.","marker":"[27]"},{"why":"SongEval provides the song-aesthetic scores used both as the primary DPO reward and as the main quality metric.","marker":"[36]"},{"why":"Audiobox-aesthetic scores instrumentals and serves as the secondary reward and metric, mapped to SongEval's scale.","marker":"[37]"},{"why":"Direct Preference Optimization gives the training objective that aligns generated samples with the winner-loser pairs.","marker":"[35]"},{"why":"Tango2 supplies the theoretical motivation for applying DPO to diffusion-based generative models.","marker":"[34]"},{"why":"Conditional flow matching defines the generative training objective that turns noise into song latents.","marker":"[33]"}],"fun_headline_variants":["Lyrics to full songs, 260x faster, style from text or audio","DiffRhythm+: balanced data and DPO refine song generation","Full-song diffusion with text or audio style control","Preference-optimized full-length songs, still 260x faster","Balanced 120k hours of data lifts song quality and speed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The premise the argument rests on is that SongEval and Audiobox-aesthetic scores are trustworthy surrogates for human listening preference, because they alone decide which outputs become DPO winners and losers and drive the headline quality improvements.","fun_headline_variants_meta":{"raw":{"variants":["Lyrics to full songs, 260x faster, style from text or audio","DiffRhythm+: balanced data and DPO refine song generation","Full-song diffusion with text or audio style control","Preference-optimized full-length songs, still 260x faster","Balanced 120k hours of data lifts song quality and speed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000453,"raw_usage":{"total_tokens":2294,"prompt_tokens":973,"completion_tokens":1321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1231}},"tokens_in":589,"tokens_out":1321,"duration_ms":10820,"temperature":1.0,"reasoning_tokens":1231,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:35:12.057543+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Play the exact winner-loser pairs that SongEval selected (e.g., 'DiffR.+ DPO' vs. 'DiffR.+ Pretrain' outputs it rated apart by more than 0.4) to a group of listeners and ask them which they prefer; if their choices agree with SongEval no more than chance, the preference-optimization gains are an artifact of the automated scorer rather than evidence of alignment with human taste.","supporting_citations":[{"cited_title":"Direct preference optimization: Your language model is secretly a reward model,","cited_arxiv_id":null,"evidence_quote":"Direct Preference Optimization gives the training objective that aligns generated samples with the winner-loser pairs."},{"cited_title":"Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,","cited_arxiv_id":null,"evidence_quote":"Tango2 supplies the theoretical motivation for applying DPO to diffusion-based generative models."}],"review_version":1}