{"id":"62fb50d3-78ac-43e4-b948-3fa1bc11bd5f","arxiv_id":"2607.11117","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Embedding watermark bits into diffusion semantic latents via a frozen-backbone adapter yields far more robust music provenance than post-hoc watermarking under codecs and cover-song attacks.","lead":"MusicMark embeds multi-bit watermarks into a music diffusion model’s semantic latents during generation, not as a post-generation patch. That design keeps provenance signals alive under neural codecs and cover-song voice swaps better than post-hoc audio watermarks, with little quality loss.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Robustness gains may be inflated by training on the same attack families used at evaluation, so the semantic-latent claim is only partially stress-tested.","rationale":"The reader correctly flags the matched train/test attack pool as the weakest assumption and keeps the verdict CONDITIONAL with moderate confidence. That is the right load-bearing concern: the paper’s evidence for a real robustness gap over post-hoc baselines under the reported protocol is solid (especially vs. AudioSeal-M trained identically), quality preservation and ablations are careful, and “first generative music watermarking” is a fair systems claim. What is least secure is the stronger causal reading that semantic-latent embedding itself guarantees robustness to real-world and adaptive manipulations outside the augmentation set. Absolute recovery is already fragile under some neural codecs; without held-out or adaptive tests (and without released artifacts), provenance guarantees should not be overstated. No internal inconsistency or experimental error is evident; the concern is about the scope of the robustness claim. Verdict stays CONDITIONAL—accept-shaped if authors add held-out/adaptive stress tests, uncertainty, and artifacts—not REJECT.","tokens_in":20284,"tokens_out":646,"duration_ms":6612,"concrete_test":"Hold out one neural codec (e.g., SNAC or a newer unseen codec) and one cover-song pipeline (different separator + non-RVC VC) entirely from training augmentations; retrain MusicMark and AudioSeal-M on the reduced pool; re-evaluate Table I Abs/Bit on the held-out attacks. If MusicMark’s Abs advantage over AudioSeal-M falls below ~0.2 or Abs under held-out neural codec drops below ~0.3, the generalization half of the strongest claim is not supported by the current protocol.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that embedding multi-bit messages into the semantic latent during diffusion (decoupled cross-attention adapter + latent consistency loss) yields substantially stronger detection/extraction than post-hoc methods under neural codecs and cover-song edits while preserving quality (Tables I–III, §IV-B/C). That comparison is load-bearing only if the reported Acc/Bit/Abs gap reflects the embedding locus rather than matched attack-augmentation training. §III-F and Appendix C/Table VII show the joint adapter–detector objective is trained with EnCodec, MP3, cover-song RVC, cut, speed, etc.—the same families used at test time (including Cover Song + Cut). AudioSeal-M is trained on the same pool, which helps, but does not rule out that both systems overfit the shared pool while MusicMark’s latent coupling simply memorizes those distortions better. Absolute accuracy already collapses under DAC/SNAC (0.342/0.261) and Cover Song + Cut (0.813), and no held-out codec, alternative VC stack, or detector-aware removal is reported. If the gap shrinks under truly unseen transforms, the “semantic latent is inherently more robust” interpretation weakens even if the systems contribution remains useful.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"MusicMark proposes the first generative watermarking framework for lyrics- and text-conditioned music generation. It freezes a diffusion backbone (primarily ACE-Step) and inserts a lightweight watermark adapter that embeds multi-bit messages into the semantic latent via a decoupled cross-attention branch during denoising, then trains a detector jointly with generation, fidelity, and watermark losses, including a latent consistency regularizer that keeps watermarked latents close to stop-gradient unwatermarked references. Attack augmentations (including neural codecs and a cover-song pipeline) are used during training. Empirically, MusicMark substantially outperforms post-hoc baselines (WavMark, AudioSeal, AudioSeal-M trained on the same data/attacks, SilentCipher) on detection and message extraction under 20 attacks—especially neural codec re-synthesis and cover-song + cut—while preserving objective quality (FAD/CLAP/PER/aesthetics) and human MOS relative to the unwatermarked backbone (Tables I–III), with ablations on stage, injection method, layer position, latent loss, and transfer to Stable Audio 3 (Tables IV–V).","tokens_in":20758,"tokens_out":1069,"duration_ms":9851,"significance":"If the results hold under broader evaluation, this is a solid systems contribution for provenance of AI-generated music: generative semantic-latent watermarking is a natural response to the known fragility of post-hoc audio watermarks under neural codecs, and the cover-song attack is a useful music-specific stress test. Strengths include a fair same-data/same-attack re-training of AudioSeal-M, multi-metric quality evaluation with human MOS, systematic ablations (stage, shared vs. decoupled attention, layer position, latent loss), and a backbone-transfer experiment on SA3. The work is timely given commercial AI music platforms and the documented failure modes of residual-signal watermarks under codec re-synthesis.","major_comments":[{"comment":"§III-F, Appendix C, Table VII, and Table I: robustness is optimized under the same attack families later reported as wins (EnCodec/DAC/SNAC, MP3, cover-song RVC, cut, speed, etc.). AudioSeal-M matches the training pool, which helps isolate the embedding locus, but absolute accuracy still collapses under DAC/SNAC (0.342/0.261) and Cover Song + Cut (0.813), and no held-out codec family, alternative voice-conversion stack, or detector-aware removal is reported. The central claim that semantic-latent embedding is inherently more robust than post-hoc residual insertion is only partially stress-tested; at least one truly unseen transform suite (or adaptive attack) is needed to support the interpretation beyond matched-augmentation superiority.","section":null},{"comment":"§IV-B and Table I (Neural Codec / Cover Song rows): the paper’s strongest narrative is robustness under neural codecs and music-specific edits, yet Abs under DAC and SNAC remains low (0.342 and 0.261) despite perfect Acc and high Bit. The abstract and conclusion should more carefully distinguish detection reliability from full-message recovery under the hardest codecs, and clarify what message capacity / payload reliability is claimed for practical provenance use.","section":null},{"comment":"§V and experimental scope: message capacity is fixed at N=16 bits and evaluation is mainly text/lyrics-conditioned generation (plus one SA3 style-only transfer). For a provenance framework, the manuscript should either demonstrate higher-capacity settings or more explicitly bound the claim (e.g., short identifiers only) and discuss collision / multi-user attribution limits under that capacity.","section":null}],"minor_comments":[{"comment":"SilentCipher Acc is marked “–” in Table I because it lacks a detection head; state this once in the table caption and avoid implying detection parity in prose summaries of “all metrics.”","section":null},{"comment":"Fig. 3 difference visualizations are informative but would benefit from a fixed color scale / gain normalization so amplitude attenuation (SilentCipher) and band artifacts (AudioSeal-M) are comparable across columns.","section":null},{"comment":"Notation: α is both the learnable watermark scale (Eq. 7) and appears as a training “Scale Factor α=0.1” in Table VI; disambiguate initialization vs. learned scale.","section":null},{"comment":"Appendix B filtering (style–music similarity ≥0.5, English-only Muse clips) should note possible domain shift relative to the Suno-metadata evaluation prompts.","section":null},{"comment":"Straight-through estimator for non-differentiable attacks is stated but not analyzed; a short note on whether STE bias affects codec vs. cover-song gradients would help reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"The novelty claim (“first generative watermarking for music”) appears plausible relative to the cited speech/generative audio literature, but the core technical idea (adapter + latent consistency + attack-augmented joint training) is incremental relative to recent generative speech watermarking and latent-injection methods. The paper is still a good fit for a systems/audio venue if the unseen-attack and capacity points are tightened; I would not reject for lack of theory alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean systems paper that actually moves the music-watermarking problem. The new piece is not “generative watermarking” in the abstract—that already exists for speech and some audio generators—but a practical music-specific setup: multi-bit messages injected into the semantic latent of a frozen diffusion backbone (ACE-Step) via a decoupled watermark cross-attention adapter, plus a latent-consistency regularizer that keeps watermarked latents near unwatermarked ones. They also define a cover-song attack (MDX-Net + RVC vocal swap, optional cut) that matches how music actually gets redistributed.\n\nWhat they do well is the comparison. Table I is the load-bearing result: against WavMark, AudioSeal, a same-data/same-attack AudioSeal-M, and SilentCipher, MusicMark holds detection and bit accuracy under neural codecs and cover-song+cut where post-hoc methods collapse. Quality tables (FAD/CLAP/PER/aesthetics + MOS) show they do not trash the backbone. Ablations on stage (latent vs vocoder), shared vs decoupled attention, layer position, and the latent loss are the right ones, and the SA3 transfer check is a nice extra. Citations are fair; they own the speech/post-hoc literature instead of ignoring it.\n\nSoft spots, in proportion: the robustness claim is trained on the same attack families used at test (EnCodec/DAC/SNAC, MP3, RVC cover-song, cuts, etc.). AudioSeal-M matching that pool helps, but it does not fully separate “semantic latent is inherently robust” from “this adapter+detector memorizes the shared pool better.” Absolute message recovery is already soft under DAC/SNAC (0.34/0.26) and cover+cut (0.81). No held-out codec, alternate VC stack, or detector-aware removal. Capacity is 16 bits; no public code/checkpoints. Those are real limits, not fatal ones for a first systems result.\n\nWho it’s for: anyone working on AI-music provenance, generative audio watermarking, or platform attribution. Worth a serious referee. I would engage, cite the cover-song protocol and the latent-adapter design, and push for unseen-attack and artifact release.","headline":"Solid first generative music watermarking system with real codec/cover-song gains over post-hoc baselines; the semantic-latent story is useful but only partly stress-tested because training and eval share the same attack families.","tokens_in":21368,"tokens_out":565,"would_cite":true,"duration_ms":6475,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"MusicMark embeds multi-bit watermarks into music’s semantic latents during diffusion so they ride with the music itself and survive codecs, cuts, and cover-song voice swaps without spoiling quality.","keywords":["music watermarking","generative watermarking","semantic latent embedding","diffusion music generation","neural codec robustness","cover-song attack","provenance verification"],"falsifier":"Take a held-out set of MusicMark tracks, re-encode them with a neural codec or voice-conversion stack never used in training or evaluation, then measure whether absolute message accuracy collapses toward the post-hoc baseline levels reported in Table I.","tokens_in":21153,"feed_emoji":"🎵","tokens_out":647,"duration_ms":10682,"temperature":0.7,"pith_summary":"Commercial AI music is flooding the web, but existing audio watermarks were built for speech and are usually painted on after generation as faint noise. That residual signal is easy to strip—especially with modern neural codecs that keep the song’s meaning and throw away everything else—and the watermark step can simply be skipped. MusicMark claims a different route: inject the message into the semantic latent while the diffusion model is still denoising, so the watermark becomes part of the musical content rather than a post-production overlay. A lightweight adapter and detector are trained on a frozen generator with a joint objective that keeps watermarked latents close to their clean twins, plus heavy attack training that includes codec re-synthesis and a new cover-song pipeline. If the claim holds, provenance for generated music can no longer be bypassed by re-encoding or re-voicing, while the songs still sound like the original model’s output.","feed_headline":"Watermarks baked into music latents survive codecs and covers","feed_subtitle":"MusicMark ties the message to the song’s semantics so re-encoding and voice swaps no longer erase provenance","key_machinery":"The watermark adapter: a zero-initialized, decoupled cross-attention branch that injects a projected multi-bit message embedding into selected late layers of a frozen diffusion music model, combined with a latent consistency loss that keeps watermarked latents near their unwatermarked references.","core_discovery":"MusicMark is presented as the first generative watermarking framework for lyrics- and text-conditioned music: by conditioning a diffusion latent denoiser on multi-bit messages through a decoupled watermark cross-attention adapter, the watermark is written into the semantic representation itself. Joint training with latent-consistency, fidelity, and watermark losses plus attack augmentations yields far higher detection and message-recovery rates than post-hoc baselines under neural codecs and cover-song transformations, while objective and human quality scores stay comparable to the unwatermarked backbone.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["MusicMark bakes watermarks into music latents during generation","Generative music watermarks survive codecs and cover-song attacks","First generative watermark embeds messages in music semantic space","MusicMark adapter writes provenance into diffusion latents","Watermarks in music generation latents resist neural re-synthesis"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That the strong robustness numbers will still hold when real attackers use codecs, voice converters, or removal methods that were never seen in the training attack pool.","fun_headline_variants_meta":{"raw":{"variants":["MusicMark bakes watermarks into music latents during generation","Generative music watermarks survive codecs and cover-song attacks","First generative watermark embeds messages in music semantic space","MusicMark adapter writes provenance into diffusion latents","Watermarks in music generation latents resist neural re-synthesis"]},"model":"grok-4.5","effort":"low","cost_usd":0.005438,"raw_usage":{"total_tokens":1555,"prompt_tokens":877,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":54380000,"prompt_tokens_details":{"text_tokens":877,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":593,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":877,"tokens_out":85,"duration_ms":5664,"temperature":1.0,"reasoning_tokens":593,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T06:56:02.123224+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Take a held-out set of MusicMark tracks, re-encode them with a neural codec or voice-conversion stack never used in training or evaluation, then measure whether absolute message accuracy collapses toward the post-hoc baseline levels reported in Table I.","supporting_citations":[],"review_version":1}