{"id":"6a1a4189-4e17-403d-a7e1-727b39ff2834","arxiv_id":"2508.15521","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A single audio watermark that simultaneously identifies the generative model and its training dataset, with reported robustness to compression, noise, pruning, and resampling.","lead":"DualMark embeds two hidden signatures into AI-generated audio, one identifying the model and one identifying its training data. It aims to answer both 'who made this audio' and 'what data trained it' with a single watermark.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract asserts seamless embedding but reports no audio quality metrics; dual watermark may degrade fidelity, a load-bearing gap.","rationale":"The reader's weakest assumption already identified the load-bearing premise that dual watermarking does not degrade generated audio quality. Our stress-test sharpens this into a concrete, checkable gap: the abstract provides no audio-quality evidence whatsoever. This is the single most load-bearing concern because every downstream claim (practical robustness, accountability, copyright protection) presupposes that the watermarked generator still produces high-quality audio. We do not raise concerns about novelty or internal consistency because the abstract alone cannot adjudicate those. The verdict remains UNVERDICTED/unchanged: a full-text review is necessary to see whether the experimental section includes the missing quality evaluation.","tokens_in":653,"tokens_out":2496,"duration_ms":29690,"concrete_test":"Inspect the full paper (or, if code is released, run a reproduction) for a direct audio-quality comparison between the DualMark watermarked model and an identical unwatermarked baseline. Compute objective metrics (PESQ, STOI, MCD, or ViSQOL) and/or conduct a listening test on generated samples from the same prompts. If these metrics are absent or show a statistically significant degradation (e.g., PESQ drop > 0.5), the 'seamlessly embed' claim is refuted and the central contribution is materially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DualMark is a usable dual-provenance framework depends on the watermarks being embedded 'seamlessly'—i.e., without degrading generated audio quality. The abstract reports only attribution accuracy (97.01% F1, 91.51% AUC) and robustness results, with no objective or subjective audio quality evaluation. If the dual watermark embedding into Mel-spectrogram representations or the Watermark Consistency Loss (WCL) noticeably degrades fidelity (e.g., PESQ/STOI drops, audible artifacts), the method's practical utility collapses: users would not deploy a generator whose output quality suffers for the sake of attribution. The risk is acute because DWE and WCL alter the training objective, potentially trading reconstruction quality for watermark extractability. Without evidence that fidelity is preserved, the high attribution scores are insufficient to substantiate the claim of a 'foundational step' toward accountable audio generation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DualMark, a dual-provenance watermarking framework for audio generative models. It claims to be the first to embed two distinct attribution signatures—model identity and dataset origin—into Mel-spectrogram representations during training, using a Dual Watermark Embedding (DWE) module and a Watermark Consistency Loss (WCL). The paper also introduces the Dual Attribution Benchmark (DAB) for evaluating joint model-data attribution. The abstract reports high attribution accuracy (97.01% F1 for model attribution, 91.51% AUC for dataset attribution) and robustness against pruning, compression, noise, and sampling attacks.","tokens_in":893,"tokens_out":1864,"duration_ms":22787,"significance":"If the claims hold, DualMark addresses a real gap in audio provenance: current watermarking methods trace the generating model but not the training data. Joint attribution would strengthen copyright enforcement and accountability. The proposed benchmark could also be a useful community resource. The significance is tempered by the fact that the abstract provides no experimental details, no audio quality evaluation, and no comparison to prior work; the novelty and utility depend on evidence not visible in the abstract.","major_comments":[{"comment":"The central quantitative claims (97.01% F1, 91.51% AUC, robustness to pruning/compression/noise/sampling) are presented without any experimental setup: no model architecture, dataset, training details, baseline methods, or error bars. Single-point numbers without variance or statistical significance cannot substantiate the central claim of a 'foundational step'. Please provide full evaluation details, including number of runs and standard deviations, and describe how the benchmark splits are constructed and why they are non-trivial.","section":"Abstract, third paragraph"},{"comment":"The abstract asserts that dual watermarks are embedded 'seamlessly' into Mel-spectrogram representations, but reports no audio quality metrics (e.g., PESQ, STOI, MUSHRA, or a listening test). Since the DWE module and WCL alter the training objective, they may trade generation fidelity for extractability. A dual-provenance watermark that degrades perceived quality would not be deployable. This is a load-bearing omission: without evidence of fidelity preservation, the high attribution scores alone are insufficient to establish practical utility.","section":"Abstract, second paragraph"},{"comment":"The robustness claims are underspecified. 'Aggressive pruning, lossy compression, additive noise, and sampling attacks' are not defined: what pruning ratios, codecs, bitrates, SNRs, or sample rates were used? How were prior methods configured for comparison? The statement that these conditions 'severely compromise prior methods' needs direct quantitative comparison. Without this, the claimed advantage over prior art is not verifiable.","section":"Abstract, third paragraph"},{"comment":"The Dual Attribution Benchmark (DAB) is not described. A benchmark's value depends on its protocol, dataset composition, metric definitions, and how it avoids overfitting to a single watermarking scheme. Please specify the size and diversity of DAB, the evaluation protocol, and how model identity and dataset origin are controlled for confounds. As it stands, the 'first' claim for the benchmark cannot be assessed.","section":"Abstract, third paragraph"}],"minor_comments":[{"comment":"The adjective 'seamlessly' is vague and appears to be a qualitative assertion rather than a measured property. Please replace with objective fidelity results or remove.","section":"Abstract, second paragraph"},{"comment":"To support 'the first dual-provenance watermarking framework', please cite recent model-level watermarking methods and clearly state what distinguishes dual provenance from prior multi-bit or multi-task watermarking approaches.","section":"Abstract, first paragraph"},{"comment":"The metrics F1-score and AUC are used for different tasks (model attribution and dataset attribution). Please define both tasks and explain why AUC is appropriate for dataset attribution; also report the corresponding positive/negative class balance and decision-threshold methodology.","section":"Abstract, third paragraph"}],"recommendation":"uncertain","confidential_remarks":"This review is based solely on the abstract because the full text was not available. The central idea is plausible and timely, but none of the supporting evidence is present in the manuscript portion I could review. If the full paper provides the missing experimental details, fidelity measurements, and benchmark description, the contribution may be strong; however, on the current evidence I cannot reach a confident verdict. I recommend the editor either obtain the full text for a second-stage review or ask the authors to resubmit with the complete details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [colleague],\n\nIf this paper delivers what the abstract promises, it's the first audio watermarking scheme that asks and answers two provenance questions at once—which model generated a clip, and on which dataset the model was trained. That is a genuine capability extension, not just a tweak. The proposed DWE module and WCL loss are plausible mechanisms, and the attribution numbers (97% F1 for model, 91.5% AUC for dataset) plus robustness to pruning/compression/noise/sampling are consistent with a well-engineered system. The DAB benchmark is also a useful community artifact if it becomes public.\n\nWhat I can't tell from the abstract, and what worries me more than the stress-test note suggests, is whether the dual embedding actually leaves generation quality intact. The abstract says 'seamlessly' but gives no PESQ/STOI or listening test. That's a load-bearing claim: watermarking that degrades output by a noticeable amount won't be adopted, no matter how accurate the attribution. This isn't a fatal flaw; it's an omission in the evidence presented. I'd expect the full paper to include at least objective quality metrics and ideally a subjective test. If they're absent, the reviewers should push on it.\n\nOther soft spots: the 'first' claim needs a careful prior-art check against simultaneous or sequential watermarks in other modalities. And the reported numbers come without explicit baselines or error bars in the abstract—standard stuff, but let's see the full comparison. The stress-test note's anxiety about training-objective interference is legitimate but not damning; WCL and DWE are constrained modifications, and the robustness gains suggest the optimization isn't obviously broken.\n\nNet: worth a serious refereeing. The novelty is moderate-to-high, the application is timely, and the core idea is testable. I'd want the fidelity data and a head-to-head against model-only watermarking plus a dataset-attribution baseline. If those exist in the manuscript, this is a solid venue-level acceptance; if not, major revision.\n\nI'd bring it to reading group once the full text is out; not just the abstract. Cite it if it pans out, but I'm not citing an abstract.\n\nRecommendation: accept for peer review, conditional on the full paper containing fidelity evaluation and comparison to prior art.","headline":"Abstract-only read, but the dual-provenance idea is a real step up from model-only watermarking; the missing fidelity numbers are the one thing I'd want before trusting the 'seamless' claim.","tokens_in":1280,"tokens_out":1991,"would_cite":false,"duration_ms":21355,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DualMark claims a single watermarking framework can trace both the generating model and the training dataset of audio output.","keywords":["audio watermarking","generative audio","provenance attribution","model identity","dataset origin","dual watermark","robustness","Mel-spectrogram"],"falsifier":"Train a DualMark-augmented audio generative model, then apply a distortion not in the reported robustness set—such as low-bitrate MP3 encoding (e.g., 32 kbps), band-pass filtering, or audio re-recording through a speaker and microphone—and measure whether both watermarks remain extractable while the generated audio still matches the unwatermarked baseline in listening quality. If either signature drops below chance-level extraction or the audio quality degradation is clearly audible, the central claim of joint, robust attribution is falsified.","tokens_in":657,"feed_emoji":"🔊","tokens_out":1454,"duration_ms":17500,"temperature":0.7,"pith_summary":"DualMark is a watermarking framework for audio generative models that aims to solve a provenance gap: existing methods can only identify which model generated an audio clip, but not which dataset trained it. The paper proposes embedding two distinct signatures into the model during training—one for model identity, one for dataset origin—so that both can be extracted from generated audio. If successful, this gives creators and platforms a way to hold the right party accountable when generated audio is misused, and to verify whether a model was trained on protected data. The paper reports high attribution accuracy and robustness to common attacks, positioning the method as a first step toward fully accountable audio generative models.","feed_headline":"Two watermarks trace both model and training data in audio","feed_subtitle":"DualMark reports ~97% model attribution F1 and ~91% dataset AUC, surviving compression, noise, and pruning.","key_machinery":"The central mechanism is the Dual Watermark Embedding (DWE) module, which injects two distinct watermark signatures into the Mel-spectrogram representations used during training, combined with a Watermark Consistency Loss (WCL) that enforces both signatures are reliably reproduced when the model generates audio. The DWE acts as the carrier of the two attribution identities, and the WCL is the training signal that binds them to the generative process, enabling joint extraction at inference time.","core_discovery":"The paper introduces DualMark, described as the first dual-provenance watermarking framework that simultaneously encodes model identity and dataset origin into audio generative models during training. The central proposal is a Dual Watermark Embedding (DWE) module that injects two watermarks into Mel-spectrogram representations, paired with a Watermark Consistency Loss (WCL) that trains the model to reproduce both watermarks in generated audio. The authors also construct the Dual Attribution Benchmark (DAB) for evaluating joint model–data attribution. Experiments on this benchmark report a 97.01% F1-score for model attribution, 91.51% AUC for dataset attribution, and robustness against pruni","pith_inferences":["A natural extension the authors leave implicit is testing DualMark across distribution shifts beyond the listed attacks—for example, real-world microphone re-recording, streaming codecs, or adversarial watermark-removal attempts—since the stated robustness set may not cover all deployment conditions.","The dual-signature idea could be combined with dataset inference or membership inference techniques to strengthen data-origin claims, but the paper does not compare against such baselines.","A testable follow-up is whether the two watermarks interfere with each other or with generation quality when embedded at higher capacities or when the model is fine-tuned on new data, a scenario closer to real-world model updates."],"forward_implications":["If DualMark works as reported, generated audio could be traced back to both the specific generative model and the specific training dataset that produced it, enabling more precise copyright enforcement and accountability.","A unified benchmark (DAB) for joint model–data attribution would allow future watermarking methods to be compared on a consistent robustness standard.","The approach could be extended to other generative modalities (images, video, text) by adapting the embedding and consistency-loss mechanism to their respective representation spaces.","Watermark consistency during training may serve as a regularizer that affects generation quality, an effect the paper must characterize for practical deployment."],"supporting_citations":[],"fun_headline_variants":["DualMark: dual watermarks trace audio to model and dataset","First dual-provenance watermark for audio: model + dataset","DualMark: 97% model, 91% dataset attribution in audio","Two watermarks, one audio file: model and data origin","Audio provenance: dual watermarks ID model and dataset"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The method's central claim rests on the premise that two watermarks can be embedded into Mel-spectrogram representations during training without noticeably degrading the generated audio, and that both signatures remain extractable independently under real-world distortions.","fun_headline_variants_meta":{"raw":{"variants":["DualMark: dual watermarks trace audio to model and dataset","First dual-provenance watermark for audio: model + dataset","DualMark: 97% model, 91% dataset attribution in audio","Two watermarks, one audio file: model and data origin","Audio provenance: dual watermarks ID model and dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001145,"raw_usage":{"total_tokens":4591,"prompt_tokens":753,"completion_tokens":3838,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":3749}},"tokens_in":497,"tokens_out":3838,"duration_ms":27795,"temperature":1.0,"reasoning_tokens":3749,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:48:55.811921+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a DualMark-augmented audio generative model, then apply a distortion not in the reported robustness set—such as low-bitrate MP3 encoding (e.g., 32 kbps), band-pass filtering, or audio re-recording through a speaker and microphone—and measure whether both watermarks remain extractable while the generated audio still matches the unwatermarked baseline in listening quality. If either signature drops below chance-level extraction or the audio quality degradation is clearly audible, the central claim of joint, robust attribution is falsified.","supporting_citations":[],"review_version":1}