{"id":"2e6cc0e2-5337-4b41-ae1d-0ec3136ba88a","arxiv_id":"2505.19663","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Neural audio codecs, particularly Descript Audio Codec, are the hardest attack class for current deep audio watermarking, and retraining with those codecs does not restore reliable watermark extraction.","lead":"This paper introduces RAW-Bench, a standardized test for audio watermarking that simulates 20 real-world distortions. Evaluating four deep-learning watermarking methods, it finds that modern neural audio codecs, especially Descript Audio Codec, erase watermarks even when the methods are trained to resist such compression.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'even when trained with such compressions' claim (Abstract, §4) rests on a single underspecified retraining recipe; without ablations over DA exposure and training budget, 'fundamental limitation' and 'insufficient training' remain indistinguishable.","rationale":"The paper does solid benchmark work: the clean-table, domain table, and the DA attack results are stark and internally consistent. The central claim, however, includes the qualifier 'even when algorithms are trained with such compressions' (Abstract). That qualifier is what turns a benchmark observation into a claim about a fundamental limitation. The only evidence for it is the retraining experiment, whose protocol is described in a few sentences in Section 3: uniform category weighting, a proprietary 1250-hour music dataset, a lowered SDR bound for SilentCipher, and no stated epochs, optimizer, loss, or schedule. With uniform category weighting, the two neural codecs (EN and DA) receive a small share of training batches, and no ablation isolates codec exposure. The observed post-retraining DA results—bitwise accuracy around 0.42–0.60 and full-message accuracy 0.00—are exactly what an undertrained codec defense would look like; they do not by themselves demonstrate in-principle failure. The reader identified this same weakest assumption. Other concerns, such as the comparability of full-message accuracy across different message lengths (16 bits for AS/WM, 23.8 for SC, 30 for TI), are secondary because the DA full-message results are all near zero; the retraining concern is the load-bearing one. A single DA-heavy retraining experiment with the released code and a public corpus would settle whether the headline should be softened from 'even when trained' to 'under uniform attack augmentation, training did not confer DA robustness.' Since my concern does not change the reader's verdict, 'UNCHANGED' is appropriate.","tokens_in":9341,"tokens_out":4349,"duration_ms":50816,"concrete_test":"Using the released RAW-Bench pipeline, retrain AudioSeal and SilentCipher on a public music-plus-speech corpus (e.g., MUSDB18 plus VCTK) with a DA-heavy curriculum: sample DA in 50% of training batches, always use the full 9-codebook DA configuration, and train for at least three times the reported budget; then evaluate DA full-message accuracy under strict settings. If either model exceeds 0.5 full-message accuracy, the 'even when trained' claim is an artifact of the training schedule rather than a fundamental codec incompatibility. If both still yield 0.00, the concern is resolved and the headline claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—that neural codecs defeat watermarking even when models are trained with such compressions (Abstract; Section 4, 'Will Watermarks Survive Neural Codecs?')—depends entirely on the retraining protocol described in Section 3 ('Retraining') and Section 4 ('Re-training'). That protocol uses uniform weighting across attack categories, a proprietary 1250-hour music dataset plus 40-hour VCTK and BBC subsets, a lowered SDR bound for SilentCipher, and unspecified epochs, optimizer, loss, and learning-rate schedule. With uniform category weighting, the neural-compression category receives only about one-sixth of the training batches, split between EN and DA. No ablation varies the fraction of codec examples, the number of codebooks used during training, or the training duration. The post-retraining DA numbers are consistent with undertraining: AS* reaches 0.60 bitwise and 0.00 full-message accuracy on DA, and SC* reaches 0.42 bitwise and 0.00 full-message. These are near-chance bitwise scores with no message integrity, which is exactly what a weakly represented or prematurely stopped codec defense would produce. Without ablations isolating codec exposure, the paper cannot distinguish 'fundamental incompatibility' from 'insufficient or unrepresentative training.' This is load-bearing because the 'even when trained' qualifier elevates the result from a benchmark observation about four checkpoints to a general limitation of the training approach.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RAW-Bench, a standardized benchmark for deep learning-based audio watermarking, together with a 20-attack robustness pipeline and a diverse test set of raw 44.1 kHz recordings spanning speech, music, and environmental sounds. It evaluates four publicly available pre-trained watermarking models (AudioSeal, SilentCipher, Timbre, WavMark) under loose and strict attack settings, and additionally retrains AudioSeal and SilentCipher with the proposed attack pipeline. The main findings are that neural codecs (Encodec and Descript Audio Codec) are the most damaging distortions, that even retraining with such codec attacks leaves full-message accuracy at or near zero for Descript Audio Codec, and that attack-augmented training generally helps but does not fix all vulnerabilities.","tokens_in":9643,"tokens_out":3632,"duration_ms":38511,"significance":"If the results hold, the paper makes a valuable and timely contribution: it provides a reproducible benchmark, a public test dataset, an out-of-sample verification protocol, and the first systematic cross-model comparison that includes SilentCipher and retrained variants. The negative result on neural codecs is important for the audio watermarking community, and the 'compete for the same space' argument is a thought-provoking design principle. The release of the evaluation code (github.com/SonyResearch/raw_bench) and the careful construction of the test set from multiple public corpora are concrete strengths that support the paper's utility.","major_comments":[{"comment":"The abstract and Section 4 claim that neural codecs pose the most significant challenge 'even when algorithms are trained with such compressions.' This claim rests on a single retraining configuration: uniform per-category attack weighting (so the two neural-codec attacks receive only a small fraction of training batches), a proprietary 1250-hour music dataset plus 40-hour VCTK and BBC subsets, and no reported epochs, optimizer, loss, or learning-rate schedule. The post-retraining DA numbers (AS* bitwise 0.60, full-message 0.00; SC* bitwise 0.42, full-message 0.00) are consistent with undertraining rather than a fundamental incompatibility. I request ablations that vary the fraction of codec-distorted examples, the codebook configuration, and the training budget, and that report the training hyperparameters, before the 'even when trained' conclusion is stated as a general limitation.","section":"Section 3 (Retraining) and Section 4 (Table 5, rows AS* and SC*)"},{"comment":"All robustness results are reported as point estimates without confidence intervals, error bars, or significance tests. For a benchmark explicitly intended to enable systematic comparison, this is a limitation for the finer-grained ranking claims (for example, AS versus TI on many attack columns, or the small AS* versus AS improvements). I request at least standard deviations or confidence intervals across test segments or independent runs, or a statistical test for the main comparisons, so that the reader can distinguish meaningful differences from noise.","section":"Table 5"},{"comment":"Full-message accuracy is compared across models with different message lengths (AudioSeal 16 bits, SilentCipher 23.8 bits, Timbre 30 bits, WavMark 16 bits). Since the probability that every bit decodes correctly depends on message length even at equal per-bit accuracy, cross-model comparisons of full-message accuracy are biased. The capacity is approximately matched, but the message lengths are not; I recommend reporting per-bit equivalent accuracy or fixing the payload length whenever full-message accuracy is used for cross-model conclusions.","section":"Table 1 vs. Table 5"}],"minor_comments":[{"comment":"The paragraph arguing that watermarking and neural codecs 'compete for the same space' is speculative; it is labeled as a belief, but it should be explicitly presented as a hypothesis for future work rather than a conclusion of the benchmark.","section":"Section 4, 'Will Watermarks Survive Neural Codecs?'"},{"comment":"The row for DA lists 'Descript Audio Codec [16] (at 44.1 kHz)', but the abbreviations EN and DA are not expanded in the table caption; please define 'EN' as Encodec and 'DA' as Descript Audio Codec in the caption.","section":"Table 2"},{"comment":"The phrase 'the architecture of AS is based on EN' is imprecise; AudioSeal uses an Encodec-based architecture, and the sentence would benefit from stating that explicitly rather than using the abbreviation EN alone.","section":"Section 4, Robustness paragraph"},{"comment":"The paper would benefit from a figure summarizing the Table 5 results (for example, grouped bar charts per attack category), since the current dense table is hard to read and the main qualitative findings (e.g., the DA collapse) are less visually salient than they deserve.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The benchmark itself is a solid, useful contribution, and the public release of code and data is commendable. The main risk is the overstatement of the 'even when trained with such compressions' conclusion given the single, underspecified retraining recipe and the lack of ablations. If the authors add the requested ablations, report training details, and temper the abstract accordingly, this would be a strong paper. The use of a proprietary dataset for retraining also limits reproducibility, and I would encourage the authors to either release the retraining data or clearly document the exact training configuration so others can replicate the procedure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a solid, useful benchmark paper with a real negative result. Across four deep watermarking methods, Descript Audio Codec (DA) destroys full-message accuracy for every model, and retraining with their attack pipeline barely moves the needle. That is worth knowing. The weaker part is the \"even when trained with such compressions\" headline: it rests on a single retraining recipe with no ablations, so the paper can support \"current attack-augmented training doesn't fix DA\" but not \"no reasonable training could.\"\n\nWhat's genuinely new: it extends AudioMarkBench by adding SilentCipher, adding Encodec and DAC as attack classes, using a raw 44.1 kHz test set across music, speech, and environmental sounds, and testing retraining. The evaluation framework is public, the test sets are public collections, and capacity is roughly matched. The authors also verify their test data isn't in the models' training sets. That's the right way to build a benchmark.\n\nSoft spots, in order of seriousness. The retraining section uses a proprietary 1250-hour music dataset, uniform weighting across six attack categories, a lowered SDR bound for SilentCipher, and unspecified epochs, optimizer, and loss. No ablation varies codec exposure or training budget. So the \"even when trained\" claim is an observation about one protocol, not a demonstrated fundamental limit. The Section 4 \"fundamental issue\" language goes beyond the evidence; the evidence is about current checkpoints. Second, Table 5 reports point estimates with no error bars or significance tests. Some gaps are huge and need no statistics, but the small differences (e.g., AS vs SC on EN) are best treated as noise. Third, full-message accuracy is compared across different message lengths (16/23.8/30 bits); bitwise accuracy is the safer headline metric.\n\nThe citation pattern is fine. SilentCipher is the authors' own model, but it performs worse than Timbre on most attacks and worse than AS after retraining, so self-preference isn't driving the story.\n\nWho it's for: watermarking researchers and anyone building provenance or forensics tools on top of neural codecs. It deserves a proper peer review, with the main request being ablations around retraining and release of the retrained weights or at least a reproducible training recipe.\n\nRecommendation: send it to review. It's a legitimate benchmark contribution with a findable weakness, not a flawed paper.","headline":"Useful benchmark with a real DAC-negative result, but the 'even when trained' claim outruns the evidence.","tokens_in":10180,"tokens_out":2562,"would_cite":true,"duration_ms":27309,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Neural codecs beat all tested audio watermarks, even retrained ones","keywords":["audio watermarking","neural codecs","robustness benchmark","adversarial training","Encodec","Descript Audio Codec","imperceptibility","RAW-Bench"],"falsifier":"Retrain one of the tested models with Descript Audio Codec as the dominant attack (for example, applying DAC distortion to a large share of each batch, with hyperparameters tuned for this single attack) and evaluate on the RAW-Bench DA attack at strict settings. If full-message accuracy rises well above the reported near-zero values, the paper's claim that neural codecs present a fundamental limitation would be refuted; if it stays near zero, the claim is strengthened.","tokens_in":9173,"feed_emoji":"🎧","tokens_out":6752,"duration_ms":64810,"temperature":0.7,"pith_summary":"This paper introduces RAW-Bench, a standardized benchmark for testing deep-learning audio watermarking systems against real-world distortions. The authors evaluate four published watermarking methods under a pipeline of twenty attacks, including conventional and neural compression, noise, filtering, and time-domain modifications, on a new test set spanning speech, music, and environmental sounds. The central finding is that neural compression, specifically Encodec and Descript Audio Codec, is the most damaging attack, and that retraining the watermarking models with these attacks improves bitwise accuracy but never brings full-message accuracy to acceptable levels. In fact, none of the considered methods survives the Descript Audio Codec attack. If this holds, it means current deep-learning watermarking cannot protect audio that passes through modern neural codecs, which are increasingly the final stage of real-world audio pipelines.","feed_headline":"Neural codecs beat all tested audio watermarks, even retrained ones","feed_subtitle":"A 20-attack benchmark shows Descript Audio Codec erases full messages from every deep-learning watermarking system tested.","key_machinery":"The load-bearing object is the Robust Audio Watermarking Benchmark (RAW-Bench), built from three components: a diverse test dataset of raw 44.1 kHz recordings across music, speech, and environmental sounds; an attack pipeline of twenty distortions organized into six categories (mixing, dynamics, filtering, low-level modifications, neural compression, and conventional compression) with loose and strict parameter settings; and a retraining protocol that feeds the strict attacks into the models with uniform per-category weighting. This machinery turns the question 'do watermarks survive neural codecs?' into measurable bitwise and full-message accuracies, and it is what allows the paper to attribute failures to specific attack types rather than to dataset quirks.","core_discovery":"On the paper's own terms, the discovery is that neural codecs and deep-learning audio watermarking compete for the same imperceptible information, and the codecs currently win. Across four pre-trained watermarking models, the Descript Audio Codec at 44.1 kHz reduces full-message extraction accuracy to essentially zero in every case, and Encodec is nearly as destructive. Retraining two of the models on a pipeline that includes these codec distortions improves robustness on some attacks but leaves full-message accuracy near zero for both neural codecs. The paper concludes that this is not a tuning failure but a fundamental tension: a codec that succeeds at removing imperceptible components will remove watermarks that are designed to be imperceptible.","pith_inferences":["A test the paper leaves implicit: retraining with the neural codec as the sole attack, or with a curriculum that starts on mild codec settings and hardens, would separate 'the attack is fundamentally destructive' from 'the retraining recipe was too diluted'.","The same RAW-Bench pipeline could be pointed at latent-based watermarks for generative audio, which may survive codecs better because they are embedded in a space the codec already preserves; that would test whether the competition-for-imperceptible-information story extends beyond carrier-signal methods.","If codec makers and watermark designers iterate against each other on this benchmark, the specific result about Descript Audio Codec could become a moving target, but the general trade-off between imperceptible embedding and lossy compression is a structural constraint that would remain."],"forward_implications":["Audio that passes through Encodec or Descript Audio Codec will lose embedded watermarks from all four tested deep-learning methods, making current watermarking unreliable for distribution chains that use neural codecs.","Retraining with a broad attack pipeline improves robustness on some distortions but not on neural compression, reverb, or phase shift, so attack augmentation alone is not a sufficient fix.","The advantage one model gains from using Encodec's architecture does not transfer to a different neural codec, indicating that codec-specific robustness does not generalize.","The benchmark's strict and loose attack thresholds provide a common yardstick for future watermarking systems to report robustness at matched perceptual impact."],"supporting_citations":[{"why":"supplies the AudioSeal baseline, the only model whose architecture matches Encodec and which retains some robustness to it.","marker":"[8]"},{"why":"supplies the SilentCipher baseline, which the paper retrains with its attack pipeline.","marker":"[9]"},{"why":"supplies the Timbre baseline, the most robust model on conventional distortions.","marker":"[10]"},{"why":"supplies the WavMark baseline, which fails completely on both neural codecs.","marker":"[11]"},{"why":"provides the Encodec neural codec used as one of the two critical neural-compression attacks.","marker":"[15]"},{"why":"provides the Descript Audio Codec, the attack that defeats every tested watermarking method.","marker":"[16]"},{"why":"prior result showing codec-augmented training improves robustness, which this paper confirms is insufficient.","marker":"[17]"},{"why":"prior speech-only watermark benchmark that this paper extends with a larger attack set and broader audio domains.","marker":"[20]"}],"fun_headline_variants":["Neural codecs erase audio watermarks even after retraining","Watermarking fails against neural codecs in real-world test","Neural codecs destroy watermarks; retraining is no fix","Audio watermarks can't survive neural codecs in benchmark","Neural codecs beat watermarks even when retrained on them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that retraining with neural-codec attacks is insufficient rests on one specific retraining setup—a balanced mix of many attacks, a proprietary music dataset, relaxed quality constraints for one of the models, and undisclosed training details—so a more focused or intensive codec-training regime could in principle overturn the claim.","fun_headline_variants_meta":{"raw":{"variants":["Neural codecs erase audio watermarks even after retraining","Watermarking fails against neural codecs in real-world test","Neural codecs destroy watermarks; retraining is no fix","Audio watermarks can't survive neural codecs in benchmark","Neural codecs beat watermarks even when retrained on them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000587,"raw_usage":{"total_tokens":2706,"prompt_tokens":846,"completion_tokens":1860,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":1773}},"tokens_in":462,"tokens_out":1860,"duration_ms":13901,"temperature":1.0,"reasoning_tokens":1773,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:08:41.477388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain one of the tested models with Descript Audio Codec as the dominant attack (for example, applying DAC distortion to a large share of each batch, with hyperparameters tuned for this single attack) and evaluate on the RAW-Bench DA attack at strict settings. If full-message accuracy rises well above the reported near-zero values, the paper's claim that neural codecs present a fundamental limitation would be refuted; if it stays near zero, the claim is strengthened.","supporting_citations":[{"cited_title":"Simple and controllable music gen- eration,","cited_arxiv_id":null,"evidence_quote":"supplies the AudioSeal baseline, the only model whose architecture matches Encodec and which retains some robustness to it."},{"cited_title":"Music ControlNet: Multiple time-varying controls for music genera- tion,","cited_arxiv_id":null,"evidence_quote":"supplies the SilentCipher baseline, which the paper retrains with its attack pipeline."},{"cited_title":"A comparative study on recent neu- ral spoofing countermeasures for synthetic speech detection,","cited_arxiv_id":null,"evidence_quote":"supplies the Timbre baseline, the most robust model on conventional distortions."},{"cited_title":"Towards assessing data replication in music generation with mu- sic similarity metrics on raw audio,","cited_arxiv_id":null,"evidence_quote":"supplies the WavMark baseline, which fails completely on both neural codecs."},{"cited_title":"De- tecting voice cloning attacks via Timbre Watermarking,","cited_arxiv_id":null,"evidence_quote":"provides the Encodec neural codec used as one of the two critical neural-compression attacks."},{"cited_title":"Maskmark: Robust neural watermarking for real and synthetic speech,","cited_arxiv_id":null,"evidence_quote":"prior result showing codec-augmented training improves robustness, which this paper confirms is insufficient."},{"cited_title":"High fidelity neural audio compression,","cited_arxiv_id":null,"evidence_quote":"prior speech-only watermark benchmark that this paper extends with a larger attack set and broader audio domains."}],"review_version":1}