REVIEW 3 major objections 5 minor 45 references
MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A training-free watermark for diffusion images can pack 8,384 bits and still survive rotation, scaling, and translation attacks by using an X-shaped correction template.
desk verdict Novel X-shaped template makes training-free diffusion watermarking genuinely robust to rotation, but the 8,384-bit capacity claim is information-theoretic sleight of hand. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is an X-shaped binary template: two lines of eight points each crossing at the image center, injected into the Fourier domain of the predicted latent at every timestep via Eq. (8), with magnitude scaled by the standard deviation of the spectrum. At decoding, a greedy search over angles finds the line orientation that maximizes mean Fourier magnitude, and a check on the template's outermost points detects scaling; the detected transform is undone directly on the inverted latent, avoiding a second DDIM inversion. The watermark itself is a normalized Gaussian vector duplicated and shuffled with private keys, and similarity is scored by Pearson correlation. A Shannon-entropy capacity formula, 2.0471 bits per Gaussian element versus 1 bit per Bernoulli bit, is what turns the high-dimensional watermark into a quantitative capacity advantage.
What would settle it
Take a MaXsive-watermarked image, rotate it by 135 degrees, and run the decoder: at 135 degrees the scaling factor is $\gamma = 1/(\sin 135^\circ + \cos 135^\circ) = 1/0$, so the correction is undefined, while at angles where the denominator is negative the rescaling would invert the image size. Observing either failure for rotations inside the claimed RST range would falsify the robustness claim as stated.
Extended reading notes
Core claim
The paper's central discovery is that RST robustness in training-free diffusion watermarking does not require coupling the watermark to a repetitive pattern. Instead, a separate X-shaped template is injected into the Fourier spectrum of the predicted latent at each DDIM sampling step; at decoding, the template's lines are located by maximizing the mean Fourier magnitude over candidate angles, and the detected angle plus a scaling check drive a correction of the recovered initial noise before watermark extraction. Because the template and the watermark are independent, the watermark can be a high-dimensional vector sampled from the standard normal distribution, duplicated and privately shuffled to form the initial noise. Using Shannon entropy as the capacity measure, the paper computes 8,384 bits for MaXsive's 4,096 normal elements versus 11 bits for RingID and 256 bits for Gaussian Shading, and reports higher verification and identification robustness than existing algorithms on Stirmark 3.1 and WAVES.
Load-bearing premise
Appendix A derives the scaling factor $\gamma = 1/(\sin\theta + \cos\theta)$ only for rotations in $[0,90^\circ]$ before central cropping, yet Section 4.3.2 applies it to any detected angle in $[0,360^\circ]$, where $\sin\theta+\cos\theta$ can vanish or go negative and produce an invalid correction.
Editorial extensions
If this is right
- With 8,384 bits per watermark, a service with thousands of users can assign unique identifiers without the identity collisions that plague 11-bit ring-based watermarks.
- RST attacks no longer force a capacity-robustness trade-off, so the same watermark can remain usable after rotation, scaling, and geometric distortions.
- Because no training or fine-tuning is needed, the method applies to an already-deployed diffusion model by changing only the initial noise and sampling guidance.
- In the reported benchmarks, rotation and rotation-plus-scaling verification reach full success under WAVES rotation and 0.70 on Stirmark RST, while cropping remains a partial weakness the paper explicitly flags.
Reading between the lines
- Editorial extension: the template-decoupled design could be tested on other generative architectures, such as text-to-image transformers or video diffusion models, since the template operates on latent Fourier spectra rather than on a specific sampler.
- Editorial extension: the entropy-based capacity yardstick suggests a direct search over watermark distributions with higher differential entropy under a fixed image-quality budget, which could push capacity beyond 8,384 bits.
- Editorial extension: the rotation-correction formula is derived only for 0-90 degree rotations, so a full 360-degree robustness audit would clarify whether angles outside that range fail and whether a universal correction rule is needed.
- Editorial extension: because the watermark is reconstructed by averaging many shuffled copies, error-correcting codes layered on the Gaussian payload could likely lower the false-positive rate or the required template strength.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MaXsive proposes a training-free watermarking method for diffusion models. It embeds a Gaussian watermark into the initial noise of a DDIM generation process via shuffling and duplication, and separately injects an X-shaped template in the Fourier domain to estimate and correct rotation/scaling distortions during detection. The watermark is detected by DDIM inversion, template-based geometric correction, deshuffling, aggregation, and Pearson correlation. The paper claims a watermark capacity of 8,384 bits based on multiplying the number of Gaussian elements by the differential entropy of a standard Gaussian, and reports verification and identification results on the WAVES and Stirmark benchmarks, claiming superiority over Tree-Rings, RingID, and Gaussian Shading.
Significance. The template design is a genuinely interesting idea: decoupling the geometric-correction template from the watermark pattern is more elegant than the repetitive-ring approaches of Tree-Rings and RingID, and the Stirmark RST result (0.70 verification TPR versus 0.34 for RingID, Table 3) is a real experimental improvement. The exhaustive per-distortion tables (Tables 6-8) are valuable for the community. The main weakness is that the headline capacity number (8,384 bits) is not an information-theoretic capacity in any operational sense: it is the product of an arbitrarily chosen vector length and the differential entropy of a Gaussian, which is not the number of reliably recoverable bits under attack. The rotation-scaling formula also lacks a valid derivation for the full 0-360 degree detection range used in Eq. (10). If these two issues are fixed, the paper's contribution would be solid, but as written the central 'high-capacity' claim is overstated.
major comments (3)
- [Sec. 4.4, Eq. (14)] The capacity measure in Eq. (14) equates differential entropy with watermark capacity. Differential entropy is not the number of bits that can be reliably recovered from an attacked image; a continuous random variable sampled from a standard Gaussian has infinite Shannon entropy if measured with unbounded precision, and the right quantity is the capacity of the channel from the embedded watermark to the extracted watermark under the specific attacks considered. The paper never defines a decoding model for the 8,384 bits; detection is performed by Pearson correlation between the entire extracted vector and candidate watermarks, so the 8,384-bit number does not correspond to an embedded bitstream. Since the title, abstract, and introduction advertise high capacity as the primary contribution, this needs to be reworked: either remove the 8,384-bit claim and report the size of the identification space supported by the correlation detector at the target FPR, or estimate the per-coordinate SNR after attacks and compute the resulting channel capacity. As it stands, the comparison in Table 1 between Bernoulli and Gaussian methods is not a valid capacity comparison.
- [Sec. 4.3.2 and Appendix A] The scaling parameter gamma = 1/(sin(theta)+cos(theta)) is derived in Appendix A only for a square image rotated by theta in [0, 90] degrees before central cropping, where sin(theta)+cos(theta) is positive. In Sec. 4.3.2 this same formula is applied to any detected angle, and Eq. (10) searches over [0, 360) degrees. For angles outside the first quadrant, sin(theta)+cos(theta) can be zero or negative, making gamma invalid. The authors should either restrict the applicability statement to the angle range covered by the derivation, or replace the formula with gamma = 1/(|sin(theta)|+|cos(theta)|) and verify that the template detection still works with the corrected expression. This is load-bearing for the claimed RST robustness beyond the narrow benchmark angles.
- [Abstract and Tables 7-8] The abstract states that MaXsive 'surpasses all existing algorithms on the robustness benchmarks, Stirmark 3.1 and WAVES, in both verification and identification settings,' but Table 7 shows that on individual WAVES distortions MaXsive is not the best: on C&R (crop-and-resize) verification it achieves 0.20 vs. 0.47 for Gaussian Shading, and on blurring it achieves 0.85 vs. 0.92 for Gaussian Shading. The aggregate results in Table 2 are indeed strong, but the unqualified 'surpasses all existing algorithms' claim is contradicted by the paper's own per-distortion data. Please qualify the claim to refer to average performance or to specific distortion categories, and explicitly acknowledge the C&R gap, which is also consistent with the limitation stated in Sec. 7.
minor comments (5)
- [Sec. 6.2] There is a typo: 'desgined' should be 'designed'.
- [Eq. (8)] The notation M eta[std(|F(z_t^0)|)] is unclear; M is a binary mask and eta is a scalar, so the intended product is M * eta * std(|F(z_t^0)|). Please clarify the operator precedence.
- [Sec. 5.1] The reported FPR is 1e-34, which is extremely stringent and not typical for watermark evaluation. Please clarify how this threshold is derived for each method (e.g., from the theoretical distribution of the Pearson correlation under the null hypothesis) and whether the same threshold is used for all compared methods.
- [Table 1] The table lists Tree-Rings capacity as 20.471 bits, which is 10*2.0471, but the original Tree-Rings paper does not report capacity in bits; this is the authors' reinterpretation. A footnote or explanation of how the column was computed from the original settings would improve reproducibility.
- [Sec. 4.2.2] The template pattern design describes a line composed of 8 points 'located in the range of 0.2w/2 to 0.5w/2 with an interval of 0.1w/2.' A figure or explicit coordinates would help readers understand the exact mask geometry, especially since the template is central to the method.
Circularity Check
No circular derivation found; the 8,384-bit figure is an explicit definitional calculation, and the self-citation is peripheral.
full rationale
MaXsive's derivation chain is self-contained. The watermark is sampled from N(0,1), normalized, duplicated with private-key shuffles, and recovered by DDIM inversion; the template is an independent X-shape mask injected via Eq. (8) and detected via Eq. (10). None of these components is fitted to the benchmark numbers it later claims. The 8,384.9216-bit capacity is the literal product of L=4096 and H_g≈2.0471 in Eq. (14), i.e., an explicit definitional calculation rather than a measured or fitted result; whether differential entropy is a valid capacity measure is a soundness question (and arguably an overstatement), not a circular one. The only author-overlap citation [30] supports a peripheral normalization decision and is not load-bearing: the method would not reduce to that citation. No step in the paper defines an input in terms of its advertised output or fits a parameter and then calls it a prediction. Honest finding: no significant circularity.
Assumptions & free parameters
free parameters (3)
- Template strength eta =
5
- X-template geometry =
8 points per line; radial range 0.2 to 0.5 of w/2; theta_d unspecified
- Watermark dimension L =
4096
assumptions (5)
- domain assumption DDIM inversion (Eq. 3) recovers the initial noise of a generated image accurately enough for watermark extraction.
- domain assumption Shuffling duplicated watermark blocks keeps the initial noise statistically close to N(0,I).
- ad hoc to paper The X-shaped template creates local magnitude extrema in the Fourier domain that survive rotation, scaling, and common image distortions.
- ad hoc to paper The geometric scaling formula gamma = 1/(sin(theta)+cos(theta)) is valid for the detected angle.
- standard math Stable Diffusion 2.1 and DPM-Solver behave as described in their original publications.
Cite this review
Pith. "Pith review of MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models." pith.science (2026). https://pith.science/paper/LCKHGKXR
@misc{pith2026250721195,
author = {Pith},
title = {Pith review of: MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/LCKHGKXR}},
note = {Machine review of arXiv:2507.21195}
}
read the original abstract
The great success of the diffusion model in image synthesis led to the release of gigantic commercial models, raising the issue of copyright protection and inappropriate content generation. Training-free diffusion watermarking provides a low-cost solution for these issues. However, the prior works remain vulnerable to rotation, scaling, and translation (RST) attacks. Although some methods employ meticulously designed patterns to mitigate this issue, they often reduce watermark capacity, which can result in identity (ID) collusion. To address these problems, we propose MaXsive, a training-free diffusion model generative watermarking technique that has high capacity and robustness. MaXsive best utilizes the initial noise to watermark the diffusion model. Moreover, instead of using a meticulously repetitive ring pattern, we propose injecting the X-shape template to recover the RST distortions. This design significantly increases robustness without losing any capacity, making ID collusion less likely to happen. The effectiveness of MaXsive has been verified on two well-known watermarking benchmarks under the scenarios of verification and identification.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Bang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal, Yuancheng Xu, Chenghao Deng, Sicheng Zhu, Abdirisak Mohamed, Yuxin Wen, Tom Goldstein, and Furong Huang. 2024. WAVES: Benchmarking the Robustness of Image Wa- termarks. In Proceedings of the 41st International Conference on Machine Learning , Vol. 235. PMLR, Vienna, Austria, 1456–1492
work page 2024
-
[2]
Teal Witter, Chinmay Hegde, and Niv Co- hen
Kasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde, and Niv Co- hen. 2025. Hidden in the Noise: Two-Stage Robust Watermarking for Images. arXiv:2412.04653 [cs.CV] https://arxiv.org/abs/2412.04653
arXiv 2025
-
[3]
Mohammad Awrangjeb, Manzur Murshed, and Guojun Lu. 2006. Global Geomet- ric Distortion Correction in Images. In 2006 IEEE Workshop on Multimedia Signal Processing. IEEE, Victoria, BC, Canada, 435–440
work page 2006
- [4]
-
[5]
Provisions on the Administration of Deep Synthesis Internet Information Services
Cyberspace Administration of China, Ministry of Industry and Information Technology of the People’s Republic of China, and Ministry of Public Security of the People’s Republic of China 2022. Provisions on the Administration of Deep Synthesis Internet Information Services . Cyberspace Administration of China, Ministry of Industry and Information Technology...
work page 2022
-
[6]
Hai Ci, Pei Yang, Yiren Song, and Mike Zheng Shou. 2024. RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-key Identification. In Computer Vision – ECCV 2024 , Vol. 28. Springer Nature Switzerland, Milan, Italy, 338–354
work page 2024
-
[7]
Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion Models Beat GANs on Image Synthesis. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., Virtual, 8780–8794
work page 2021
-
[8]
Ping Dong, J.G. Brankov, N.P. Galatsanos, Yongyi Yang, and F. Davoine. 2005. Digital watermarking robust to geometric distortions.IEEE Transactions on Image Processing 14, 12 (2005), 2140–2150. doi:10.1109/TIP.2005.857263
arXiv 2005
Show all 45 references
-
[9]
Brankov, Nikolas P
Ping Dong, Jovan G. Brankov, Nikolas P. Galatsanos, Yongyi Yang, and Franck Davoine. 2005. Digital Watermarking Robust to Geometric Distortions. IEEE Transactions on Image Processing 14, 12 (2005), 2140–2150. doi:10.1109/TIP.2005. 857263
2005 doi
-
[10]
Jean-Luc Dugelay, Stéphane Roche, Christian Rey, and Gwenaël Doërr. 2006. Still- Image Watermarking Robust to Local Geometric Distortions. IEEE Transactions on Image Processing 15, 9 (2006), 2831–2842. doi:10.1109/TIP.2006.877311
2006
-
[11]
Weitao Feng, Wenbo Zhou, Jiyan He, Jie Zhang, Tianyi Wei, Guanlin Li, Tianwei Zhang, Weiming Zhang, and Nenghai Yu. 2024. AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA. In Proceedings of the 41st International Conference on Mac...
2024
-
[12]
Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, and Teddy Furon. 2023. The Stable Signature: Rooting Watermarks in Latent Diffusion Models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Paris, France, 22409–22420
2023
-
[13]
Xinbo Gao, Cheng Deng, Xuelong Li, and Dacheng Tao. 2010. Geometric Dis- tortion Insensitive Image Watermarking in Affine Covariant Regions. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 40, 3 (2010), 278–286. doi:10.1109/TSMCC.2009.2037512
2010
-
[14]
Ton Kalker, Geert Depovere, Jaap Haitsma, and Maurice J.J.J.B. Maes. 1999. Video Watermarking System for Broadcast Monitoring. InProceedings of SPIE, Vol. 3657. SPIE, San Jose, CA, USA, 103–112
1999
-
[15]
Shi, and Yan Lin
Xiangui Kang, Jiwu Huang, Yun Q. Shi, and Yan Lin. 2003. A DWT-DFT Composite Watermarking Scheme Robust to Both Affine Transform and JPEG Compression. IEEE Transactions on Circuits and Systems for Video Technology 13, 8 (2003), 776–
2003
-
[16]
Xiangui Kang, Jiwu Huang, and Wenjun Zeng. 2010. Efficient General Print- Scanning Resilient Data Hiding Based on Uniform Log-Polar Mapping. IEEE Transactions on Information Forensics and Security 5, 1 (2010), 1–12. doi:10.1109/ TIFS.2009.2039604
2010
-
[17]
Bhattacharjee, and Touradj Ebrahimi
Martin Kutter, Sushil K. Bhattacharjee, and Touradj Ebrahimi. 1999. Towards Second Generation Watermarking Schemes. In Proceedings 1999 International Conference on Image Processing (Cat. 99CH36348) , Vol. 1. IEEE, Kobe, Japan, 320– 323
1999
-
[18]
Bloom, Ingemar J
Ching-Yung Lin, Min Wu, Jeffrey A. Bloom, Ingemar J. Cox, Matt L. Miller, and Yui Man Lui. 2001. Rotation, Scale, and Translation Resilient Watermarking for Images. IEEE Transactions on Image Processing 10, 5 (2001), 767–782. doi:10.1109/ 83.918569
2001
-
[19]
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps. In Advances in Neural Information Processing Systems , Vol. 35. Curran Associates, Inc., New Orleans, LA, ...
2022
-
[20]
Chun-Shien Lu, Shih-Wei Sun, Chao-Yong Hsu, and Pao-Chi Chang. 2006. Media Hash-Dependent Image Watermarking Resilient Against Both Geometric At- tacks and Estimation Attacks Based on False Positive-Oriented Detection. IEEE Transactions on Multimedia 8, 4 (2006), 668–685. doi:...
2006
-
[21]
Zhiyuan Ma, Guoli Jia, Biqing Qi, and Bowen Zhou. 2024. Safe-sd: Safe and traceable stable diffusion with text prompt trigger for invisible generative water- marking. In Proceedings of the 32nd ACM International Conference on Multimedia . 7113–7122
2024
-
[22]
Tambiama Madiega. 2023. Generative AI and watermarking. European Parliamen- tary Research Service. https://www.europarl.europa.eu/RegData/etudes/BRIE/ 2023/757583/EPRS_BRI(2023)757583_EN.pdf
2023
-
[23]
National Assembly of South Korea [n. d.]. Content Industry Promotion Act
-
[24]
Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved Denoising Diffusion Probabilistic Models. In Proceedings of the 38th International Conference on Machine Learning, Vol. 139. PMLR, Virtual, 8162–8171
2021
-
[25]
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. 2022. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In Proceedings of the 39th International Confe...
2022
-
[26]
Ò Ruanaidh and Thierry Pun
Joseph J.K. Ò Ruanaidh and Thierry Pun. 1997. Rotation, Translation and Scale Invariant Digital Image Watermarking. In Proceedings of International Conference on Image Processing, Vol. 1. IEEE, Santa Barbara, CA, USA, 536–539
1997
-
[27]
Shelby Pereira and Thierry Pun. 2000. Robust Template Matching for Affine Resistant Image Watermarks. IEEE Transactions on Image Processing 9, 6 (2000), 1123–1129. doi:10.1109/83.846253
2000 doi
-
[28]
Petitcolas
Fabien A.P. Petitcolas. 2000. Watermarking Schemes Evaluation. IEEE Signal Processing Magazine 17, 5 (2000), 58–64. doi:10.1109/79.879339
2000 doi
-
[29]
Petitcolas, Ross J
Fabien A.P. Petitcolas, Ross J. Anderson, and Markus G. Kuhn. 1998. Attacks on Copyright Marking Systems. In Information Hiding, Vol. 1. Springer Berlin Heidelberg, Portland, OR, USA, 218–238
1998
-
[30]
Mao Po-Yuan, Shashank Kotyan, Tham Yik Foong, and Danilo Vasconcellos Vargas. 2023. Synthetic Shifts to Initial Seed Vector Exposes the Brittle Nature of Latent-Based Diffusion Models. arXiv:2312.11473 [cs.CV] https://arxiv.org/abs/ 2312.11473
2023 arXiv
-
[31]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, New Orleans, LA, USA, 10674–10685
2022
-
[32]
Gustavo Santana. 2022. Stable-Diffusion-Prompts. Hugging Face. https: //huggingface.co/datasets/Gustavosta/Stable-Diffusion-Prompts
2022
-
[33]
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2022. Denoising Diffusion Implicit Models. arXiv:2010.02502 [cs.LG] https://arxiv.org/abs/2010.02502
2022 arXiv
-
[34]
Chih-Wei Tang and Hsueh-Ming Hang. 2003. A Feature-Based Robust Digital Image Watermarking Scheme. IEEE Transactions on Signal Processing 51, 4 (2003), 950–959. doi:10.1109/TSP.2003.809367
2003
-
[35]
Huawei Tian, Yao Zhao, Rongrong Ni, Lunming Qin, and Xuelong Li. 2013. LDFT-Based Watermarking Resilient to Local Desynchronization Attacks. IEEE Transactions on Cybernetics 43, 6 (2013), 2190–2201. doi:10.1109/TCYB.2013. 2245415
2013 doi
-
[36]
Sviatoslav Voloshynovskiy, Frédéric Deguillaume, and Thierry Pun. 2001. Multi- bit digital watermarking robust against local nonlinear geometrical distor- tions. In Proceedings 2001 International Conference on Image Processing (Cat. No. 01CH37205), Vol. 3. IEEE, Thessaloniki, ...
2001
-
[37]
Yuxin Wen, John Kirchenbauer, Jonas Geiping, and Tom Goldstein. 2023. Tree- Rings Watermarks: Invisible Fingerprints for Diffusion Images. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36...
2023
-
[38]
Shijun Xiang, Hyoung Joong Kim, and Jiwu Huang. 2008. Invariant Image Watermarking Based on Statistical Features in the Low-Frequency Domain. IEEE Transactions on Circuits and Systems for Video Technology 18, 6 (2008), 777–790. doi:10.1109/TCSVT.2008.918843
2008
-
[39]
Rui Xu, Mengya Hu, Deren Lei, Yaxi Li, David Lowe, Alex Gorevski, Mingyu Wang, Emily Ching, and Alex Deng. 2024. InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance. arXiv:2411.07795 [cs.CR] https://arxiv.org/abs/2411.07795
2024 arXiv
-
[40]
Zijin Yang, Kai Zeng, Kejiang Chen, Han Fang, Weiming Zhang, and Nenghai Yu
-
[41]
Kevin Alex Zhang, Lei Xu, Alfredo Cuesta-Infante, and Kalyan Veera- machaneni. 2019. Robust Invisible Video Watermarking with Attention. arXiv:1909.01285 [cs.MM] https://arxiv.org/abs/1909.01285
2019 arXiv
-
[42]
Zheng, J
D. Zheng, J. Zhao, and A. El Saddik. 2003. RST-invariant digital image wa- termarking based on log-polar mapping and phase correlation. IEEE Trans- actions on Circuits and Systems for Video Technology 13, 8 (2003), 753–765. doi:10.1109/TCSVT.2003.815959
2003
-
[43]
Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. 2018. HiDDeN: Hiding Data With Deep Networks. In Computer Vision – ECCV 2018 , Vol. 15. Springer International Publishing, Munich, Germany, 682–697. MaXsive: High-Capacity and Robust Training-Free Generative Image Wate...
2018
-
[46]
Similarly, the top boundary of 𝑆 is at 𝑦 = 𝑛; thus, for 𝑆 to be contained in𝑅, we must have 𝑛≤ 𝑁 2 sin𝜃−𝑛 cos𝜃 2 sin𝜃 +𝑛 2, which implies that 𝑛≤ 𝑁 sin𝜃+ cos𝜃
sin𝜃 = 𝑁 2 =⇒ 𝑦 = 𝑁 2 sin𝜃−𝑛 cos𝜃 2 sin𝜃 +𝑛 2. Similarly, the top boundary of 𝑆 is at 𝑦 = 𝑛; thus, for 𝑆 to be contained in𝑅, we must have 𝑛≤ 𝑁 2 sin𝜃−𝑛 cos𝜃 2 sin𝜃 +𝑛 2, which implies that 𝑛≤ 𝑁 sin𝜃+ cos𝜃. Since a similar derivation holds when considering the constraints from...
2025
-
[2024]
In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Gaussian Shading: Provable Performance-Lossless Image Watermarking for MM ’25, October 27–31, 2025, Dublin, Ireland Po-Yuan Mao, Cheng-Chang Tsai, and Chun-Shien Lu Diffusion Models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, Seattle, ...
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.