A theoretical framework decouples diffusion model generation from watermark decisions, enabling SSB to reach any security-robustness-fidelity regime without model-specific empirical tests.
Video seal: Open and efficient video watermarking
10 Pith papers cite this work. Polarity classification is still indexing.
abstract
The proliferation of AI-generated content and sophisticated video editing tools has made it both important and challenging to moderate digital platforms. Video watermarking addresses these challenges by embedding imperceptible signals into videos, allowing for identification. However, the rare open tools and methods often fall short on efficiency, robustness, and flexibility. To reduce these gaps, this paper introduces Video Seal, a comprehensive framework for neural video watermarking and a competitive open-sourced model. Our approach jointly trains an embedder and an extractor, while ensuring the watermark robustness by applying transformations in-between, e.g., video codecs. This training is multistage and includes image pre-training, hybrid post-training and extractor fine-tuning. We also introduce temporal watermark propagation, a technique to convert any image watermarking model to an efficient video watermarking model without the need to watermark every high-resolution frame. We present experimental results demonstrating the effectiveness of the approach in terms of speed, imperceptibility, and robustness. Video Seal achieves higher robustness compared to strong baselines especially under challenging distortions combining geometric transformations and video compression. Additionally, we provide new insights such as the impact of video compression during training, and how to compare methods operating on different payloads. Contributions in this work - including the codebase, models, and a public demo - are open-sourced under permissive licenses to foster further research and development in the field.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
GIFGuard is the first spatiotemporal watermarking framework for proactive deepfake forensics in facial GIFs, using a 3D adaptive residual encoder and hourglass decoder plus a new GIFfaces dataset.
LAVA is a layered audio-visual watermarking system using cross-modal fusion and calibration-aware alignment to achieve robust deepfake tamper detection and localization under compression and asynchrony.
RAW benchmark shows existing watermark methods fail on avatar post-processing like background removal; new WALT method reaches 92.4% robustness on zoom and 95.6% on background removal.
CAT trains watermark detectors against adaptive compositional adversaries using differentiable attack selection, yielding up to 63.5% capacity gains on hard attacks versus random-augmentation baselines.
ISTS watermarking dynamically controls injection based on prompt semantics and uses two-sided detection to resist removal and forgery attacks in diffusion models.
Guidance watermarking steers diffusion denoising steps via gradients from an off-the-shelf watermark decoder to embed marks during generation, converting post-hoc schemes into in-generation ones while remaining complementary to VAE modifications.
FlowMark learns content-adaptive spatial masks for video watermark embedding, achieving 50+ dB PSNR, 128-bit capacity, and robustness to compression, temporal edits, and social media pipelines.
Watermark removal leaves detectable forensic artifacts, so no current method balances attack success, perceptual quality, and undetectability.
The paper analyzes evolving security and safety threats in generative AI from content generation to agentic actions, noting that attack surfaces expand faster than defenses and that many safeguards require institutional coordination not yet in place.
citing papers explorer
-
Secure Seed-Based Multi-bit Watermarking for Diffusion Models from First Principles
A theoretical framework decouples diffusion model generation from watermark decisions, enabling SSB to reach any security-robustness-fidelity regime without model-specific empirical tests.
-
GIFGuard: Proactive Forensics against Deepfakes in Facial GIFs via Spatiotemporal Watermarking
GIFGuard is the first spatiotemporal watermarking framework for proactive deepfake forensics in facial GIFs, using a 3D adaptive residual encoder and hourglass decoder plus a new GIFfaces dataset.
-
LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization
LAVA is a layered audio-visual watermarking system using cross-modal fusion and calibration-aware alignment to achieve robust deepfake tamper detection and localization under compression and asynchrony.
-
RAW: Robust Avatar Watermarking -- Benchmarking and Baseline
RAW benchmark shows existing watermark methods fail on avatar post-processing like background removal; new WALT method reaches 92.4% robustness on zoom and 95.6% on background removal.
-
Compositional Adversarial Training for Robust Visual Watermarking
CAT trains watermark detectors against adaptive compositional adversaries using differentiable attack selection, yielding up to 63.5% capacity gains on hard attacks versus random-augmentation baselines.
-
Towards Robust Content Watermarking Against Removal and Forgery Attacks
ISTS watermarking dynamically controls injection based on prompt semantics and uses two-sided detection to resist removal and forgery attacks in diffusion models.
-
Guidance Watermarking for Diffusion Models
Guidance watermarking steers diffusion denoising steps via gradients from an off-the-shelf watermark decoder to embed marks during generation, converting post-hoc schemes into in-generation ones while remaining complementary to VAE modifications.
-
FlowMark: Mask-Guided Video Watermarking
FlowMark learns content-adaptive spatial masks for video watermark embedding, achieving 50+ dB PSNR, 128-bit capacity, and robustness to compression, temporal edits, and social media pipelines.
-
The Forensic Cost of Watermark Removal: From Dedicated Attacks to Image Editing
Watermark removal leaves detectable forensic artifacts, so no current method balances attack success, perceptual quality, and undetectability.
-
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
The paper analyzes evolving security and safety threats in generative AI from content generation to agentic actions, noting that attack surfaces expand faster than defenses and that many safeguards require institutional coordination not yet in place.