A distribution-level watermark for categorical data, embedded by secret hashing and detected by inverting the hash mixture and comparing total variation distance to the original distribution.
Adaptive and Robust Watermark for Generative Tabular Data
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In recent years, watermarking generative tabular data has become a prominent framework to protect against the misuse of synthetic data. However, while most prior work in watermarking methods for tabular data demonstrate a wide variety of desirable properties (e.g., high fidelity, detectability, robustness), the findings often emphasize empirical guarantees against common oblivious and adversarial attacks. In this paper, we study a flexible and robust watermarking algorithm for generative tabular data. Specifically, we demonstrate theoretical guarantees on the performance of the algorithm on metrics like fidelity, detectability, robustness, and hardness of decoding. The proof techniques introduced in this work may be of independent interest and may find applicability in other areas of machine learning. Finally, we validate our theoretical findings on synthetic and real-world tabular datasets.
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Watermarking Generative Categorical Data
A distribution-level watermark for categorical data, embedded by secret hashing and detected by inverting the hash mixture and comparing total variation distance to the original distribution.